In Amazon SQS, a redrive policy moves a message to the dead-letter queue (DLQ) once its receive count exceeds maxReceiveCount, and a redrive moves messages back to the source queue after a fix (Amazon SQS dead-letter queues). Set the DLQ’s retention longer than the source queue’s: in a standard queue a moved message keeps its original enqueue timestamp, so it can expire before anyone looks.
A DLQ answers the question a job-queue or streaming problem ends with: what happens to the message that keeps failing? Separate transient failures, retried with backoff, from permanent ones such as a malformed payload, which go to the DLQ at once with the error, the attempt count and the original payload attached. If order matters, as in a FIFO queue or a Kafka partition, dead-lettering one message lets later messages for the same key overtake it, so park the rest of that key too or make the handler tolerate reordering.
Decide who is alerted when the DLQ is not empty, and how a replay avoids repeating side effects, with an idempotency key. In a customer integration, the DLQ and a replay command are what let the customer’s operations team fix bad records without calling you.
Related: backpressure, exponential backoff.