DLQs
Dead-Letter Queues (DLQs)
Learn how dead-letter queues isolate messages that cannot be processed successfully, how DLQs work with bounded retries, and how failure classification, metadata preservation, monitoring, investigation, correction, redrive, idempotency, ordering, retention, security, and platform-specific behaviour in Kafka, RabbitMQ, and Amazon SQS support reliable asynchronous processing.
Introduction
Asynchronous consumers can fail while processing messages. Some failures are temporary and can succeed after a delayed retry. Other failures are permanent until the message, consumer, schema, configuration, or business data is corrected.
Producer publishes message
|
v
Primary queue or topic
|
v
Consumer receives message
|
v
Processing fails
Retrying every failed message indefinitely can consume processing capacity, increase costs, delay healthy messages, and create an endless failure loop.
Silently deleting the failed message is also unsafe when the operation is important.
A dead-letter queue, commonly abbreviated as DLQ, provides a separate destination for messages that cannot be completed through the normal processing and retry policy.
Primary destination
|
v
Consumer processing
|
+-- Success:
| complete message
|
+-- Temporary failure:
| delayed retry
|
+-- Attempts exhausted
or permanent failure:
send to DLQ
Core idea: A DLQ isolates terminally failed messages so healthy work can continue. A DLQ does not solve the failure automatically. Every DLQ needs monitoring, ownership, investigation, correction, controlled redrive, and retention policies.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Queues vs logs | A DLQ can be implemented as a queue or a dedicated failure topic. |
| 2 | Kafka, RabbitMQ, and Amazon SQS | Each platform handles failed-message routing differently. |
| 3 | Delivery semantics | At-least-once processing can deliver the same message repeatedly. |
| 4 | Retries | Retryable messages should normally use bounded retries before terminal handling. |
| 5 | Idempotency | Redriving a message can repeat a previously completed side effect. |
| 6 | Ordering | Removing one failed message from an ordered stream can allow later messages to overtake it. |
| 7 | Observability | DLQ depth, message age, repeated failures, and redrive outcomes require monitoring. |
What Is a Dead-Letter Queue?
A dead-letter queue is a separate message destination used to preserve messages that cannot be processed successfully through the approved primary and retry paths.
Message enters primary queue
|
v
Consumer attempts processing
|
v
Processing fails
|
v
Retry count increases
|
v
Maximum attempts reached
|
v
Message moves to DLQ
The destination can be called:
- Dead-letter queue
- Dead-letter topic
- Failure queue
- Error queue
- Parking-lot queue
- Quarantine destination
The exact term and behaviour depend on the messaging platform and application architecture.
Primary Queue, Retry Queue, and DLQ
| Destination | Purpose | Expected Message State |
|---|---|---|
| Primary queue or topic | Normal business processing | The message is expected to succeed under ordinary conditions. |
| Retry queue or topic | Delay another attempt after a temporary failure | The message can plausibly succeed without manual correction. |
| Dead-letter queue | Isolate a terminal or exhausted failure | The message requires investigation, correction, controlled replay, or disposal. |
Separation rule: A retry queue is an automated recovery path. A DLQ is a terminal investigation path. Do not use the DLQ as an unbounded additional retry tier.
Why Messages Enter a DLQ
| Failure | Example | Likely Direction |
|---|---|---|
| Malformed payload | The message is not valid JSON | DLQ without repeated unchanged retries |
| Schema incompatibility | The consumer does not support the message version | DLQ and contract investigation |
| Missing required field | A business identifier is absent | DLQ or producer correction |
| Permanent business-rule failure | The referenced account does not permit the operation | DLQ, rejection, or business exception workflow |
| Temporary dependency failure | A database or API is briefly unavailable | Delayed retry before DLQ |
| Retry exhaustion | A temporary-looking failure remains unresolved | DLQ after the configured attempt limit |
| Deserialization failure | The consumer cannot construct the expected object | DLQ with original payload preserved safely |
| Consumer defect | A valid message triggers an unhandled exception | DLQ until the consumer is corrected |
| Message expiration | The operation is no longer useful to the requester | Drop or DLQ according to business policy |
| Routing failure | No valid destination is available | Dead-letter routing where supported |
Poison Messages
A poison message is a message that repeatedly fails for a deterministic reason.
Consumer receives message
|
v
Message causes deterministic failure
|
v
Consumer retries unchanged message
|
v
Same deterministic failure occurs
|
v
Partition or queue progress is delayed
Common poison-message causes include:
- Invalid JSON or binary encoding
- Unsupported schema version
- Incorrect data type
- Missing mandatory identifier
- Unexpected null value
- Corrupted payload
- Permanent business validation failure
- Unhandled consumer code path
A poison message should not be retried indefinitely. The message should be isolated while retaining enough safe information for investigation.
Failure Classification
Message processing fails
|
v
Classify failure
|
+-- Transient:
| delayed retry
|
+-- Permanent:
| DLQ
|
+-- Unknown:
bounded retry
then DLQ if unresolved
Accurate classification prevents permanent errors from consuming retry capacity and prevents temporary errors from being sent to manual investigation too early.
Bounded Retry before DLQ
Temporary errors should normally receive a limited number of retries using backoff and jitter.
Attempt 1 fails
|
v
Delayed retry
|
v
Attempt 2 fails
|
v
Longer delayed retry
|
v
Attempt limit reached
|
v
Send to DLQ
A simplified decision can be expressed as:
\[ Destination = \begin{cases} RetryPath, & AttemptCount < MaximumAttempts \\ DeadLetterQueue, & AttemptCount \geq MaximumAttempts \end{cases} \]
The maximum attempt limit must be selected from recovery probability, business urgency, operation cost, and downstream capacity rather than from an arbitrary universal number.
Delayed Retry
Immediate retries can create a tight loop when the dependency remains unavailable.
Failure
|
v
Immediate retry
|
v
Same failure
|
v
Immediate retry
|
v
Capacity is consumed
without progress.
A delayed retry gives temporary conditions time to recover and protects the consumer and downstream service.
Information to Preserve
A DLQ message should preserve enough context to understand and safely recover the failure.
| Field | Purpose |
|---|---|
| Message ID | Identifies the original logical message |
| Correlation ID | Connects the failure with the broader workflow |
| Original destination | Identifies the source queue, topic, or route |
| Original message key | Preserves entity routing and ordering context |
| Schema version | Identifies the payload contract used by the producer |
| Failure category | Classifies the error as transient, permanent, or unknown |
| Error code | Provides a stable machine-readable failure identifier |
| Safe error summary | Supports investigation without exposing unnecessary sensitive data |
| Attempt count | Shows how many processing attempts occurred |
| Failure timestamp | Supports age and incident analysis |
| Consumer version | Identifies the application version that failed |
| Trace reference | Connects the message to distributed tracing where available |
DLQ Envelope
{
"deadLetterId": "stable-dead-letter-id",
"originalMessageId": "stable-original-message-id",
"correlationId": "stable-correlation-id",
"originalDestination": "certificate-generation",
"originalMessageKey": "learner-1042",
"messageType": "GenerateCertificate",
"schemaVersion": 1,
"failureCategory": "schema-validation",
"errorCode": "MISSING_COURSE_ID",
"attemptCount": 4,
"failedAt": "failure-timestamp",
"consumerVersion": "approved-consumer-version",
"payloadReference": "protected-payload-reference"
}
When the payload contains sensitive information, storing a protected reference can be safer than copying the complete original payload into several destinations.
Preserve the Original Message
Editing the only stored copy of a failed message destroys evidence about the original failure.
Original failed message
|
v
Preserve immutable original
|
v
Create corrected replay message
with correction metadata
|
v
Redrive corrected copy
Corrections should be auditable. Preserve the relationship between the original message and the corrected or replayed message.
DLQ Investigation
Investigation should determine whether the failure originated in the producer, message contract, consumer, dependency, or business data.
- Identify the original message and correlation ID.
- Identify the primary destination and consumer.
- Review the failure code and category.
- Validate the message schema and required fields.
- Check whether the consumer version supports the schema.
- Check downstream dependency health during the failed attempts.
- Determine whether any business side effect already occurred.
- Identify whether the failure affects one message or a complete message class.
- Correct the producer, consumer, configuration, or business data.
- Verify the correction using a controlled test.
- Redrive only eligible messages.
- Confirm successful processing and close the failure record.
Correction Strategies
| Root Cause | Possible Correction |
|---|---|
| Producer created an invalid payload | Correct the producer and create a validated replacement message |
| Consumer does not support the schema | Deploy compatible consumer logic or transform through an approved process |
| Reference data was temporarily unavailable | Restore the dependency and redrive the original message |
| Consumer defect | Deploy the correction and replay affected messages |
| Permanent business rejection | Resolve through a business exception workflow rather than blind replay |
| Message is no longer useful | Dispose of it through the approved retention and audit policy |
What Is Redrive?
Redrive moves or republishes a failed message from the DLQ to a processing destination after the underlying problem has been corrected.
DLQ message
|
v
Validate redrive eligibility
|
v
Confirm consumer correction
|
v
Check duplicate side effects
|
v
Redrive to retry or source path
|
v
Monitor processing result
Redrive should be a controlled operation rather than an automatic loop from the DLQ back to the primary destination.
Redrive rule: Do not bulk-redrive a DLQ until the root cause is corrected and the consumer is safe to execute the messages again.
Redrive and Idempotency
A failed message can already have produced part or all of its intended side effect.
Consumer updates database
|
v
Database commit succeeds
|
v
Consumer fails before
acknowledgment
|
v
Message reaches failure path
|
v
Blind replay can repeat update.
Before redrive, determine whether:
- The original business transaction committed
- An external API accepted the request
- An email or notification was sent
- A file or certificate was generated
- A downstream event was published
Use stable message IDs, idempotency keys, database constraints, and status reconciliation to prevent duplicate outcomes.
DLQs and Ordering
Moving one failed record to a DLQ allows later records to continue. This can change business ordering.
Original ordered events:
Version 1
Version 2
Version 3
Version 2 fails and moves to DLQ.
Consumer processes:
Version 1
Version 3
Later redrive:
Version 2
Version 2 now arrives after Version 3.
Possible controls include:
- Version-aware conditional updates
- Per-entity sequence checks
- Blocking later events for the same entity
- Parking the complete entity stream
- Rebuilding state from an authoritative event history
- Rejecting stale messages during redrive
Ordering rule: A DLQ improves pipeline availability by isolating failed messages, but that isolation can weaken ordering. Define the expected behaviour for later messages belonging to the same entity.
Kafka Dead-letter Topics
Kafka applications commonly implement the DLQ pattern using a dedicated Kafka topic for failed records.
Source Topic
|
v
Kafka Consumer
|
+-- Success:
| commit processing progress
|
+-- Retryable failure:
| publish to retry topic
|
+-- Terminal failure:
publish to dead-letter topic
The consumer application, connector, or stream-processing framework must define when and how the record is published to the dead-letter topic.
The application must coordinate publication to the dead-letter topic with progress on the original partition. Otherwise, the failed event can be lost or repeatedly dead-lettered.
Kafka DLQ Envelope
{
"originalTopic": "learning-domain-events",
"originalPartition": 2,
"originalOffset": 8500,
"originalKey": "learner-1042",
"eventId": "stable-event-id",
"eventType": "CourseCompleted",
"failureCategory": "consumer-processing",
"errorCode": "UNSUPPORTED_EVENT_VERSION",
"attemptCount": 3,
"consumerGroup": "certificate-eligibility"
}
Kafka-specific Concerns
- Publishing to the DLQ and advancing the source offset must be coordinated safely.
- Moving the record out of the source partition can change per-key ordering.
- The DLQ topic requires its own retention and access policy.
- The DLQ schema must remain readable even when the original schema is invalid.
- Redrive must preserve the original key and event identity where required.
RabbitMQ Dead-letter Exchanges
RabbitMQ can route dead-lettered messages through a configured dead-letter exchange. Bindings on that exchange determine the destination queue.
Primary Exchange
|
v
Primary Queue
|
v
Consumer rejects or
negatively acknowledges
according to policy
|
v
Dead-letter Exchange
|
v
Dead-letter Queue
RabbitMQ retry designs commonly use a retry queue with delayed return to the primary exchange, followed by a terminal dead-letter queue after bounded attempts.
Primary Queue
|
v
Consumer failure
|
v
Retry Queue
|
v
Delay expires
|
v
Primary Exchange
|
v
Primary Queue
Attempts exhausted
|
v
Terminal DLQ
Immediate requeueing without a delay can create a rapid redelivery loop.
Amazon SQS Dead-letter Queues
Amazon SQS can associate a source queue with a dead-letter queue through a redrive policy.
Source SQS Queue
|
v
Consumer receives message
|
v
Processing fails
or message is not deleted
|
v
Visibility timeout expires
|
v
Message becomes available again
|
v
Receive count reaches limit
|
v
Message moves to DLQ
The source queue's redrive policy defines the destination and the receive count at which the message becomes eligible for movement.
The DLQ should be compatible with the source queue type and business ordering requirements.
Conceptual SQS Policy
sqs:
sourceQueue:
name: certificate-generation
type: standard-or-fifo
processing:
visibilityTimeout: approved-processing-boundary
deleteAfterDurableSuccess: true
redrive:
deadLetterQueue: certificate-generation-failures
maximumReceiveCount: approved-limit
deadLetterMonitoring:
messageCount: enabled
oldestMessageAge: enabled
notifications: enabled
redriveOperation:
authorization: restricted
rootCauseCorrectionRequired: true
idempotencyValidationRequired: true
Platform Comparison
| Area | Kafka | RabbitMQ | Amazon SQS |
|---|---|---|---|
| Failure destination | Dedicated dead-letter topic pattern | Dead-letter exchange routing to a queue | Associated SQS dead-letter queue |
| Routing decision | Consumer, connector, or processing framework | Consumer disposition and broker topology | Receive count and redrive policy |
| Retry path | Consumer retry or retry topics | Requeue or delayed retry queues | Visibility timeout and repeated receive |
| Ordering risk | Later partition events can continue without the failed record | Requeueing and multiple consumers can affect completion order | Standard queues are not strictly ordered; FIFO workflows require message-group analysis |
| Progress control | Source offset and DLQ publication | Acknowledgment, rejection, and dead-letter routing | Message receive count, visibility, and deletion |
| Redrive | Republish from DLQ topic through controlled tooling or application logic | Republish or route through approved recovery topology | Use an approved SQS redrive process |
DLQ Tracking Table
CREATE TABLE dead_letter_records
(
dead_letter_id VARCHAR(150) PRIMARY KEY,
original_message_id VARCHAR(150) NOT NULL,
correlation_id VARCHAR(150) NULL,
original_destination VARCHAR(200) NOT NULL,
failure_category VARCHAR(100) NOT NULL,
error_code VARCHAR(100) NOT NULL,
attempt_count INTEGER NOT NULL,
dead_letter_status VARCHAR(30) NOT NULL,
failed_at TIMESTAMP NOT NULL,
resolved_at TIMESTAMP NULL,
redriven_at TIMESTAMP NULL
);
This type of tracking table can support ownership, investigation, resolution, and redrive auditing. It should not replace the actual durable message destination unless explicitly designed to do so.
Java DLQ-routing Example
public final class MessageProcessor {
private final BusinessHandler businessHandler;
private final RetryPublisher retryPublisher;
private final DeadLetterPublisher deadLetterPublisher;
private final FailureClassifier failureClassifier;
public void process(
MessageEnvelope message
) {
try {
businessHandler.processIdempotently(
message
);
} catch (Exception failure) {
FailureCategory category =
failureClassifier.classify(
failure
);
if (
category.isRetryable()
&& message.getAttemptCount()
< message.getMaximumAttempts()
) {
retryPublisher.publish(
message.nextAttempt(),
failure
);
return;
}
deadLetterPublisher.publish(
DeadLetterEnvelope.from(
message,
failure,
category
)
);
}
}
}
This educational example omits platform-specific offset, acknowledgment, transaction, publication-confirmation, authorization, and error-handling details. Production code must ensure that routing to the retry or DLQ path cannot silently lose the original message.
PHP DLQ Decision Example
<?php
declare(strict_types=1);
final class FailureRouter
{
public function __construct(
private RetryPublisher $retryPublisher,
private DeadLetterPublisher $deadLetterPublisher
) {
}
public function route(
MessageEnvelope $message,
ProcessingFailure $failure
): void {
if (
$failure->isRetryable()
&&
$message->getAttemptCount()
< $message->getMaximumAttempts()
) {
$this->retryPublisher->publish(
$message->nextAttempt(),
$failure
);
return;
}
$this->deadLetterPublisher->publish(
DeadLetterEnvelope::fromFailure(
$message,
$failure
)
);
}
}
Redrive Authorization
Redriving DLQ messages can repeat business operations and expose sensitive data. Redrive access should therefore be restricted.
Use:
- Least-privilege access
- Authenticated operator or service identity
- Environment-specific permissions
- Approval for high-risk message classes
- Audited redrive operations
- Per-message or bounded-batch controls
- Separation between investigation and production execution permissions
Security and Privacy
DLQs can contain messages that ordinary processing failed to sanitize or transform.
- Encrypt DLQ data according to organizational policy.
- Restrict read and redrive access.
- Do not copy reusable credentials or tokens into failure metadata.
- Redact sensitive error details.
- Apply tenant isolation to DLQ access.
- Define retention and deletion requirements.
- Audit downloads, edits, exports, and redrives.
- Use protected payload references when appropriate.
DLQ Retention
A DLQ requires an explicit retention period.
The retention policy should consider:
- Time required to detect the failure
- Time required to investigate and deploy a correction
- Business recovery obligations
- Privacy and regulatory requirements
- Message-system retention limits
- Cost of retaining large payloads
- Required audit evidence
A short retention can delete failures before resolution. Unlimited retention creates uncontrolled storage and privacy risk.
DLQ Ownership
Every DLQ should have a documented operational owner.
The owner should be responsible for:
- Monitoring DLQ alerts
- Classifying failures
- Coordinating producer or consumer corrections
- Approving redrive
- Verifying successful recovery
- Managing retention and disposal
- Documenting recurring failure patterns
A DLQ without ownership becomes an unmonitored storage location rather than a recovery mechanism.
DLQ Metrics
Useful metrics include:
- Messages entering the DLQ
- Current DLQ depth
- Age of the oldest unresolved message
- Messages by failure category
- Messages by original destination
- Messages by producer and consumer version
- Messages awaiting investigation
- Messages corrected
- Messages redriven
- Redrive success and failure counts
- Messages disposed of
- Recurring message or error signatures
Monitoring rule: Monitor the age of unresolved DLQ messages as well as message count. A small DLQ containing one old critical message can require more attention than a larger recent batch.
Structured DLQ Event
{
"deadLetterId": "stable-dead-letter-id",
"originalMessageId": "stable-message-id",
"originalDestination": "learning-domain-events",
"failureCategory": "schema-incompatible",
"errorCode": "UNSUPPORTED_VERSION",
"attemptCount": 3,
"deadLetterStatus": "awaiting-investigation",
"redriveEligible": false
}
Routine logs should avoid the complete failed payload when identifiers, error codes, and protected references are sufficient.
Alert Conditions
Alert when:
- A new message enters a critical DLQ
- DLQ depth continues growing
- The oldest unresolved-message age exceeds the response objective
- One failure category increases suddenly
- One producer version generates repeated invalid messages
- One consumer version creates repeated processing failures
- Retry messages move rapidly to the DLQ
- Redrive failures increase
- Redriven messages return to the DLQ
- DLQ retention approaches message expiry
- No owner acknowledges a DLQ incident
- DLQ storage approaches capacity
Troubleshooting Workflow
- Identify the DLQ destination and affected environment.
- Identify the oldest and highest-priority unresolved messages.
- Identify the original queue, topic, partition, or route.
- Check the message ID and correlation ID.
- Review the failure category and stable error code.
- Validate the original schema and required message fields.
- Check producer and consumer application versions.
- Check retry attempts and delay history.
- Check downstream dependency health during the attempts.
- Determine whether the business side effect already occurred.
- Correct the message, producer, consumer, data, or dependency.
- Test the correction using a controlled message.
- Redrive a small approved batch.
- Monitor successful completion and duplicate detection.
- Complete the remaining redrive or approved disposal.
Common DLQ Mistakes
Not Configuring a DLQ
Failed messages remain in endless retry loops or disappear without an investigation path.
Using the DLQ as another Automatic Retry Queue
Messages cycle repeatedly without root-cause correction.
Sending Every First Failure Directly to the DLQ
Temporary failures create unnecessary manual work instead of using an appropriate delayed retry path.
Retrying Permanent Failures before DLQ
Invalid data consumes retry capacity even though repetition cannot correct the problem.
Monitoring Count but Not Message Age
Old critical failures remain unresolved while the overall DLQ depth looks small.
Bulk-redriving without Fixing the Root Cause
The same messages fail again and return to the DLQ.
Redriving without Idempotency
A message whose earlier side effect succeeded can produce a duplicate business result.
Ignoring Ordering during Redrive
An older event can overwrite or conflict with newer state.
Losing the Original Payload or Metadata
The team cannot reproduce, diagnose, or safely correct the failure.
Editing the Original DLQ Message
Failure evidence is destroyed and the correction is not auditable.
Allowing Unrestricted Redrive Access
An unauthorized or accidental action can repeat high-risk business operations.
Having No Operational Owner
Failed messages accumulate without investigation, correction, or disposal.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Temporary dependency failure | The message uses delayed retry and succeeds when the dependency recovers. |
| Malformed payload | The message reaches the DLQ without unlimited retries. |
| Unsupported schema | The DLQ preserves the original schema and failure details. |
| Retry exhaustion | The message moves to the DLQ at the approved attempt boundary. |
| Missing acknowledgment | Duplicate-safe processing prevents a repeated business effect. |
| DLQ publication failure | The original message is not silently lost. |
| Kafka DLQ routing | The failed record and original partition metadata are preserved. |
| RabbitMQ delayed retry | The message does not enter a tight immediate-requeue loop. |
| SQS receive-count exhaustion | The message moves to the configured SQS DLQ. |
| Consumer correction | A controlled redrive succeeds after the code fix. |
| Blind redrive attempt | Authorization or workflow controls block unapproved replay. |
| Out-of-order redrive | Version checks prevent stale state from replacing newer state. |
| Repeated redrive failure | The message remains visible and does not loop indefinitely. |
| Retention expiry | Alerts occur before an unresolved message is deleted. |
DLQ Best Practices
Recommended Practices
- Create a DLQ for every critical asynchronous processing path.
- Separate retry destinations from terminal DLQs.
- Classify failures before deciding whether to retry or dead-letter.
- Use bounded retries with backoff and jitter.
- Send permanent errors to the DLQ without wasteful repetition.
- Preserve the stable message ID and correlation ID.
- Preserve the original destination, key, schema version, and attempt count.
- Use stable machine-readable error codes.
- Protect sensitive payloads and failure metadata.
- Keep the original failed message immutable.
- Create corrected replay messages with audit references.
- Assign an operational owner to every DLQ.
- Monitor both DLQ depth and oldest unresolved-message age.
- Alert when critical messages enter the DLQ.
- Correct the root cause before redrive.
- Redrive a small controlled batch before bulk processing.
- Make consumers idempotent before replay.
- Check whether an external side effect already occurred.
- Protect ordering using versions or sequence checks.
- Define retention, disposal, audit, and access policies.
Practice Exercise
Design DLQ handling for the asynchronous workflows in your online learning platform.
Requirements
- Create a certificate-generation primary queue.
- Create a delayed retry path for temporary failures.
- Create a terminal certificate-generation DLQ.
- Add stable message and correlation identifiers.
- Preserve failure category, error code, and attempt count.
- Route malformed messages directly to controlled failure handling.
- Retry temporary storage failures with backoff and jitter.
- Move exhausted messages to the DLQ.
- Add a unique learner-course certificate constraint.
- Monitor DLQ depth and oldest-message age.
- Create alerts for new critical DLQ messages.
- Deploy a correction for one simulated consumer defect.
- Redrive one test message.
- Redrive a controlled batch.
- Verify that no duplicate certificates are created.
- Record redrive authorization and results.
DLQ-design Template
| Decision | Selected Direction | Primary Risk Controlled |
|---|---|---|
| Transient failure | Bounded delayed retry | Prevents temporary failures from requiring immediate manual work |
| Permanent failure | Route to terminal DLQ | Prevents infinite retry loops |
| Message identity | Stable message and correlation IDs | Supports tracing and duplicate detection |
| Failure metadata | Category, code, attempt count, source, and version | Supports accurate investigation |
| Monitoring | Depth, ingress rate, and oldest unresolved age | Prevents invisible failure accumulation |
| Redrive | Restricted and auditable controlled workflow | Prevents accidental re-execution |
| Duplicate protection | Idempotent consumer and database constraints | Prevents repeated business effects |
| Ordering | Version-aware conditional application | Prevents older redriven messages from overwriting newer state |
Frequently Asked Questions
What is a DLQ?
A dead-letter queue is a separate destination for messages that cannot be processed through the approved primary and retry paths.
When should a message enter a DLQ?
A message should enter a DLQ when it has a permanent failure or exhausts the approved bounded retry policy.
Is a DLQ the same as a retry queue?
No. A retry queue delays another automated attempt, while a DLQ isolates a terminal failure for investigation and controlled recovery.
What is a poison message?
A poison message is a message that repeatedly fails for a deterministic reason such as malformed data, an unsupported schema, or a consumer defect.
Should every failed message go directly to the DLQ?
No. Failures that are plausibly temporary should normally use an approved bounded retry path first.
What information should a DLQ preserve?
Preserve the original message identity, correlation information, destination, key, schema version, failure code, attempt count, timestamp, and safe diagnostic context.
What is redrive?
Redrive is the controlled movement or republication of corrected, eligible DLQ messages back into a processing path.
Why is idempotency required during redrive?
A previous processing attempt can have completed its business operation before the message entered the failure path.
Can a DLQ affect ordering?
Yes. Later events can continue while an earlier failed event remains in the DLQ, causing the earlier event to be processed later during redrive.
Does Kafka have a DLQ?
Kafka applications commonly implement the pattern using a dedicated dead-letter topic and consumer or framework error-handling logic.
How does Amazon SQS use a DLQ?
A source SQS queue can use a redrive policy that moves repeatedly received messages to an associated dead-letter queue after the configured receive boundary.
What should be monitored?
Monitor DLQ message count, ingress rate, oldest unresolved-message age, failure categories, redrive outcomes, and messages approaching retention expiry.
Key Takeaway
A dead-letter queue isolates messages that cannot be processed successfully through the normal and bounded retry paths. Temporary failures should use delayed retries with backoff and jitter, while permanent or exhausted failures should be preserved in the DLQ. Store enough safe metadata to identify the original message, destination, schema, attempt count, error category, and consumer version. Treat the DLQ as an operational recovery workflow rather than a message graveyard or another automatic retry tier. Assign an owner, monitor both depth and message age, correct the root cause, test the correction, and redrive a controlled batch before processing the complete backlog. Make consumers idempotent, verify whether external side effects already occurred, and protect ordering with version or sequence checks. Finally, restrict and audit redrive access, define retention and deletion policies, and test poison messages, DLQ-routing failure, repeated redrive failure, and retention expiry.