retries
Retries
Learn how retries recover from temporary failures, why immediate or unlimited retries can overload a distributed system, how exponential backoff and jitter spread repeated attempts, and how retry classification, idempotency, timeouts, deadlines, retry budgets, dead-letter handling, circuit breakers, observability, and platform-specific behaviour make retries safe.
Introduction
Distributed systems communicate through networks, databases, APIs, queues, event logs, and external services. Any request can fail temporarily even when the request and the target service are normally valid.
Application sends request
|
v
Temporary network interruption
|
v
Request fails
|
v
A later attempt can succeed
A retry repeats a failed operation because another attempt has a reasonable chance of succeeding.
Retries are useful for temporary failures such as short network interruptions, temporary service unavailability, rate limiting, broker leadership changes, and transient database conflicts.
Retries are dangerous when the failure is permanent, the operation is not safe to repeat, or many clients retry simultaneously.
Core idea: A retry is another request that consumes processing time, connections, memory, network capacity, and downstream resources. Retry only failures that can plausibly succeed later, and keep every retry policy bounded, delayed, randomized, observable, and safe to execute again.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Timeouts and deadlines | A retry should begin only after the current attempt has ended or become unusable. |
| 2 | Delivery semantics | Retries can turn one logical operation into several physical deliveries. |
| 3 | Idempotency | A completed operation can be repeated when its response is lost. |
| 4 | Queues and consumer groups | Message redelivery and offset recovery are forms of retry. |
| 5 | Backpressure | Retries add traffic to a dependency that can already be overloaded. |
| 6 | Circuit breakers | Repeated failures can require temporary rejection rather than more attempts. |
| 7 | Observability | Original calls and retry attempts must be measured separately. |
What Is a Retry?
A retry is a new attempt to perform an operation after an earlier attempt failed or returned an uncertain result.
Attempt 1
|
v
Temporary failure
|
v
Wait according to retry policy
|
v
Attempt 2
|
v
Success
The retry should normally preserve the same logical operation identity. Generating a new business identifier for every attempt can bypass duplicate protection.
Retryable vs Non-Retryable Failures
| Failure Category | Typical Direction | Reasoning |
|---|---|---|
| Temporary network interruption | Retry within the remaining deadline | A later connection can succeed. |
| Temporary service unavailability | Retry with backoff and jitter | The service can recover between attempts. |
| Rate limiting | Respect server guidance and retry later | An immediate attempt can remain throttled. |
| Transient database deadlock | Retry the complete transaction using bounded logic | The conflicting transaction can complete before the next attempt. |
| Invalid request | Do not retry unchanged | The same invalid input will fail again. |
| Authentication failure | Do not blindly retry | Repeated use of invalid credentials does not correct the problem. |
| Authorization failure | Do not retry without a relevant state change | The caller still lacks permission. |
| Schema-validation failure | Reject or dead-letter | The message needs correction rather than immediate repetition. |
| Permanent business rule failure | Return or record the failure | Time does not make the operation valid. |
| Unknown outcome | Investigate, query status, or retry idempotently | The original operation might already have succeeded. |
Classification rule: Do not retry every exception. A permanent validation, authentication, authorization, or configuration error normally requires correction rather than repetition.
Immediate Retry
An immediate retry starts another attempt without a meaningful delay.
Request fails
|
v
Retry immediately
|
v
Retry immediately
|
v
Retry immediately
Immediate retry can be useful for a rare connection race or another narrowly defined transient failure. Repeated immediate retries against an overloaded dependency usually make the failure worse.
Fixed-delay Retry
A fixed-delay policy waits for the same duration before every retry.
Attempt 1 fails
Wait D
Attempt 2 fails
Wait D
Attempt 3 fails
Fixed delays are simple, but many clients failing together can retry together at the same fixed interval.
Exponential Backoff
Exponential backoff increases the delay after every failure.
A conceptual formula is:
\[ Delay_n = \min \left( MaximumDelay, BaseDelay \times Multiplier^n \right) \]
Where:
- \(n\) is the retry number
- \(BaseDelay\) is the initial delay
- \(Multiplier\) controls delay growth
- \(MaximumDelay\) caps the delay
Conceptual retry schedule:
First retry:
Base delay
Second retry:
Longer delay
Third retry:
Longer delay again
Later retries:
Delay remains capped
at the configured maximum.
Backoff reduces repeated pressure and gives the dependency time to recover.
Jitter
Jitter introduces controlled randomness into retry delays.
Without jitter:
Client A retries at Time X
Client B retries at Time X
Client C retries at Time X
With jitter:
Client A retries near Time X1
Client B retries near Time X2
Client C retries near Time X3
Jitter reduces synchronization among clients after a shared failure.
A conceptual full-jitter formula is:
\[ RetryDelay_n = Random \left( 0, \min \left( MaximumDelay, BaseDelay \times Multiplier^n \right) \right) \]
Jitter rule: Exponential backoff reduces retry frequency. Jitter prevents many clients from following the same retry schedule.
Retry Storm
A retry storm occurs when repeated client attempts add enough load to delay or prevent dependency recovery.
Service becomes overloaded
|
v
Requests fail
|
v
Every client retries immediately
|
v
Service receives more traffic
|
v
More requests fail
|
v
More retries are generated
Retries can amplify the original workload.
If one original operation performs \(r\) retries, the maximum number of physical attempts is:
\[ TotalAttempts = 1 + r \]
If several application layers retry independently, amplification can multiply.
Retry Multiplication
API Gateway retries
Application Service retries
Database Client retries
Result:
One original request can create
many dependency calls.
Choose one appropriate retry layer for each dependency boundary and understand retries already provided by SDKs, proxies, service meshes, cloud services, and message brokers.
Retry Amplification
When \(L\) layers each make up to \(A_i\) total attempts, the maximum end-to-end call multiplication can be approximated as:
\[ MaximumAttemptMultiplication = \prod_{i=1}^{L} A_i \]
This upper bound illustrates why independent retries at every layer can rapidly increase traffic.
Retry Budget
A retry budget limits how much additional work retries are allowed to create.
A retry budget can constrain:
- Maximum retry attempts per operation
- Maximum elapsed retry time
- Maximum retry traffic relative to original traffic
- Maximum retries per tenant
- Maximum concurrent retrying operations
- Maximum cost of an external operation
Retry requested
|
v
Is failure retryable?
|
+-- No:
| stop
|
+-- Yes:
is retry budget available?
|
+-- No:
| stop or dead-letter
|
+-- Yes:
wait with jitter
retry
Deadline and Time Budget
Every retry must fit within the caller's useful deadline.
A conceptual total retry budget is:
\[ TotalElapsedTime = \sum AttemptDurations + \sum RetryDelays \]
Starting another attempt is not useful when insufficient time remains for that attempt to complete and return a usable result.
Define:
- Connection timeout
- Per-attempt timeout
- Maximum delay
- Maximum attempts
- Total operation deadline
Idempotency
A retry can repeat an operation that already succeeded when the response was lost.
Client sends enrollment request
|
v
Server commits enrollment
|
v
Response is lost
|
v
Client retries
|
v
Without idempotency:
Duplicate enrollment can be created.
A stable idempotency key lets the server identify repeated attempts for the same logical operation.
{
"operationId": "stable-unique-operation-id",
"tenantId": 17,
"learnerId": 1042,
"courseId": 42,
"operation": "enroll"
}
Idempotency rule: Every retry of one logical write should use the same operation identity. Generating a new identity on every attempt defeats duplicate detection.
Database-backed Idempotency
CREATE TABLE idempotent_operations
(
operation_id VARCHAR(150) PRIMARY KEY,
operation_type VARCHAR(100) NOT NULL,
request_hash VARCHAR(150) NOT NULL,
result_reference VARCHAR(150) NULL,
operation_status VARCHAR(30) NOT NULL,
created_at TIMESTAMP NOT NULL,
completed_at TIMESTAMP NULL
);
The operation record and business update should use an appropriate transactional boundary. A simple check followed by an unrelated insert can still race with another attempt.
Idempotent Processing Flow
Begin transaction
|
v
Insert stable operation ID
using unique constraint
|
+-- Operation already exists:
| return stored outcome
|
+-- New operation:
perform business update
store result
|
v
Commit transaction
Unknown Outcomes
A timeout does not prove that an operation failed.
Client calls payment service
|
v
Payment service completes charge
|
v
Response is lost
|
v
Client receives timeout
|
v
Outcome is unknown,
not necessarily failed.
Possible controls include:
- Retry with the same provider-supported idempotency key
- Query operation status using the stable business identifier
- Record an uncertain state
- Run reconciliation before another charge
- Require manual review for unresolved high-risk operations
Retry Strategies Comparison
| Strategy | Behaviour | Main Risk |
|---|---|---|
| Immediate retry | Retry without a meaningful wait | Increases pressure on a failing dependency |
| Fixed delay | Wait the same amount between attempts | Clients can retry in synchronized waves |
| Linear backoff | Increase delay by a fixed amount | Can remain too aggressive during sustained failures |
| Exponential backoff | Increase delay multiplicatively | Deterministic clients can remain synchronized |
| Exponential backoff with jitter | Increase delays and randomize retry times | Still requires limits, classification, and idempotency |
| Server-directed retry | Use valid server timing guidance | Requires correct interpretation and an overall deadline |
Circuit Breaker and Retries
A circuit breaker prevents repeated calls to a dependency that is failing consistently.
Closed:
Calls are allowed
Failure threshold reached
|
v
Open:
Calls fail fast
without dependency call
Recovery interval passes
|
v
Half-open:
Limited probe calls
Dependency succeeds
|
v
Closed
Retries help with isolated transient failures. A circuit breaker limits repeated load during persistent failures.
Retries and Backpressure
A dependency experiencing overload does not need more immediate traffic.
Combine retries with:
- Rate limiting
- Concurrency limiting
- Bounded queues
- Load shedding
- Circuit breakers
- Caller deadlines
- Dependency health and saturation signals
Message Queue Retries
A queue consumer can reject, abandon, or fail to acknowledge a message, causing later delivery according to the platform policy.
Queue delivers message
|
v
Consumer processing fails
|
v
Classify failure
|
+-- Temporary:
| retry using delayed path
|
+-- Permanent:
| dead-letter
|
+-- Success:
acknowledge or delete
Immediate requeueing can create a tight failure loop. Prefer a delayed retry path, bounded attempts, and a terminal failure destination.
Kafka Consumer Retries
A Kafka consumer can fail before committing its offset, causing a record to be read again after restart or reassignment.
Read event
|
v
Processing fails
|
+-- Retry in consumer:
| partition progress waits
|
+-- Publish to retry topic:
| primary consumer can continue
|
+-- Publish to dead-letter topic:
preserve terminal failure
Moving events to another topic can alter ordering. Define whether later events for the same key are allowed to proceed before the failed event.
RabbitMQ Retries
RabbitMQ consumers can negatively acknowledge or reject deliveries according to the queue topology and retry policy.
RabbitMQ delivery
|
v
Consumer fails
|
+-- Requeue immediately:
| risk of tight retry loop
|
+-- Route to delayed retry queue:
| message returns later
|
+-- Route to dead-letter queue:
terminal failure handling
Publisher confirms and consumer acknowledgments cover separate parts of the message path. Retrying publication does not prove business processing failed.
Amazon SQS Retries
An SQS message becomes available again when the visibility timeout expires without successful deletion.
Consumer receives message
|
v
Message becomes invisible
|
v
Consumer fails or does not delete
|
v
Visibility timeout expires
|
v
Message becomes available again
Configure the visibility timeout to match the processing design. A timeout that is too short can cause concurrent duplicate processing. A timeout that is much too long can delay recovery.
Dead-letter Handling
A dead-letter destination stores messages that cannot be processed within the approved retry policy.
Primary processing
|
v
Bounded attempts exhausted
|
v
Dead-letter destination
|
v
Investigate failure
|
v
Correct data or consumer
|
v
Controlled redrive
Preserve:
- Stable message ID
- Original destination
- Attempt count
- Failure classification
- Relevant schema version
- Correlation information
- Safe redrive status
Poison Messages
A poison message fails consistently because of invalid data, incompatible schema, missing business context, or a deterministic consumer defect.
Poison message
|
v
Consumer fails
|
v
Retry without correction
|
v
Consumer fails again
|
v
Useful processing is delayed.
Retrying a permanent error indefinitely wastes capacity and can block ordered partition or queue progress.
Conceptual Retry Policy
retryPolicy:
classification:
retryable:
- temporary-network-failure
- timeout-with-safe-replay
- service-unavailable
- rate-limited
- transient-concurrency-conflict
nonRetryable:
- authentication-failure
- authorization-failure
- validation-failure
- incompatible-schema
- permanent-business-rule-failure
attempts:
maximumAttempts: approved-limit
timing:
strategy: capped-exponential-backoff
jitter: enabled
maximumDelay: approved-delay
totalDeadline: approved-time-budget
respectServerGuidance: true
safety:
idempotencyKey: required-for-writes
retryAtOneLayer: true
terminalFailure:
deadLetterDestination: approved-failure-path
observability:
originalRequests: enabled
retryAttempts: enabled
retrySuccesses: enabled
exhaustedRetries: enabled
duplicateDetections: enabled
Java Retry Example
public final class RetryExecutor {
private final int maximumAttempts;
private final long baseDelayMillis;
private final long maximumDelayMillis;
public RetryExecutor(
int maximumAttempts,
long baseDelayMillis,
long maximumDelayMillis
) {
this.maximumAttempts = maximumAttempts;
this.baseDelayMillis = baseDelayMillis;
this.maximumDelayMillis = maximumDelayMillis;
}
public <T> T execute(
RetryableOperation<T> operation,
RetryClassifier classifier
) throws Exception {
Exception lastFailure = null;
for (
int attempt = 1;
attempt <= maximumAttempts;
attempt++
) {
try {
return operation.run();
} catch (Exception failure) {
lastFailure = failure;
boolean lastAttempt =
attempt == maximumAttempts;
if (
lastAttempt
|| !classifier.isRetryable(failure)
) {
throw failure;
}
long delayCap =
Math.min(
maximumDelayMillis,
baseDelayMillis
* (1L << (attempt - 1))
);
long delayWithJitter =
java.util.concurrent.ThreadLocalRandom
.current()
.nextLong(
delayCap + 1
);
Thread.sleep(
delayWithJitter
);
}
}
throw lastFailure;
}
}
This is an educational example. Production code should also enforce a total deadline, honor cancellation and server guidance, prevent overflow, use approved resilience libraries where appropriate, classify provider-specific failures, and propagate observability context.
PHP Retry Example
<?php
declare(strict_types=1);
final class RetryExecutor
{
public function execute(
callable $operation,
callable $isRetryable,
int $maximumAttempts,
int $baseDelayMilliseconds,
int $maximumDelayMilliseconds
): mixed {
$lastFailure = null;
for (
$attempt = 1;
$attempt <= $maximumAttempts;
$attempt++
) {
try {
return $operation();
} catch (Throwable $failure) {
$lastFailure = $failure;
$isLastAttempt =
$attempt === $maximumAttempts;
if (
$isLastAttempt
|| !$isRetryable($failure)
) {
throw $failure;
}
$delayCap =
min(
$maximumDelayMilliseconds,
$baseDelayMilliseconds
* (2 ** ($attempt - 1))
);
$delayWithJitter =
random_int(
0,
$delayCap
);
usleep(
$delayWithJitter * 1000
);
}
}
throw $lastFailure;
}
}
Do not hold a database transaction, row lock, HTTP connection, or other scarce resource while sleeping between retries.
Transaction Retry
Retrying a transaction requires rerunning the complete state-dependent calculation.
Transaction reads state
|
v
Calculates new value
|
v
Update conflict occurs
|
v
Retry must:
- start a new transaction
- reread current state
- recalculate the result
- attempt the update again
Reusing a calculation based on stale data can apply an incorrect business result.
Transaction rule: A transaction retry must repeat all reads and calculations that depend on transactional state. Do not retry only the final write.
External Side Effects inside Retried Code
Begin transaction
Send external email
Database update conflicts
Retry transaction
Send external email again
External side effects cannot normally be rolled back with the database transaction.
Move external publication after the local commit or record it through a transactional outbox.
Transactional Outbox
Begin local transaction
|
+-- Update business data
+-- Insert outbox record
|
v
Commit
|
v
Outbox publisher retries
message publication safely
The outbox makes publication recoverable without repeatedly executing the original business transaction and external message send together.
Retry vs Fallback vs Circuit Breaker
| Pattern | Purpose |
|---|---|
| Retry | Repeat an operation because another attempt can succeed. |
| Fallback | Return an alternative response or use another approved dependency. |
| Circuit breaker | Temporarily stop calls to a persistently failing dependency. |
| Timeout | Bound how long one attempt can consume resources. |
| Bulkhead | Prevent one dependency or workload from exhausting shared resources. |
| Dead-letter handling | Preserve terminally failed asynchronous work for investigation. |
Learning-platform Examples
| Workflow | Retry Direction | Safety Control |
|---|---|---|
| Course-catalog read | Retry a temporary read failure within the request deadline | Capped backoff, jitter, and cancellation |
| Learner enrollment | Retry only with the same enrollment operation ID | Unique enrollment and idempotency constraints |
| Certificate generation | Retry failed queue processing | Unique learner-course certificate constraint |
| Enrollment email | Use bounded delayed retries | Stable notification ID and delivery deduplication |
| Search-index update | Retry transient indexing failures | Document version and idempotent upsert |
| Payment-backed enrollment | Retry only using the same payment idempotency key | Status lookup and reconciliation for uncertain outcomes |
| Invalid event schema | Do not repeatedly retry unchanged data | Dead-letter and controlled correction |
Security Considerations
- Do not retry invalid authentication credentials indefinitely.
- Do not retry authorization failures without a relevant permission change.
- Preserve the same trusted operation identity across attempts.
- Do not expose credentials or private payloads in retry logs.
- Apply retry limits by trusted tenant or caller identity where required.
- Protect dead-letter destinations with least-privilege access.
- Validate messages again before controlled redrive.
- Audit retries of security-sensitive operations.
Observability
Useful retry metrics include:
- Original request count
- Total attempt count
- Retry-attempt count
- Retry rate
- Success-after-retry count
- Retries exhausted
- Retry delay distribution
- Failures by retry classification
- Duplicate-operation detections
- Dead-letter message count
- Retry traffic by dependency
- Retry traffic by tenant
- Time spent in retry waits
- Circuit-breaker state changes
- Uncertain external outcomes
Structured Retry Event
{
"operationId": "stable-operation-id",
"dependency": "certificate-service",
"attempt": 3,
"maximumAttempts": 4,
"failureCategory": "temporary-unavailable",
"retryDecision": "retry",
"backoffStrategy": "capped-exponential-with-jitter"
}
Avoid logging secrets, bearer tokens, payment data, or complete sensitive request payloads.
Alert Conditions
Alert when:
- Retry rate increases unexpectedly
- Retries remain elevated after the original incident ends
- Retry attempts exceed the approved budget
- Retries-exhausted events increase
- One dependency receives concentrated retry traffic
- Duplicate-operation detections increase
- Dead-letter volume grows
- Circuit breakers remain open
- Retry delays consume most of the request deadline
- Several application layers retry the same dependency
- Rate-limit responses continue during retries
- Retry traffic causes dependency saturation
Troubleshooting Workflow
- Identify the original logical operation and stable operation ID.
- Identify every layer performing retries.
- Check the maximum attempts and total deadline.
- Check whether the failure is transient, permanent, or uncertain.
- Check whether the operation is safe to repeat.
- Check whether all attempts use the same idempotency key.
- Check backoff, jitter, and maximum-delay behaviour.
- Check server-provided retry guidance.
- Check circuit-breaker and rate-limit state.
- Check downstream CPU, connections, latency, and saturation.
- Check retries already provided by SDKs or infrastructure.
- Check dead-letter and poison-message handling.
- Check duplicate and reconciliation records.
- Reduce, disable, or relocate retries through the approved configuration.
Common Retry Mistakes
Retrying Every Error
Validation, authentication, authorization, and deterministic business errors continue failing while consuming more resources.
Immediate Retry without Backoff
The retry adds load while the dependency is still failing or overloaded.
Exponential Backoff without Jitter
Clients failing together can continue retrying in synchronized waves.
Unlimited Retry
A permanent failure creates an endless loop, growing backlog and cost.
Retrying at Every Layer
Nested retry policies multiply the number of physical dependency calls.
Retrying Non-Idempotent Writes
The first attempt can succeed while the retry creates another real-world effect.
Generating a New Operation ID per Attempt
The server treats every retry as a new business operation.
Ignoring the Overall Deadline
A retry begins when insufficient time remains to produce a useful result.
Retrying while Holding Locks
The waiting client continues blocking other work and increases contention.
Sending External Side Effects inside a Retried Transaction
A transaction retry repeats an email, message, payment, or HTTP request that cannot be rolled back with the database.
Retrying Poison Messages in Place Forever
One invalid message consumes capacity and can prevent useful progress.
Monitoring Only Final Failures
Successful final responses can hide a growing dependence on retries and increasing load.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Temporary network failure | A later bounded attempt succeeds. |
| Permanent validation failure | The operation fails without repeated attempts. |
| Rate-limited response | The retry respects valid server timing guidance. |
| Shared dependency outage | Jitter spreads retry attempts over time. |
| Maximum attempts reached | The operation stops and follows the terminal failure policy. |
| Overall deadline reached | No additional attempt begins after the useful budget is exhausted. |
| Response lost after successful write | The repeated operation returns the existing outcome. |
| Concurrent duplicate attempts | A unique constraint allows one business effect. |
| Transaction deadlock | The complete transaction rereads and recalculates state. |
| Poison message | Bounded processing moves it to the approved failure path. |
| Nested retry configuration | Total attempt multiplication remains within the approved budget. |
| Dependency saturation | Backoff, circuit breaking, and concurrency limits protect recovery. |
Retry Best Practices
Recommended Practices
- Retry only failures that are plausibly transient.
- Do not retry unchanged validation, authentication, or authorization failures.
- Use capped exponential backoff.
- Add jitter to prevent synchronized retry waves.
- Set a maximum number of attempts.
- Set a total elapsed-time or deadline budget.
- Honor valid server retry guidance.
- Use one stable idempotency key across all attempts.
- Make message consumers and write operations idempotent.
- Use database constraints to protect business uniqueness.
- Retry at one deliberate architectural layer.
- Understand retry behaviour already supplied by SDKs and infrastructure.
- Do not sleep while holding locks or scarce resources.
- Reread and recalculate state when retrying transactions.
- Move external side effects outside retried database transactions.
- Use a transactional outbox for reliable message publication.
- Use bounded retries and dead-letter handling for messages.
- Combine retries with backpressure and circuit breakers.
- Measure original requests and retry attempts separately.
- Test uncertain outcomes, retry exhaustion, and retry storms.
Practice Exercise
Design retry policies for the asynchronous and synchronous workflows in your online learning platform.
Requirements
- List every operation currently retried.
- Identify retries already supplied by SDKs and infrastructure.
- Classify errors as transient, permanent, or uncertain.
- Define maximum attempts for every dependency.
- Add capped exponential backoff and jitter.
- Define a total operation deadline.
- Add stable idempotency keys to enrollment writes.
- Add a unique constraint preventing duplicate enrollment.
- Move notification publication to an outbox.
- Add bounded retry queues for email delivery.
- Add dead-letter handling for invalid messages.
- Retry database deadlocks by rerunning the complete transaction.
- Introduce a shared dependency outage.
- Verify that retry traffic remains within the approved budget.
- Monitor success-after-retry and retries-exhausted separately.
Retry-design Template
| Operation | Retryable Failure | Replay Protection | Terminal Action |
|---|---|---|---|
| Course-catalog read | Temporary connection or service failure | Read is naturally repeatable | Return controlled fallback or failure |
| Learner enrollment | Temporary infrastructure failure | Stable operation ID and unique enrollment constraint | Return the known result or unresolved status |
| Certificate generation | Temporary worker or storage failure | Unique certificate identity | Dead-letter after bounded attempts |
| Email notification | Temporary provider failure | Stable notification ID | Failure queue and operator review |
| Database deadlock | Transient concurrency conflict | New transaction with reread and recalculation | Return conflict after bounded attempts |
| Payment request | Timeout with uncertain outcome | Provider idempotency key and status reconciliation | Record uncertain state for controlled resolution |
Frequently Asked Questions
What is a retry?
A retry is a new attempt to perform an operation after an earlier attempt failed or returned an uncertain result.
Which failures should be retried?
Retry only failures that are plausibly temporary and operations that are safe to repeat within the remaining deadline.
What is exponential backoff?
Exponential backoff increases the delay between later retry attempts, normally with a configured maximum delay.
What is jitter?
Jitter adds controlled randomness to retry delays so clients do not all retry at the same time.
What is a retry storm?
A retry storm occurs when repeated attempts add enough load to delay or prevent dependency recovery.
What is a retry budget?
A retry budget limits attempts, elapsed time, concurrency, traffic, or cost created by retries.
Why is idempotency required?
The original operation can succeed while its response is lost, causing a later attempt to repeat an already completed operation.
Should every application layer retry?
No. Retries at several layers can multiply calls. Choose one deliberate layer for each dependency boundary.
Should a database deadlock be retried?
A transient deadlock can be retried using bounded logic, but the complete transaction must reread current state and recalculate dependent results.
Should invalid messages be retried?
Not repeatedly without correction. Invalid or incompatible messages should be rejected or moved to a controlled failure path.
How do retries interact with circuit breakers?
Retries recover from isolated transient failures, while circuit breakers temporarily stop calls during persistent failure.
What is the safest general retry policy?
Classify the failure, verify idempotency, enforce a deadline and retry budget, use capped exponential backoff with jitter, and stop through a controlled terminal failure path.
Key Takeaway
Retries recover from temporary failures by repeating an operation, but every retry adds load and can repeat a business effect. Retry only errors that can plausibly succeed later, and do not retry permanent validation, authentication, authorization, or schema failures unchanged. Use capped exponential backoff to reduce retry frequency and jitter to prevent synchronized retry waves. Bound the number of attempts, individual delays, total elapsed time, concurrency, and aggregate retry traffic. Preserve one stable idempotency key across all attempts, because a timeout can mean that the original operation succeeded and only its response was lost. Avoid retries at several architectural layers, do not hold locks while waiting, and rerun all state-dependent calculations when retrying a transaction. For asynchronous messages, use delayed retries, bounded attempts, and dead-letter handling. Finally, combine retries with timeouts, backpressure, circuit breakers, observability, and tested recovery procedures.