Table of Contents

    retries

    MESSAGING & ASYNCHRONOUS PROCESSING

    Retries

    Learn how retries recover from temporary failures, why immediate or unlimited retries can overload a distributed system, how exponential backoff and jitter spread repeated attempts, and how retry classification, idempotency, timeouts, deadlines, retry budgets, dead-letter handling, circuit breakers, observability, and platform-specific behaviour make retries safe.

    Introduction

    Distributed systems communicate through networks, databases, APIs, queues, event logs, and external services. Any request can fail temporarily even when the request and the target service are normally valid.

    Application sends request
          |
          v
    Temporary network interruption
          |
          v
    Request fails
          |
          v
    A later attempt can succeed

    A retry repeats a failed operation because another attempt has a reasonable chance of succeeding.

    Retries are useful for temporary failures such as short network interruptions, temporary service unavailability, rate limiting, broker leadership changes, and transient database conflicts.

    Retries are dangerous when the failure is permanent, the operation is not safe to repeat, or many clients retry simultaneously.

    Core idea: A retry is another request that consumes processing time, connections, memory, network capacity, and downstream resources. Retry only failures that can plausibly succeed later, and keep every retry policy bounded, delayed, randomized, observable, and safe to execute again.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Timeouts and deadlines A retry should begin only after the current attempt has ended or become unusable.
    2 Delivery semantics Retries can turn one logical operation into several physical deliveries.
    3 Idempotency A completed operation can be repeated when its response is lost.
    4 Queues and consumer groups Message redelivery and offset recovery are forms of retry.
    5 Backpressure Retries add traffic to a dependency that can already be overloaded.
    6 Circuit breakers Repeated failures can require temporary rejection rather than more attempts.
    7 Observability Original calls and retry attempts must be measured separately.

    What Is a Retry?

    A retry is a new attempt to perform an operation after an earlier attempt failed or returned an uncertain result.

    Attempt 1
          |
          v
    Temporary failure
          |
          v
    Wait according to retry policy
          |
          v
    Attempt 2
          |
          v
    Success

    The retry should normally preserve the same logical operation identity. Generating a new business identifier for every attempt can bypass duplicate protection.

    Safe Retry Flow
    operation fails → classify failure → verify replay safety → check remaining budget → wait with backoff and jitter → retry with same operation identity

    Retryable vs Non-Retryable Failures

    Failure Category Typical Direction Reasoning
    Temporary network interruption Retry within the remaining deadline A later connection can succeed.
    Temporary service unavailability Retry with backoff and jitter The service can recover between attempts.
    Rate limiting Respect server guidance and retry later An immediate attempt can remain throttled.
    Transient database deadlock Retry the complete transaction using bounded logic The conflicting transaction can complete before the next attempt.
    Invalid request Do not retry unchanged The same invalid input will fail again.
    Authentication failure Do not blindly retry Repeated use of invalid credentials does not correct the problem.
    Authorization failure Do not retry without a relevant state change The caller still lacks permission.
    Schema-validation failure Reject or dead-letter The message needs correction rather than immediate repetition.
    Permanent business rule failure Return or record the failure Time does not make the operation valid.
    Unknown outcome Investigate, query status, or retry idempotently The original operation might already have succeeded.

    Classification rule: Do not retry every exception. A permanent validation, authentication, authorization, or configuration error normally requires correction rather than repetition.

    Immediate Retry

    An immediate retry starts another attempt without a meaningful delay.

    Request fails
          |
          v
    Retry immediately
          |
          v
    Retry immediately
          |
          v
    Retry immediately

    Immediate retry can be useful for a rare connection race or another narrowly defined transient failure. Repeated immediate retries against an overloaded dependency usually make the failure worse.

    Fixed-delay Retry

    A fixed-delay policy waits for the same duration before every retry.

    Attempt 1 fails
    
    Wait D
    
    Attempt 2 fails
    
    Wait D
    
    Attempt 3 fails

    Fixed delays are simple, but many clients failing together can retry together at the same fixed interval.

    Exponential Backoff

    Exponential backoff increases the delay after every failure.

    A conceptual formula is:

    \[ Delay_n = \min \left( MaximumDelay, BaseDelay \times Multiplier^n \right) \]

    Where:

    • \(n\) is the retry number
    • \(BaseDelay\) is the initial delay
    • \(Multiplier\) controls delay growth
    • \(MaximumDelay\) caps the delay
    Conceptual retry schedule:
    
    First retry:
    
    Base delay
    
    
    Second retry:
    
    Longer delay
    
    
    Third retry:
    
    Longer delay again
    
    
    Later retries:
    
    Delay remains capped
    at the configured maximum.

    Backoff reduces repeated pressure and gives the dependency time to recover.

    Jitter

    Jitter introduces controlled randomness into retry delays.

    Without jitter:
    
    Client A retries at Time X
    Client B retries at Time X
    Client C retries at Time X
    
    
    With jitter:
    
    Client A retries near Time X1
    Client B retries near Time X2
    Client C retries near Time X3

    Jitter reduces synchronization among clients after a shared failure.

    A conceptual full-jitter formula is:

    \[ RetryDelay_n = Random \left( 0, \min \left( MaximumDelay, BaseDelay \times Multiplier^n \right) \right) \]

    Jitter rule: Exponential backoff reduces retry frequency. Jitter prevents many clients from following the same retry schedule.

    Retry Storm

    A retry storm occurs when repeated client attempts add enough load to delay or prevent dependency recovery.

    Service becomes overloaded
          |
          v
    Requests fail
          |
          v
    Every client retries immediately
          |
          v
    Service receives more traffic
          |
          v
    More requests fail
          |
          v
    More retries are generated

    Retries can amplify the original workload.

    If one original operation performs \(r\) retries, the maximum number of physical attempts is:

    \[ TotalAttempts = 1 + r \]

    If several application layers retry independently, amplification can multiply.

    Retry Multiplication

    API Gateway retries
    
    Application Service retries
    
    Database Client retries
    
    
    Result:
    
    One original request can create
    many dependency calls.

    Choose one appropriate retry layer for each dependency boundary and understand retries already provided by SDKs, proxies, service meshes, cloud services, and message brokers.

    Retry Amplification

    When \(L\) layers each make up to \(A_i\) total attempts, the maximum end-to-end call multiplication can be approximated as:

    \[ MaximumAttemptMultiplication = \prod_{i=1}^{L} A_i \]

    This upper bound illustrates why independent retries at every layer can rapidly increase traffic.

    Retry Budget

    A retry budget limits how much additional work retries are allowed to create.

    A retry budget can constrain:

    • Maximum retry attempts per operation
    • Maximum elapsed retry time
    • Maximum retry traffic relative to original traffic
    • Maximum retries per tenant
    • Maximum concurrent retrying operations
    • Maximum cost of an external operation
    Retry requested
          |
          v
    Is failure retryable?
          |
          +-- No:
          |      stop
          |
          +-- Yes:
                 is retry budget available?
                        |
                        +-- No:
                        |      stop or dead-letter
                        |
                        +-- Yes:
                               wait with jitter
                               retry

    Deadline and Time Budget

    Every retry must fit within the caller's useful deadline.

    A conceptual total retry budget is:

    \[ TotalElapsedTime = \sum AttemptDurations + \sum RetryDelays \]

    Starting another attempt is not useful when insufficient time remains for that attempt to complete and return a usable result.

    Define:

    • Connection timeout
    • Per-attempt timeout
    • Maximum delay
    • Maximum attempts
    • Total operation deadline

    Idempotency

    A retry can repeat an operation that already succeeded when the response was lost.

    Client sends enrollment request
          |
          v
    Server commits enrollment
          |
          v
    Response is lost
          |
          v
    Client retries
          |
          v
    Without idempotency:
    
    Duplicate enrollment can be created.

    A stable idempotency key lets the server identify repeated attempts for the same logical operation.

    {
      "operationId": "stable-unique-operation-id",
      "tenantId": 17,
      "learnerId": 1042,
      "courseId": 42,
      "operation": "enroll"
    }

    Idempotency rule: Every retry of one logical write should use the same operation identity. Generating a new identity on every attempt defeats duplicate detection.

    Database-backed Idempotency

    CREATE TABLE idempotent_operations
    (
        operation_id VARCHAR(150) PRIMARY KEY,
        operation_type VARCHAR(100) NOT NULL,
        request_hash VARCHAR(150) NOT NULL,
        result_reference VARCHAR(150) NULL,
        operation_status VARCHAR(30) NOT NULL,
        created_at TIMESTAMP NOT NULL,
        completed_at TIMESTAMP NULL
    );

    The operation record and business update should use an appropriate transactional boundary. A simple check followed by an unrelated insert can still race with another attempt.

    Idempotent Processing Flow

    Begin transaction
          |
          v
    Insert stable operation ID
    using unique constraint
          |
          +-- Operation already exists:
          |      return stored outcome
          |
          +-- New operation:
                 perform business update
                 store result
          |
          v
    Commit transaction

    Unknown Outcomes

    A timeout does not prove that an operation failed.

    Client calls payment service
          |
          v
    Payment service completes charge
          |
          v
    Response is lost
          |
          v
    Client receives timeout
          |
          v
    Outcome is unknown,
    not necessarily failed.

    Possible controls include:

    • Retry with the same provider-supported idempotency key
    • Query operation status using the stable business identifier
    • Record an uncertain state
    • Run reconciliation before another charge
    • Require manual review for unresolved high-risk operations

    Retry Strategies Comparison

    Strategy Behaviour Main Risk
    Immediate retry Retry without a meaningful wait Increases pressure on a failing dependency
    Fixed delay Wait the same amount between attempts Clients can retry in synchronized waves
    Linear backoff Increase delay by a fixed amount Can remain too aggressive during sustained failures
    Exponential backoff Increase delay multiplicatively Deterministic clients can remain synchronized
    Exponential backoff with jitter Increase delays and randomize retry times Still requires limits, classification, and idempotency
    Server-directed retry Use valid server timing guidance Requires correct interpretation and an overall deadline

    Circuit Breaker and Retries

    A circuit breaker prevents repeated calls to a dependency that is failing consistently.

    Closed:
    
    Calls are allowed
    
    
    Failure threshold reached
          |
          v
    Open:
    
    Calls fail fast
    without dependency call
    
    
    Recovery interval passes
          |
          v
    Half-open:
    
    Limited probe calls
    
    
    Dependency succeeds
          |
          v
    Closed

    Retries help with isolated transient failures. A circuit breaker limits repeated load during persistent failures.

    Retries and Backpressure

    A dependency experiencing overload does not need more immediate traffic.

    Combine retries with:

    • Rate limiting
    • Concurrency limiting
    • Bounded queues
    • Load shedding
    • Circuit breakers
    • Caller deadlines
    • Dependency health and saturation signals

    Message Queue Retries

    A queue consumer can reject, abandon, or fail to acknowledge a message, causing later delivery according to the platform policy.

    Queue delivers message
          |
          v
    Consumer processing fails
          |
          v
    Classify failure
          |
          +-- Temporary:
          |      retry using delayed path
          |
          +-- Permanent:
          |      dead-letter
          |
          +-- Success:
                 acknowledge or delete

    Immediate requeueing can create a tight failure loop. Prefer a delayed retry path, bounded attempts, and a terminal failure destination.

    Kafka Consumer Retries

    A Kafka consumer can fail before committing its offset, causing a record to be read again after restart or reassignment.

    Read event
          |
          v
    Processing fails
          |
          +-- Retry in consumer:
          |      partition progress waits
          |
          +-- Publish to retry topic:
          |      primary consumer can continue
          |
          +-- Publish to dead-letter topic:
                 preserve terminal failure

    Moving events to another topic can alter ordering. Define whether later events for the same key are allowed to proceed before the failed event.

    RabbitMQ Retries

    RabbitMQ consumers can negatively acknowledge or reject deliveries according to the queue topology and retry policy.

    RabbitMQ delivery
          |
          v
    Consumer fails
          |
          +-- Requeue immediately:
          |      risk of tight retry loop
          |
          +-- Route to delayed retry queue:
          |      message returns later
          |
          +-- Route to dead-letter queue:
                 terminal failure handling

    Publisher confirms and consumer acknowledgments cover separate parts of the message path. Retrying publication does not prove business processing failed.

    Amazon SQS Retries

    An SQS message becomes available again when the visibility timeout expires without successful deletion.

    Consumer receives message
          |
          v
    Message becomes invisible
          |
          v
    Consumer fails or does not delete
          |
          v
    Visibility timeout expires
          |
          v
    Message becomes available again

    Configure the visibility timeout to match the processing design. A timeout that is too short can cause concurrent duplicate processing. A timeout that is much too long can delay recovery.

    Dead-letter Handling

    A dead-letter destination stores messages that cannot be processed within the approved retry policy.

    Primary processing
          |
          v
    Bounded attempts exhausted
          |
          v
    Dead-letter destination
          |
          v
    Investigate failure
          |
          v
    Correct data or consumer
          |
          v
    Controlled redrive

    Preserve:

    • Stable message ID
    • Original destination
    • Attempt count
    • Failure classification
    • Relevant schema version
    • Correlation information
    • Safe redrive status

    Poison Messages

    A poison message fails consistently because of invalid data, incompatible schema, missing business context, or a deterministic consumer defect.

    Poison message
          |
          v
    Consumer fails
          |
          v
    Retry without correction
          |
          v
    Consumer fails again
          |
          v
    Useful processing is delayed.

    Retrying a permanent error indefinitely wastes capacity and can block ordered partition or queue progress.

    Conceptual Retry Policy

    retryPolicy:
      classification:
        retryable:
          - temporary-network-failure
          - timeout-with-safe-replay
          - service-unavailable
          - rate-limited
          - transient-concurrency-conflict
    
        nonRetryable:
          - authentication-failure
          - authorization-failure
          - validation-failure
          - incompatible-schema
          - permanent-business-rule-failure
    
      attempts:
        maximumAttempts: approved-limit
    
      timing:
        strategy: capped-exponential-backoff
        jitter: enabled
        maximumDelay: approved-delay
        totalDeadline: approved-time-budget
        respectServerGuidance: true
    
      safety:
        idempotencyKey: required-for-writes
        retryAtOneLayer: true
    
      terminalFailure:
        deadLetterDestination: approved-failure-path
    
      observability:
        originalRequests: enabled
        retryAttempts: enabled
        retrySuccesses: enabled
        exhaustedRetries: enabled
        duplicateDetections: enabled

    Java Retry Example

    public final class RetryExecutor {
    
        private final int maximumAttempts;
        private final long baseDelayMillis;
        private final long maximumDelayMillis;
    
        public RetryExecutor(
            int maximumAttempts,
            long baseDelayMillis,
            long maximumDelayMillis
        ) {
            this.maximumAttempts = maximumAttempts;
            this.baseDelayMillis = baseDelayMillis;
            this.maximumDelayMillis = maximumDelayMillis;
        }
    
        public <T> T execute(
            RetryableOperation<T> operation,
            RetryClassifier classifier
        ) throws Exception {
    
            Exception lastFailure = null;
    
            for (
                int attempt = 1;
                attempt <= maximumAttempts;
                attempt++
            ) {
                try {
                    return operation.run();
                } catch (Exception failure) {
                    lastFailure = failure;
    
                    boolean lastAttempt =
                        attempt == maximumAttempts;
    
                    if (
                        lastAttempt
                        || !classifier.isRetryable(failure)
                    ) {
                        throw failure;
                    }
    
                    long delayCap =
                        Math.min(
                            maximumDelayMillis,
                            baseDelayMillis
                                * (1L << (attempt - 1))
                        );
    
                    long delayWithJitter =
                        java.util.concurrent.ThreadLocalRandom
                            .current()
                            .nextLong(
                                delayCap + 1
                            );
    
                    Thread.sleep(
                        delayWithJitter
                    );
                }
            }
    
            throw lastFailure;
        }
    }

    This is an educational example. Production code should also enforce a total deadline, honor cancellation and server guidance, prevent overflow, use approved resilience libraries where appropriate, classify provider-specific failures, and propagate observability context.

    PHP Retry Example

    <?php
    
    declare(strict_types=1);
    
    final class RetryExecutor
    {
        public function execute(
            callable $operation,
            callable $isRetryable,
            int $maximumAttempts,
            int $baseDelayMilliseconds,
            int $maximumDelayMilliseconds
        ): mixed {
            $lastFailure = null;
    
            for (
                $attempt = 1;
                $attempt <= $maximumAttempts;
                $attempt++
            ) {
                try {
                    return $operation();
                } catch (Throwable $failure) {
                    $lastFailure = $failure;
    
                    $isLastAttempt =
                        $attempt === $maximumAttempts;
    
                    if (
                        $isLastAttempt
                        || !$isRetryable($failure)
                    ) {
                        throw $failure;
                    }
    
                    $delayCap =
                        min(
                            $maximumDelayMilliseconds,
                            $baseDelayMilliseconds
                                * (2 ** ($attempt - 1))
                        );
    
                    $delayWithJitter =
                        random_int(
                            0,
                            $delayCap
                        );
    
                    usleep(
                        $delayWithJitter * 1000
                    );
                }
            }
    
            throw $lastFailure;
        }
    }

    Do not hold a database transaction, row lock, HTTP connection, or other scarce resource while sleeping between retries.

    Transaction Retry

    Retrying a transaction requires rerunning the complete state-dependent calculation.

    Transaction reads state
          |
          v
    Calculates new value
          |
          v
    Update conflict occurs
          |
          v
    Retry must:
    
    - start a new transaction
    - reread current state
    - recalculate the result
    - attempt the update again

    Reusing a calculation based on stale data can apply an incorrect business result.

    Transaction rule: A transaction retry must repeat all reads and calculations that depend on transactional state. Do not retry only the final write.

    External Side Effects inside Retried Code

    Begin transaction
    
    Send external email
    
    Database update conflicts
    
    Retry transaction
    
    Send external email again

    External side effects cannot normally be rolled back with the database transaction.

    Move external publication after the local commit or record it through a transactional outbox.

    Transactional Outbox

    Begin local transaction
          |
          +-- Update business data
          +-- Insert outbox record
          |
          v
    Commit
          |
          v
    Outbox publisher retries
    message publication safely

    The outbox makes publication recoverable without repeatedly executing the original business transaction and external message send together.

    Retry vs Fallback vs Circuit Breaker

    Pattern Purpose
    Retry Repeat an operation because another attempt can succeed.
    Fallback Return an alternative response or use another approved dependency.
    Circuit breaker Temporarily stop calls to a persistently failing dependency.
    Timeout Bound how long one attempt can consume resources.
    Bulkhead Prevent one dependency or workload from exhausting shared resources.
    Dead-letter handling Preserve terminally failed asynchronous work for investigation.

    Learning-platform Examples

    Workflow Retry Direction Safety Control
    Course-catalog read Retry a temporary read failure within the request deadline Capped backoff, jitter, and cancellation
    Learner enrollment Retry only with the same enrollment operation ID Unique enrollment and idempotency constraints
    Certificate generation Retry failed queue processing Unique learner-course certificate constraint
    Enrollment email Use bounded delayed retries Stable notification ID and delivery deduplication
    Search-index update Retry transient indexing failures Document version and idempotent upsert
    Payment-backed enrollment Retry only using the same payment idempotency key Status lookup and reconciliation for uncertain outcomes
    Invalid event schema Do not repeatedly retry unchanged data Dead-letter and controlled correction

    Security Considerations

    • Do not retry invalid authentication credentials indefinitely.
    • Do not retry authorization failures without a relevant permission change.
    • Preserve the same trusted operation identity across attempts.
    • Do not expose credentials or private payloads in retry logs.
    • Apply retry limits by trusted tenant or caller identity where required.
    • Protect dead-letter destinations with least-privilege access.
    • Validate messages again before controlled redrive.
    • Audit retries of security-sensitive operations.

    Observability

    Useful retry metrics include:

    • Original request count
    • Total attempt count
    • Retry-attempt count
    • Retry rate
    • Success-after-retry count
    • Retries exhausted
    • Retry delay distribution
    • Failures by retry classification
    • Duplicate-operation detections
    • Dead-letter message count
    • Retry traffic by dependency
    • Retry traffic by tenant
    • Time spent in retry waits
    • Circuit-breaker state changes
    • Uncertain external outcomes

    Structured Retry Event

    {
      "operationId": "stable-operation-id",
      "dependency": "certificate-service",
      "attempt": 3,
      "maximumAttempts": 4,
      "failureCategory": "temporary-unavailable",
      "retryDecision": "retry",
      "backoffStrategy": "capped-exponential-with-jitter"
    }

    Avoid logging secrets, bearer tokens, payment data, or complete sensitive request payloads.

    Alert Conditions

    Alert when:

    • Retry rate increases unexpectedly
    • Retries remain elevated after the original incident ends
    • Retry attempts exceed the approved budget
    • Retries-exhausted events increase
    • One dependency receives concentrated retry traffic
    • Duplicate-operation detections increase
    • Dead-letter volume grows
    • Circuit breakers remain open
    • Retry delays consume most of the request deadline
    • Several application layers retry the same dependency
    • Rate-limit responses continue during retries
    • Retry traffic causes dependency saturation

    Troubleshooting Workflow

    1. Identify the original logical operation and stable operation ID.
    2. Identify every layer performing retries.
    3. Check the maximum attempts and total deadline.
    4. Check whether the failure is transient, permanent, or uncertain.
    5. Check whether the operation is safe to repeat.
    6. Check whether all attempts use the same idempotency key.
    7. Check backoff, jitter, and maximum-delay behaviour.
    8. Check server-provided retry guidance.
    9. Check circuit-breaker and rate-limit state.
    10. Check downstream CPU, connections, latency, and saturation.
    11. Check retries already provided by SDKs or infrastructure.
    12. Check dead-letter and poison-message handling.
    13. Check duplicate and reconciliation records.
    14. Reduce, disable, or relocate retries through the approved configuration.

    Common Retry Mistakes

    1

    Retrying Every Error

    Validation, authentication, authorization, and deterministic business errors continue failing while consuming more resources.

    2

    Immediate Retry without Backoff

    The retry adds load while the dependency is still failing or overloaded.

    3

    Exponential Backoff without Jitter

    Clients failing together can continue retrying in synchronized waves.

    4

    Unlimited Retry

    A permanent failure creates an endless loop, growing backlog and cost.

    5

    Retrying at Every Layer

    Nested retry policies multiply the number of physical dependency calls.

    6

    Retrying Non-Idempotent Writes

    The first attempt can succeed while the retry creates another real-world effect.

    7

    Generating a New Operation ID per Attempt

    The server treats every retry as a new business operation.

    8

    Ignoring the Overall Deadline

    A retry begins when insufficient time remains to produce a useful result.

    9

    Retrying while Holding Locks

    The waiting client continues blocking other work and increases contention.

    10

    Sending External Side Effects inside a Retried Transaction

    A transaction retry repeats an email, message, payment, or HTTP request that cannot be rolled back with the database.

    11

    Retrying Poison Messages in Place Forever

    One invalid message consumes capacity and can prevent useful progress.

    12

    Monitoring Only Final Failures

    Successful final responses can hide a growing dependence on retries and increasing load.

    Recommended Test Cases

    Test Expected Evidence
    Temporary network failure A later bounded attempt succeeds.
    Permanent validation failure The operation fails without repeated attempts.
    Rate-limited response The retry respects valid server timing guidance.
    Shared dependency outage Jitter spreads retry attempts over time.
    Maximum attempts reached The operation stops and follows the terminal failure policy.
    Overall deadline reached No additional attempt begins after the useful budget is exhausted.
    Response lost after successful write The repeated operation returns the existing outcome.
    Concurrent duplicate attempts A unique constraint allows one business effect.
    Transaction deadlock The complete transaction rereads and recalculates state.
    Poison message Bounded processing moves it to the approved failure path.
    Nested retry configuration Total attempt multiplication remains within the approved budget.
    Dependency saturation Backoff, circuit breaking, and concurrency limits protect recovery.

    Retry Best Practices

    Recommended Practices

    • Retry only failures that are plausibly transient.
    • Do not retry unchanged validation, authentication, or authorization failures.
    • Use capped exponential backoff.
    • Add jitter to prevent synchronized retry waves.
    • Set a maximum number of attempts.
    • Set a total elapsed-time or deadline budget.
    • Honor valid server retry guidance.
    • Use one stable idempotency key across all attempts.
    • Make message consumers and write operations idempotent.
    • Use database constraints to protect business uniqueness.
    • Retry at one deliberate architectural layer.
    • Understand retry behaviour already supplied by SDKs and infrastructure.
    • Do not sleep while holding locks or scarce resources.
    • Reread and recalculate state when retrying transactions.
    • Move external side effects outside retried database transactions.
    • Use a transactional outbox for reliable message publication.
    • Use bounded retries and dead-letter handling for messages.
    • Combine retries with backpressure and circuit breakers.
    • Measure original requests and retry attempts separately.
    • Test uncertain outcomes, retry exhaustion, and retry storms.

    Practice Exercise

    Design retry policies for the asynchronous and synchronous workflows in your online learning platform.

    Requirements

    1. List every operation currently retried.
    2. Identify retries already supplied by SDKs and infrastructure.
    3. Classify errors as transient, permanent, or uncertain.
    4. Define maximum attempts for every dependency.
    5. Add capped exponential backoff and jitter.
    6. Define a total operation deadline.
    7. Add stable idempotency keys to enrollment writes.
    8. Add a unique constraint preventing duplicate enrollment.
    9. Move notification publication to an outbox.
    10. Add bounded retry queues for email delivery.
    11. Add dead-letter handling for invalid messages.
    12. Retry database deadlocks by rerunning the complete transaction.
    13. Introduce a shared dependency outage.
    14. Verify that retry traffic remains within the approved budget.
    15. Monitor success-after-retry and retries-exhausted separately.

    Retry-design Template

    Operation Retryable Failure Replay Protection Terminal Action
    Course-catalog read Temporary connection or service failure Read is naturally repeatable Return controlled fallback or failure
    Learner enrollment Temporary infrastructure failure Stable operation ID and unique enrollment constraint Return the known result or unresolved status
    Certificate generation Temporary worker or storage failure Unique certificate identity Dead-letter after bounded attempts
    Email notification Temporary provider failure Stable notification ID Failure queue and operator review
    Database deadlock Transient concurrency conflict New transaction with reread and recalculation Return conflict after bounded attempts
    Payment request Timeout with uncertain outcome Provider idempotency key and status reconciliation Record uncertain state for controlled resolution

    Frequently Asked Questions

    1

    What is a retry?

    A retry is a new attempt to perform an operation after an earlier attempt failed or returned an uncertain result.

    2

    Which failures should be retried?

    Retry only failures that are plausibly temporary and operations that are safe to repeat within the remaining deadline.

    3

    What is exponential backoff?

    Exponential backoff increases the delay between later retry attempts, normally with a configured maximum delay.

    4

    What is jitter?

    Jitter adds controlled randomness to retry delays so clients do not all retry at the same time.

    5

    What is a retry storm?

    A retry storm occurs when repeated attempts add enough load to delay or prevent dependency recovery.

    6

    What is a retry budget?

    A retry budget limits attempts, elapsed time, concurrency, traffic, or cost created by retries.

    7

    Why is idempotency required?

    The original operation can succeed while its response is lost, causing a later attempt to repeat an already completed operation.

    8

    Should every application layer retry?

    No. Retries at several layers can multiply calls. Choose one deliberate layer for each dependency boundary.

    9

    Should a database deadlock be retried?

    A transient deadlock can be retried using bounded logic, but the complete transaction must reread current state and recalculate dependent results.

    10

    Should invalid messages be retried?

    Not repeatedly without correction. Invalid or incompatible messages should be rejected or moved to a controlled failure path.

    11

    How do retries interact with circuit breakers?

    Retries recover from isolated transient failures, while circuit breakers temporarily stop calls during persistent failure.

    12

    What is the safest general retry policy?

    Classify the failure, verify idempotency, enforce a deadline and retry budget, use capped exponential backoff with jitter, and stop through a controlled terminal failure path.

    Key Takeaway

    Retries recover from temporary failures by repeating an operation, but every retry adds load and can repeat a business effect. Retry only errors that can plausibly succeed later, and do not retry permanent validation, authentication, authorization, or schema failures unchanged. Use capped exponential backoff to reduce retry frequency and jitter to prevent synchronized retry waves. Bound the number of attempts, individual delays, total elapsed time, concurrency, and aggregate retry traffic. Preserve one stable idempotency key across all attempts, because a timeout can mean that the original operation succeeded and only its response was lost. Avoid retries at several architectural layers, do not hold locks while waiting, and rerun all state-dependent calculations when retrying a transaction. For asynchronous messages, use delayed retries, bounded attempts, and dead-letter handling. Finally, combine retries with timeouts, backpressure, circuit breakers, observability, and tested recovery procedures.