Table of Contents

    sagas

    SYSTEM DESIGN • CHAPTER 15.7

    Sagas

    Understand how long-running business transactions maintain consistency without distributed locks, why compensation is not the same as rollback, and what it costs to make partial failure visible rather than hidden.

    Learning objective: By the end of this article, you will understand the saga pattern, why compensation differs fundamentally from rollback, choreography versus orchestration, isolation anomalies and countermeasures, failure handling when compensation itself fails, and how to model a workflow that tolerates every step failing.

    Prerequisites

    Recommended Knowledge

    • Two-phase commit and its blocking behaviour
    • Transactions, atomicity, and isolation
    • Idempotency and duplicate handling
    • Eventual consistency and its anomalies
    • Message delivery semantics and retries
    • Event-driven communication
    • State machines and legal transitions
    • Partial failure and timeout ambiguity

    The Problem Sagas Address

    Two-phase commit holds locks across a network for the duration of a transaction. For a business process spanning seconds, minutes, or days, that is untenable. Sagas abandon atomicity to keep each step short and local.

    THE STRUCTURAL SHIFT
    One Distributed Transaction Sequence of Local Transactions
    Property Two-Phase Commit Saga
    Atomicity Guaranteed Approximated by compensation
    Isolation Provided by locks None across steps
    Lock duration Entire transaction Single local step
    Availability Compounds downward Each step independent
    Intermediate state Invisible Visible to other readers
    Failure handling Automatic rollback Explicit compensation
    Duration tolerated Milliseconds Days

    Simple Analogy

    Booking flights, a hotel, and a car separately. If the car is unavailable, you cannot un-book the flight as though it never happened. You cancel it, which may incur a fee and appears on your statement.

    Compensation Is Not Rollback

    This distinction is the source of most saga design errors. A rollback erases history; a compensation adds to it.

    Rollback

    • State returns to exactly as before
    • No trace remains
    • Nobody observed the intermediate state
    • Performed by the database engine

    Compensation

    • State becomes semantically equivalent
    • Both actions remain in history
    • Others may have already reacted
    • Written by you, as business logic
    Forward Action Compensation Residual Effect
    Reserve inventory Release reservation None, cleanly reversible
    Charge a card Issue a refund Two entries on the statement
    Create a shipment Cancel the shipment Cancellation record persists
    Send a notification Send a correction Recipient saw the original
    Award loyalty points Deduct the points Balance may have gone negative
    Dispatch a courier Recall the courier Fuel and time already spent
    Publish content Unpublish It may have been copied
    THE DESIGN QUESTION
    For each step, ask what the world looks like after compensation. If the answer involves a customer noticing, that step belongs later in the sequence or behind a reservation.

    Irreversible Steps

    Some actions cannot be compensated at all. Their placement determines whether a saga is viable.

    Step Type Property Placement Rule
    Compensatable Can be semantically undone Anywhere before the pivot
    Pivot Last point where abort is possible After all compensatable steps
    Retriable Must eventually succeed Only after the pivot
    SAGA STRUCTURE
    Compensatable Steps Pivot Retriable Steps
    Why the Pivot Matters Before it, failure means compensate backwards. After it, failure means retry forwards until success. Everything irreversible must sit after the pivot, or the saga has no safe failure path.
    Misplaced Irreversibility Sending a physical package before verifying payment places an uncompensatable step before the pivot. Any later failure leaves no correct recovery action.

    Modelling a Saga

    const ORDER_SAGA = {
        name: "place-order",
        steps: [
            {
                name: "reserve-inventory",
                type: "compensatable",
                action: (ctx) => inventory.reserve(ctx.orderId, ctx.items),
                compensation: (ctx) => inventory.release(ctx.orderId)
            },
            {
                name: "authorize-payment",
                type: "compensatable",
                action: (ctx) => payments.authorize(ctx.orderId, ctx.amount),
                compensation: (ctx) => payments.voidAuthorization(ctx.orderId)
            },
            {
                name: "capture-payment",
                type: "pivot",
                action: (ctx) => payments.capture(ctx.orderId),
                compensation: (ctx) => payments.refund(ctx.orderId)
            },
            {
                name: "commit-inventory",
                type: "retriable",
                action: (ctx) => inventory.commit(ctx.orderId)
            },
            {
                name: "create-shipment",
                type: "retriable",
                action: (ctx) => shipping.create(ctx.orderId, ctx.address)
            },
            {
                name: "notify-customer",
                type: "retriable",
                action: (ctx) => notifications.orderConfirmed(ctx.orderId)
            }
        ]
    };
    async function executeSaga(definition, context, log) {
        const completed = [];
    
        for (const step of definition.steps) {
            try {
                await log.record(context.sagaId, step.name, "started");
    
                const result = await step.action(context);
    
                await log.record(context.sagaId, step.name, "completed", result);
                completed.push(step);
            } catch (error) {
                await log.record(context.sagaId, step.name, "failed", error);
    
                if (step.type === "retriable") {
                    await log.markForRetry(context.sagaId, step.name);
                    return { status: "awaiting-retry", failedAt: step.name };
                }
    
                await compensate(completed, context, log);
                return { status: "compensated", failedAt: step.name };
            }
        }
    
        await log.record(context.sagaId, null, "saga-completed");
        return { status: "completed" };
    }
    
    async function compensate(completedSteps, context, log) {
        for (const step of completedSteps.slice().reverse()) {
            if (!step.compensation) continue;
    
            await retryWithBackoff(async () => {
                await step.compensation(context);
                await log.record(context.sagaId, step.name, "compensated");
            }, { maxAttempts: 10 });
        }
    }
    Compensate in reverse order: Later steps may depend on earlier ones. Releasing inventory before voiding the payment authorization that referenced it can leave dangling references.

    Choreography Versus Orchestration

    Two coordination styles exist, and the choice determines where workflow knowledge lives.

    Aspect Choreography Orchestration
    Control Distributed across services Centralized in a coordinator
    Communication Events published and consumed Commands issued directly
    Workflow visibility Implicit, spread across code Explicit in one definition
    Coupling Loose between services Services coupled to coordinator
    Adding a step Modify multiple services Modify the definition
    Debugging Trace events across systems Inspect coordinator state
    Cyclic dependency risk High as the saga grows Low
    Single point of failure None Coordinator, unless replicated

    Choreography Flow

    EVENT CHAIN
    Order Created Inventory Reserved Payment Captured Shipment Created
    // Each service reacts to events and emits its own
    inventoryService.on("OrderCreated", async (event) => {
        try {
            await inventory.reserve(event.orderId, event.items);
            await publish("InventoryReserved", { orderId: event.orderId });
        } catch (error) {
            await publish("InventoryReservationFailed", {
                orderId: event.orderId,
                reason: error.message
            });
        }
    });
    
    inventoryService.on("PaymentFailed", async (event) => {
        await inventory.release(event.orderId);
        await publish("InventoryReleased", { orderId: event.orderId });
    });
    Choreography at Scale With eight services and several failure branches, no single place describes the workflow. Answering "what happens if step five fails" requires reading every service.

    Orchestration Flow

    class SagaOrchestrator {
        async advance(sagaId) {
            const state = await this.store.load(sagaId);
    
            if (state.status === "compensating") {
                return this.continueCompensation(state);
            }
    
            const next = this.definition.steps[state.currentStep];
    
            if (!next) {
                return this.store.complete(sagaId);
            }
    
            const outcome = await this.dispatch(next, state.context);
    
            if (outcome.ok) {
                await this.store.advanceStep(sagaId, next.name);
                return this.advance(sagaId);
            }
    
            if (next.type === "retriable") {
                return this.store.scheduleRetry(sagaId, next.name, outcome.error);
            }
    
            await this.store.beginCompensation(sagaId, next.name);
            return this.continueCompensation(await this.store.load(sagaId));
        }
    }
    Orchestration Scales Better Beyond three or four steps, centralized workflow definition becomes substantially easier to reason about, monitor, and modify.

    The Isolation Problem

    Sagas provide no isolation. Every intermediate state is visible to concurrent readers and writers, producing anomalies that traditional transactions prevent.

    Anomaly What Happens Example
    Lost update Concurrent saga overwrites a change Two orders modify the same record
    Dirty read Reader sees state later compensated Report includes a cancelled order
    Fuzzy read Value changes mid-saga Price differs between steps
    Phantom New rows appear during the saga Stock arrives after the check

    Countermeasures

    Technique Mechanism Cost
    Semantic lock Mark records as pending Readers must handle pending state
    Commutative updates Use deltas rather than absolutes Requires operation redesign
    Pessimistic ordering Place risky steps after the pivot Constrains workflow design
    Re-read value Verify assumptions before committing Extra reads, possible late abort
    Version file Record operations and reorder on apply Significant complexity
    By value Route high-risk cases to 2PC Two mechanisms to maintain
    -- Semantic lock: the record advertises an in-progress saga
    UPDATE orders
    SET status      = 'payment_pending',
        saga_id     = :saga_id,
        locked_until = CURRENT_TIMESTAMP + INTERVAL '5 minutes'
    WHERE order_id  = :order_id
      AND (saga_id IS NULL OR locked_until < CURRENT_TIMESTAMP);
    
    -- Readers decide how to treat pending records
    SELECT order_id, status, total
    FROM orders
    WHERE customer_id = :customer_id
      AND status NOT IN ('payment_pending', 'compensating');
    Semantic locks need expiry: A saga that crashes mid-flight leaves records marked pending forever. Attach a timeout and a recovery process that resolves abandoned sagas.

    Idempotency Is Mandatory

    Every step and every compensation will be retried. Without idempotency, retries duplicate effects and compensations over-correct.

    CREATE TABLE saga_step_execution (
        saga_id      VARCHAR(80) NOT NULL,
        step_name    VARCHAR(80) NOT NULL,
        attempt_key  VARCHAR(80) NOT NULL,
        status       VARCHAR(20) NOT NULL,
        result       JSONB       NULL,
        executed_at  TIMESTAMP   NOT NULL,
        PRIMARY KEY (saga_id, step_name)
    );
    
    -- Claim execution atomically; conflict means already done
    INSERT INTO saga_step_execution (
        saga_id, step_name, attempt_key, status, executed_at
    )
    VALUES (:saga_id, :step_name, :attempt_key, 'in_progress', CURRENT_TIMESTAMP)
    ON CONFLICT (saga_id, step_name) DO NOTHING
    RETURNING attempt_key;
    Without Idempotency Consequence
    Payment step retried Customer charged twice
    Refund compensation retried Refunded twice
    Inventory release retried Stock inflated
    Notification retried Customer emailed repeatedly
    COMPENSATIONS TOO
    Compensations are retried under exactly the same conditions as forward steps. A non-idempotent compensation is as dangerous as a non-idempotent action.

    When Compensation Fails

    This is the hardest case in the pattern. The forward step succeeded, the saga must abort, and the undo will not complete.

    Cause Response
    Transient service failure Retry with backoff indefinitely
    Downstream permanently rejects Escalate for manual resolution
    Business rule prevents undo Apply an alternative remedy
    Compensation logic defect Halt, alert, fix, and replay
    Physical action already taken No technical remedy exists
    async function compensateWithEscalation(step, context, log) {
        const outcome = await retryWithBackoff(
            () => step.compensation(context),
            {
                maxAttempts: 20,
                maxDurationMs: 6 * 60 * 60 * 1000,
                isRetriable: (e) => !e.permanent
            }
        );
    
        if (outcome.succeeded) {
            return await log.record(context.sagaId, step.name, "compensated");
        }
    
        await log.record(context.sagaId, step.name, "compensation-failed", {
            attempts: outcome.attempts,
            lastError: outcome.error
        });
    
        await escalation.raise({
            severity: "high",
            sagaId: context.sagaId,
            step: step.name,
            message: "Compensation exhausted, manual intervention required",
            context: context.summary()
        });
    
        return { status: "requires-intervention" };
    }
    Do Not Silently Give Up Abandoning a failed compensation leaves the system in a state no code accounts for. It must be visible, queued, and owned by someone.

    Saga State Persistence

    CREATE TABLE saga_instance (
        saga_id        VARCHAR(80)  PRIMARY KEY,
        definition     VARCHAR(80)  NOT NULL,
        status         VARCHAR(30)  NOT NULL,
        current_step   VARCHAR(80)  NULL,
        context        JSONB        NOT NULL,
        started_at     TIMESTAMP    NOT NULL,
        updated_at     TIMESTAMP    NOT NULL,
        completed_at   TIMESTAMP    NULL,
        failure_reason TEXT         NULL
    );
    
    CREATE INDEX idx_saga_stuck
        ON saga_instance (status, updated_at)
        WHERE status IN ('running', 'compensating');
    
    -- Recovery sweep for abandoned sagas
    SELECT saga_id, definition, current_step, updated_at
    FROM saga_instance
    WHERE status IN ('running', 'compensating')
      AND updated_at < CURRENT_TIMESTAMP - INTERVAL '10 minutes'
    ORDER BY updated_at;
    Status Meaning Next Action
    running Executing forward steps Advance to next step
    awaiting-retry Retriable step failed Retry after backoff
    compensating Undoing completed steps Continue reverse execution
    compensated Fully undone Terminal
    completed All steps succeeded Terminal
    requires-intervention Compensation exhausted Human resolution
    Persist before dispatching: Record the intent to execute a step before calling the service. A crash between call and record leaves an executed step with no log entry, and recovery will execute it again.

    A Worked Example

    Step Type Compensation Visible Residue
    Validate order Compensatable Mark order cancelled Cancelled order visible
    Reserve inventory Compensatable Release reservation None
    Authorize payment Compensatable Void authorization Possible pending hold
    Apply discount code Compensatable Restore code usage None
    Capture payment Pivot Refund Two statement entries
    Commit inventory Retriable None Must succeed
    Create shipment Retriable None Must succeed
    Award points Retriable None Must succeed
    Send confirmation Retriable None Must succeed
    Why Capture Is the Pivot Before capture, everything can be voided with minimal customer impact. After it, money has moved and the business is committed to fulfilment, so remaining steps must be driven to completion.

    Monitoring

    Signals Worth Tracking

    • Saga completion rate by definition
    • Compensation rate and which step triggered it
    • Sagas stuck past their expected duration
    • Compensation failures requiring intervention
    • Retry counts per retriable step
    • End-to-end saga duration percentiles
    • Step-level failure rates
    • Semantic locks held past expiry
    • Idempotency conflicts detected
    • Sagas awaiting manual resolution
    A compensation rate that is zero usually means failures are being silently swallowed, not that nothing fails. Track which step triggers compensation, not merely how often.

    Verification

    describe("order saga", function () {
        it("compensates all completed steps when a step fails", async function () {
            paymentService.failNext("authorize");
    
            const result = await saga.execute(orderContext);
    
            expect(result.status).toBe("compensated");
            expect(await inventory.reservationFor(orderContext.orderId)).toBeNull();
        });
    
        it("retries forward rather than compensating after the pivot", async function () {
            shippingService.failNext("create");
    
            const result = await saga.execute(orderContext);
    
            expect(result.status).toBe("awaiting-retry");
            expect(await payments.captured(orderContext.orderId)).toBe(true);
        });
    
        it("is idempotent under duplicate step execution", async function () {
            await saga.executeStep("capture-payment", orderContext);
            await saga.executeStep("capture-payment", orderContext);
    
            const charges = await payments.chargesFor(orderContext.orderId);
            expect(charges.length).toBe(1);
        });
    
        it("escalates when compensation cannot complete", async function () {
            inventoryService.failPermanently("release");
            paymentService.failNext("authorize");
    
            const result = await saga.execute(orderContext);
    
            expect(result.status).toBe("requires-intervention");
            expect(escalation.raised()).toHaveLength(1);
        });
    
        it("resumes correctly after orchestrator restart", async function () {
            const started = saga.execute(orderContext);
            await faultInjector.crashOrchestrator();
            await orchestrator.restart();
            await orchestrator.recoverStuckSagas();
    
            expect(await saga.status(orderContext.sagaId)).toBe("completed");
        });
    });

    Common Design Mistakes

    Weak Design

    • Treating compensation as rollback
    • Irreversible steps before the pivot
    • Non-idempotent steps or compensations
    • No persisted saga state
    • Ignoring isolation anomalies
    • Choreography for complex workflows
    • Silently abandoning failed compensation
    • Semantic locks without expiry
    • No recovery sweep for stuck sagas

    Strong Design

    • Designs compensations as business actions
    • Places the pivot deliberately
    • Makes every step idempotent
    • Persists state before dispatching
    • Applies countermeasures for anomalies
    • Orchestrates beyond a few steps
    • Escalates exhausted compensation
    • Expires semantic locks
    • Sweeps for abandoned sagas

    System Design Interview Discussion

    Question What Your Answer Should Cover
    Why a saga over 2PC? Lock duration and compounding availability
    How does compensation differ from rollback? Semantic undo leaving visible history
    What if a step cannot be undone? Pivot placement and forward recovery
    What about isolation? Anomalies and semantic lock countermeasures
    Choreography or orchestration? Complexity threshold and visibility
    What if compensation fails? Retry, escalate, never silently abandon
    How do you handle retries? Idempotency on steps and compensations
    What if the orchestrator crashes? Persisted state and recovery sweep

    Design Checklist

    Production Checklist

    • Classify each step as compensatable, pivot, or retriable
    • Place all irreversible steps after the pivot
    • Write a compensation for every compensatable step
    • Document the visible residue of each compensation
    • Make every step and compensation idempotent
    • Persist saga state before dispatching any call
    • Compensate in reverse completion order
    • Retry retriable steps indefinitely with backoff
    • Escalate compensations that cannot complete
    • Apply semantic locks where anomalies matter
    • Expire semantic locks and recover abandoned sagas
    • Use orchestration beyond three or four steps
    • Run a recovery sweep for stalled instances
    • Monitor compensation rate by triggering step
    • Test every step failing, including compensation failure
    • Test orchestrator crash and resumption

    Knowledge Check

    1

    How does compensation differ from rollback?

    Rollback erases the change as though it never occurred. Compensation performs a new action producing an equivalent outcome, leaving both in history.

    2

    What is the pivot step?

    The last point at which the saga can still abort. Before it failure means compensating backwards; after it, retrying forwards.

    3

    Why do sagas lack isolation?

    Each step commits locally and immediately, so intermediate state is visible to concurrent readers and writers.

    4

    Why must compensations be idempotent?

    They are retried under the same failure conditions as forward steps, so a non-idempotent refund can issue twice.

    5

    When should you prefer orchestration?

    Once the workflow exceeds a few steps or acquires branching failure paths, since choreography then hides the workflow across many services.

    Summary

    A saga replaces one distributed transaction with a sequence of local transactions, each committing immediately. Atomicity is abandoned and approximated through compensating actions that semantically undo completed work.

    Compensation is not rollback. It is a new business action with its own visible consequences, and some actions cannot be compensated at all. This makes step ordering critical: irreversible work must sit after the pivot, where failure means retrying forwards rather than undoing backwards.

    Sagas provide no isolation, so intermediate state is visible and anomalies such as dirty reads and lost updates become possible. Semantic locks, commutative operations, and careful ordering are the practical countermeasures.

    Every step and compensation will be retried, making idempotency mandatory. Saga state must be persisted before dispatching, and the hardest case — compensation that cannot complete — must be escalated rather than silently abandoned.

    Key Takeaway

    A saga trades atomicity for availability and makes partial failure your problem to model. Classify every step, place the pivot deliberately, write a compensation for everything before it, make all of it idempotent, persist state before you dispatch, and treat a compensation that cannot complete as an incident rather than an edge case.