Table of Contents

    exactly-once claims

    SYSTEM DESIGN • CHAPTER 15.10

    Exactly-Once Claims

    Understand why exactly-once delivery is impossible, what vendors actually mean when they advertise it, and how to evaluate the claim against the guarantee your system genuinely needs.

    Learning objective: By the end of this article, you will understand why exactly-once delivery cannot exist, the difference between delivery and processing semantics, what the scope boundaries of real exactly-once implementations are, and how to interrogate a claim before depending on it.

    Prerequisites

    Recommended Knowledge

    • Partial failure and timeout ambiguity
    • The two generals problem
    • Delivery semantics and retries
    • Idempotency and deduplication
    • Transactional outbox and inbox patterns
    • Two-phase commit
    • Consumer offsets and checkpointing
    • Stream processing fundamentals

    Why Exactly-Once Delivery Cannot Exist

    This is not an engineering limitation awaiting a better implementation. It follows directly from the two generals problem, which establishes that certain agreement over an unreliable channel is impossible.

    THE IMPOSSIBILITY
    Send Message No Acknowledgement Resend or Not?
    Sender's Choice If Message Was Lost If Acknowledgement Was Lost
    Resend Correct, delivered once Delivered twice
    Do not resend Never delivered Correct, delivered once
    THE FORCED CHOICE
    Since the sender cannot distinguish the two cases, it must choose a policy. Resending risks duplication; not resending risks loss. There is no third option that avoids both.

    Simple Analogy

    You place an order by phone and the line drops before confirmation. Calling back risks ordering twice. Not calling back risks not ordering at all. Only the shop's own records can resolve it, not your caution.

    Exactly-once delivery is not hard. It is impossible. What is achievable is exactly-once effect, and the distinction determines whether a vendor claim means anything.

    Delivery Versus Processing

    The confusion almost always stems from conflating two different things: how many times a message arrives, and how many times its effect is applied.

    Term Meaning Achievable
    Exactly-once delivery The message arrives precisely once No, provably impossible
    Exactly-once processing The handler runs precisely once No, the handler may be retried
    Exactly-once effect The observable outcome occurs once Yes, with deduplication
    Effectively-once Same as exactly-once effect Yes
    WHAT IS ACTUALLY BUILT
    At-Least-Once Delivery + Deduplication Exactly-Once Effect
    The Honest Framing Every production exactly-once system delivers at least once and discards duplicates somewhere. The engineering question is where that deduplication happens and what it covers.

    What Vendors Actually Guarantee

    Exactly-once features are real and useful. They are also scoped far more narrowly than the marketing suggests.

    Typical Claim Actual Scope
    Exactly-once producer Deduplicates retries within one session
    Exactly-once semantics Read, process, write atomically inside the platform
    Exactly-once processing Checkpointed state is consistent with offsets
    Exactly-once sink Transactional write to a supported destination
    End-to-end exactly-once Within the vendor's boundary only
    THE GUARANTEE BOUNDARY
    Source in Platform Processing Sink in Platform
    The Guarantee Ends at the Boundary The moment your handler calls a payment gateway, sends an email, or writes to a database outside the platform, the guarantee no longer applies to that effect.

    How Platform Exactly-Once Works

    Mechanism Purpose
    Producer sequence numbers Broker discards duplicate retries
    Producer epoch Fences out a zombie producer instance
    Transactional writes Output and offset commit atomically
    Read-committed isolation Consumers never see aborted output
    Checkpointed state State and position restored together
    It is deduplication plus atomic commit: The producer deduplicates by sequence number, and the offset commits in the same transaction as the output. That is genuinely valuable, and it is not exactly-once delivery.

    Where Claims Break Down

    Situation Why the Guarantee Does Not Hold
    Handler calls an external API That call is outside the transaction
    Handler writes to another database No shared transaction exists
    Handler sends a notification Side effect cannot be rolled back
    Producer restarts with a new identity Deduplication state is lost
    Non-deterministic processing logic Replay produces different output
    Deduplication window expires A late duplicate is no longer recognized
    Manual offset manipulation Operator reprocesses deliberately
    Topic replayed after an incident Everything is reprocessed by design
    // The guarantee covers the offset commit and the output topic.
    // It does not cover the payment call.
    async function handleOrderPlaced(message, context) {
        const order = message.value;
    
        // Outside any transaction the platform controls
        await paymentGateway.charge({
            customerId: order.customerId,
            amount: order.total
        });
    
        // Covered by the platform transaction
        await context.produce("OrderCharged", { orderId: order.orderId });
    
        // Offset commits atomically with the produce, not with the charge
    }
    A Replay Charges Twice If the process fails after charging but before committing, the message is reprocessed. The platform correctly avoids duplicating the output event, and the customer is still charged twice.

    Non-Determinism Breaks Replay

    Exactly-once processing guarantees frequently assume that reprocessing the same input produces the same output. Ordinary code violates this constantly.

    Source of Non-Determinism Consequence on Replay Remedy
    Reading the current time Different timestamp in output Take time from the event
    Generating a random identifier Different key, duplicate record Derive deterministically from input
    Calling an external service Different response Cache the response by input
    Reading mutable shared state Different computed result Include the state version in input
    Iterating an unordered collection Different output ordering Sort explicitly
    // Non-deterministic: replay produces a different record
    function buildRecord(event) {
        return {
            id: crypto.randomUUID(),
            processedAt: new Date().toISOString(),
            orderId: event.orderId
        };
    }
    
    // Deterministic: replay produces an identical record
    function buildRecord(event) {
        return {
            id: deriveId(event.orderId, event.eventId),
            processedAt: event.occurredAt,
            orderId: event.orderId
        };
    }
    Derive, Do Not Generate Any identifier that must survive replay should be a deterministic function of the input. A random value creates a new record every time the message is reprocessed.

    Building Exactly-Once Effects

    Since the platform guarantee stops at its boundary, the application must extend it. Three techniques cover most cases.

    1

    Natural Idempotency

    Design the operation so repetition changes nothing. Setting a status to shipped twice is harmless; incrementing a counter twice is not.

    2

    Deduplication Table

    Record processed message identifiers in the same transaction as the effect. A duplicate finds the record and stops.

    3

    Idempotency Keys Downstream

    Pass a deterministic key to external services that support them, so the external system performs the deduplication.

    CREATE TABLE processed_messages (
        consumer     VARCHAR(80)  NOT NULL,
        message_id   VARCHAR(120) NOT NULL,
        processed_at TIMESTAMP    NOT NULL DEFAULT CURRENT_TIMESTAMP,
        PRIMARY KEY (consumer, message_id)
    );
    
    CREATE INDEX idx_processed_cleanup
        ON processed_messages (processed_at);
    async function processExactlyOnce(db, message, handler) {
        return await db.transaction(async (tx) => {
            const claim = await tx.query(`
                INSERT INTO processed_messages (consumer, message_id)
                VALUES ($1, $2)
                ON CONFLICT (consumer, message_id) DO NOTHING
                RETURNING message_id
            `, [handler.name, message.id]);
    
            if (claim.rowCount === 0) {
                return { status: "duplicate", messageId: message.id };
            }
    
            // Effect and deduplication record share one commit
            await handler.apply(tx, message);
    
            return { status: "applied", messageId: message.id };
        });
    }
    The two writes must share a transaction: Recording the message as processed separately from applying its effect recreates the dual-write problem. Either both commit or neither does.

    When the Effect Is External

    async function chargeExactlyOnce(db, message, gateway) {
        const idempotencyKey = `charge:${message.orderId}:${message.eventId}`;
    
        // The gateway deduplicates; we do not need to
        const result = await gateway.charge({
            amount: message.total,
            customerId: message.customerId,
            idempotencyKey
        });
    
        await db.query(`
            INSERT INTO charge_records (order_id, gateway_ref, idempotency_key)
            VALUES ($1, $2, $3)
            ON CONFLICT (idempotency_key) DO NOTHING
        `, [message.orderId, result.reference, idempotencyKey]);
    
        return result;
    }
    External System Deduplication Support Fallback
    Payment gateway Usually idempotency keys Query before charging
    Email provider Rarely Local deduplication table
    SMS gateway Rarely Local deduplication table
    Third-party API Varies Check documentation, assume none
    Physical dispatch None possible Guard before, reconcile after

    Deduplication Windows

    Deduplication state cannot be retained forever. Every implementation has a window, and duplicates arriving after it are not recognized.

    SAFE WINDOW
    \[ T_{\text{dedup}} > T_{\text{max retry delay}} + T_{\text{max consumer downtime}} \]
    Window Too Short Window Too Long
    Late duplicates reprocessed Deduplication table grows large
    Guarantee silently lapses Lookup latency increases
    Replay after long outage duplicates Storage cost rises
    -- Retain well beyond the longest plausible redelivery
    DELETE FROM processed_messages
    WHERE processed_at < CURRENT_TIMESTAMP - INTERVAL '30 days'
      AND ctid IN (
          SELECT ctid FROM processed_messages
          WHERE processed_at < CURRENT_TIMESTAMP - INTERVAL '30 days'
          LIMIT 10000
      );
    The Silent Window Expiry A deduplication window shorter than your longest outage means a post-incident replay reprocesses everything. Nothing reports an error; effects simply happen twice.

    The Cost of Exactly-Once

    Cost Impact
    Transactional coordination Higher latency per message
    Deduplication lookups Extra read on every message
    Deduplication storage Table growth and cleanup burden
    Reduced batching Lower throughput
    Transaction timeouts Additional failure modes
    Operational complexity More states to reason about
    Often At-Least-Once Is Correct If the operation is naturally idempotent, exactly-once machinery adds cost for no benefit. Setting a value, upserting a record, or indexing a document are all safe to repeat.
    Operation Duplicate-Safe Needs Deduplication
    Set status to shipped Yes No
    Upsert a search document Yes No
    Invalidate a cache entry Yes No
    Increment a counter No Yes
    Charge a payment No Yes
    Send a notification No Yes
    Append a ledger entry No Yes

    Interrogating a Claim

    When a platform advertises exactly-once, these questions reveal what is actually covered.

    Questions to Ask

    • Does this cover delivery or effect?
    • Where exactly does the guarantee boundary sit?
    • Does it survive a consumer restart?
    • Does it survive a producer restart with a new instance?
    • How long is the deduplication window?
    • Does it hold if my handler calls an external service?
    • Does it assume deterministic processing?
    • What happens during a manual offset reset?
    • What happens when I replay a topic deliberately?
    • What is the throughput cost of enabling it?
    THE DIAGNOSTIC QUESTION
    Ask what happens if the consumer crashes immediately after calling an external service but before committing. Any honest answer will concede that the external call may repeat.

    A Worked Example

    An order pipeline where each step needs a different guarantee.

    Step Guarantee Needed Mechanism
    Update order status At-least-once Naturally idempotent write
    Index for search At-least-once Upsert by document key
    Decrement inventory Exactly-once effect Deduplication table in transaction
    Charge payment Exactly-once effect Gateway idempotency key
    Award loyalty points Exactly-once effect Deduplication table in transaction
    Send confirmation email Exactly-once effect Local deduplication before sending
    Emit analytics event At-least-once Downstream tolerates duplicates
    Guarantee Per Operation, Not Per Pipeline Applying exactly-once machinery uniformly wastes throughput on steps that are already safe. Classify each effect and pay only where duplication actually matters.

    Monitoring

    Signals Worth Tracking

    • Duplicate messages detected per consumer
    • Duplicate rate as a proportion of throughput
    • Deduplication table size and growth
    • Deduplication lookup latency
    • Messages arriving outside the window
    • Transaction abort rate
    • External calls rejected as duplicates
    • Reprocessing events after offset resets
    • Throughput with and without the guarantee enabled
    A duplicate detection rate of zero usually means your deduplication is not working, not that duplicates never occur. At-least-once delivery guarantees they will.

    Verification

    describe("exactly-once effects", function () {
        it("applies an effect once despite duplicate delivery", async function () {
            const message = buildMessage({ orderId: "o-1", points: 100 });
    
            await processExactlyOnce(db, message, loyaltyHandler);
            await processExactlyOnce(db, message, loyaltyHandler);
            await processExactlyOnce(db, message, loyaltyHandler);
    
            const balance = await db.loyalty.balanceFor("customer-1");
            expect(balance).toBe(100);
        });
    
        it("does not record the message when the effect fails", async function () {
            loyaltyHandler.failNext();
    
            await expect(
                processExactlyOnce(db, message, loyaltyHandler)
            ).rejects.toThrow();
    
            const recorded = await db.processedMessages.exists(message.id);
            expect(recorded).toBe(false);
        });
    
        it("reprocesses after the deduplication window expires", async function () {
            await processExactlyOnce(db, message, loyaltyHandler);
            await db.processedMessages.expireAll();
    
            await processExactlyOnce(db, message, loyaltyHandler);
    
            const balance = await db.loyalty.balanceFor("customer-1");
            expect(balance).toBe(200); // Documents the real limitation
        });
    
        it("produces identical output on replay", async function () {
            const first = buildRecord(message);
            const second = buildRecord(message);
    
            expect(first).toEqual(second);
        });
    
        it("passes a stable idempotency key to the gateway", async function () {
            await chargeExactlyOnce(db, message, gateway);
            await chargeExactlyOnce(db, message, gateway);
    
            expect(gateway.distinctKeysReceived()).toBe(1);
            expect(gateway.chargesApplied()).toBe(1);
        });
    });
    Test the window expiry explicitly: A test asserting that reprocessing occurs after expiry documents the guarantee's actual boundary, rather than leaving the team assuming it is unconditional.

    Common Design Mistakes

    Weak Design

    • Believing delivery can be exactly-once
    • Assuming the platform covers external calls
    • Non-deterministic identifiers in handlers
    • Deduplication written outside the transaction
    • Window shorter than plausible downtime
    • Enabling the guarantee everywhere uniformly
    • No deduplication cleanup job
    • Never measuring the duplicate rate
    • Treating duplicates as an exceptional case

    Strong Design

    • Designs for at-least-once by default
    • Knows exactly where the boundary ends
    • Derives identifiers deterministically
    • Commits deduplication with the effect
    • Window exceeds worst-case redelivery
    • Applies the guarantee per operation
    • Cleans up in bounded batches
    • Monitors duplicates as normal traffic
    • Tests window expiry deliberately

    System Design Interview Discussion

    Question What Your Answer Should Cover
    Can you guarantee exactly-once? Delivery no, effect yes
    Why is delivery impossible? The two generals problem
    What do platforms actually provide? Deduplication plus atomic offset commit
    Where does the guarantee end? At the first external side effect
    How do you extend it? Deduplication table or idempotency keys
    How long do you retain dedup state? Beyond worst-case redelivery delay
    What does it cost? Latency, throughput, storage, complexity
    When is at-least-once enough? Naturally idempotent operations

    Design Checklist

    Production Checklist

    • Assume at-least-once delivery everywhere
    • Classify each effect as duplicate-safe or not
    • Make naturally idempotent operations the default
    • Commit deduplication records with the effect
    • Derive all identifiers deterministically
    • Take timestamps from the event, not the clock
    • Pass idempotency keys to external services
    • Set the deduplication window above worst-case downtime
    • Clean up deduplication state in bounded batches
    • Document where the guarantee boundary ends
    • Avoid enabling the guarantee where it adds nothing
    • Monitor duplicate rate as expected traffic
    • Alert on messages arriving outside the window
    • Test duplicate delivery, window expiry, and replay
    • Verify handler output is identical on replay

    Knowledge Check

    1

    Why is exactly-once delivery impossible?

    A sender cannot distinguish a lost message from a lost acknowledgement, so it must choose between risking loss and risking duplication.

    2

    What is achievable instead?

    Exactly-once effect: at-least-once delivery combined with deduplication, so the observable outcome occurs once.

    3

    Where does a platform guarantee end?

    At the platform boundary. Any external call, notification, or write to another system is outside the transaction it controls.

    4

    Why does non-determinism break replay?

    Reprocessing generates different identifiers or timestamps, producing a new record rather than recognizing the existing one.

    5

    Why must the deduplication window be generous?

    A duplicate arriving after expiry is not recognized, so the guarantee lapses silently during exactly the long outages where replay is most likely.

    Summary

    Exactly-once delivery is provably impossible. A sender that receives no acknowledgement cannot determine whether the message or the acknowledgement was lost, so it must choose between risking duplication and risking loss.

    What is achievable is exactly-once effect: at-least-once delivery combined with deduplication somewhere in the path. Every real exactly-once feature works this way, using producer sequence numbers to discard retries and committing output atomically with the consumer offset.

    Those guarantees are genuine but narrowly scoped. They hold within the platform boundary and end at the first external side effect. They also typically assume deterministic processing, which ordinary code violates through clocks, random identifiers, and external calls.

    The practical response is to assume at-least-once everywhere, classify each effect as duplicate-safe or not, and extend the guarantee only where it matters through deduplication tables committed with the effect or idempotency keys passed downstream.

    Key Takeaway

    Treat exactly-once as a claim to interrogate, not a guarantee to depend on. Ask where the boundary sits, what happens across a restart, and how long the deduplication window is. Then design for at-least-once anyway, make identifiers deterministic, commit deduplication alongside the effect, and pay for the guarantee only where duplication would actually cause harm.