exactly-once claims
Exactly-Once Claims
Understand why exactly-once delivery is impossible, what vendors actually mean when they advertise it, and how to evaluate the claim against the guarantee your system genuinely needs.
Prerequisites
Recommended Knowledge
- Partial failure and timeout ambiguity
- The two generals problem
- Delivery semantics and retries
- Idempotency and deduplication
- Transactional outbox and inbox patterns
- Two-phase commit
- Consumer offsets and checkpointing
- Stream processing fundamentals
Why Exactly-Once Delivery Cannot Exist
This is not an engineering limitation awaiting a better implementation. It follows directly from the two generals problem, which establishes that certain agreement over an unreliable channel is impossible.
| Sender's Choice | If Message Was Lost | If Acknowledgement Was Lost |
|---|---|---|
| Resend | Correct, delivered once | Delivered twice |
| Do not resend | Never delivered | Correct, delivered once |
Simple Analogy
You place an order by phone and the line drops before confirmation. Calling back risks ordering twice. Not calling back risks not ordering at all. Only the shop's own records can resolve it, not your caution.
Delivery Versus Processing
The confusion almost always stems from conflating two different things: how many times a message arrives, and how many times its effect is applied.
| Term | Meaning | Achievable |
|---|---|---|
| Exactly-once delivery | The message arrives precisely once | No, provably impossible |
| Exactly-once processing | The handler runs precisely once | No, the handler may be retried |
| Exactly-once effect | The observable outcome occurs once | Yes, with deduplication |
| Effectively-once | Same as exactly-once effect | Yes |
What Vendors Actually Guarantee
Exactly-once features are real and useful. They are also scoped far more narrowly than the marketing suggests.
| Typical Claim | Actual Scope |
|---|---|
| Exactly-once producer | Deduplicates retries within one session |
| Exactly-once semantics | Read, process, write atomically inside the platform |
| Exactly-once processing | Checkpointed state is consistent with offsets |
| Exactly-once sink | Transactional write to a supported destination |
| End-to-end exactly-once | Within the vendor's boundary only |
How Platform Exactly-Once Works
| Mechanism | Purpose |
|---|---|
| Producer sequence numbers | Broker discards duplicate retries |
| Producer epoch | Fences out a zombie producer instance |
| Transactional writes | Output and offset commit atomically |
| Read-committed isolation | Consumers never see aborted output |
| Checkpointed state | State and position restored together |
Where Claims Break Down
| Situation | Why the Guarantee Does Not Hold |
|---|---|
| Handler calls an external API | That call is outside the transaction |
| Handler writes to another database | No shared transaction exists |
| Handler sends a notification | Side effect cannot be rolled back |
| Producer restarts with a new identity | Deduplication state is lost |
| Non-deterministic processing logic | Replay produces different output |
| Deduplication window expires | A late duplicate is no longer recognized |
| Manual offset manipulation | Operator reprocesses deliberately |
| Topic replayed after an incident | Everything is reprocessed by design |
// The guarantee covers the offset commit and the output topic.
// It does not cover the payment call.
async function handleOrderPlaced(message, context) {
const order = message.value;
// Outside any transaction the platform controls
await paymentGateway.charge({
customerId: order.customerId,
amount: order.total
});
// Covered by the platform transaction
await context.produce("OrderCharged", { orderId: order.orderId });
// Offset commits atomically with the produce, not with the charge
}
Non-Determinism Breaks Replay
Exactly-once processing guarantees frequently assume that reprocessing the same input produces the same output. Ordinary code violates this constantly.
| Source of Non-Determinism | Consequence on Replay | Remedy |
|---|---|---|
| Reading the current time | Different timestamp in output | Take time from the event |
| Generating a random identifier | Different key, duplicate record | Derive deterministically from input |
| Calling an external service | Different response | Cache the response by input |
| Reading mutable shared state | Different computed result | Include the state version in input |
| Iterating an unordered collection | Different output ordering | Sort explicitly |
// Non-deterministic: replay produces a different record
function buildRecord(event) {
return {
id: crypto.randomUUID(),
processedAt: new Date().toISOString(),
orderId: event.orderId
};
}
// Deterministic: replay produces an identical record
function buildRecord(event) {
return {
id: deriveId(event.orderId, event.eventId),
processedAt: event.occurredAt,
orderId: event.orderId
};
}
Building Exactly-Once Effects
Since the platform guarantee stops at its boundary, the application must extend it. Three techniques cover most cases.
Natural Idempotency
Design the operation so repetition changes nothing. Setting a status to shipped twice is harmless; incrementing a counter twice is not.
Deduplication Table
Record processed message identifiers in the same transaction as the effect. A duplicate finds the record and stops.
Idempotency Keys Downstream
Pass a deterministic key to external services that support them, so the external system performs the deduplication.
CREATE TABLE processed_messages (
consumer VARCHAR(80) NOT NULL,
message_id VARCHAR(120) NOT NULL,
processed_at TIMESTAMP NOT NULL DEFAULT CURRENT_TIMESTAMP,
PRIMARY KEY (consumer, message_id)
);
CREATE INDEX idx_processed_cleanup
ON processed_messages (processed_at);
async function processExactlyOnce(db, message, handler) {
return await db.transaction(async (tx) => {
const claim = await tx.query(`
INSERT INTO processed_messages (consumer, message_id)
VALUES ($1, $2)
ON CONFLICT (consumer, message_id) DO NOTHING
RETURNING message_id
`, [handler.name, message.id]);
if (claim.rowCount === 0) {
return { status: "duplicate", messageId: message.id };
}
// Effect and deduplication record share one commit
await handler.apply(tx, message);
return { status: "applied", messageId: message.id };
});
}
When the Effect Is External
async function chargeExactlyOnce(db, message, gateway) {
const idempotencyKey = `charge:${message.orderId}:${message.eventId}`;
// The gateway deduplicates; we do not need to
const result = await gateway.charge({
amount: message.total,
customerId: message.customerId,
idempotencyKey
});
await db.query(`
INSERT INTO charge_records (order_id, gateway_ref, idempotency_key)
VALUES ($1, $2, $3)
ON CONFLICT (idempotency_key) DO NOTHING
`, [message.orderId, result.reference, idempotencyKey]);
return result;
}
| External System | Deduplication Support | Fallback |
|---|---|---|
| Payment gateway | Usually idempotency keys | Query before charging |
| Email provider | Rarely | Local deduplication table |
| SMS gateway | Rarely | Local deduplication table |
| Third-party API | Varies | Check documentation, assume none |
| Physical dispatch | None possible | Guard before, reconcile after |
Deduplication Windows
Deduplication state cannot be retained forever. Every implementation has a window, and duplicates arriving after it are not recognized.
| Window Too Short | Window Too Long |
|---|---|
| Late duplicates reprocessed | Deduplication table grows large |
| Guarantee silently lapses | Lookup latency increases |
| Replay after long outage duplicates | Storage cost rises |
-- Retain well beyond the longest plausible redelivery
DELETE FROM processed_messages
WHERE processed_at < CURRENT_TIMESTAMP - INTERVAL '30 days'
AND ctid IN (
SELECT ctid FROM processed_messages
WHERE processed_at < CURRENT_TIMESTAMP - INTERVAL '30 days'
LIMIT 10000
);
The Cost of Exactly-Once
| Cost | Impact |
|---|---|
| Transactional coordination | Higher latency per message |
| Deduplication lookups | Extra read on every message |
| Deduplication storage | Table growth and cleanup burden |
| Reduced batching | Lower throughput |
| Transaction timeouts | Additional failure modes |
| Operational complexity | More states to reason about |
| Operation | Duplicate-Safe | Needs Deduplication |
|---|---|---|
| Set status to shipped | Yes | No |
| Upsert a search document | Yes | No |
| Invalidate a cache entry | Yes | No |
| Increment a counter | No | Yes |
| Charge a payment | No | Yes |
| Send a notification | No | Yes |
| Append a ledger entry | No | Yes |
Interrogating a Claim
When a platform advertises exactly-once, these questions reveal what is actually covered.
Questions to Ask
- Does this cover delivery or effect?
- Where exactly does the guarantee boundary sit?
- Does it survive a consumer restart?
- Does it survive a producer restart with a new instance?
- How long is the deduplication window?
- Does it hold if my handler calls an external service?
- Does it assume deterministic processing?
- What happens during a manual offset reset?
- What happens when I replay a topic deliberately?
- What is the throughput cost of enabling it?
A Worked Example
An order pipeline where each step needs a different guarantee.
| Step | Guarantee Needed | Mechanism |
|---|---|---|
| Update order status | At-least-once | Naturally idempotent write |
| Index for search | At-least-once | Upsert by document key |
| Decrement inventory | Exactly-once effect | Deduplication table in transaction |
| Charge payment | Exactly-once effect | Gateway idempotency key |
| Award loyalty points | Exactly-once effect | Deduplication table in transaction |
| Send confirmation email | Exactly-once effect | Local deduplication before sending |
| Emit analytics event | At-least-once | Downstream tolerates duplicates |
Monitoring
Signals Worth Tracking
- Duplicate messages detected per consumer
- Duplicate rate as a proportion of throughput
- Deduplication table size and growth
- Deduplication lookup latency
- Messages arriving outside the window
- Transaction abort rate
- External calls rejected as duplicates
- Reprocessing events after offset resets
- Throughput with and without the guarantee enabled
Verification
describe("exactly-once effects", function () {
it("applies an effect once despite duplicate delivery", async function () {
const message = buildMessage({ orderId: "o-1", points: 100 });
await processExactlyOnce(db, message, loyaltyHandler);
await processExactlyOnce(db, message, loyaltyHandler);
await processExactlyOnce(db, message, loyaltyHandler);
const balance = await db.loyalty.balanceFor("customer-1");
expect(balance).toBe(100);
});
it("does not record the message when the effect fails", async function () {
loyaltyHandler.failNext();
await expect(
processExactlyOnce(db, message, loyaltyHandler)
).rejects.toThrow();
const recorded = await db.processedMessages.exists(message.id);
expect(recorded).toBe(false);
});
it("reprocesses after the deduplication window expires", async function () {
await processExactlyOnce(db, message, loyaltyHandler);
await db.processedMessages.expireAll();
await processExactlyOnce(db, message, loyaltyHandler);
const balance = await db.loyalty.balanceFor("customer-1");
expect(balance).toBe(200); // Documents the real limitation
});
it("produces identical output on replay", async function () {
const first = buildRecord(message);
const second = buildRecord(message);
expect(first).toEqual(second);
});
it("passes a stable idempotency key to the gateway", async function () {
await chargeExactlyOnce(db, message, gateway);
await chargeExactlyOnce(db, message, gateway);
expect(gateway.distinctKeysReceived()).toBe(1);
expect(gateway.chargesApplied()).toBe(1);
});
});
Common Design Mistakes
Weak Design
- Believing delivery can be exactly-once
- Assuming the platform covers external calls
- Non-deterministic identifiers in handlers
- Deduplication written outside the transaction
- Window shorter than plausible downtime
- Enabling the guarantee everywhere uniformly
- No deduplication cleanup job
- Never measuring the duplicate rate
- Treating duplicates as an exceptional case
Strong Design
- Designs for at-least-once by default
- Knows exactly where the boundary ends
- Derives identifiers deterministically
- Commits deduplication with the effect
- Window exceeds worst-case redelivery
- Applies the guarantee per operation
- Cleans up in bounded batches
- Monitors duplicates as normal traffic
- Tests window expiry deliberately
System Design Interview Discussion
| Question | What Your Answer Should Cover |
|---|---|
| Can you guarantee exactly-once? | Delivery no, effect yes |
| Why is delivery impossible? | The two generals problem |
| What do platforms actually provide? | Deduplication plus atomic offset commit |
| Where does the guarantee end? | At the first external side effect |
| How do you extend it? | Deduplication table or idempotency keys |
| How long do you retain dedup state? | Beyond worst-case redelivery delay |
| What does it cost? | Latency, throughput, storage, complexity |
| When is at-least-once enough? | Naturally idempotent operations |
Design Checklist
Production Checklist
- Assume at-least-once delivery everywhere
- Classify each effect as duplicate-safe or not
- Make naturally idempotent operations the default
- Commit deduplication records with the effect
- Derive all identifiers deterministically
- Take timestamps from the event, not the clock
- Pass idempotency keys to external services
- Set the deduplication window above worst-case downtime
- Clean up deduplication state in bounded batches
- Document where the guarantee boundary ends
- Avoid enabling the guarantee where it adds nothing
- Monitor duplicate rate as expected traffic
- Alert on messages arriving outside the window
- Test duplicate delivery, window expiry, and replay
- Verify handler output is identical on replay
Knowledge Check
Why is exactly-once delivery impossible?
A sender cannot distinguish a lost message from a lost acknowledgement, so it must choose between risking loss and risking duplication.
What is achievable instead?
Exactly-once effect: at-least-once delivery combined with deduplication, so the observable outcome occurs once.
Where does a platform guarantee end?
At the platform boundary. Any external call, notification, or write to another system is outside the transaction it controls.
Why does non-determinism break replay?
Reprocessing generates different identifiers or timestamps, producing a new record rather than recognizing the existing one.
Why must the deduplication window be generous?
A duplicate arriving after expiry is not recognized, so the guarantee lapses silently during exactly the long outages where replay is most likely.
Summary
Exactly-once delivery is provably impossible. A sender that receives no acknowledgement cannot determine whether the message or the acknowledgement was lost, so it must choose between risking duplication and risking loss.
What is achievable is exactly-once effect: at-least-once delivery combined with deduplication somewhere in the path. Every real exactly-once feature works this way, using producer sequence numbers to discard retries and committing output atomically with the consumer offset.
Those guarantees are genuine but narrowly scoped. They hold within the platform boundary and end at the first external side effect. They also typically assume deterministic processing, which ordinary code violates through clocks, random identifiers, and external calls.
The practical response is to assume at-least-once everywhere, classify each effect as duplicate-safe or not, and extend the guarantee only where it matters through deduplication tables committed with the effect or idempotency keys passed downstream.
Key Takeaway
Treat exactly-once as a claim to interrogate, not a guarantee to depend on. Ask where the boundary sits, what happens across a restart, and how long the deduplication window is. Then design for at-least-once anyway, make identifiers deterministic, commit deduplication alongside the effect, and pay for the guarantee only where duplication would actually cause harm.