Table of Contents

    read-your-writes

    SYSTEM DESIGN • CHAPTER 14.8

    Read-Your-Writes

    Understand the session guarantee that eliminates the most jarring user-visible inconsistency, why it costs far less than strong consistency, and how to implement it without coordinating every read.

    Learning objective: By the end of this article, you will understand the four session guarantees, why read-your-writes matters more to users than global consistency, implementation approaches from sticky routing to version tokens, the scope boundaries that break it, and how to verify it holds.

    Prerequisites

    Recommended Knowledge

    • Replication and replica lag
    • Consistency models and their guarantees
    • Logical clocks and version tracking
    • Sessions and identity propagation
    • Load balancing and request routing
    • Caching and cache invalidation
    • Quorum reads and writes
    • PACELC latency trade-offs

    The Anomaly Users Actually Notice

    A user updates their profile, the interface confirms success, and the page reloads showing the old value. Nothing has failed. The write was durable and will propagate. Yet the user's confidence in the system has been damaged.

    THE FAILURE SEQUENCE
    Write to Leader Success Returned Read from Lagging Replica Old Value Shown

    Simple Analogy

    You submit a form at a counter, receive a receipt, then ask a second clerk who has not yet received the paperwork. They tell you nothing was submitted. The record exists, but your experience says otherwise.

    Users do not perceive global consistency. They perceive whether their own action took effect. Those are very different requirements, and only one of them is expensive.

    Why This Anomaly Is Disproportionately Damaging

    User Behavior Resulting Problem
    Resubmits the form Duplicate records created
    Assumes the system is broken Support burden and lost trust
    Retries a payment Potential double charge
    Re-uploads a file Storage waste and confusion
    Reports a bug Engineering time spent on expected behavior
    Abandons the workflow Direct business loss

    The Four Session Guarantees

    Read-your-writes belongs to a family of guarantees scoped to a single client session rather than to the system as a whole. Each addresses a distinct perceived anomaly.

    Guarantee Promise Anomaly Prevented
    Read-your-writes You see your own updates Your change appears lost
    Monotonic reads You never see older data than before Time appearing to move backwards
    Writes follow reads Your write is ordered after what you read Replying before the original appears
    Monotonic writes Your writes apply in issue order Later edit overwritten by earlier one
    Why session scope is the key insight: Guaranteeing coherence for one client requires tracking only what that client has seen. Guaranteeing it globally requires agreement across all nodes, which is orders of magnitude more expensive.
    COST COMPARISON
    \[ C_{\text{session}} \ll C_{\text{linearizable}} \]

    Implementation Approaches

    Approach Mechanism Principal Weakness
    Read from leader Route all reads to the primary Discards read scalability entirely
    Sticky routing Pin a session to one replica Breaks on failover or rebalance
    Recency window Use the leader briefly after writing Guessed window may be insufficient
    Version token Client carries the version it wrote Requires client cooperation
    Replica version check Select a replica caught up enough Needs replica progress metadata
    Write-through cache Serve the writer's own value locally Only covers the same process

    Reading from the Leader

    Advantages

    • Trivially correct
    • No client-side state
    • No version tracking required
    • Easy to reason about

    Costs

    • Leader handles all read traffic
    • Replicas serve no purpose
    • Higher latency for distant clients
    • Leader failure affects all reads
    Selective Application Route only reads of recently written keys to the leader. Most reads touch data the session never modified and can safely use a replica.

    Version Token Tracking

    The most flexible approach records the version produced by each write and requires subsequent reads to meet or exceed it. Most reads still hit a local replica.

    class SessionVersionTracker {
        constructor(sessionId) {
            this.sessionId = sessionId;
            this.versions = new Map();
        }
    
        recordWrite(key, version) {
            const existing = this.versions.get(key) ?? 0;
            this.versions.set(key, Math.max(existing, version));
        }
    
        requiredFor(key) {
            return this.versions.get(key) ?? 0;
        }
    
        prune(maxAgeMs) {
            const cutoff = Date.now() - maxAgeMs;
    
            for (const [key, entry] of this.versions) {
                if (entry.recordedAt < cutoff) {
                    this.versions.delete(key);
                }
            }
        }
    }
    
    async function sessionRead(key, session) {
        const required = session.requiredFor(key);
    
        if (required === 0) {
            const value = await datastore.readLocalReplica(key);
            return { value, path: "local", escalated: false };
        }
    
        const local = await datastore.readLocalReplica(key);
    
        if (local.version >= required) {
            return { value: local.value, path: "local", escalated: false };
        }
    
        const authoritative = await datastore.readLeader(key);
    
        return {
            value: authoritative.value,
            path: "leader",
            escalated: true,
            lagVersions: required - local.version
        };
    }
    The efficiency argument: Escalation occurs only when the local replica has genuinely not caught up. In a healthy system with low lag, the overwhelming majority of reads remain local.

    Propagating the Token

    Version state must travel with the client. Several mechanisms exist, each with different trust and scope properties.

    Carrier Suitability Consideration
    Response header API clients Client must echo it back
    Session cookie Browser applications Size limits constrain key count
    Server-side session Backend-managed sessions Requires a session store lookup
    Signed opaque token Untrusted clients Prevents tampering with versions
    Request context Service-to-service calls Must propagate across every hop
    HTTP/1.1 200 OK
    Content-Type: application/json
    X-Session-Version: eyJwcm9maWxlOjQ4MjEiOjE3fQ
    X-Consistency-Path: leader
    {
        "data": { "displayName": "R. Ansari" },
        "consistency": {
            "guarantee": "read-your-writes",
            "servedFrom": "local-replica",
            "sessionVersion": 17,
            "replicaVersion": 17,
            "escalated": false
        }
    }
    Unsigned Token Risk A client that can edit the version claim could demand an arbitrarily high version, forcing every read to escalate to the leader. Sign or encrypt tokens exposed to untrusted clients.

    Sticky Routing and Its Fragility

    Pinning a session to one replica appears to solve the problem cheaply. The guarantee holds only while the pin does, and pins break for ordinary operational reasons.

    Event Effect on the Guarantee
    Replica failure Session moves to a possibly behind replica
    Deployment or restart All pinned sessions redistributed
    Autoscaling event Routing recalculated, pins lost
    Client network change New connection may route elsewhere
    Multi-device usage Each device pins independently
    Cross-region movement Geographic routing overrides the pin
    Combine Rather Than Choose Use sticky routing as an optimization that keeps most reads local, with version tokens as the correctness mechanism that survives when the pin breaks.

    Where the Guarantee Breaks

    Read-your-writes is scoped to a session, and every boundary crossing is a place the guarantee can silently fail.

    Boundary Failure Remedy
    Second device Version state is not shared Track versions per user, not per device
    Downstream service Context not propagated Pass version in request headers
    Search index Indexed asynchronously Merge recent writes into results
    Cache layer Serves a pre-write entry Invalidate on write, version cache keys
    Read model projection Built from an event stream Wait for projection to reach the version
    Aggregated views Computed from multiple sources Track the slowest contributing source
    Session expiry Version state discarded Retention exceeding maximum lag
    SCOPE RULE
    The guarantee extends exactly as far as the version context travels. Anywhere the context stops, the guarantee stops with it, without any visible error.

    The Search Index Problem

    Search is the most common place this guarantee fails visibly. A user creates a document, searches for it immediately, and finds nothing, because indexing is asynchronous by design.

    async function searchWithRecentWrites(query, session) {
        const indexResults = await searchIndex.query(query);
    
        const recentWrites = session.recentWrites({
            withinMs: 60_000
        });
    
        if (recentWrites.length === 0) {
            return indexResults;
        }
    
        const indexedIds = new Set(
            indexResults.hits.map(hit => hit.id)
        );
    
        const missing = recentWrites.filter(
            write => !indexedIds.has(write.entityId)
        );
    
        if (missing.length === 0) {
            return indexResults;
        }
    
        const fetched = await datastore.readMany(
            missing.map(write => write.entityId)
        );
    
        const matching = fetched.filter(
            doc => matchesQuery(doc, query)
        );
    
        return {
            hits: [...matching, ...indexResults.hits],
            pendingIndexCount: matching.length
        };
    }
    Merging affects more than presence: Injected results disturb ranking, pagination counts, and facet totals. Decide deliberately whether they appear first, are marked as pending, or are excluded from counts.

    Caching Interaction

    A cache in front of a replicated store adds a second staleness layer. Even a fully caught-up replica does not help if the cache returns a pre-write entry.

    Strategy Behavior Trade-off
    Invalidate on write Delete the entry immediately Invalidation may itself be delayed
    Write-through Update cache during the write Only helps the writing node's cache
    Version in cache key New version yields a new key Old entries linger until eviction
    Bypass after write Skip cache briefly for written keys Load spike on the origin
    Per-session overlay Session-local view of own writes Additional state per session
    async function cachedSessionRead(key, session) {
        const required = session.requiredFor(key);
    
        if (required > 0 && session.recentlyWrote(key, 30_000)) {
            const fresh = await datastore.readLeader(key);
            await cache.set(key, fresh, { ttlSeconds: 60 });
            return fresh;
        }
    
        const cached = await cache.get(key);
    
        if (cached && cached.version >= required) {
            return cached;
        }
    
        const value = required > 0
            ? await datastore.readLeader(key)
            : await datastore.readLocalReplica(key);
    
        await cache.set(key, value, { ttlSeconds: 60 });
        return value;
    }

    Client-Side Alternatives

    Some anomalies can be hidden at the presentation layer without any backend guarantee, though these approaches have clear limits.

    Reasonable Uses

    • Optimistic UI updates after success
    • Local overlay of pending changes
    • Indicating an update is still propagating
    • Deferring a refetch briefly

    Where They Fail

    • A different device shows stale data
    • Page reload discards local state
    • Server-rendered pages bypass the overlay
    • Derived values computed server-side differ
    Correct Positioning Treat client-side overlays as a latency improvement layered on top of a backend guarantee, never as a replacement for one.

    Applying It Selectively

    Operation Guarantee Needed Reasoning
    Profile after editing Read-your-writes User expects their change immediately
    Cart after adding Read-your-writes Missing item triggers re-adding
    Order list after placing Read-your-writes Absence suggests failure
    Document after saving Read-your-writes Data loss fear is acute
    Another user's profile Eventual No expectation of recency
    Product catalogue Eventual User made no change
    Recommendations Eventual Approximate by nature
    Aggregate counters Eventual Precision is not expected
    const SESSION_GUARANTEE = {
        "profile.read":       "read-your-writes",
        "cart.read":          "read-your-writes",
        "orders.list":        "read-your-writes",
        "document.read":      "read-your-writes",
        "user.publicProfile": "eventual",
        "catalog.read":       "eventual",
        "recommendations":    "eventual",
        "metrics.counters":   "eventual"
    };
    
    function guaranteeFor(operation) {
        return SESSION_GUARANTEE[operation] ?? "eventual";
    }
    The selectivity principle: Apply the guarantee only to data the session actually modified. Reads of untouched data have no expectation to violate and should stay on the fast local path.

    Verification

    This guarantee fails intermittently under lag, which means it passes ordinary testing and breaks in production. Deliberate verification is required.

    describe("read-your-writes guarantee", function () {
        it("returns the written value immediately", async function () {
            const session = await createSession();
    
            await client.write(session, "profile:4821", {
                displayName: "Updated Name"
            });
    
            const result = await client.read(session, "profile:4821");
    
            expect(result.value.displayName).toBe("Updated Name");
        });
    
        it("holds when the local replica lags", async function () {
            const session = await createSession();
    
            await faultInjector.pauseReplication("replica-west", 5_000);
    
            await client.write(session, "profile:4821", {
                displayName: "Lag Test"
            });
    
            const result = await client.readFrom(
                session,
                "replica-west",
                "profile:4821"
            );
    
            expect(result.value.displayName).toBe("Lag Test");
            expect(result.escalated).toBe(true);
        });
    
        it("survives replica failover", async function () {
            const session = await createSession();
    
            await client.write(session, "profile:4821", {
                displayName: "Failover Test"
            });
    
            await faultInjector.failReplica("replica-east");
    
            const result = await client.read(session, "profile:4821");
    
            expect(result.value.displayName).toBe("Failover Test");
        });
    });

    Conditions Worth Testing

    • Read immediately after write
    • Read while replication is artificially paused
    • Read after replica failover
    • Read through a populated cache
    • Read via a downstream service
    • Search immediately after creation
    • Read after sticky routing is disrupted
    • Read after session token expiry

    Monitoring

    Signals Worth Tracking

    • Escalation rate from replica to leader
    • Version gap at the point of escalation
    • Replica lag distribution
    • Reads served locally versus escalated
    • Session version map size
    • Guarantee violations detected in testing
    • Latency difference between paths
    • Leader read load attributable to escalation
    • Search results injected from recent writes
    • Cache bypasses triggered by recent writes
    A rising escalation rate is an early warning that replica lag is growing. The guarantee is compensating for a problem that will soon affect everything else.

    Common Design Mistakes

    Weak Design

    • Relying on sticky routing alone
    • Guessing a fixed post-write window
    • Forgetting cache and search layers
    • Not propagating context downstream
    • Using unsigned client-supplied versions
    • Tracking versions per device only
    • Applying the guarantee to every read
    • Treating optimistic UI as the solution
    • Never testing under injected lag

    Strong Design

    • Tracks versions per session and user
    • Escalates only when the replica is behind
    • Covers cache, search, and projections
    • Propagates context across service hops
    • Signs tokens exposed to clients
    • Applies the guarantee selectively
    • Bounds version map size and age
    • Discloses which path served the read
    • Tests under lag and failover

    System Design Interview Discussion

    Question What Your Answer Should Cover
    Why not use strong consistency? Session scope costs far less
    How is it implemented? Version tracking with selective escalation
    Why is sticky routing insufficient? Failover, scaling, and multi-device use
    What about the search index? Merging recent writes into results
    What about caching? Invalidation or bypass after write
    Does it cross service boundaries? Context propagation requirements
    Which reads need it? Only data the session modified
    How would you verify it? Lag injection and failover tests

    Design Checklist

    Production Checklist

    • Identify reads where users expect their own writes
    • Return a version with every write response
    • Track written versions per session and user
    • Compare replica version before serving
    • Escalate only when the replica is behind
    • Sign version tokens exposed to clients
    • Propagate version context across service hops
    • Invalidate or bypass cache after writes
    • Merge recent writes into search results
    • Bound version map size and entry age
    • Set retention above maximum replica lag
    • Combine sticky routing with version checks
    • Disclose the serving path in responses
    • Monitor escalation rate and version gaps
    • Test under injected lag and failover

    Knowledge Check

    1

    What does read-your-writes guarantee?

    That a client always observes its own completed updates, even when reads are served from replicas that have not yet caught up.

    2

    Why is it cheaper than linearizability?

    It requires tracking only what one session wrote, rather than coordinating agreement about recency across all nodes for all clients.

    3

    Why is sticky routing insufficient?

    Pins break during failover, deployment, scaling, and network changes, and a second device is never pinned to the same replica.

    4

    Where does the guarantee commonly fail?

    At boundaries the version context does not cross: caches, search indexes, downstream services, projections, and additional devices.

    5

    Why apply it selectively?

    Reads of data the session never modified carry no expectation to violate, so forcing them through the leader wastes the read capacity replicas provide.

    Summary

    Read-your-writes guarantees that a client observes its own completed updates. It addresses the single most damaging user-visible inconsistency: a change that was accepted and durable yet appears to have vanished.

    It belongs to a family of session guarantees alongside monotonic reads, writes-follow-reads, and monotonic writes. Because these are scoped to one client rather than the whole system, they cost a fraction of what linearizability demands while removing the anomalies users actually perceive.

    Implementation approaches range from leader reads through sticky routing to version tokens. Version tracking is the most robust, serving most reads locally and escalating only when the replica has genuinely not caught up. Sticky routing is a useful optimization but breaks on failover, scaling, and multi-device use, so it cannot be the correctness mechanism.

    The guarantee extends exactly as far as the version context travels. Caches, search indexes, downstream services, and projections all break it silently, which is why applying it selectively to the data a session actually modified, and testing under injected lag, both matter.

    Key Takeaway

    Guarantee what users perceive, not what theory maximizes. Track the versions each session wrote, escalate only when a replica is behind, carry that context through caches, search, and downstream services, and verify the guarantee under injected lag rather than assuming it holds.