Table of Contents

    fencing tokens

    SYSTEM DESIGN • CHAPTER 15.5

    Fencing Tokens

    Understand the mechanism that makes exclusion actually enforceable, why the resource must be the arbiter rather than the client, and how a single monotonic number closes a gap that no amount of timing can.

    Learning objective: By the end of this article, you will understand what a fencing token guarantees, why monotonicity is the essential property, where tokens must be generated and validated, how to apply them to databases, storage, queues, and external systems, and what to do when a resource cannot be fenced.

    Prerequisites

    Recommended Knowledge

    • Distributed locks and their limitations
    • Leases and expiry semantics
    • Leader election and the stale holder problem
    • Consensus and monotonic sequencing
    • Process pauses and timeout ambiguity
    • Conditional writes and compare-and-swap
    • Idempotency and duplicate handling
    • Version guards and optimistic concurrency

    What a Fencing Token Is

    A fencing token is a strictly increasing number issued with every grant of exclusive access. The protected resource records the highest token it has accepted and refuses anything lower.

    THE FENCING RULE
    \[ \text{Accept} \iff token_{\text{presented}} \geq token_{\text{highest accepted}} \]

    Simple Analogy

    A service counter calls ticket numbers in order. Someone arriving with ticket forty-one is turned away once forty-two has been served, without the counter needing to know where they have been or why they are late.

    The resource does not need to know who currently holds the lock. It only needs to know that it has already served someone newer.

    The Gap Tokens Close

    Every coordination mechanism shares one structural weakness: the interval between being granted permission and using it. Nothing the client does can eliminate that interval.

    THE UNAVOIDABLE INTERVAL
    Permission Granted Unbounded Delay Write Arrives
    Attempted Fix Why It Fails
    Shorter lease duration Pauses can exceed any duration chosen
    Check the lock before writing A pause can occur after the check
    Faster failure detection Detects the failure, not the resumption
    More reliable network Pauses are local, not network-caused
    Better clock synchronization The paused process observes no elapsed time
    Larger safety margin Reduces probability, never reaches zero
    THE STRUCTURAL INSIGHT
    The client is the wrong place to enforce exclusion, because it cannot observe its own suspension. Only the resource observes the true arrival order of writes.

    The Sequence Tokens Prevent

    Step Without Fencing With Fencing
    Client A acquires Holds the lock Receives token 41
    Client A pauses Still believes it holds Still believes it holds
    Lease expires Lock released Lock released
    Client B acquires Holds the lock Receives token 42
    Client B writes Write applied Applied, resource records 42
    Client A resumes and writes Overwrites B silently Rejected, 41 is below 42

    Monotonicity Is the Whole Property

    The token's only requirement is that it strictly increases across grants. Everything else about it is incidental.

    Property Required Reason
    Strictly increasing Essential Ordering is the entire mechanism
    Durable across restarts Essential Restart must not reissue a used value
    Issued by a single authority Essential Two issuers can produce duplicates
    Contiguous with no gaps Not required Only relative order matters
    Related to wall-clock time Not required Clocks are the problem being avoided
    Meaningful to humans Not required Useful for debugging only
    Timestamps Are Not Tokens Clock skew means two nodes can generate identical or inverted timestamps. A token generated from the local clock provides no ordering guarantee at all.
    Valid Token Sources A consensus log index, a database sequence, a lock service counter, or a leader term number. Each is issued by a single authority and survives restart.

    Generating Tokens

    CREATE TABLE lock_grants (
        resource       VARCHAR(160) PRIMARY KEY,
        holder_id      VARCHAR(80)  NOT NULL,
        fencing_token  BIGINT       NOT NULL DEFAULT 0,
        granted_at     TIMESTAMP    NOT NULL,
        expires_at     TIMESTAMP    NOT NULL
    );
    
    -- Atomic grant with token increment
    UPDATE lock_grants
    SET holder_id     = :holder_id,
        fencing_token = fencing_token + 1,
        granted_at    = CURRENT_TIMESTAMP,
        expires_at    = CURRENT_TIMESTAMP + (:ttl_ms * INTERVAL '1 millisecond')
    WHERE resource = :resource
      AND (expires_at < CURRENT_TIMESTAMP OR holder_id = :holder_id)
    RETURNING fencing_token, expires_at;
    Source Mechanism Durability
    Consensus log index Position in the replicated log Strong, survives any minority failure
    Leader term number Incremented on each election Strong, persisted before voting
    Database sequence Monotonic counter Strong if the database is durable
    Coordination service version Node modification counter Strong, consensus-backed
    In-memory counter Incremented locally None, resets on restart
    The counter must outlive the process: An in-memory counter that resets to zero on restart issues tokens the resource has already seen and rejected. Every subsequent write from that issuer is silently dropped.

    Validating at the Resource

    Validation must happen where the write lands, in the same atomic operation as the write. A separate check-then-write introduces the very gap the token exists to close.

    -- Correct: comparison and write are one atomic operation
    UPDATE account_balances
    SET balance       = balance + :delta,
        fencing_token = :token,
        updated_at    = CURRENT_TIMESTAMP
    WHERE account_id    = :account_id
      AND fencing_token <= :token;
    
    -- Zero rows affected means a stale writer was rejected
    Check-Then-Write Defeats the Purpose Reading the stored token, comparing it in application code, and then writing reopens the window. A newer writer can commit between your read and your write.
    class FencedResource {
        constructor(store) {
            this.store = store;
        }
    
        async apply(resourceId, mutation, token) {
            const result = await this.store.conditionalUpdate({
                id: resourceId,
                set: { ...mutation, fencingToken: token },
                condition: { fencingToken: { lessThanOrEqual: token } }
            });
    
            if (result.rowsAffected === 0) {
                const current = await this.store.readToken(resourceId);
    
                return {
                    accepted: false,
                    reason: "stale-token",
                    presented: token,
                    required: current
                };
            }
    
            return { accepted: true, token };
        }
    }

    Greater-Than or Greater-Than-Or-Equal

    Comparison Effect Suitable When
    Strictly greater Each token permits one write Single decisive action per grant
    Greater or equal One token permits many writes Holder performs ongoing work
    Usually Greater-Or-Equal A leader typically writes repeatedly under one term. Requiring a new token per write would force a coordination round trip for every operation.

    Fencing Different Resource Types

    Resource Fencing Mechanism Strength
    Relational database Token column in the where clause Strong and natural
    Key-value store Compare-and-swap on a version Strong where supported
    Object storage Conditional write on generation Strong where supported
    Message topic Producer epoch rejecting older Strong in modern brokers
    Internal service Token header validated server-side Strong if state is tracked
    Distributed file system Client identifier revocation Varies by implementation
    External partner API Usually none available Weak, mitigate differently

    Fencing a Downstream Service

    async function callFencedService(client, request, token) {
        const response = await client.post("/apply", {
            headers: {
                "X-Fencing-Token": String(token),
                "X-Resource-Id": request.resourceId
            },
            body: request.payload
        });
    
        if (response.status === 409) {
            return {
                accepted: false,
                reason: "fenced by downstream",
                requiredToken: Number(response.headers.get("X-Required-Token"))
            };
        }
    
        return { accepted: true, result: response.body };
    }
    function fencingMiddleware(store) {
        return async function (request, response, next) {
            const token = Number(request.headers["x-fencing-token"]);
            const resourceId = request.headers["x-resource-id"];
    
            if (!Number.isFinite(token)) {
                return response.status(400).json({
                    error: "missing or invalid fencing token"
                });
            }
    
            const highest = await store.highestToken(resourceId);
    
            if (token < highest) {
                response.setHeader("X-Required-Token", String(highest));
                return response.status(409).json({
                    error: "stale fencing token",
                    presented: token,
                    required: highest
                });
            }
    
            request.fencingToken = token;
            return next();
        };
    }
    Tokens must propagate across every hop: A fenced entry point calling an unfenced downstream service leaves the write path unprotected at its final destination.

    When Fencing Is Impossible

    Many external systems accept any authenticated write with no version check. Exclusion cannot be guaranteed for these, so the goal shifts to limiting damage.

    Mitigation Mechanism Residual Risk
    Idempotency key from token Duplicate call recognized Different tokens still both apply
    Fence a local proxy record Guard before calling externally Pause between guard and call
    Short lease plus margin Reduce overlap probability Never eliminated
    Reconciliation after the fact Detect and correct duplicates Some effects are irreversible
    Make the call commutative Repetition becomes harmless Not always possible
    async function callUnfenceableExternal(external, request, token) {
        const guard = await localStore.conditionalUpdate({
            id: request.resourceId,
            set: { pendingToken: token, pendingAt: Date.now() },
            condition: { pendingToken: { lessThanOrEqual: token } }
        });
    
        if (guard.rowsAffected === 0) {
            return { accepted: false, reason: "fenced locally before external call" };
        }
    
        return await external.submit(request.payload, {
            idempotencyKey: `${request.resourceId}:${token}`
        });
    }
    The Honest Limitation A local guard narrows the window but does not close it. A pause between the guard and the external call still permits two submissions with different idempotency keys.

    Token Scope

    A token is meaningful only within the scope its issuer covers. Comparing tokens from different issuers or different resources produces nonsense.

    Scope Token Meaning Consideration
    Per resource Ordering of grants for that resource Most precise, more counters to maintain
    Per lock name Ordering for that lock Natural when locks map to resources
    Per cluster term Ordering of leadership epochs Coarse, one token covers all writes
    Global sequence Ordering across everything Simple, but a sequencing bottleneck
    Cross-Scope Comparison Storing a token from resource A in resource B's column produces a comparison with no meaning. Either value may be higher for unrelated reasons.
    Record the Scope Store the issuer and scope alongside the token so a mismatched comparison is detectable rather than silently wrong.
    {
        "resourceId": "shard-0042",
        "fencing": {
            "token": 1847,
            "issuer": "lock-service",
            "scope": "resource:shard-0042",
            "issuedAt": "2026-09-24T06:14:22Z"
        }
    }

    Fencing and Version Guards

    Fencing tokens resemble optimistic concurrency version guards, but they answer different questions and can coexist.

    Aspect Fencing Token Version Guard
    Question answered Is this writer still authorized? Has the data changed since I read it?
    Increments on Grant of exclusive access Every write to the data
    Issued by Coordination authority The data store itself
    Comparison Greater or equal Exactly equal
    Rejects Deposed holders Concurrent modifications
    -- Both guards applied together
    UPDATE documents
    SET content       = :content,
        version       = version + 1,
        fencing_token = :fencing_token
    WHERE document_id    = :document_id
      AND version        = :expected_version
      AND fencing_token <= :fencing_token;
    COMPLEMENTARY, NOT ALTERNATIVE
    The version guard prevents lost updates between concurrent authorized writers. The fencing token prevents writes from writers who are no longer authorized at all.

    Failure Modes

    Failure Consequence Mitigation
    Counter resets on restart All writes permanently rejected Durable, replicated counter
    Two issuers for one scope Duplicate tokens, no ordering Single authority per scope
    Token not propagated downstream Final write unprotected Carry the token across every hop
    Check separated from write Race reopens the window Single atomic conditional write
    Rejections silently swallowed Work believed done was dropped Surface and alert on rejections
    Token column missing on a table That path remains unfenced Audit all write paths
    Client supplies its own token Arbitrary high value bypasses fencing Issue server-side only
    A rejection is a successful defence: Fencing rejections should be logged and alerted, not suppressed. Each one is evidence a stale writer was stopped, and a rising rate indicates an instability worth investigating.

    Monitoring

    Signals Worth Tracking

    • Fencing rejections per resource
    • Gap between presented and required token
    • Token issuance rate per scope
    • Current token value and its progression
    • Writes accepted without a token present
    • Token propagation failures across hops
    • Time between token issue and first use
    • Resources with no token column configured
    • Issuer availability and latency
    A large gap between the presented and required token indicates a writer was absent for many grants. That is a paused or partitioned node worth investigating on its own.

    Verification

    describe("fencing token enforcement", function () {
        it("rejects a lower token after a higher one is accepted", async function () {
            await resource.apply("res-1", { value: "b" }, 42);
    
            const stale = await resource.apply("res-1", { value: "a" }, 41);
    
            expect(stale.accepted).toBe(false);
            expect(stale.required).toBe(42);
        });
    
        it("accepts repeated writes with the same token", async function () {
            await resource.apply("res-1", { value: "x" }, 50);
            const again = await resource.apply("res-1", { value: "y" }, 50);
    
            expect(again.accepted).toBe(true);
        });
    
        it("issues strictly increasing tokens across grants", async function () {
            const tokens = [];
    
            for (let i = 0; i < 20; i += 1) {
                const grant = await lockService.acquire("res-1", `client-${i}`);
                tokens.push(grant.fencingToken);
                await lockService.forceExpire("res-1");
            }
    
            for (let i = 1; i < tokens.length; i += 1) {
                expect(tokens[i]).toBeGreaterThan(tokens[i - 1]);
            }
        });
    
        it("preserves monotonicity across issuer restart", async function () {
            const before = await lockService.acquire("res-1", "client-a");
    
            await faultInjector.restartIssuer();
    
            const after = await lockService.acquire("res-1", "client-b");
    
            expect(after.fencingToken).toBeGreaterThan(before.fencingToken);
        });
    });

    Common Design Mistakes

    Weak Design

    • Deriving tokens from wall-clock time
    • Using an in-memory counter
    • Validating in application code before writing
    • Accepting a client-supplied token value
    • Fencing one write path but not others
    • Dropping the token at a service boundary
    • Comparing tokens across different scopes
    • Swallowing rejections as ordinary errors
    • Assuming the lock alone is sufficient

    Strong Design

    • Issues from a single durable authority
    • Persists the counter before responding
    • Validates atomically with the write
    • Generates tokens server-side only
    • Audits every write path for fencing
    • Propagates tokens across all hops
    • Records issuer and scope with the token
    • Alerts on rejection rate and token gaps
    • Tests issuer restart and stale writers

    System Design Interview Discussion

    Question What Your Answer Should Cover
    Why are fencing tokens needed? The unbounded gap between grant and write
    What property must the token have? Strict monotonicity from one durable issuer
    Why not use a timestamp? Clock skew breaks the ordering guarantee
    Where is validation performed? At the resource, atomically with the write
    What if the resource cannot fence? Idempotency keys and accepted residual risk
    How do tokens relate to versions? Authorization versus concurrent modification
    What happens on issuer restart? Durable counter preserves monotonicity
    How would you verify it works? Stale-writer and restart tests

    Design Checklist

    Production Checklist

    • Issue tokens from one durable authority per scope
    • Persist the counter before returning a grant
    • Never derive tokens from wall-clock time
    • Reject tokens supplied by clients
    • Store the token alongside the protected state
    • Validate in the same atomic operation as the write
    • Use greater-or-equal for repeated writes per grant
    • Record issuer and scope with every token
    • Propagate tokens across every service hop
    • Audit all write paths for missing fencing
    • Return a distinct status code on rejection
    • Include the required token in rejection responses
    • Use idempotency keys where fencing is unavailable
    • Alert on rejection rate and token gaps
    • Test stale writers, issuer restart, and monotonicity

    Knowledge Check

    1

    What gap do fencing tokens close?

    The unbounded interval between being granted access and the write arriving, during which authorization may have been revoked without the client knowing.

    2

    Why must tokens be monotonic?

    Ordering is the entire mechanism. Without strict increase, the resource cannot determine which grant is newer.

    3

    Why must validation be atomic with the write?

    A separate check and write reopens the race, because a newer writer can commit in between the two operations.

    4

    Why must the counter be durable?

    A counter resetting on restart reissues values the resource has already exceeded, causing every subsequent write to be rejected.

    5

    How do tokens differ from version guards?

    Tokens establish whether the writer is still authorized; version guards establish whether the data changed since it was read.

    Summary

    A fencing token is a strictly increasing number issued with each grant of exclusive access, recorded by the resource and used to reject any writer presenting a lower value. It closes the gap that no coordination mechanism can close on its own.

    That gap exists because a client cannot observe its own suspension. A paused process perceives no elapsed time, so it cannot detect that its lease expired and another holder has since acted. Shorter leases and larger margins reduce the probability without ever reaching zero.

    Monotonicity is the only property that matters, and it requires a single durable authority. Timestamps fail because of clock skew, and in-memory counters fail because restart reissues consumed values. Consensus log indices, leader terms, and database sequences all work.

    Validation must occur at the resource, in the same atomic operation as the write, and the token must propagate across every hop in the write path. Where a resource cannot be fenced, the goal shifts from prevention to damage limitation through idempotency keys, local guards, and reconciliation.

    Key Takeaway

    Move enforcement to where the write lands. Issue strictly increasing tokens from one durable authority, validate them atomically with every write, carry them across every service hop, audit for write paths that lack them, and treat each rejection as evidence the mechanism worked rather than as an error to suppress.