Table of Contents

    leases

    SYSTEM DESIGN • CHAPTER 15.3

    Leases

    Understand the mechanism that converts an unreliable grant into a self-expiring one, why time-bounded permission removes the need for explicit release, and the clock assumptions that determine whether a lease is safe.

    Learning objective: By the end of this article, you will understand what a lease guarantees, why expiry solves the failed-holder problem, the two clocks involved and their asymmetric risk, renewal strategy, revocation, and how leases apply beyond leadership to caches, locks, and resource ownership.

    Prerequisites

    Recommended Knowledge

    • Leader election and fencing tokens
    • Partial failure and timeout ambiguity
    • Clock skew and synchronization error
    • Consensus and quorum-backed stores
    • Caching and invalidation
    • Failure detection limits
    • Process pauses and scheduling delays
    • Idempotency and duplicate handling

    What a Lease Is

    A lease is a grant of permission that is valid only for a bounded period. Unlike a lock, it does not require the holder to release it. If the holder disappears, the grant lapses on its own.

    THE LEASE CONTRACT
    Grant Valid Until Expiry Renew or Lapse

    Simple Analogy

    A parking meter grants a bay for an hour. Nobody must confirm you have left; the claim simply expires. The attendant does not need to reach you to reclaim the space.

    Lease Versus Lock

    Property Lock Lease
    Duration Until explicitly released Bounded by expiry
    Holder crashes Held indefinitely Reclaimed automatically
    Requires liveness Yes, to release No, expiry is passive
    Requires clocks No Yes, on both sides
    Reclaim cost Manual intervention Waiting for expiry
    Failure mode Deadlock Premature reclaim
    THE TRADE MADE
    A lease exchanges a dependency on the holder being alive for a dependency on clocks being approximately right. That trade is usually worthwhile, but it is a trade, not a free improvement.

    The Problem Leases Solve

    A plain lock assumes the holder will eventually release it. In a distributed system, holders crash, get partitioned, and pause. None of those states produce a release.

    The Stuck Lock A node acquires a lock and is terminated. The lock remains held by a process that no longer exists. Every other node waits forever, and recovery requires a human.
    Holder Failure Lock Outcome Lease Outcome
    Process killed Held forever Expires and is reclaimed
    Host powered off Held forever Expires and is reclaimed
    Network partitioned Held forever Expires and is reclaimed
    Long pause Held, then released late Expires, holder must stop acting
    Graceful shutdown Released cleanly Released early, better than waiting
    Why Expiry Is the Key Property Reclaiming does not require contacting the holder. The grantor waits out the clock, which works regardless of whether the holder is dead, unreachable, or frozen.

    Two Clocks, Asymmetric Risk

    Every lease involves two independent measurements of elapsed time: the grantor deciding when to reclaim, and the holder deciding when to stop. Disagreement between them determines safety.

    Disagreement Effect Severity
    Holder stops too early Unnecessary handover Availability cost only
    Holder stops too late Two holders simultaneously Correctness violation
    Grantor reclaims too early Overlap with a valid holder Correctness violation
    Grantor reclaims too late Extended outage Availability cost only
    SAFE RECLAIM CONDITION
    \[ T_{\text{grantor waits}} > T_{\text{lease}} + \epsilon_{\text{skew}} + \delta_{\text{network}} \]
    Err in opposite directions: The holder should stop early and the grantor should reclaim late. Both errors cost availability. Erring the other way costs correctness.

    Monotonic Time Is Mandatory

    Wall-clock time can move backwards during synchronization, which makes a lease appear to have more remaining validity than it does. Monotonic clocks never regress.

    The Backward Jump A holder acquires a lease and records the wall-clock expiry. Time synchronization then moves the clock back by ten seconds. The holder now believes it has ten extra seconds of validity that have already elapsed in reality.
    class Lease {
        constructor(durationMs, safetyMarginMs) {
            this.durationMs = durationMs;
            this.safetyMarginMs = safetyMarginMs;
            this.acquiredAt = null;
            this.token = null;
        }
    
        grant(token) {
            this.acquiredAt = performance.now();
            this.token = token;
        }
    
        remainingMs() {
            if (this.acquiredAt === null) return 0;
    
            const elapsed = performance.now() - this.acquiredAt;
            return Math.max(0, this.durationMs - elapsed);
        }
    
        isSafelyHeld() {
            return this.remainingMs() > this.safetyMarginMs;
        }
    
        assertHeld() {
            if (!this.isSafelyHeld()) {
                throw new LeaseExpiredError({
                    remainingMs: this.remainingMs(),
                    requiredMarginMs: this.safetyMarginMs
                });
            }
        }
    }
    The Safety Margin Stopping when a margin remains, rather than at nominal expiry, absorbs scheduling delay and clock error. The margin is the difference between believing you are safe and being safe.

    Renewal Strategy

    A holder that needs continued permission renews before expiry. How early it renews determines how many consecutive failures it can absorb.

    FAILURE TOLERANCE
    \[ N_{\text{tolerable failures}} = \left\lfloor \frac{T_{\text{lease}}}{T_{\text{renew interval}}} \right\rfloor - 1 \]
    Renewal Point Failures Absorbed Renewal Traffic
    At expiry None Minimal
    At half the lease One Low
    At one third Two Moderate
    At one quarter Three Higher
    class LeaseHolder {
        constructor(grantor, config) {
            this.grantor = grantor;
            this.resource = config.resource;
            this.holderId = config.holderId;
            this.lease = new Lease(config.leaseMs, config.marginMs);
            this.renewIntervalMs = config.leaseMs / 3;
            this.consecutiveFailures = 0;
        }
    
        async acquire() {
            const result = await this.grantor.grant({
                resource: this.resource,
                holder: this.holderId,
                durationMs: this.lease.durationMs
            });
    
            if (!result.granted) return { acquired: false };
    
            this.lease.grant(result.token);
            this.scheduleRenewal();
    
            return { acquired: true, token: result.token };
        }
    
        scheduleRenewal() {
            this.timer = setInterval(async () => {
                try {
                    const result = await this.grantor.renew({
                        resource: this.resource,
                        holder: this.holderId,
                        token: this.lease.token,
                        durationMs: this.lease.durationMs
                    });
    
                    if (!result.renewed) {
                        return this.surrender(result.reason);
                    }
    
                    this.lease.grant(result.token);
                    this.consecutiveFailures = 0;
                } catch (error) {
                    this.consecutiveFailures += 1;
    
                    if (!this.lease.isSafelyHeld()) {
                        this.surrender("margin exhausted during renewal failures");
                    }
                }
            }, this.renewIntervalMs);
        }
    
        surrender(reason) {
            clearInterval(this.timer);
            this.lease.acquiredAt = null;
            this.onLeaseLost(reason);
        }
    }
    Surrender on margin exhaustion, not on failure count: What matters is remaining validity, not how many renewals failed. A single failure late in the lease is more dangerous than three early ones.

    Choosing the Duration

    Short Lease

    • Fast reclaim after genuine failure
    • Smaller window for a stale holder
    • Heavy renewal traffic
    • Lost during ordinary latency spikes

    Long Lease

    • Stable under transient disruption
    • Low renewal overhead
    • Extended outage after real failure
    • Longer overlap risk window
    Factor Pushes Toward
    Latency variance to the grantor Longer leases
    Observed process pause duration Longer leases
    Cost of holder warmup Longer leases
    Sensitivity to outage duration Shorter leases
    Absence of fencing at the resource Shorter leases
    Grantor capacity for renewals Longer leases
    MINIMUM VIABLE LEASE
    \[ T_{\text{lease}} > 3 \times (T_{\text{renew RTT}} + T_{\text{pause p99}}) \]

    Leases Alone Are Not Sufficient

    A lease bounds how long a holder should believe it holds permission. It cannot prevent a holder from acting after expiry, because the holder may be paused and unaware.

    THE RESIDUAL RISK
    Lease Expires Holder Still Paused Resumes and Writes
    Why Timing Alone Fails No margin is large enough to guarantee safety, because process pauses have no upper bound. A virtual machine can be suspended for arbitrarily long.
    Pair Leases With Fencing The lease reduces how often overlap occurs. The fencing token, validated at the resource, ensures overlap is harmless when it does occur.
    -- Grantor issues a monotonic token with each grant
    UPDATE lease_registry
    SET holder_id  = :holder_id,
        token      = token + 1,
        granted_at = CURRENT_TIMESTAMP,
        expires_at = CURRENT_TIMESTAMP + (:duration_ms * INTERVAL '1 millisecond')
    WHERE resource = :resource
      AND (expires_at < CURRENT_TIMESTAMP OR holder_id = :holder_id)
    RETURNING token, expires_at;
    
    -- Resource validates the token before accepting work
    UPDATE protected_state
    SET payload     = :payload,
        last_token  = :incoming_token
    WHERE resource   = :resource
      AND last_token <= :incoming_token;
    DIVISION OF RESPONSIBILITY
    The lease is a liveness optimization that limits how long a failed holder blocks progress. The fencing token is the safety mechanism. Do not conflate them.

    Revocation

    Sometimes a grant must end before its natural expiry. Because the holder may be unreachable, revocation has the same fundamental limitation as everything else.

    Approach Mechanism Limitation
    Cooperative return Ask the holder to release Requires the holder to be reachable
    Wait for expiry Let the clock run out Bounded by lease duration
    Token invalidation Advance the resource token Effective immediately at the resource
    Forced reclaim Grant to another before expiry Unsafe without fencing
    Token Invalidation Is Immediate Advancing the resource's expected token instantly invalidates the current holder, without needing to reach it. Reclaim is effective at the point that matters.

    Leases Beyond Leadership

    Leadership is the most discussed application, but the pattern appears throughout distributed systems wherever a grant must survive holder failure.

    Application What Is Leased On Expiry
    Leadership The right to coordinate Another node campaigns
    Distributed lock Exclusive access to a resource Lock becomes available
    Shard ownership Responsibility for a data range Range is reassigned
    Cache entry Permission to serve cached data Entry is revalidated
    Read lease Right to serve reads locally Reads escalate to the leader
    Work claim Right to process a queue item Item is redelivered
    Session Authenticated state Re-authentication required
    Service registration Presence in service discovery Instance is deregistered

    Read Leases

    A particularly useful application lets a leader serve linearizable reads without a round trip. Holding a read lease means no other node can have become leader within the lease window.

    async function leasedRead(leader, key) {
        if (!leader.readLease.isSafelyHeld()) {
            const confirmed = await leader.confirmLeadershipViaQuorum();
    
            if (!confirmed) {
                return { ok: false, redirect: leader.knownLeaderId };
            }
    
            leader.readLease.grant(confirmed.token);
        }
    
        return {
            ok: true,
            value: leader.stateMachine.get(key),
            servedFrom: "lease"
        };
    }
    The read lease payoff: Reads within a valid lease skip the quorum round trip entirely. This is often the single largest latency improvement available in a consensus-backed store.

    Work Claims

    -- Claim an item with a processing lease
    UPDATE work_queue
    SET claimed_by = :worker_id,
        claim_token = claim_token + 1,
        claim_expires_at = CURRENT_TIMESTAMP + INTERVAL '60 seconds'
    WHERE item_id = (
        SELECT item_id
        FROM work_queue
        WHERE status = 'pending'
          AND (claim_expires_at IS NULL
               OR claim_expires_at < CURRENT_TIMESTAMP)
        ORDER BY created_at
        LIMIT 1
        FOR UPDATE SKIP LOCKED
    )
    RETURNING item_id, claim_token;
    Why Queues Use Leases A worker that crashes mid-item does not strand it. The claim expires, and another worker picks it up, without any dead-worker detection mechanism.

    Where the Grantor Lives

    Grantor Strength Weakness
    Consensus service Safe under partition Additional dependency
    Relational database Already present, transactional Availability tied to the database
    Single cache node Low latency Unsafe, grants can be lost on failover
    Object storage Durable, conditional writes Higher latency per operation
    Self-managed peers No external dependency You are implementing consensus
    The Single-Node Grantor Hazard A grantor without replicated durable state can lose its record of who holds what during failover, then grant the same lease twice. The lease mechanism is only as safe as its grantor.

    Failure Modes

    Failure Symptom Mitigation
    Grantor unavailable No renewals possible Surrender at margin, degrade gracefully
    Holder pause beyond lease Acts after expiry Fencing tokens at the resource
    Clock jump on holder Believes lease is longer Monotonic clock, not wall clock
    Renewal storm Grantor overloaded Longer leases, jittered renewal
    Lease thrashing Rapid holder changes Increase duration and margin
    Duplicate grant Two valid holders Consensus-backed grantor
    Orphaned work Item claimed but never completed Claim expiry and redelivery
    Jitter renewal timing: Many holders renewing on identical intervals create synchronized load spikes on the grantor. Randomize the interval slightly to spread them.

    Monitoring

    Signals Worth Tracking

    • Lease acquisition and loss rate
    • Renewal success and failure counts
    • Renewal latency against the lease period
    • Remaining margin at renewal time
    • Surrenders caused by margin exhaustion
    • Time a resource spends unleased
    • Holder changes per resource per hour
    • Fencing rejections following expiry
    • Observed process pause durations
    • Clock skew across holders
    • Grantor request volume and availability
    Track the margin remaining at each renewal, not just whether renewal succeeded. A margin trending toward zero predicts the outage before it happens.

    Verification

    describe("lease safety", function () {
        it("surrenders when renewal cannot reach the grantor", async function () {
            const holder = await acquireLease("resource-a");
    
            await faultInjector.blockGrantor(holder.id);
            await sleep(leaseMs);
    
            expect(holder.lease.isSafelyHeld()).toBe(false);
            expect(holder.surrendered).toBe(true);
        });
    
        it("does not grant twice within one lease period", async function () {
            const first = await acquireLease("resource-a");
            const second = await tryAcquireLease("resource-a", "other-node");
    
            expect(second.acquired).toBe(false);
            expect(first.lease.isSafelyHeld()).toBe(true);
        });
    
        it("rejects work from a holder that paused past expiry", async function () {
            const holder = await acquireLease("resource-a");
            const oldToken = holder.lease.token;
    
            await faultInjector.pauseProcess(holder.id, leaseMs * 2);
            const replacement = await acquireLease("resource-a", "other-node");
    
            const staleWrite = await resource.write("value", oldToken);
    
            expect(staleWrite.accepted).toBe(false);
            expect(replacement.lease.token).toBeGreaterThan(oldToken);
        });
    });

    Common Design Mistakes

    Weak Design

    • Treating a lease as a safety guarantee
    • Using wall-clock time for expiry
    • Renewing at the last moment
    • Acting with zero remaining margin
    • Granting from a non-durable store
    • Leases shorter than observed pauses
    • Synchronized renewal across all holders
    • Not releasing on graceful shutdown
    • Never testing grantor unavailability

    Strong Design

    • Pairs leases with fencing tokens
    • Measures elapsed time monotonically
    • Renews at a fraction of the lease
    • Surrenders on margin exhaustion
    • Uses a consensus-backed grantor
    • Sizes leases above pause percentiles
    • Jitters renewal intervals
    • Releases explicitly on shutdown
    • Drills grantor outage and holder pause

    System Design Interview Discussion

    Question What Your Answer Should Cover
    Why a lease rather than a lock? Reclaim without holder cooperation
    What does a lease depend on? Bounded clock error on both sides
    Is a lease sufficient for safety? No, pauses require fencing
    How long should the lease be? Above pause and latency percentiles
    When do you renew? Fraction of lease, absorbing failures
    How do you revoke early? Token advance at the resource
    Where does the grantor live? Consensus-backed durable store
    What else uses leases? Read leases, work claims, registration

    Design Checklist

    Production Checklist

    • Grant leases from a consensus-backed store
    • Issue a monotonic token with every grant
    • Validate tokens at the protected resource
    • Measure lease elapsed time monotonically
    • Define a safety margin and stop before expiry
    • Renew at one third of the lease period
    • Jitter renewal intervals across holders
    • Surrender on margin exhaustion, not failure count
    • Size leases above the pause and latency percentiles
    • Have the grantor wait beyond nominal expiry
    • Release explicitly on graceful shutdown
    • Support revocation via token advance
    • Re-check the lease inside long operations
    • Monitor margin at renewal and surrender causes
    • Test grantor outage, holder pause, and clock jumps

    Knowledge Check

    1

    What problem does a lease solve that a lock does not?

    Reclaim without holder cooperation. A crashed or partitioned holder cannot release a lock, but a lease expires on its own.

    2

    Why use monotonic rather than wall-clock time?

    Wall-clock time can jump backwards during synchronization, making an expired lease appear still valid. Monotonic clocks never regress.

    3

    Why is a lease alone insufficient for safety?

    Process pauses have no upper bound, so a holder may resume after expiry and act. Only resource-side fencing makes that harmless.

    4

    Why renew at a fraction of the lease?

    It allows several consecutive renewal failures to be absorbed before the lease lapses, rather than losing it on any single transient error.

    5

    What is a read lease used for?

    Serving linearizable reads without a quorum round trip, because holding it means no other node can have become leader within the window.

    Summary

    A lease is a time-bounded grant. Its defining property is that reclaim requires no cooperation from the holder, which solves the problem locks cannot: a crashed or partitioned holder blocking progress indefinitely.

    This convenience comes from trading a dependency on holder liveness for a dependency on clocks. Two independent measurements are involved, and their errors are asymmetric: the holder stopping early and the grantor reclaiming late both cost only availability, while the reverse costs correctness.

    Safe implementation requires monotonic time, a margin before acting, and renewal at a fraction of the lease so transient failures are absorbed. Duration balances fast reclaim against thrashing, and must exceed observed pause and latency percentiles.

    Critically, a lease is a liveness optimization rather than a safety mechanism. Process pauses are unbounded, so overlap cannot be eliminated by timing alone. Pairing leases with fencing tokens validated at the resource is what makes the overlap harmless, and the pattern extends well beyond leadership to read leases, work claims, shard ownership, and service registration.

    Key Takeaway

    A lease bounds how long you wait, not how long someone acts. Use monotonic clocks with a safety margin, renew at a fraction of the period, surrender when margin runs out rather than when renewal fails, grant from a consensus-backed store, and always pair the lease with fencing at the resource it protects.