Table of Contents

    secrets and key rotation

    SYSTEM DESIGN • CHAPTER 13.6

    Secrets and Key Rotation

    Understand how systems store, distribute, and renew credentials without embedding them in code, and how to design rotation that completes reliably instead of stalling halfway.

    Learning objective: By the end of this article, you will understand secret types and sprawl, secret management architecture, dynamic and short-lived credentials, workload identity, dual-credential rotation, blast radius, and how to detect and respond to a leaked secret.

    Prerequisites

    Recommended Knowledge

    • Symmetric and asymmetric key concepts
    • Envelope encryption and key hierarchies
    • TLS certificates and their lifecycle
    • Token signing and verification
    • Environment configuration and deployment pipelines
    • Containers and orchestration basics
    • Access control and least privilege
    • Audit logging and monitoring

    What Counts as a Secret

    A secret is any value that grants access or proves identity, and whose disclosure would allow an unauthorized party to act as a legitimate component of the system.

    Secret Type Purpose Typical Rotation Difficulty
    Database credentials Service authenticates to a datastore Moderate, connection pools must refresh
    API keys Calling an internal or third-party service Depends on provider dual-key support
    Signing keys Producing tokens and signatures Low, with key identifiers and overlap
    Encryption keys Protecting data at rest High if data must be re-encrypted
    TLS private keys Proving server identity Low when renewal is automated
    OAuth client secrets Authenticating a registered client Moderate, coordinate with the provider
    SSH keys Administrative and automated access Moderate, distribution is often manual
    Webhook signing secrets Verifying inbound callbacks Moderate, requires accepting both values
    Service account tokens Workload-to-workload authentication Low when issued dynamically
    If losing a value would let someone impersonate part of your system, it is a secret, regardless of what the configuration file calls it.

    Secret Sprawl

    Secrets multiply quietly. A value created once for convenience gets copied into a second environment, pasted into a message, and referenced by a script, until nobody can enumerate where it exists.

    Where Secrets Accumulate Source repositories, configuration files, container images, continuous integration variables, chat messages, ticket attachments, developer machines, log output, shell history, crash dumps, and infrastructure state files.
    ROTATION FEASIBILITY
    \[ T_{\text{rotation}} \propto N_{\text{copies}} \]

    The practical cost of rotating a secret scales with the number of places it was copied to. This is why sprawl is not merely untidy: it is the reason rotation gets deferred indefinitely.

    CENTRALIZATION RULE
    A secret should have exactly one authoritative source. Everything else should fetch it at runtime rather than hold a copy.

    Secret Management Architecture

    A secret manager centralizes storage, access control, auditing, and rotation. Applications request secrets at startup or on demand rather than receiving them through deployment artifacts.

    RUNTIME RETRIEVAL
    Workload Identity Secret Manager Secret Value In-Memory Use
    1

    Encrypted Storage

    Secrets are encrypted at rest under keys the manager controls, so the underlying storage layer alone does not expose them.

    2

    Identity-Based Access

    Each workload authenticates as itself and receives only the secrets its function requires, rather than a shared application password.

    3

    Versioning

    Multiple versions coexist, so a rotation can introduce a new value while consumers still resolve the previous one during the transition.

    4

    Audit Logging

    Every read, write, and rotation is recorded with the requesting identity, making unusual access patterns detectable.

    5

    Dynamic Issuance

    For supported backends, the manager creates a short-lived credential on request and revokes it on expiry, so no long-lived value exists to leak.

    The Bootstrap Problem

    If secrets live in a manager, the application needs a credential to authenticate to that manager. Storing that credential in configuration simply relocates the original problem.

    Approach Mechanism Assessment
    Static bootstrap token A long-lived token in configuration Recreates the problem it aimed to solve
    Platform-issued identity Runtime attests the workload automatically Preferred, no secret to distribute
    Short-lived bootstrap Single-use token injected at launch Acceptable where platform identity is unavailable
    Hardware attestation Trusted module proves machine identity Strong, requires supporting infrastructure
    Workload Identity The orchestration platform vouches for the workload, and the secret manager trusts that attestation. No credential is written into an image, repository, or configuration file at all.

    Static Versus Dynamic Secrets

    Static Secrets

    • Long lifetime increases exposure
    • Shared across instances and environments
    • Difficult to attribute usage to a caller
    • Rotation requires coordinated updates
    • A leak stays valuable indefinitely

    Dynamic Secrets

    • Created per request with a short lease
    • Unique per workload or session
    • Usage traceable to a specific consumer
    • Expire automatically without intervention
    • A leak has limited practical value
    EXPOSURE WINDOW
    \[ W_{\text{exposure}} = T_{\text{revocation}} - T_{\text{compromise}} \]

    Short leases bound this window automatically. A credential valid for an hour limits damage even when compromise goes undetected, which is the common case.

    Why Rotation Matters

    Rotation replaces a credential with a new value and invalidates the old one. It serves several distinct purposes that are often conflated.

    Driver Reasoning
    Bounding undetected compromise Limits how long a silently leaked value remains useful
    Limiting cryptographic exposure Reduces the volume of data protected by one key
    Personnel changes Removes access held by departing staff
    Exercising the procedure Proves rotation works before an emergency requires it
    Incident response Invalidates credentials after suspected exposure
    Regulatory expectation Demonstrates lifecycle control over credentials
    Underrated benefit: Routine rotation keeps the process working. Organizations that never rotate discover during an incident that their procedure is undocumented, untested, and breaks production.

    The Dual-Credential Pattern

    Naive rotation causes an outage: the old credential is revoked before every consumer has picked up the new one. The dual pattern eliminates that gap by allowing two valid credentials during the transition.

    ROTATION PHASES
    Create New Both Valid Switch Consumers Revoke Old
    1

    Create

    Generate the new credential and register it at the provider without disturbing the existing one. Both are now accepted.

    2

    Publish

    Store the new value as the current version in the secret manager while retaining the previous version.

    3

    Propagate

    Allow consumers to refresh on their normal cadence, or trigger a rolling restart. Verify that traffic has migrated.

    4

    Verify

    Confirm through provider telemetry that the old credential receives no further use. This step is the one most commonly skipped.

    5

    Revoke

    Disable the old credential, retaining the ability to re-enable briefly before permanent destruction.

    CREATE TABLE managed_secrets (
        secret_id        VARCHAR(80)  PRIMARY KEY,
        secret_name      VARCHAR(160) NOT NULL,
        secret_type      VARCHAR(40)  NOT NULL,
        current_version  INT          NOT NULL,
        previous_version INT          NULL,
        owner_team       VARCHAR(80)  NOT NULL,
        rotation_days    INT          NOT NULL,
        last_rotated_at  TIMESTAMP    NOT NULL,
        next_rotation_at TIMESTAMP    NOT NULL,
        last_verified_at TIMESTAMP    NULL
    );
    
    SELECT secret_name,
           secret_type,
           owner_team,
           last_rotated_at,
           next_rotation_at
    FROM managed_secrets
    WHERE next_rotation_at < CURRENT_TIMESTAMP
    ORDER BY next_rotation_at;

    Orchestrating the Rotation

    async function rotateSecret(secretName) {
        const record = await secretStore.getMetadata(secretName);
    
        const newValue = await provider.createCredential({
            secretName: secretName,
            keepExisting: true
        });
    
        const newVersion = await secretStore.addVersion(
            secretName,
            newValue,
            { stage: "current" }
        );
    
        await secretStore.setVersionStage(
            secretName,
            record.currentVersion,
            "previous"
        );
    
        await notifyConsumers(secretName, newVersion);
    
        const migrated = await waitForMigration(secretName, {
            timeoutMs: record.propagationWindowMs
        });
    
        if (!migrated.complete) {
            await alerting.raise({
                severity: "high",
                message: `Rotation stalled for ${secretName}`,
                stillUsingOld: migrated.laggingConsumers
            });
            return { status: "PENDING_MIGRATION" };
        }
    
        await provider.disableCredential(record.currentVersion);
    
        await secretStore.updateMetadata(secretName, {
            previousVersion: record.currentVersion,
            currentVersion: newVersion,
            lastRotatedAt: new Date().toISOString()
        });
    
        return { status: "ROTATED", version: newVersion };
    }
    Why Verification Matters Revoking on a timer rather than on confirmed migration is how rotation causes outages. Proceed only when the old credential demonstrably has no remaining traffic.

    Rotation by Secret Type

    Secret Type Rotation Approach Principal Risk
    Database password Create a second user, migrate, drop the first Pooled connections holding old credentials
    Token signing key Publish new key, sign with it, accept both Verifiers caching an outdated key set
    Data encryption key Encrypt new data with the new key Re-encrypting historical data at scale
    TLS certificate Renew and reload before expiry Expiry causing sudden total failure
    Third-party API key Use provider dual-key support if offered Providers permitting only one active key
    Webhook secret Accept both during an overlap window Rejecting valid inbound callbacks
    SSH key Add new key, remove old after confirmation Losing access to unreachable hosts

    Signing Key Overlap

    {
        "keys": [
            {
                "kid": "signing-2026-09",
                "status": "active",
                "use": "sig",
                "activatedAt": "2026-09-01T00:00:00Z"
            },
            {
                "kid": "signing-2026-06",
                "status": "deprecated",
                "use": "sig",
                "acceptUntil": "2026-09-30T00:00:00Z"
            }
        ]
    }

    Because tokens signed by the previous key remain in circulation until they expire, the overlap window must exceed the maximum token lifetime. Removing the old key sooner invalidates valid tokens.

    MINIMUM OVERLAP
    \[ T_{\text{overlap}} \geq T_{\text{max token lifetime}} + T_{\text{key cache TTL}} \]

    Connection Pools and Cached Credentials

    Long-lived connections are the most frequent cause of rotation failures. A pool established at startup keeps using the old credential until connections are recycled.

    Handling Cached Credentials

    • Set a maximum connection lifetime in the pool
    • Fetch credentials through a provider, not a fixed variable
    • Cache with a TTL shorter than the overlap window
    • Refresh on authentication failure before failing the request
    • Trigger a rolling restart if refresh is unsupported
    • Expose the loaded credential version as a metric
    class CredentialProvider {
        constructor(secretName, options) {
            this.secretName = secretName;
            this.ttlMs = options.ttlMs;
            this.cached = null;
            this.fetchedAt = 0;
        }
    
        async get({ forceRefresh = false } = {}) {
            const age = Date.now() - this.fetchedAt;
    
            if (!forceRefresh && this.cached && age < this.ttlMs) {
                return this.cached;
            }
    
            this.cached = await secretManager.getCurrent(this.secretName);
            this.fetchedAt = Date.now();
    
            return this.cached;
        }
    }
    
    async function queryWithRetry(sql, params) {
        try {
            return await database.query(sql, params);
        } catch (error) {
            if (!isAuthenticationError(error)) {
                throw error;
            }
    
            await credentialProvider.get({ forceRefresh: true });
            await database.reconnect();
    
            return await database.query(sql, params);
        }
    }

    Blast Radius

    The impact of a compromised secret depends on how much it unlocks and how many systems share it. Partitioning credentials limits what a single leak achieves.

    Wide Blast Radius

    • One credential shared by every service
    • Same value across all environments
    • Administrative permissions by default
    • No expiry on the credential
    • Usage cannot be attributed to a caller

    Narrow Blast Radius

    • A distinct credential per workload
    • Fully separate values per environment
    • Minimum permissions for the function
    • Short lease with automatic expiry
    • Every use attributable to one consumer
    Environment Mixing Reusing production credentials in staging or development means a compromise of the least protected environment yields production access. Environments must never share secret values.

    Detecting Leaked Secrets

    Secrets leak through routine developer activity far more often than through sophisticated attack. Detection must operate continuously rather than at review time.

    Control When It Acts Coverage
    Pre-commit hook Before code is committed Prevents local introduction
    Pipeline scanning During continuous integration Blocks merges containing secrets
    History scanning Across past commits Finds previously committed values
    Image scanning On container build Detects baked-in credentials
    Log inspection Continuously in production Catches accidental logging
    Anomaly detection On secret access patterns Flags unusual retrieval behavior
    Deleting the commit is insufficient: Once a secret reaches a repository, it persists in history, clones, forks, caches, and mirrors. The only reliable response is to rotate the value.

    Emergency Rotation

    Scheduled rotation prioritizes availability. Emergency rotation prioritizes containment, and accepts disruption as the cost of closing exposure quickly.

    Aspect Scheduled Emergency
    Trigger Policy interval Confirmed or suspected exposure
    Overlap window Generous Minimal or none
    Old credential Retired after verification Revoked immediately
    Disruption tolerance None expected Accepted deliberately
    Approval path Standard change process Incident authority
    Follow-up Routine record Access review and root cause analysis

    Emergency Response Sequence

    • Revoke the exposed credential at the provider
    • Issue a replacement and publish it
    • Restart or refresh affected workloads
    • Review logs for use of the exposed value
    • Determine what the credential could reach
    • Rotate anything the credential could have unlocked
    • Remove the value from the leak source
    • Record the timeline and correct the root cause

    Threats and Mitigations

    Threat Description Mitigation
    Committed credentials Secret pushed into version control Scanning at commit, pipeline, and history
    Secrets in images Values baked into container layers Runtime injection and image scanning
    Environment variable exposure Values visible in process listings and dumps Fetch at runtime, avoid crash dump inclusion
    Logged secrets Credentials written during debugging Structured redaction and log review
    Shared credentials One value used across many services Per-workload identity and credentials
    Never-expiring secrets Indefinitely valid long-lived values Mandatory expiry and scheduled rotation
    Failed rotation Old value revoked before migration completed Dual credentials with verified cutover
    Certificate expiry Renewal missed, service becomes unreachable Automated renewal and early alerting
    Over-privileged secret Credential grants more than required Least privilege scoping per workload
    Secret manager compromise Central store becomes a single point of failure Strict access control, audit, and anomaly alerts

    Monitoring and Auditing

    Signals Worth Tracking

    • Age of every secret against its rotation policy
    • Count of secrets overdue for rotation
    • Secrets lacking an assigned owner
    • Rotation success and failure rates
    • Consumers still using a deprecated version
    • Retrieval volume by identity and secret
    • Access from unexpected workloads or regions
    • Authentication failures following a rotation
    • Certificate expiry countdown
    • Scanning alerts from repositories and images
    • Secret manager availability and latency
    A rotation that completes without verifying that the old credential stopped being used has not finished. It has simply stopped being observed.

    Common Design Mistakes

    Weak Design

    • Committing secrets and planning to remove them later
    • Treating base64 encoding as protection
    • Sharing one credential across all services
    • Reusing production values in lower environments
    • Revoking the old value on a fixed timer
    • Rotating manually with no runbook
    • Granting administrative permissions by default
    • Renewing certificates by hand
    • Deleting a leaked commit without rotating

    Strong Design

    • Fetches secrets at runtime from one source
    • Uses platform workload identity for bootstrap
    • Issues a distinct credential per workload
    • Keeps environments fully isolated
    • Revokes only after verified migration
    • Automates rotation end to end
    • Scopes each credential to least privilege
    • Automates certificate renewal with alerting
    • Rotates immediately on any suspected exposure

    System Design Interview Discussion

    Question What Your Answer Should Cover
    Where do secrets live? Central manager with runtime retrieval
    How does a service authenticate to it? Workload identity and the bootstrap problem
    How is rotation performed? Dual credentials, propagation, and verified cutover
    What breaks during rotation? Connection pools and cached credentials
    How long is the overlap? Token lifetime plus cache TTL
    What if a secret leaks? Emergency rotation and downstream impact review
    How is blast radius limited? Per-workload scoping and environment isolation
    What if the manager is unavailable? Cache windows and degradation behavior

    Implementation Checklist

    Production Checklist

    • Inventory every secret and assign an owner
    • Store secrets in one authoritative manager
    • Retrieve at runtime rather than at build time
    • Use platform workload identity where available
    • Issue a distinct credential per workload
    • Scope permissions to the minimum required
    • Keep environments completely isolated
    • Prefer dynamic short-lived credentials
    • Support two valid versions during rotation
    • Verify migration before revoking the old value
    • Bound connection lifetime in pools
    • Refresh credentials on authentication failure
    • Automate certificate renewal and alert early
    • Scan commits, history, images, and logs
    • Redact secrets from logs and error output
    • Audit every retrieval and rotation event
    • Document and rehearse emergency rotation
    • Track rotation age and overdue secrets

    Knowledge Check

    1

    What is the bootstrap problem?

    A service needs a credential to reach the secret manager, which recreates the original storage problem unless the platform issues the identity automatically.

    2

    Why use dual credentials during rotation?

    Consumers update at different times. Accepting both values during the transition prevents an outage between publishing the new credential and revoking the old one.

    3

    Why do connection pools break rotation?

    They authenticate once at connection creation and reuse that connection, so the old credential remains in use until the connection is recycled.

    4

    Why must a leaked secret be rotated rather than deleted?

    Once exposed, it persists in history, clones, mirrors, and caches. Only invalidating the value removes its usefulness.

    5

    Why prefer dynamic secrets?

    They expire automatically, are unique per consumer, make usage attributable, and sharply limit the value of any leak.

    Summary

    Secrets are the credentials that let components prove identity and obtain access. They accumulate copies across repositories, images, pipelines, and machines, and that sprawl is the reason rotation gets postponed.

    A secret manager provides one authoritative source with encrypted storage, identity-based access, versioning, and auditing. Platform-issued workload identity resolves the bootstrap problem without writing a credential into any artifact.

    Rotation bounds undetected compromise and, equally importantly, keeps the procedure exercised. The dual-credential pattern avoids outages by allowing two valid values, but only if the old value is revoked after verified migration rather than on a timer.

    Connection pools and cached credentials are the usual cause of rotation failures. Narrow blast radius through per-workload credentials and strict environment isolation, scan continuously for leaks, and treat any exposure as grounds for immediate rotation.

    Key Takeaway

    Design rotation before you need it. Keep one authoritative source, fetch at runtime, prefer short-lived dynamic credentials, support two valid versions during cutover, verify migration before revoking, and rehearse the emergency path so it works when exposure is real.