Table of Contents

    audit logs

    SYSTEM DESIGN • CHAPTER 13.7

    Audit Logs

    Understand how audit logs establish an accountable record of who did what, when, and to which resource, and how to design them so the evidence survives the incident it was meant to explain.

    Learning objective: By the end of this article, you will understand how audit logs differ from application logs, what a complete audit event contains, where to emit them reliably, how to make them tamper-evident, and how to retain and query them without creating a new data exposure.

    Prerequisites

    Recommended Knowledge

    • Authentication and identity propagation
    • Authorization models and enforcement points
    • Hashing and digital signatures
    • Transactions and durability guarantees
    • Event logs, queues, and delivery semantics
    • Object storage and retention policies
    • Structured logging and observability basics
    • Trust boundaries from threat modeling

    What an Audit Log Is

    An audit log is a durable, ordered record of security-relevant actions, capturing the acting identity, the operation performed, the resource affected, the outcome, and the time. It exists to answer questions after the fact, under conditions where memory and assumption are insufficient.

    THE AUDIT QUESTION
    Who Did What To What When and With What Result

    Simple Analogy

    A bank ledger records every transaction in sequence, in ink, and is reviewed by someone who did not write it. Its value comes not from being readable but from being difficult for the person it records to alter.

    An audit log is written for a reader who does not trust the system that produced it. That assumption shapes every design decision that follows.

    Audit Logs Versus Application Logs

    Conflating these two is the most common structural mistake. They serve different readers, carry different guarantees, and have different lifetimes.

    Dimension Application Log Audit Log
    Primary reader Engineers debugging behavior Investigators, auditors, compliance
    Content Diagnostic detail and stack traces Business and security-relevant actions
    Loss tolerance Acceptable, sampling is common Loss is itself an incident
    Mutability Freely rotated and deleted Append-only, tamper-evident
    Retention Days to weeks Months to years, policy driven
    Schema Informal and evolving Defined, versioned, stable
    Access control Broad engineering access Restricted and itself audited
    Volume driver Code paths executed Security-relevant actions taken
    Common Failure Treating a debug line as the audit trail. When log levels change, retention expires, or output is sampled, the evidence disappears precisely when an investigation needs it.
    Correct Separation Emit audit events through a dedicated path with its own schema, storage, retention, and access controls, independent of diagnostic logging configuration.

    What to Record

    Volume is not the objective. Recording everything produces noise that obscures the events that matter and inflates cost without improving investigation.

    Category Examples Why It Matters
    Authentication Login, logout, failure, step-up, factor change Establishes account compromise timelines
    Authorization Permission denied, elevation, role assignment Reveals probing and privilege changes
    Sensitive data access Reading regulated or classified records Supports breach scope determination
    Data modification Create, update, delete on protected entities Reconstructs how state changed
    Administrative action Configuration, user management, policy change Highest-impact operations in the system
    Credential lifecycle Key creation, rotation, revocation Links credential state to observed activity
    Bulk operations Exports, mass deletes, large queries Primary signal for data exfiltration
    Audit system events Log access, retention change, export Detects attempts to inspect or alter evidence
    SELECTION RULE
    Record an action if its absence would leave an investigator unable to answer a question they will predictably need to ask.

    Anatomy of an Audit Event

    A well-formed event is self-contained. An investigator reading it years later, without access to the original system, should still understand what occurred.

    {
        "eventId": "evt-8f41c290",
        "schemaVersion": 2,
        "occurredAt": "2026-09-24T04:18:42.317Z",
        "recordedAt": "2026-09-24T04:18:42.492Z",
        "eventType": "report.exported",
        "category": "data-access",
        "severity": "high",
        "outcome": "success",
        "actor": {
            "type": "user",
            "id": "user-4821",
            "displayName": "Example User",
            "authMethod": "oidc",
            "authLevel": "mfa",
            "sessionId": "sess-2b91f0",
            "onBehalfOf": null
        },
        "action": {
            "operation": "export",
            "method": "POST",
            "endpoint": "/api/reports/export"
        },
        "resource": {
            "type": "financial-report",
            "id": "report-7734",
            "tenantId": "tenant-1042",
            "classification": "confidential",
            "recordCount": 4820
        },
        "context": {
            "sourceIp": "203.0.113.47",
            "userAgent": "Mozilla/5.0",
            "requestId": "req-cc71e4",
            "traceId": "trace-9a2d5f",
            "region": "ap-south-1",
            "deviceManaged": true
        },
        "decision": {
            "policyId": "POL-0042",
            "effect": "Permit",
            "reason": "tenant-scoped export permitted for finance role"
        }
    }
    Field Group Purpose Consequence If Missing
    Event identity Unique reference and schema version Duplicates cannot be reconciled
    Timestamps When it happened and when it was recorded Delays are indistinguishable from gaps
    Actor Who performed the action No accountability, the core purpose fails
    Action What operation was attempted Intent cannot be determined
    Resource What was affected, and how much Breach scope cannot be established
    Outcome Success, failure, or denial Attempts and successes become indistinguishable
    Context Network, device, and correlation identifiers Cross-system correlation becomes impossible
    Decision Which rule produced the outcome Cannot explain why access was allowed

    Two Timestamps, Not One

    INGESTION LAG
    \[ L_{\text{audit}} = T_{\text{recorded}} - T_{\text{occurred}} \]

    Recording both reveals pipeline delay. A single timestamp makes a backlogged collector look identical to an ordered sequence, which distorts any timeline reconstruction.

    Capturing the Real Actor

    In distributed systems, the service performing the write is rarely the party that initiated it. Recording only the immediate caller produces an audit trail that names your own services.

    Confused Deputy in the Log Every entry attributes the action to the reporting service account, because the originating user identity was dropped at the first internal hop.
    Identity Propagation Carry the originating principal through every internal call and record both it and the acting service, so delegation is visible rather than hidden.
    {
        "actor": {
            "type": "service",
            "id": "report-service",
            "onBehalfOf": {
                "type": "user",
                "id": "user-4821",
                "sessionId": "sess-2b91f0"
            }
        }
    }

    Actor Types Worth Distinguishing

    • End user acting directly
    • Service acting on a user's behalf
    • Service acting autonomously in a scheduled job
    • Administrator performing privileged operations
    • Support staff impersonating a user with approval
    • External system authenticated by API credential
    Impersonation requires care: When support staff act as a customer, record both identities explicitly. An entry showing only the customer misrepresents who actually performed the action.

    Where to Emit Events

    Emission Point Strength Weakness
    API gateway Central, hard to bypass Lacks business meaning and record detail
    Service layer Knows intent and resource context Must be applied on every code path
    Authorization point Captures permits and denials uniformly Misses actions that bypass the check
    Database triggers Very difficult to circumvent Often lacks the originating user identity
    Change data capture Reliable record of actual mutations Describes effect, not intent or actor
    PLACEMENT RULE
    Emit at the authorization boundary, where both the identity and the resource are known and every path must pass through. Supplement with data-layer capture for high-value tables.

    Durability and Consistency

    An audit event that is lost when the process crashes provides no assurance. The relationship between the business write and the audit write must be deliberate.

    Pattern Behavior Trade-off
    Fire and forget Emit asynchronously, ignore failure Fast, but events are silently lost
    Same transaction Write audit row with the business row Atomic, but couples audit to the datastore
    Transactional outbox Commit locally, relay asynchronously No loss, adds a relay component
    Synchronous to collector Confirm before completing the operation Strong guarantee, collector becomes critical
    Write-ahead Record the attempt before executing Captures crashes, needs outcome reconciliation
    BEGIN;
    
    UPDATE financial_reports
    SET status = 'archived',
        updated_at = CURRENT_TIMESTAMP
    WHERE report_id = :report_id;
    
    INSERT INTO audit_outbox (
        event_id,
        event_type,
        actor_id,
        on_behalf_of_id,
        resource_type,
        resource_id,
        tenant_id,
        outcome,
        payload,
        occurred_at,
        published_at
    )
    VALUES (
        :event_id,
        'report.archived',
        :actor_id,
        :on_behalf_of_id,
        'financial-report',
        :report_id,
        :tenant_id,
        'success',
        :payload_json,
        CURRENT_TIMESTAMP,
        NULL
    );
    
    COMMIT;

    The outbox makes the audit record share the fate of the business change. Either both commit or neither does, and a relay process forwards unpublished rows to durable audit storage.

    Failure policy decision: If audit writing fails, does the operation proceed? For high-sensitivity actions, refusing to act without an audit record is often the correct answer, and it must be a conscious choice.

    Tamper Evidence

    An attacker who gains sufficient access will attempt to remove traces. Audit logs cannot always prevent modification, but they can make modification detectable.

    1

    Append-Only Storage

    Grant write permission without update or delete. The application that produces events should have no ability to alter them afterwards.

    2

    Hash Chaining

    Each entry includes a hash of the previous entry, so removing or editing any record breaks the chain at that point and everywhere after it.

    3

    Separate Trust Domain

    Ship events to an account or system that application credentials cannot administer, so compromising the application does not grant control of its record.

    4

    Immutable Retention

    Apply storage-level policies that prevent deletion before the retention period elapses, even by privileged operators.

    5

    Periodic Anchoring

    Publish a signed digest of the chain at intervals to an external location, bounding how far back undetected rewriting could reach.

    HASH CHAIN
    \[ H_n = \text{SHA-256}(H_{n-1} \parallel E_n) \]
    function buildChainedEntry(event, previousHash) {
        const canonical = canonicalizeJson({
            eventId: event.eventId,
            occurredAt: event.occurredAt,
            eventType: event.eventType,
            actor: event.actor,
            action: event.action,
            resource: event.resource,
            outcome: event.outcome
        });
    
        const entryHash = sha256Hex(previousHash + canonical);
    
        return {
            sequence: event.sequence,
            payload: canonical,
            previousHash: previousHash,
            entryHash: entryHash
        };
    }
    
    function verifyChain(entries, genesisHash) {
        let expectedPrevious = genesisHash;
    
        for (const entry of entries) {
            if (entry.previousHash !== expectedPrevious) {
                return {
                    valid: false,
                    brokenAt: entry.sequence,
                    reason: "previous hash mismatch"
                };
            }
    
            const recomputed = sha256Hex(
                entry.previousHash + entry.payload
            );
    
            if (recomputed !== entry.entryHash) {
                return {
                    valid: false,
                    brokenAt: entry.sequence,
                    reason: "entry hash mismatch"
                };
            }
    
            expectedPrevious = entry.entryHash;
        }
    
        return { valid: true, verifiedCount: entries.length };
    }
    Chaining does not stop deletion. It ensures that deletion cannot happen quietly, which is frequently the more achievable and more useful guarantee.

    Audit Logs as a Data Exposure

    Audit logs concentrate sensitive information and are often retained far longer than the data they describe. They routinely become the least protected copy of regulated data.

    Should Not Appear

    • Passwords or credential values
    • Authentication tokens in full
    • Complete payment card numbers
    • Full record contents after an edit
    • Encryption keys or key material
    • Search query text containing personal data

    Safer Alternatives

    • Stable identifiers rather than values
    • Field names changed, not their contents
    • Hashes for correlation without disclosure
    • Masked fragments where partial display helps
    • Record counts instead of the records
    • Classification labels rather than payloads
    {
        "eventType": "customer.updated",
        "resource": {
            "type": "customer",
            "id": "customer-9021"
        },
        "changes": [
            { "field": "email", "changed": true, "previousHash": "a3f9...", "newHash": "7c2b..." },
            { "field": "phone", "changed": true },
            { "field": "displayName", "changed": false }
        ]
    }

    Recording that a field changed, without recording its value, preserves accountability while avoiding a long-lived copy of the data in a system with different access controls.

    Access must itself be audited: Reading audit logs is a privileged action. Record who queried them, what filters they used, and what they exported.

    Storage and Query Design

    Audit volume grows steadily and is queried rarely but urgently. The access pattern is narrow and predictable: a specific actor, resource, or time range during an investigation.

    CREATE TABLE audit_events (
        sequence        BIGINT       PRIMARY KEY,
        event_id        VARCHAR(64)  NOT NULL UNIQUE,
        occurred_at     TIMESTAMP    NOT NULL,
        recorded_at     TIMESTAMP    NOT NULL,
        event_type      VARCHAR(80)  NOT NULL,
        category        VARCHAR(40)  NOT NULL,
        outcome         VARCHAR(20)  NOT NULL,
        actor_type      VARCHAR(20)  NOT NULL,
        actor_id        VARCHAR(80)  NOT NULL,
        on_behalf_of_id VARCHAR(80)  NULL,
        resource_type   VARCHAR(80)  NULL,
        resource_id     VARCHAR(80)  NULL,
        tenant_id       VARCHAR(80)  NULL,
        source_ip       VARCHAR(45)  NULL,
        trace_id        VARCHAR(64)  NULL,
        payload         JSONB        NOT NULL,
        previous_hash   VARCHAR(64)  NOT NULL,
        entry_hash      VARCHAR(64)  NOT NULL
    );
    
    CREATE INDEX idx_audit_actor_time
        ON audit_events (actor_id, occurred_at DESC);
    
    CREATE INDEX idx_audit_resource_time
        ON audit_events (resource_type, resource_id, occurred_at DESC);
    
    CREATE INDEX idx_audit_tenant_time
        ON audit_events (tenant_id, occurred_at DESC);
    SELECT occurred_at,
           event_type,
           actor_id,
           on_behalf_of_id,
           resource_id,
           outcome,
           source_ip
    FROM audit_events
    WHERE actor_id = :actor_id
      AND occurred_at BETWEEN :window_start AND :window_end
      AND category IN ('data-access', 'administrative')
    ORDER BY occurred_at;
    Tier Typical Age Access Characteristic
    Hot Recent weeks Indexed, queried during active investigation
    Warm Recent months Queryable with higher latency
    Cold Beyond a year Archived, retrieved on request

    Retention

    Retention sits between two opposing pressures. Keeping logs too briefly destroys evidence before slow-moving compromises are discovered. Keeping them indefinitely accumulates liability and cost.

    MINIMUM USEFUL RETENTION
    \[ R_{\text{min}} \geq T_{\text{typical detection delay}} + T_{\text{investigation}} \]
    Consideration Pressure Design Response
    Detection delay Toward longer retention Retain beyond realistic discovery time
    Regulatory obligation Toward defined minimums Set policy per data category
    Storage cost Toward shorter retention Tier and compress older data
    Privacy obligation Toward minimization Store identifiers, not personal content
    Legal hold Suspends deletion Support selective retention override
    Deletion tension: A request to erase personal data may conflict with an obligation to retain audit records. Designing logs around identifiers rather than personal content substantially reduces this conflict.

    From Recording to Detection

    Logs that are only read after an incident provide forensics without prevention. The same stream can drive detection when specific patterns are monitored.

    Patterns Worth Alerting On

    • Repeated authorization denials from one identity
    • Privilege escalation outside a change window
    • Bulk export exceeding a normal threshold
    • Access to unusually many distinct records
    • Administrative action from an unfamiliar network
    • Activity from a dormant account
    • Concurrent sessions from distant locations
    • Changes to audit configuration or retention
    • A sudden drop in expected event volume
    • Hash chain verification failure
    SILENCE IS A SIGNAL
    A collapse in audit volume deserves an alert. Absence of events more often indicates a broken pipeline or suppressed logging than a genuinely quiet period.

    Threats and Mitigations

    Threat Description Mitigation
    Log deletion Attacker removes entries covering their activity Append-only storage in a separate trust domain
    Log alteration Entries edited to misrepresent events Hash chaining with periodic anchoring
    Silent loss Events dropped without anyone noticing Outbox delivery and volume monitoring
    Injection Crafted input forges or breaks log entries Structured fields and strict encoding
    Sensitive data leakage Regulated content stored in the trail Redaction, hashing, and field-level review
    Unauthorized reading Logs browsed to gather information Restricted access, itself audited
    Clock manipulation Timestamps altered to distort sequence Server-assigned time and monotonic sequence
    Volume flooding Noise generated to bury real events Rate awareness and anomaly baselines
    Missing identity Only the service is recorded as actor End-to-end identity propagation

    Monitoring the Audit System

    Signals Worth Tracking

    • Event volume by type against a baseline
    • Ingestion lag between occurrence and recording
    • Unpublished outbox depth and oldest entry age
    • Emission failure rate by service
    • Chain verification results and break location
    • Storage growth against retention projections
    • Query latency for investigation patterns
    • Audit log read and export activity
    • Retention policy and configuration changes
    • Events lacking a resolved actor identity

    Common Design Mistakes

    Weak Design

    • Using application logs as the audit trail
    • Recording only successful operations
    • Attributing actions to service accounts
    • Storing full record contents in events
    • Allowing the application to delete entries
    • Emitting asynchronously with no delivery guarantee
    • Granting broad read access to the trail
    • Leaving retention undefined
    • Never verifying that logging still works

    Strong Design

    • Maintains a separate, versioned audit schema
    • Records denials and failures alongside successes
    • Propagates the originating identity end to end
    • Stores identifiers and change indicators
    • Writes append-only to a separate trust domain
    • Uses the outbox pattern for guaranteed delivery
    • Restricts and audits access to the trail
    • Defines retention per data category
    • Alerts on volume collapse and chain breaks

    System Design Interview Discussion

    Question What Your Answer Should Cover
    What gets audited? Security-relevant categories, not every code path
    Where is the event emitted? Authorization boundary, plus data-layer capture
    How is loss prevented? Transactional outbox and delivery monitoring
    Who is recorded as the actor? Propagated identity and delegation chain
    How is tampering detected? Append-only storage, hash chaining, anchoring
    How long is data kept? Detection delay, regulation, and tiering
    How is it queried? Actor, resource, and time-range access patterns
    Is the log itself sensitive? Redaction and audited access to the trail

    Implementation Checklist

    Production Checklist

    • Define a versioned audit event schema
    • Separate audit output from diagnostic logging
    • Record occurrence and recording timestamps
    • Propagate the originating identity end to end
    • Capture both the actor and any delegating principal
    • Log denials and failures, not only successes
    • Include the decision reason or matched policy
    • Record resource identifiers and affected counts
    • Emit at the authorization boundary consistently
    • Use a transactional outbox for delivery
    • Decide whether audit failure blocks the operation
    • Write append-only to a separate trust domain
    • Chain entries and verify the chain periodically
    • Redact credentials and sensitive values
    • Restrict and audit access to the audit store
    • Define retention and legal hold behavior
    • Index for actor, resource, and time queries
    • Alert on volume collapse and verification failure
    • Rehearse an investigation against real data

    Knowledge Check

    1

    Why separate audit logs from application logs?

    They have different readers, loss tolerance, retention, and access controls. Diagnostic logs are rotated, sampled, and reconfigured, which destroys evidence.

    2

    Why record two timestamps?

    The difference reveals pipeline delay. With one timestamp, a backlogged collector is indistinguishable from an accurate sequence of events.

    3

    Why log denied attempts?

    Failed authorization is the clearest signal of probing or a compromised account. Logging only successes hides reconnaissance entirely.

    4

    What does hash chaining achieve?

    It makes alteration or removal detectable, because each entry depends on the one before it and any break propagates forward.

    5

    Why can audit logs become a liability?

    They concentrate sensitive data, are retained longer than the source records, and are often protected less carefully than the systems they describe.

    Summary

    Audit logs record who performed which action on which resource, when, and with what outcome. They are distinct from application logs in reader, guarantees, retention, and access control, and conflating the two destroys evidence at the moment it is needed.

    A complete event is self-contained, carrying actor, action, resource, outcome, context, decision reason, and both occurrence and recording timestamps. In distributed systems, the originating identity must be propagated so the trail names the real actor rather than an internal service.

    Reliability comes from emitting at the authorization boundary and using the transactional outbox pattern so audit records share the fate of the business change. Tamper evidence comes from append-only storage in a separate trust domain, hash chaining, and periodic anchoring.

    Because audit logs concentrate sensitive data and outlive the records they describe, they must store identifiers and change indicators rather than values, restrict and audit their own access, and follow a deliberate retention policy.

    Key Takeaway

    Design the audit trail for an investigator who does not trust the system that wrote it. Emit at the authorization boundary, guarantee delivery, propagate the real identity, record denials alongside successes, make tampering detectable, and store identifiers rather than the sensitive data itself.