Table of Contents

    Threat modeling

    SYSTEM DESIGN • CHAPTER 13.1

    Threat Modeling

    Learn how to systematically identify what you are protecting, where trust boundaries exist, what can go wrong, and which preventive, detective, and recovery controls your architecture actually needs.

    Learning objective: By the end of this article, you will be able to define trust boundaries, build a data-flow view of a system, apply a structured threat-identification method, prioritize risks, and map each risk to a defensive control.

    Prerequisites

    Threat modeling is a design activity rather than a coding exercise, but the following background makes it far more effective:

    Recommended Knowledge

    • Client-server request flow and API design
    • Authentication and authorization concepts
    • Sessions, tokens, and identity propagation
    • Databases, caches, queues, and object storage
    • Transport security and encryption at rest
    • Service-to-service communication
    • Logging, metrics, and audit trails
    • Rate limiting and overload control
    • Backups, recovery, and incident response basics

    What Is Threat Modeling?

    Threat modeling is a structured process for examining a system design, identifying how it could be attacked or misused, and deciding which defenses are justified. It is a reasoning activity performed on the architecture itself, ideally before the system is built and repeated as the design evolves.

    The goal is not to imagine every conceivable attack. The goal is to find realistic weaknesses early, while changing the design is still inexpensive.

    FOUR CORE QUESTIONS
    What are we building? What can go wrong? What will we do about it? Did we do it well?

    Simple Analogy

    Before constructing a building, an architect reviews the entrances, windows, corridors, and restricted areas, and decides where doors, locks, cameras, and guards belong. Threat modeling applies the same reasoning to software architecture.

    A design review that never asks how the system could be abused is an incomplete design review.

    Why Threat Modeling Matters

    Without Threat Modeling

    • Security is added late and inconsistently
    • Trust boundaries are undocumented
    • Controls are chosen by habit rather than risk
    • Sensitive data spreads into unintended systems
    • Detection and recovery are unplanned
    • Design flaws surface only after an incident

    With Threat Modeling

    • Assets and trust boundaries are explicit
    • Risks are prioritized before implementation
    • Controls map to identified threats
    • Logging supports investigation
    • Recovery expectations are defined
    • Security decisions become reviewable
    Key idea: Design flaws are usually cheaper to fix on a diagram than in a deployed distributed system carrying production data.

    Step 1: Identify Assets

    An asset is anything worth protecting. Before discussing attacks, the team should agree on what actually matters in the system.

    Asset Category Examples Primary Concern
    User data Profiles, contact details, documents, messages Confidentiality and integrity
    Credentials Passwords, tokens, API keys, certificates Unauthorized access
    Business data Orders, pricing, inventory, contracts Integrity and correctness
    Derived data Search indexes, rankings, analytics, models Manipulation and leakage
    Availability APIs, checkout flow, search, notifications Service disruption
    Audit records Access logs, administrative actions Tampering and deletion
    Infrastructure Build systems, deployment pipelines, secrets stores Privileged compromise
    Reputation Brand trust, platform integrity Abuse, spam, and fraud
    ASSET RULE
    If the team cannot name what is most valuable in the system, it cannot decide where the strongest controls belong.

    Step 2: Model the System

    The second step is to describe the system accurately enough to reason about it. A data-flow view is usually more useful than a deployment diagram because attacks follow data and requests.

    1

    External Entities

    Users, administrators, partner systems, third-party APIs, and automated clients that interact with the system but are not controlled by it.

    2

    Processes

    Services, workers, jobs, functions, and components that receive, transform, or act on data.

    3

    Data Stores

    Databases, caches, search indexes, object storage, queues, logs, and configuration or secret stores.

    4

    Data Flows

    The paths along which data moves, including protocol, direction, authentication method, and sensitivity of the content.

    5

    Trust Boundaries

    The points where the level of trust changes, such as between the internet and the API gateway, between services, or between an application and an administrative tool.

    EXAMPLE DATA FLOW
    Browser API Gateway Order Service Database

    Trust Boundaries in Detail

    A trust boundary is any place where data or a request moves from a less trusted context into a more trusted one. Every crossing requires validation, authentication, or authorization decisions.

    Boundary Typical Risk Expected Control
    Internet to API gateway Unauthenticated or malicious requests Transport security, authentication, input validation, rate limits
    Gateway to internal service Forged identity or unauthorized calls Service authentication and verified identity propagation
    Service to database Excess privilege and injection Least-privilege accounts and parameterized queries
    Service to third party Data exposure and untrusted responses Scoped credentials, response validation, timeouts
    Application to admin tooling Privileged misuse Strong authentication, approval flows, audit logging
    Pipeline to derived store Loss of permission context Propagated access rules and filtered indexing
    Build system to production Supply-chain compromise Signed artifacts, restricted deploy credentials, review gates
    Common design gap: Search indexes, caches, analytics stores, and exports frequently inherit data without inheriting the permission model that protected the original source.

    Step 3: Identify Threats with STRIDE

    STRIDE is a widely used mnemonic that helps teams reason about categories of threats rather than relying on memory or intuition. It is applied to each component, data store, and data flow.

    Category Meaning Property Violated Example Concern
    Spoofing Pretending to be another identity Authentication Reusing a stolen session token
    Tampering Unauthorized modification of data Integrity Altering a stored order amount
    Repudiation Denying an action without evidence to the contrary Non-repudiation Missing audit trail for an admin change
    Information disclosure Exposing data to unauthorized parties Confidentiality Private documents appearing in search results
    Denial of service Making the system unavailable or degraded Availability Expensive queries exhausting capacity
    Elevation of privilege Gaining permissions beyond what was granted Authorization A normal user reaching an admin endpoint

    Applying STRIDE Systematically

    Walk through each element of the model and ask the six questions. A structured pass produces more complete coverage than an open-ended discussion.

    Per-Element Questions

    • Who can claim this identity, and how is it verified?
    • Who can modify this data, and how is change detected?
    • What evidence proves who performed this action?
    • Who can read this data, including indirectly?
    • What makes this component expensive or fragile?
    • How could permissions be exceeded on this path?
    • What happens if this dependency is malicious or compromised?
    • What is the impact if this component fails completely?

    Step 4: Consider Attacker Profiles

    Different adversaries have different capabilities, motives, and persistence. Modeling them prevents a design that defends only against the least sophisticated case.

    Profile Typical Motive Design Implication
    Opportunistic external actor Easy financial gain Remove trivially exploitable weaknesses
    Automated bot traffic Scraping, credential stuffing, spam Rate limiting, bot controls, anomaly detection
    Abusive legitimate user Gaming the platform or harming others Quotas, content controls, reporting workflows
    Compromised account Attacker using valid credentials Session monitoring, step-up authentication, anomaly alerts
    Malicious or careless insider Data access beyond job need Least privilege, approvals, immutable audit logs
    Compromised dependency Supply-chain foothold Artifact verification, isolation, restricted credentials
    A valid credential does not prove a legitimate intent. Systems should be able to detect unusual behavior from authenticated identities.

    Step 5: Prioritize the Risks

    Not every identified threat deserves equal investment. Risk prioritization directs effort toward issues with meaningful likelihood and impact.

    RISK MODEL
    \[ Risk = Likelihood \times Impact \]

    Impact may include data exposure, financial loss, service disruption, regulatory consequences, and erosion of user trust. Likelihood depends on exposure, attacker effort, and the presence of existing controls.

    Likelihood Low Impact Medium Impact High Impact
    High Monitor and schedule Address soon Address immediately
    Medium Accept or defer Plan mitigation Address soon
    Low Accept and document Accept or monitor Plan mitigation

    Response Options

    Valid Responses

    • Mitigate by adding or strengthening a control
    • Eliminate by removing the risky capability
    • Transfer through contractual or architectural separation
    • Accept explicitly, with an owner and review date

    Invalid Responses

    • Assuming nobody will find the weakness
    • Assuming internal networks are inherently safe
    • Deferring indefinitely without an owner
    • Relying on undocumented tribal knowledge

    Step 6: Select Layered Controls

    Effective designs layer three kinds of controls. Preventive controls stop an action, detective controls reveal that something happened, and recovery controls restore a safe state.

    1

    Preventive Controls

    Reduce the probability that an attack succeeds.

    Examples include authentication, authorization checks, input validation, output encoding, parameterized queries, encryption, least privilege, network segmentation, quotas, and secure defaults.

    2

    Detective Controls

    Reveal suspicious or unauthorized activity.

    Examples include audit logs, access monitoring, anomaly detection, alerting on privilege changes, integrity checks, and reconciliation between systems.

    3

    Recovery Controls

    Restore correct operation after an incident.

    Examples include backups with tested restores, credential and key rotation, session revocation, rollback procedures, data repair pipelines, and documented incident runbooks.

    DEFENSE IN DEPTH
    Prevent Detect Respond Recover
    LAYERING RULE
    Assume any single control may fail. A design that depends entirely on one check has no margin for error.

    Recording the Threat Model

    A threat model is only useful if it is written down, reviewed, and kept current. A structured, versioned document allows the model to evolve alongside the architecture.

    {
        "system": "product-search-platform",
        "version": "2026.09",
        "owner": "search-platform-team",
        "assets": [
            "product-catalog",
            "customer-accounts",
            "search-index",
            "audit-logs"
        ],
        "trustBoundaries": [
            "internet-to-gateway",
            "gateway-to-services",
            "service-to-database",
            "pipeline-to-search-index"
        ],
        "threats": [
            {
                "id": "THR-014",
                "component": "search-api",
                "category": "InformationDisclosure",
                "description": "Restricted documents could be returned if index permission filters are not applied.",
                "likelihood": "Medium",
                "impact": "High",
                "response": "Mitigate",
                "controls": [
                    "index-level permission filters",
                    "authorization check at query time",
                    "automated permission regression tests"
                ],
                "status": "Planned",
                "reviewDate": "2026-12-01"
            }
        ]
    }

    Each entry should identify the affected component, the threat category, the assessed risk, the chosen response, the concrete controls, the current status, and the next review date.

    Tracking Threats in a Database

    CREATE TABLE threat_model_entries (
        threat_id        VARCHAR(20) PRIMARY KEY,
        system_name      VARCHAR(100) NOT NULL,
        component        VARCHAR(100) NOT NULL,
        stride_category  VARCHAR(30)  NOT NULL,
        description      TEXT         NOT NULL,
        likelihood       VARCHAR(10)  NOT NULL,
        impact           VARCHAR(10)  NOT NULL,
        response_type    VARCHAR(20)  NOT NULL,
        control_summary  TEXT         NOT NULL,
        owner            VARCHAR(100) NOT NULL,
        status           VARCHAR(20)  NOT NULL,
        created_at       TIMESTAMP    NOT NULL,
        review_due_at    TIMESTAMP    NOT NULL
    );
    
    SELECT
        threat_id,
        component,
        stride_category,
        likelihood,
        impact,
        owner,
        review_due_at
    FROM threat_model_entries
    WHERE status <> 'Implemented'
      AND impact = 'High'
    ORDER BY review_due_at;

    Storing threats as tracked records allows them to be reviewed, assigned, reported, and revisited rather than forgotten after the initial design meeting.

    Turning Threats into Tests

    A threat model becomes durable when its conclusions are encoded as automated tests. This prevents a mitigated risk from silently reappearing during later refactoring.

    describe("search authorization boundary", function () {
        it("excludes documents the caller cannot access", async function () {
            const restrictedDocumentId =
                await createRestrictedDocument({ owner: "user-a" });
    
            const response = await searchAsUser("user-b", {
                query: "quarterly plan"
            });
    
            const returnedIds = response.results.map(function (item) {
                return item.documentId;
            });
    
            expect(returnedIds).not.toContain(restrictedDocumentId);
        });
    
        it("rejects requests without a valid identity", async function () {
            const response = await searchWithoutCredentials({
                query: "quarterly plan"
            });
    
            expect(response.statusCode).toBe(401);
        });
    });
    Practical benefit: Authorization regression tests convert a one-time design decision into a continuously enforced guarantee.

    Threat Modeling a Search and Data Pipeline

    Search and ranking systems deserve particular attention because they copy data across multiple stores, each with its own access model.

    Stage Representative Threat Control Direction
    Ingestion Untrusted content enters the pipeline Validate, sanitize, and constrain source input
    Event transport Unauthorized publishing or consumption Authenticated topics and scoped permissions
    Indexing Permission metadata is not carried forward Index access rules alongside document content
    Query serving Restricted documents surface in results Enforce authorization filters at query time
    Ranking signals Artificial engagement manipulates ranking Abuse detection and signal validation
    Deletion Removed records persist in derived stores Propagated deletion and reconciliation
    Analytics export Sensitive fields leave the protected boundary Field-level restriction, masking, and access review
    Expensive queries Resource exhaustion through costly requests Query limits, timeouts, and quotas

    Abuse Cases Versus Use Cases

    A use case describes what a legitimate user should accomplish. An abuse case describes how the same capability could be exploited. Writing both reveals gaps that functional requirements miss.

    Use Case

    • A user requests a password reset
    • A user uploads a product image
    • A user searches the catalog
    • A user submits a review

    Abuse Case

    • Reset requests are used to enumerate valid accounts
    • Uploads are used to distribute harmful files
    • Search is automated to scrape the full catalog
    • Reviews are mass-generated to distort ranking
    Every feature that creates value for a legitimate user also creates a capability that someone may attempt to misuse.

    When to Perform Threat Modeling

    Threat modeling is most valuable as a recurring practice rather than a one-time milestone. Architecture changes routinely introduce new boundaries and new exposure.

    Natural Trigger Points

    • During initial architecture design
    • Before adding an externally exposed endpoint
    • When introducing a new data store or derived dataset
    • When integrating a third-party dependency
    • When changing the authentication or permission model
    • When handling a new category of sensitive data
    • When splitting or merging services
    • After a security incident or near miss
    • On a scheduled periodic review

    Common Mistakes

    Weak Practice

    • Treating internal traffic as automatically trusted
    • Modeling only the happy path
    • Producing a document nobody revisits
    • Listing threats without assigning owners
    • Focusing only on prevention
    • Ignoring derived stores and exports
    • Excluding operations and data teams
    • Confusing a compliance checklist with real analysis

    Strong Practice

    • Documents explicit trust boundaries
    • Covers failure and abuse paths
    • Keeps the model versioned and current
    • Assigns an owner and review date per threat
    • Layers prevention, detection, and recovery
    • Extends analysis to every derived dataset
    • Includes cross-functional reviewers
    • Validates mitigations with automated tests

    System Design Interview Discussion

    Question What Your Answer Should Cover
    What are you protecting? Named assets and their sensitivity
    Where are the trust boundaries? Every point where trust level changes
    How is identity established? Authentication and identity propagation between services
    How is authorization enforced? Where checks occur and how they reach derived data
    What is the highest-risk threat? A prioritized risk with reasoning, not a generic list
    How would you detect an attack? Audit logging, monitoring, and anomaly signals
    How do you recover? Revocation, rotation, restore, and repair procedures
    What did you accept? Explicit accepted risks with justification

    Threat Modeling Checklist

    Review Checklist

    • List the assets worth protecting
    • Draw components, data stores, and data flows
    • Mark every trust boundary explicitly
    • Apply STRIDE to each element and flow
    • Write abuse cases beside use cases
    • Consider insider and compromised-account scenarios
    • Rate likelihood and impact for each threat
    • Choose mitigate, eliminate, transfer, or accept
    • Layer preventive, detective, and recovery controls
    • Verify permissions propagate to derived data
    • Confirm deletion reaches every downstream store
    • Ensure sensitive actions produce audit records
    • Assign an owner and review date per threat
    • Encode key mitigations as automated tests
    • Re-review after significant architecture changes

    Knowledge Check

    1

    What is the purpose of threat modeling?

    To identify realistic ways a system could be attacked or misused, and to decide which controls are justified, while the design is still inexpensive to change.

    2

    What is a trust boundary?

    A point where data or a request moves between contexts with different levels of trust, requiring validation, authentication, or authorization.

    3

    What does STRIDE represent?

    Spoofing, tampering, repudiation, information disclosure, denial of service, and elevation of privilege.

    4

    Why are detective controls necessary?

    Preventive controls can fail or be bypassed. Detection provides the evidence needed to notice, investigate, and respond to an incident.

    5

    Why do derived data stores need special attention?

    Search indexes, caches, analytics stores, and exports often copy data without copying the permission model, deletion behavior, or retention rules of the original source.

    Summary

    Threat modeling is a structured design activity that answers four questions: what is being built, what can go wrong, what will be done about it, and whether the work was done well.

    The process begins by naming assets and modeling components, data stores, data flows, and trust boundaries. STRIDE then provides a repeatable way to identify spoofing, tampering, repudiation, information disclosure, denial of service, and privilege escalation concerns across each element.

    Identified threats are prioritized by likelihood and impact, then addressed through mitigation, elimination, transfer, or explicit acceptance. Strong designs layer preventive, detective, and recovery controls rather than depending on a single check.

    In search and data-pipeline architectures, particular attention belongs on derived stores, where permission context, deletion propagation, and export boundaries are most frequently lost.

    Key Takeaway

    Threat modeling turns security from an assumption into a documented design decision. Name what you protect, mark where trust changes, reason systematically about what can go wrong, and layer prevention, detection, and recovery around the risks that actually matter.