Table of Contents

    DDoS

    SYSTEM DESIGN • CHAPTER 13.9

    Distributed Denial of Service

    Understand how volumetric, protocol, and application-layer attacks exhaust different resources, why rate limiting alone is insufficient, and how to layer defenses that absorb rather than merely detect an attack.

    Learning objective: By the end of this article, you will understand the three attack layers, amplification and reflection, application-layer attacks that mimic legitimate traffic, layered mitigation from network edge to application, and how to design a system that degrades gracefully instead of collapsing.

    Prerequisites

    Recommended Knowledge

    • TCP, UDP, and the connection handshake
    • DNS resolution and record types
    • TLS handshake cost and session resumption
    • Load balancers, proxies, and CDNs
    • Rate limiting algorithms and identity keys
    • Caching, autoscaling, and capacity planning
    • Timeouts, circuit breakers, and load shedding
    • Observability and anomaly detection

    What Makes an Attack Distributed

    A denial-of-service attack exhausts a finite resource so legitimate requests cannot be served. It becomes distributed when traffic originates from many sources simultaneously, which defeats the obvious defense of blocking the offender.

    THE SERVICE CONDITION
    \[ \text{Service degrades when } D_{\text{offered}} > C_{\text{available}} \]

    The attacker's objective is simply to push offered demand above available capacity at the narrowest point in the system. That point may be bandwidth, connection slots, CPU, memory, database connections, or a third-party dependency.

    Simple Analogy

    A crowd fills a shop, occupying every aisle without purchasing anything. Staff cannot serve genuine customers, who leave. Barring one person accomplishes nothing when hundreds are involved and each looks ordinary.

    The defining difficulty of a distributed attack is not volume. It is that the malicious traffic is indistinguishable from legitimate traffic at the point where you must decide.

    The Three Attack Layers

    Layer Resource Exhausted Measured In Where It Is Stopped
    Volumetric Network bandwidth Bits per second Upstream provider or scrubbing centre
    Protocol Connection state and tables Packets per second Network edge and load balancer
    Application CPU, memory, database, dependencies Requests per second Application edge and service logic
    PLACEMENT PRINCIPLE
    Each layer must be handled where capacity exists to absorb it. A volumetric flood cannot be filtered by application code that the packets never reach.

    Volumetric Attacks

    Volumetric attacks saturate the network path. The target's servers may be entirely healthy, but the pipe leading to them is full, so legitimate packets are dropped before arrival.

    Amplification and Reflection

    Attackers rarely possess the bandwidth to flood a large target directly. Instead they abuse protocols where a small request produces a large response, and forge the source address so the response lands on the victim.

    AMPLIFICATION FACTOR
    \[ A = \frac{S_{\text{response}}}{S_{\text{request}}} \]
    Mechanism How It Works Why It Is Possible
    Reflection Responses redirected to a forged source Connectionless protocols do not verify the sender
    Amplification Small query yields a large reply Some responses are far larger than requests
    Botnet flooding Many compromised devices send directly Aggregate bandwidth exceeds the target
    Carpet bombing Traffic spread across an address range No single address crosses a detection threshold

    Volumetric Mitigations

    • Absorb traffic across a distributed edge network
    • Route through scrubbing infrastructure during attack
    • Keep origin addresses unpublished and unreachable directly
    • Block traffic upstream at the provider level
    • Disable or restrict unnecessary connectionless services
    • Use anycast so load disperses across many locations
    • Establish provider escalation paths before an incident
    Origin exposure defeats edge protection: If the origin address is discoverable through DNS history, certificate records, or email headers, an attacker bypasses the edge entirely. Restrict origin access to the edge provider.

    Protocol Attacks

    Protocol attacks exhaust connection state rather than bandwidth. They exploit the fact that a server must allocate memory before knowing whether a client is genuine.

    Technique Mechanism Resource Consumed
    Half-open connections Handshake initiated but never completed Connection state table
    Slow request delivery Headers sent a few bytes at a time Connection slots and worker threads
    Slow response consumption Response read extremely slowly Send buffers and open sockets
    Handshake flooding Repeated expensive cryptographic negotiation CPU on the terminating server
    Connection churn Rapid open and close cycles Ephemeral ports and table entries
    Why These Are Effective They consume very little attacker bandwidth. A slow-delivery attack can occupy a server's connection capacity using a trickle of traffic that no volumetric detector would flag.
    Structural Defenses Terminate connections at a proxy built for high concurrency, enforce aggressive header and body timeouts, cap per-address concurrent connections, and use stateless handshake validation.
    client_header_timeout   8s;
    client_body_timeout     8s;
    send_timeout            10s;
    keepalive_timeout       15s;
    
    limit_conn_zone $binary_remote_addr zone=perip:10m;
    limit_conn perip 24;
    
    limit_req_zone $binary_remote_addr zone=reqs:10m rate=20r/s;
    limit_req zone=reqs burst=40 nodelay;

    Application-Layer Attacks

    These are the hardest to defend. Requests are well-formed, often authenticated, and individually indistinguishable from normal use. The attacker targets asymmetry: requests that cost little to send but a great deal to serve.

    COST ASYMMETRY
    \[ R_{\text{asymmetry}} = \frac{C_{\text{server}}}{C_{\text{client}}} \]

    The higher this ratio, the fewer attackers are needed. An endpoint where one request triggers a large scan, complex aggregation, or third-party call is far more attractive than a cached static resource.

    Target Why It Is Chosen Countermeasure
    Search with rare terms Bypasses cache, fans out across shards Cost-weighted limits and query complexity caps
    Deep pagination Large offsets force expensive scans Cursor pagination and maximum depth
    Report generation Heavy aggregation per request Asynchronous jobs with queue limits
    Login endpoint Deliberately slow password hashing Pre-authentication challenges and strict limits
    Cache-busting parameters Unique query strings force origin fetches Normalize keys and ignore unknown parameters
    File upload Consumes bandwidth, storage, and processing Size caps, quotas, and authenticated access
    Nested API queries One request expands into many operations Depth and complexity limits
    Audit your own asymmetry: Identify the endpoint where a single request consumes the most server resource. That is where an attacker will concentrate, and it should be the first place you apply cost-based controls.

    Layered Defense Architecture

    DEFENSE PIPELINE
    Provider Edge CDN and WAF Load Balancer Gateway Service
    1

    Network Edge

    Absorbs volumetric floods through distributed capacity and anycast routing, dropping traffic far from the origin. This is the only layer with enough bandwidth to matter.

    2

    Content Delivery Network

    Serves cacheable responses without contacting the origin, removing most legitimate load and shrinking the surface an attacker can reach.

    3

    Web Application Firewall

    Applies signature and behavioural rules, geographic controls, and challenges to suspicious clients before requests reach application infrastructure.

    4

    Load Balancer

    Terminates connections, enforces timeouts and concurrency caps, and isolates unhealthy instances from the rotation.

    5

    API Gateway

    Applies identity-aware rate limits, validates requests, and rejects malformed or oversized payloads cheaply.

    6

    Application Services

    Enforces cost-based limits, concurrency bounds, load shedding, and graceful degradation of non-essential functionality.

    Caching as a Defense

    Caching is among the most effective and least discussed mitigations. Every request served from cache is a request the origin never processes.

    ORIGIN LOAD REDUCTION
    \[ L_{\text{origin}} = R_{\text{total}} \times (1 - H_{\text{cache}}) \]

    A high hit ratio means attack traffic against cacheable content is largely absorbed. Attackers therefore attempt to bypass cache deliberately.

    Cache Bypass Techniques

    • Random query string parameters
    • Requests for non-existent resources
    • Headers that vary the cache key
    • Forcing revalidation on every request
    • Targeting intentionally uncacheable endpoints

    Cache Hardening

    • Normalize and allow-list cache key parameters
    • Cache negative responses briefly
    • Restrict which headers affect the key
    • Serve stale content while revalidating
    • Collapse concurrent identical origin requests
    STALE BEATS ABSENT
    During an attack, serving slightly outdated content is almost always preferable to serving an error. Configure stale delivery before you need it.

    Distinguishing Attackers from Users

    Since application-layer traffic looks legitimate, defense depends on accumulating weak signals rather than finding one decisive indicator.

    Signal Suspicious Pattern Reliability
    Request timing Perfectly regular intervals Moderate, easily randomized
    Navigation path Deep endpoint hit with no prior page Good for browser traffic
    Header fingerprint Inconsistent or minimal header set Moderate, can be forged
    TLS fingerprint Client library signature, not a browser Good, harder to spoof convincingly
    Session history No cookies or prior interaction Good for authenticated surfaces
    Client-side challenge Fails to execute required computation Strong against simple tooling
    Account maturity Newly created, immediately high volume Strong when authentication is required
    Endpoint distribution Traffic concentrated on one costly route Strong behavioural indicator

    Progressive Response

    function selectResponse(signals, systemState) {
        const score = computeRiskScore(signals);
    
        if (systemState.underAttack && score > 0.9) {
            return { action: "block", reason: "high-confidence attack" };
        }
    
        if (score > 0.7) {
            return { action: "challenge", type: "proof-of-work" };
        }
    
        if (score > 0.5) {
            return { action: "challenge", type: "lightweight" };
        }
    
        if (score > 0.3 && systemState.loadFactor > 0.8) {
            return { action: "throttle", delayMs: 400 };
        }
    
        return { action: "allow" };
    }
    False positives are real damage: Blocking legitimate users achieves the attacker's goal on their behalf. Prefer graduated friction over outright blocking unless confidence is very high.

    Proof of Work and Challenges

    Challenges invert the cost asymmetry. The client must perform work before the server commits resources, which is negligible for one user and prohibitive across a botnet.

    ATTACKER COST
    \[ C_{\text{attack}} = N_{\text{requests}} \times W_{\text{challenge}} \]
    {
        "challengeId": "chl-7f2a91",
        "algorithm": "sha256-prefix",
        "prefix": "a3f91c",
        "difficulty": 20,
        "issuedAt": "2026-09-24T05:02:11Z",
        "expiresAt": "2026-09-24T05:04:11Z"
    }
    Challenge Type User Impact Effectiveness
    Transparent script check None if scripting is supported Stops basic tooling
    Computational proof of work Brief delay Raises cost proportionally to volume
    Interactive challenge Noticeable friction Strong, but harms accessibility
    Device attestation None on supported platforms Strong, limited coverage
    Required authentication Blocks anonymous access Very strong where acceptable

    Load Shedding and Degradation

    When demand exceeds capacity despite every filter, the system must choose what to abandon. Shedding deliberately is far better than failing arbitrarily.

    Priority Traffic Class Behaviour Under Pressure
    Critical Payments, authentication, safety operations Protected until the last possible moment
    High Core reads for authenticated users Served, possibly from stale cache
    Normal Browsing and general anonymous traffic Challenged or throttled
    Low Recommendations, analytics, enrichment Disabled early
    Deferrable Exports, reports, batch operations Queued or rejected immediately
    function admissionControl(request, metrics) {
        const pressure = Math.max(
            metrics.cpuUtilization,
            metrics.queueDepth / metrics.queueCapacity,
            metrics.dbConnectionsInUse / metrics.dbConnectionLimit
        );
    
        if (pressure < 0.7) {
            return { admit: true };
        }
    
        const priority = classifyPriority(request);
        const threshold = SHED_THRESHOLDS[priority];
    
        if (pressure > threshold) {
            return {
                admit: false,
                status: 503,
                retryAfterSeconds: 30,
                reason: `shedding ${priority} traffic at ${pressure.toFixed(2)}`
            };
        }
    
        return { admit: true, degraded: pressure > 0.85 };
    }
    Reject Early and Cheaply A rejection issued at the edge costs almost nothing. The same rejection issued after a database query has already consumed the resource you were trying to protect.

    Autoscaling Is Not a Defense

    Scaling to meet attack traffic converts an availability problem into a financial one. The service stays up while costs rise without limit, which some attacks specifically intend.

    Scaling Alone

    • Costs grow with attack volume
    • Scaling lags behind sudden surges
    • Downstream dependencies do not scale alike
    • Database connection limits become the new ceiling

    Bounded Scaling

    • Hard maximum on instance count
    • Spend alerts wired to on-call
    • Filtering applied before scaling triggers
    • Shedding engaged once the ceiling is reached
    The dependency ceiling: Scaling application instances while the database connection pool stays fixed merely relocates the bottleneck. Every tier must scale together or the weakest one defines capacity.

    DNS as a Target

    DNS is frequently overlooked. If name resolution fails, the application is unreachable regardless of how well the application tier is defended.

    DNS Resilience

    • Use a provider with distributed anycast capacity
    • Deploy authoritative servers across multiple providers
    • Set TTLs balancing failover speed against query volume
    • Enable response rate limiting on authoritative servers
    • Monitor resolution from multiple external vantage points
    • Avoid a single registrar as a point of failure
    • Protect registrar accounts with strong authentication

    Detection and Monitoring

    Detection must distinguish an attack from a legitimate traffic surge, because the correct response differs entirely.

    Characteristic Legitimate Surge Likely Attack
    Onset Gradual or event-correlated Abrupt and unexplained
    Endpoint spread Distributed across the application Concentrated on one costly route
    Geography Matches the user base Unusual regional concentration
    Session depth Multiple pages per visitor Single request per source
    Cache ratio Stable or improving Sudden collapse
    Conversion signals Proportionate to traffic Near zero despite high volume
    Client diversity Varied agents and versions Narrow, repeated fingerprints

    Metrics Worth Alerting On

    • Inbound bandwidth against provisioned capacity
    • New connections per second at the edge
    • Half-open connection count
    • Request rate by endpoint and by source
    • Cache hit ratio trend
    • Origin request volume behind the CDN
    • Latency percentiles at each tier
    • Queue depth and rejection counts
    • Database connection saturation
    • Challenge issuance and pass rates
    • Autoscaling events and spend rate

    Incident Response

    1

    Confirm

    Verify that degradation stems from an attack rather than a deployment, dependency failure, or genuine popularity.

    2

    Classify

    Determine the layer under attack, since volumetric, protocol, and application attacks require different responses.

    3

    Contain

    Engage edge protections, tighten limits, enable challenges, and disable low-priority functionality.

    4

    Escalate

    Contact the upstream provider or scrubbing service when the volume exceeds what your edge can absorb.

    5

    Protect the Core

    Preserve critical paths such as authentication and payments while accepting degradation elsewhere.

    6

    Recover and Review

    Restore normal settings gradually, then analyse which control proved decisive and which thresholds need adjustment.

    PREPARATION RULE
    Every emergency control must be togglable without a deployment. If mitigation requires shipping code, it will arrive long after the damage.

    Common Design Mistakes

    Weak Design

    • Relying on application-level filtering alone
    • Leaving the origin directly reachable
    • Treating autoscaling as the mitigation
    • Blocking addresses against a distributed attack
    • Ignoring DNS in the threat model
    • No capacity limits on expensive endpoints
    • Mitigation requiring a code deployment
    • Failing arbitrarily rather than shedding
    • Never rehearsing the response

    Strong Design

    • Layers defenses from provider edge inward
    • Restricts origin access to the edge only
    • Bounds scaling and monitors spend
    • Uses behavioural signals over address blocking
    • Hardens DNS across multiple providers
    • Applies cost-weighted limits per endpoint
    • Exposes runtime toggles for emergency controls
    • Sheds by priority with graceful degradation
    • Rehearses response with documented runbooks

    System Design Interview Discussion

    Question What Your Answer Should Cover
    Which layer is under attack? Volumetric, protocol, or application distinction
    Where is traffic filtered? Edge absorption before origin exhaustion
    How is attack traffic identified? Layered weak signals and risk scoring
    What happens at capacity? Priority-based shedding and degradation
    Why not simply autoscale? Cost exposure and dependency ceilings
    Which endpoint is most vulnerable? Cost asymmetry analysis
    How is DNS protected? Anycast, multiple providers, and TTL strategy
    How do you avoid blocking real users? Graduated friction over binary blocking

    Preparedness Checklist

    Production Checklist

    • Route public traffic through a distributed edge
    • Restrict origin access to edge sources only
    • Rotate origin addresses if they become exposed
    • Maximize cacheability and normalize cache keys
    • Enable stale-while-revalidate and request collapsing
    • Set aggressive header, body, and idle timeouts
    • Cap concurrent connections per source
    • Apply cost-weighted limits to expensive endpoints
    • Bound pagination depth and query complexity
    • Implement priority-based admission control
    • Set hard autoscaling ceilings with spend alerts
    • Harden DNS across multiple providers
    • Expose emergency toggles without deployment
    • Establish provider escalation contacts in advance
    • Monitor bandwidth, connections, and cache ratio
    • Alert on cache collapse and origin surge
    • Document and rehearse the response runbook
    • Load-test to determine actual breaking points

    Knowledge Check

    1

    Why are application-layer attacks hardest to stop?

    The requests are well-formed and often authenticated, so they are indistinguishable from legitimate traffic at the point where the decision must be made.

    2

    What is amplification?

    Abusing a protocol where a small request yields a much larger response, with a forged source address directing that response at the victim.

    3

    Why is autoscaling insufficient?

    It converts an availability problem into an unbounded cost problem, lags sudden surges, and merely shifts the bottleneck to dependencies that do not scale alike.

    4

    Why does cache hit ratio matter?

    Cached responses never reach the origin, so a high hit ratio absorbs attack traffic. A sudden collapse in that ratio is a strong attack indicator.

    5

    Why prefer challenges over blocking?

    Blocking legitimate users accomplishes the attacker's objective. Graduated friction raises attacker cost while allowing genuine users through.

    Summary

    Distributed denial-of-service attacks exhaust a finite resource from many sources at once. They operate at three layers: volumetric attacks saturating bandwidth, protocol attacks consuming connection state, and application attacks exploiting the cost asymmetry between sending and serving a request.

    Each layer must be handled where capacity exists to absorb it. Volumetric floods require distributed edge capacity and upstream scrubbing, protocol attacks require connection termination with strict timeouts, and application attacks require identity-aware, cost-weighted limits at the service tier.

    Caching is an underrated defense, since every cached response is one the origin never processes. Attackers therefore target cache bypass, which makes key normalization, negative caching, and stale delivery important hardening measures.

    Because application-layer traffic resembles legitimate use, defense depends on layered weak signals feeding graduated responses rather than binary blocking. When demand still exceeds capacity, priority-based shedding preserves critical paths while sacrificing non-essential functionality deliberately.

    Key Takeaway

    Absorb at the edge, shed by priority, and never let mitigation depend on a deployment. Layer defenses from the provider network inward, keep the origin unreachable directly, maximize cacheability, apply cost-weighted limits to your most asymmetric endpoints, bound autoscaling, and rehearse the response before you need it.