DDoS
Distributed Denial of Service
Understand how volumetric, protocol, and application-layer attacks exhaust different resources, why rate limiting alone is insufficient, and how to layer defenses that absorb rather than merely detect an attack.
Prerequisites
Recommended Knowledge
- TCP, UDP, and the connection handshake
- DNS resolution and record types
- TLS handshake cost and session resumption
- Load balancers, proxies, and CDNs
- Rate limiting algorithms and identity keys
- Caching, autoscaling, and capacity planning
- Timeouts, circuit breakers, and load shedding
- Observability and anomaly detection
What Makes an Attack Distributed
A denial-of-service attack exhausts a finite resource so legitimate requests cannot be served. It becomes distributed when traffic originates from many sources simultaneously, which defeats the obvious defense of blocking the offender.
The attacker's objective is simply to push offered demand above available capacity at the narrowest point in the system. That point may be bandwidth, connection slots, CPU, memory, database connections, or a third-party dependency.
Simple Analogy
A crowd fills a shop, occupying every aisle without purchasing anything. Staff cannot serve genuine customers, who leave. Barring one person accomplishes nothing when hundreds are involved and each looks ordinary.
The Three Attack Layers
| Layer | Resource Exhausted | Measured In | Where It Is Stopped |
|---|---|---|---|
| Volumetric | Network bandwidth | Bits per second | Upstream provider or scrubbing centre |
| Protocol | Connection state and tables | Packets per second | Network edge and load balancer |
| Application | CPU, memory, database, dependencies | Requests per second | Application edge and service logic |
Volumetric Attacks
Volumetric attacks saturate the network path. The target's servers may be entirely healthy, but the pipe leading to them is full, so legitimate packets are dropped before arrival.
Amplification and Reflection
Attackers rarely possess the bandwidth to flood a large target directly. Instead they abuse protocols where a small request produces a large response, and forge the source address so the response lands on the victim.
| Mechanism | How It Works | Why It Is Possible |
|---|---|---|
| Reflection | Responses redirected to a forged source | Connectionless protocols do not verify the sender |
| Amplification | Small query yields a large reply | Some responses are far larger than requests |
| Botnet flooding | Many compromised devices send directly | Aggregate bandwidth exceeds the target |
| Carpet bombing | Traffic spread across an address range | No single address crosses a detection threshold |
Volumetric Mitigations
- Absorb traffic across a distributed edge network
- Route through scrubbing infrastructure during attack
- Keep origin addresses unpublished and unreachable directly
- Block traffic upstream at the provider level
- Disable or restrict unnecessary connectionless services
- Use anycast so load disperses across many locations
- Establish provider escalation paths before an incident
Protocol Attacks
Protocol attacks exhaust connection state rather than bandwidth. They exploit the fact that a server must allocate memory before knowing whether a client is genuine.
| Technique | Mechanism | Resource Consumed |
|---|---|---|
| Half-open connections | Handshake initiated but never completed | Connection state table |
| Slow request delivery | Headers sent a few bytes at a time | Connection slots and worker threads |
| Slow response consumption | Response read extremely slowly | Send buffers and open sockets |
| Handshake flooding | Repeated expensive cryptographic negotiation | CPU on the terminating server |
| Connection churn | Rapid open and close cycles | Ephemeral ports and table entries |
client_header_timeout 8s;
client_body_timeout 8s;
send_timeout 10s;
keepalive_timeout 15s;
limit_conn_zone $binary_remote_addr zone=perip:10m;
limit_conn perip 24;
limit_req_zone $binary_remote_addr zone=reqs:10m rate=20r/s;
limit_req zone=reqs burst=40 nodelay;
Application-Layer Attacks
These are the hardest to defend. Requests are well-formed, often authenticated, and individually indistinguishable from normal use. The attacker targets asymmetry: requests that cost little to send but a great deal to serve.
The higher this ratio, the fewer attackers are needed. An endpoint where one request triggers a large scan, complex aggregation, or third-party call is far more attractive than a cached static resource.
| Target | Why It Is Chosen | Countermeasure |
|---|---|---|
| Search with rare terms | Bypasses cache, fans out across shards | Cost-weighted limits and query complexity caps |
| Deep pagination | Large offsets force expensive scans | Cursor pagination and maximum depth |
| Report generation | Heavy aggregation per request | Asynchronous jobs with queue limits |
| Login endpoint | Deliberately slow password hashing | Pre-authentication challenges and strict limits |
| Cache-busting parameters | Unique query strings force origin fetches | Normalize keys and ignore unknown parameters |
| File upload | Consumes bandwidth, storage, and processing | Size caps, quotas, and authenticated access |
| Nested API queries | One request expands into many operations | Depth and complexity limits |
Layered Defense Architecture
Network Edge
Absorbs volumetric floods through distributed capacity and anycast routing, dropping traffic far from the origin. This is the only layer with enough bandwidth to matter.
Content Delivery Network
Serves cacheable responses without contacting the origin, removing most legitimate load and shrinking the surface an attacker can reach.
Web Application Firewall
Applies signature and behavioural rules, geographic controls, and challenges to suspicious clients before requests reach application infrastructure.
Load Balancer
Terminates connections, enforces timeouts and concurrency caps, and isolates unhealthy instances from the rotation.
API Gateway
Applies identity-aware rate limits, validates requests, and rejects malformed or oversized payloads cheaply.
Application Services
Enforces cost-based limits, concurrency bounds, load shedding, and graceful degradation of non-essential functionality.
Caching as a Defense
Caching is among the most effective and least discussed mitigations. Every request served from cache is a request the origin never processes.
A high hit ratio means attack traffic against cacheable content is largely absorbed. Attackers therefore attempt to bypass cache deliberately.
Cache Bypass Techniques
- Random query string parameters
- Requests for non-existent resources
- Headers that vary the cache key
- Forcing revalidation on every request
- Targeting intentionally uncacheable endpoints
Cache Hardening
- Normalize and allow-list cache key parameters
- Cache negative responses briefly
- Restrict which headers affect the key
- Serve stale content while revalidating
- Collapse concurrent identical origin requests
Distinguishing Attackers from Users
Since application-layer traffic looks legitimate, defense depends on accumulating weak signals rather than finding one decisive indicator.
| Signal | Suspicious Pattern | Reliability |
|---|---|---|
| Request timing | Perfectly regular intervals | Moderate, easily randomized |
| Navigation path | Deep endpoint hit with no prior page | Good for browser traffic |
| Header fingerprint | Inconsistent or minimal header set | Moderate, can be forged |
| TLS fingerprint | Client library signature, not a browser | Good, harder to spoof convincingly |
| Session history | No cookies or prior interaction | Good for authenticated surfaces |
| Client-side challenge | Fails to execute required computation | Strong against simple tooling |
| Account maturity | Newly created, immediately high volume | Strong when authentication is required |
| Endpoint distribution | Traffic concentrated on one costly route | Strong behavioural indicator |
Progressive Response
function selectResponse(signals, systemState) {
const score = computeRiskScore(signals);
if (systemState.underAttack && score > 0.9) {
return { action: "block", reason: "high-confidence attack" };
}
if (score > 0.7) {
return { action: "challenge", type: "proof-of-work" };
}
if (score > 0.5) {
return { action: "challenge", type: "lightweight" };
}
if (score > 0.3 && systemState.loadFactor > 0.8) {
return { action: "throttle", delayMs: 400 };
}
return { action: "allow" };
}
Proof of Work and Challenges
Challenges invert the cost asymmetry. The client must perform work before the server commits resources, which is negligible for one user and prohibitive across a botnet.
{
"challengeId": "chl-7f2a91",
"algorithm": "sha256-prefix",
"prefix": "a3f91c",
"difficulty": 20,
"issuedAt": "2026-09-24T05:02:11Z",
"expiresAt": "2026-09-24T05:04:11Z"
}
| Challenge Type | User Impact | Effectiveness |
|---|---|---|
| Transparent script check | None if scripting is supported | Stops basic tooling |
| Computational proof of work | Brief delay | Raises cost proportionally to volume |
| Interactive challenge | Noticeable friction | Strong, but harms accessibility |
| Device attestation | None on supported platforms | Strong, limited coverage |
| Required authentication | Blocks anonymous access | Very strong where acceptable |
Load Shedding and Degradation
When demand exceeds capacity despite every filter, the system must choose what to abandon. Shedding deliberately is far better than failing arbitrarily.
| Priority | Traffic Class | Behaviour Under Pressure |
|---|---|---|
| Critical | Payments, authentication, safety operations | Protected until the last possible moment |
| High | Core reads for authenticated users | Served, possibly from stale cache |
| Normal | Browsing and general anonymous traffic | Challenged or throttled |
| Low | Recommendations, analytics, enrichment | Disabled early |
| Deferrable | Exports, reports, batch operations | Queued or rejected immediately |
function admissionControl(request, metrics) {
const pressure = Math.max(
metrics.cpuUtilization,
metrics.queueDepth / metrics.queueCapacity,
metrics.dbConnectionsInUse / metrics.dbConnectionLimit
);
if (pressure < 0.7) {
return { admit: true };
}
const priority = classifyPriority(request);
const threshold = SHED_THRESHOLDS[priority];
if (pressure > threshold) {
return {
admit: false,
status: 503,
retryAfterSeconds: 30,
reason: `shedding ${priority} traffic at ${pressure.toFixed(2)}`
};
}
return { admit: true, degraded: pressure > 0.85 };
}
Autoscaling Is Not a Defense
Scaling to meet attack traffic converts an availability problem into a financial one. The service stays up while costs rise without limit, which some attacks specifically intend.
Scaling Alone
- Costs grow with attack volume
- Scaling lags behind sudden surges
- Downstream dependencies do not scale alike
- Database connection limits become the new ceiling
Bounded Scaling
- Hard maximum on instance count
- Spend alerts wired to on-call
- Filtering applied before scaling triggers
- Shedding engaged once the ceiling is reached
DNS as a Target
DNS is frequently overlooked. If name resolution fails, the application is unreachable regardless of how well the application tier is defended.
DNS Resilience
- Use a provider with distributed anycast capacity
- Deploy authoritative servers across multiple providers
- Set TTLs balancing failover speed against query volume
- Enable response rate limiting on authoritative servers
- Monitor resolution from multiple external vantage points
- Avoid a single registrar as a point of failure
- Protect registrar accounts with strong authentication
Detection and Monitoring
Detection must distinguish an attack from a legitimate traffic surge, because the correct response differs entirely.
| Characteristic | Legitimate Surge | Likely Attack |
|---|---|---|
| Onset | Gradual or event-correlated | Abrupt and unexplained |
| Endpoint spread | Distributed across the application | Concentrated on one costly route |
| Geography | Matches the user base | Unusual regional concentration |
| Session depth | Multiple pages per visitor | Single request per source |
| Cache ratio | Stable or improving | Sudden collapse |
| Conversion signals | Proportionate to traffic | Near zero despite high volume |
| Client diversity | Varied agents and versions | Narrow, repeated fingerprints |
Metrics Worth Alerting On
- Inbound bandwidth against provisioned capacity
- New connections per second at the edge
- Half-open connection count
- Request rate by endpoint and by source
- Cache hit ratio trend
- Origin request volume behind the CDN
- Latency percentiles at each tier
- Queue depth and rejection counts
- Database connection saturation
- Challenge issuance and pass rates
- Autoscaling events and spend rate
Incident Response
Confirm
Verify that degradation stems from an attack rather than a deployment, dependency failure, or genuine popularity.
Classify
Determine the layer under attack, since volumetric, protocol, and application attacks require different responses.
Contain
Engage edge protections, tighten limits, enable challenges, and disable low-priority functionality.
Escalate
Contact the upstream provider or scrubbing service when the volume exceeds what your edge can absorb.
Protect the Core
Preserve critical paths such as authentication and payments while accepting degradation elsewhere.
Recover and Review
Restore normal settings gradually, then analyse which control proved decisive and which thresholds need adjustment.
Common Design Mistakes
Weak Design
- Relying on application-level filtering alone
- Leaving the origin directly reachable
- Treating autoscaling as the mitigation
- Blocking addresses against a distributed attack
- Ignoring DNS in the threat model
- No capacity limits on expensive endpoints
- Mitigation requiring a code deployment
- Failing arbitrarily rather than shedding
- Never rehearsing the response
Strong Design
- Layers defenses from provider edge inward
- Restricts origin access to the edge only
- Bounds scaling and monitors spend
- Uses behavioural signals over address blocking
- Hardens DNS across multiple providers
- Applies cost-weighted limits per endpoint
- Exposes runtime toggles for emergency controls
- Sheds by priority with graceful degradation
- Rehearses response with documented runbooks
System Design Interview Discussion
| Question | What Your Answer Should Cover |
|---|---|
| Which layer is under attack? | Volumetric, protocol, or application distinction |
| Where is traffic filtered? | Edge absorption before origin exhaustion |
| How is attack traffic identified? | Layered weak signals and risk scoring |
| What happens at capacity? | Priority-based shedding and degradation |
| Why not simply autoscale? | Cost exposure and dependency ceilings |
| Which endpoint is most vulnerable? | Cost asymmetry analysis |
| How is DNS protected? | Anycast, multiple providers, and TTL strategy |
| How do you avoid blocking real users? | Graduated friction over binary blocking |
Preparedness Checklist
Production Checklist
- Route public traffic through a distributed edge
- Restrict origin access to edge sources only
- Rotate origin addresses if they become exposed
- Maximize cacheability and normalize cache keys
- Enable stale-while-revalidate and request collapsing
- Set aggressive header, body, and idle timeouts
- Cap concurrent connections per source
- Apply cost-weighted limits to expensive endpoints
- Bound pagination depth and query complexity
- Implement priority-based admission control
- Set hard autoscaling ceilings with spend alerts
- Harden DNS across multiple providers
- Expose emergency toggles without deployment
- Establish provider escalation contacts in advance
- Monitor bandwidth, connections, and cache ratio
- Alert on cache collapse and origin surge
- Document and rehearse the response runbook
- Load-test to determine actual breaking points
Knowledge Check
Why are application-layer attacks hardest to stop?
The requests are well-formed and often authenticated, so they are indistinguishable from legitimate traffic at the point where the decision must be made.
What is amplification?
Abusing a protocol where a small request yields a much larger response, with a forged source address directing that response at the victim.
Why is autoscaling insufficient?
It converts an availability problem into an unbounded cost problem, lags sudden surges, and merely shifts the bottleneck to dependencies that do not scale alike.
Why does cache hit ratio matter?
Cached responses never reach the origin, so a high hit ratio absorbs attack traffic. A sudden collapse in that ratio is a strong attack indicator.
Why prefer challenges over blocking?
Blocking legitimate users accomplishes the attacker's objective. Graduated friction raises attacker cost while allowing genuine users through.
Summary
Distributed denial-of-service attacks exhaust a finite resource from many sources at once. They operate at three layers: volumetric attacks saturating bandwidth, protocol attacks consuming connection state, and application attacks exploiting the cost asymmetry between sending and serving a request.
Each layer must be handled where capacity exists to absorb it. Volumetric floods require distributed edge capacity and upstream scrubbing, protocol attacks require connection termination with strict timeouts, and application attacks require identity-aware, cost-weighted limits at the service tier.
Caching is an underrated defense, since every cached response is one the origin never processes. Attackers therefore target cache bypass, which makes key normalization, negative caching, and stale delivery important hardening measures.
Because application-layer traffic resembles legitimate use, defense depends on layered weak signals feeding graduated responses rather than binary blocking. When demand still exceeds capacity, priority-based shedding preserves critical paths while sacrificing non-essential functionality deliberately.
Key Takeaway
Absorb at the edge, shed by priority, and never let mitigation depend on a deployment. Layer defenses from the provider network inward, keep the origin unreachable directly, maximize cacheability, apply cost-weighted limits to your most asymmetric endpoints, bound autoscaling, and rehearse the response before you need it.