overload control
Overload Control
Learn how admission control, rate limiting, concurrency limits, bounded queues, backpressure, load shedding, graceful degradation, circuit breakers, deadlines, retry budgets, prioritization, isolation, caching, and autoscaling work together to keep a system useful when incoming demand exceeds safe processing capacity.
Introduction
Every system has a finite processing capacity. CPU, memory, threads, connections, queues, database locks, storage operations, and downstream service quotas all have limits.
When incoming work exceeds the rate at which the system can safely complete it, the system enters overload.
Incoming workload:
10,000 requests per second
Safe processing capacity:
6,000 requests per second
Excess workload:
4,000 requests per second
Without explicit overload control, excess requests commonly wait in queues, consume memory, occupy threads, hold connections, miss deadlines, and trigger client retries. Those retries further increase the incoming workload.
Traffic exceeds capacity
|
v
Queues grow
|
v
Latency increases
|
v
Requests time out
|
v
Clients retry
|
v
Traffic increases further
|
v
System throughput collapses
Core idea: A resilient system does not attempt to accept unlimited work. It protects useful capacity by limiting, delaying, rejecting, isolating, or degrading work before overload causes widespread failure.
Rejecting a controlled number of requests early can be safer than accepting every request and allowing all of them to time out.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Latency and throughput | Overload appears when arriving work exceeds safe completion capacity. |
| 2 | Load balancing | Traffic must be distributed without overloading individual instances. |
| 3 | API gateways and reverse proxies | Ingress controls can reject excess traffic before it reaches expensive backend paths. |
| 4 | Queues and workers | Bounded queues and backpressure control producer-consumer imbalance. |
| 5 | Retries and idempotency | Uncontrolled retries can amplify an overload condition. |
| 6 | Autoscaling | Additional capacity can help, but it requires detection, provisioning, startup, and readiness time. |
| 7 | Observability | Overload controls require capacity, latency, rejection, queue, and dependency signals. |
What Is Overload?
Overload is a condition in which accepted work requires more of a constrained resource than the system can safely provide.
The constrained resource can be:
- CPU time
- Memory
- Application threads
- Event-loop capacity
- Database connections
- Database locks
- Storage operations
- Network bandwidth
- Queue capacity
- External API allowance
- One database partition or tenant-specific resource
A service can be overloaded even when the complete cluster has unused capacity. For example, one database partition, API route, tenant, or connection pool can be saturated while unrelated resources remain idle.
Overload vs Traffic Spike
| Condition | Meaning |
|---|---|
| Traffic spike | Incoming demand increases suddenly, but available capacity might still be sufficient. |
| Overload | Accepted demand exceeds the safe capacity of at least one required resource. |
| Degradation | The system intentionally provides reduced functionality, freshness, or detail. |
| Saturation | A resource is operating at or near its usable limit. |
| Collapse | The system spends increasing resources managing excess work while useful throughput falls. |
Arrival Rate and Service Rate
Let \(\lambda\) represent the average arrival rate and \(\mu\) represent the safe completion rate.
A simplified stability condition is:
\[ \lambda < \mu \]
When work arrives faster than it can be completed for a sustained period:
\[ \lambda > \mu \]
Queue depth and waiting time will continue increasing unless the system adds capacity, slows producers, rejects work, or reduces the work required per request.
Queueing and Latency
A bounded queue can absorb a short burst. It cannot solve sustained capacity mismatch.
Short burst:
Queue temporarily grows
|
v
Arrival rate returns below capacity
|
v
Workers drain backlog
Sustained overload:
Queue continuously grows
|
v
Waiting time continuously increases
A simplified form of Little's Law is:
\[ WorkInSystem = Throughput \times AverageTimeInSystem \]
If queue depth increases while throughput remains limited, waiting time increases. Many queued requests can become useless because their client deadlines expire before processing begins.
Queue rule: A queue provides controlled buffering. An unbounded queue hides overload until memory, latency, storage, or retention limits are exhausted.
Main Overload Controls
- Admission control
- Rate limiting
- Concurrency limiting
- Bounded queues
- Backpressure
- Load shedding
- Request prioritization
- Workload isolation
- Circuit breaking
- Deadlines and timeouts
- Retry budgets
- Graceful degradation
- Caching and request coalescing
- Autoscaling
Admission Control
Admission control decides whether new work should enter the system.
Incoming request
|
v
Is safe capacity available?
|
+-- Yes:
| accept request
|
+-- No:
reject, defer,
or redirect safely
The decision can consider:
- Active requests
- Queue depth
- Memory pressure
- Available database connections
- Tenant-specific capacity
- Dependency health
- Request priority
- Remaining request deadline
Admission control prevents work from consuming expensive resources when safe completion is unlikely.
Rate Limiting
Rate limiting controls how frequently a caller, tenant, route, or application can start requests.
Incoming request
|
v
Determine trusted limit key
|
v
Check rate-limiting policy
|
+-- Allowance available:
| continue
|
+-- Limit exceeded:
reject or delay
according to policy
Possible limit keys include:
- Authenticated account
- Tenant
- API subscription
- Client application
- API operation
- Trusted network identity
- A combination such as tenant and route
Rate-limit Response
HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 30
{
"type": "about:blank",
"title": "Request rate exceeded",
"status": 429,
"detail": "The caller exceeded the permitted request rate."
}
Return retry guidance only when the server can provide meaningful guidance. Clients should still use bounded retries and jitter.
Common Rate-limiting Algorithms
| Algorithm | General Behaviour | Main Trade-off |
|---|---|---|
| Fixed window | Counts requests inside a fixed time interval | Can permit bursts around window boundaries |
| Sliding window | Estimates activity across a moving interval | Requires more accounting than a fixed window |
| Token bucket | Tokens refill over time and requests consume them | Requires configuration of sustained rate and burst capacity |
| Leaky bucket | Shapes accepted work toward a controlled output rate | Can delay or reject burst traffic depending on implementation |
Token-bucket Model
A token bucket has a refill rate and a maximum token capacity.
Tokens refill at a controlled rate
|
v
Request arrives
|
+-- Token exists:
| consume token
| accept request
|
+-- Bucket empty:
reject or delay request
The refill rate controls sustained traffic. Bucket capacity determines the permitted burst.
Fairness and Per-tenant Limits
One caller or tenant should not consume all shared capacity.
Shared API capacity
|
+-- Tenant A allowance
+-- Tenant B allowance
+-- Tenant C allowance
+-- Protected shared reserve
Useful designs can combine:
- A fleet-wide safety limit
- A per-tenant limit
- A per-user limit
- A route-specific limit
- A protected reserve for critical operations
Use trusted identity context. A client-provided tenant or account identifier must not independently determine the applied limit.
Concurrency Limiting
Rate limits control how frequently work begins. Concurrency limits control how much work can remain active simultaneously.
Concurrency limit:
100 active report generations
Active operations:
100
New request:
Rejected, queued within a bound,
or deferred according to policy
Concurrency limiting is valuable when individual operations have variable or long duration.
Rate vs Concurrency Limit
| Control | Protects Against |
|---|---|
| Rate limit | Too many request starts over a period |
| Concurrency limit | Too many simultaneous in-progress operations |
| Queue limit | Too much waiting work |
| Payload limit | Excessive request or response size |
Bounded Queues
A bounded queue defines the maximum amount of waiting work.
Producer
|
v
Bounded queue
|
v
Consumer
When queue is full:
- Block or slow producer
- Reject new work
- Drop work according to policy
- Redirect work to an approved path
The queue size should reflect useful waiting time, memory or storage capacity, workload value, and recovery objectives.
Backpressure
Backpressure communicates downstream capacity limitations to upstream producers.
Consumer slows down
|
v
Queue approaches its bound
|
v
Producer receives pressure signal
|
v
Producer slows, pauses,
or reduces request creation
Backpressure can appear as:
- A blocked bounded channel
- Reduced consumer demand
- A full queue signal
- A retryable overload response
- Transport flow control
- A reduced concurrency allowance
Backpressure rule: Backpressure is effective only when upstream components respond to the signal. A sender that retries immediately or continues producing unlimited work defeats the control.
Load Shedding
Load shedding intentionally rejects, drops, abandons, or defers selected work when safe capacity is unavailable.
Active demand exceeds safe capacity
|
v
Classify incoming work
|
+-- Critical and feasible:
| process
|
+-- Optional:
| omit or degrade
|
+-- Expired:
| reject immediately
|
+-- Excess low-priority work:
shed
Load shedding protects useful throughput by preventing excess work from consuming resources needed by requests that can still complete successfully.
Rate Limiting vs Backpressure vs Load Shedding
| Control | Primary Question | Typical Location |
|---|---|---|
| Rate limiting | How frequently can this caller start work? | API gateway, proxy, application boundary |
| Backpressure | How does a constrained consumer slow its producer? | Queues, streams, service boundaries |
| Load shedding | Which excess work should be rejected to preserve useful service? | Ingress, application handlers, workers, fan-out paths |
| Admission control | Is enough safe capacity available to accept this work? | Entry points and expensive-operation boundaries |
Request Prioritization
Not every request has the same business or operational importance.
Priority 1:
Authentication and authorization
Priority 2:
Course access and enrollment
Priority 3:
Search suggestions
Priority 4:
Analytics refresh and optional recommendations
During overload, the system can protect essential operations and reduce lower-priority work.
Priority must come from trusted server-side policy. A public client should not be able to label every request as critical.
Workload Isolation
Separate capacity pools can prevent one workload from exhausting resources needed by another.
Interactive API pool
-> User-facing requests
Report-worker pool
-> Expensive report generation
Media-worker pool
-> Video and image processing
Administrative pool
-> Controlled operations
Isolation can be applied by:
- Tenant
- API route
- Workload type
- Priority
- Region
- Queue
- Worker pool
- Database connection pool
Circuit Breakers
A circuit breaker stops repeated calls to a dependency that is failing or too slow.
Dependency calls succeed
|
v
Circuit closed
Failures exceed policy
|
v
Circuit opens
|
v
New calls fail fast
or use approved fallback
Recovery interval passes
|
v
Limited test calls
|
v
Close circuit after recovery
A circuit breaker protects caller resources such as threads, sockets, and request deadlines. It does not repair the failed dependency.
Deadlines and Timeouts
Requests should not consume resources after their useful deadline.
Client deadline:
2 seconds
Gateway work:
100 milliseconds
Service processing:
600 milliseconds
Database work:
500 milliseconds
Response transfer:
200 milliseconds
Remaining time:
Reserved for variation and failure handling
Each downstream operation should receive a timeout that fits within the remaining end-to-end deadline.
Deadline-aware Admission
A request already too old to complete successfully should not enter an expensive queue.
Request reaches queue boundary
|
v
Remaining deadline:
50 milliseconds
Expected queue and processing time:
500 milliseconds
Decision:
Reject or abandon before
consuming full processing capacity
Retry Budgets
Retries compete with new requests for the same capacity.
Original traffic:
5,000 requests per second
Every failed request retries twice
Potential attempts:
Up to 15,000 attempts per second
A retry budget limits retry traffic to a controlled portion of total attempts or available capacity.
Safe retry design includes:
- Bounded attempts
- Exponential backoff
- Randomized jitter
- End-to-end deadlines
- Idempotency
- Retry budgets
- Awareness of overload responses
Retry rule: Do not retry an overload response immediately. Immediate retries transform a controlled rejection into additional load.
Graceful Degradation
Graceful degradation preserves core functionality while reducing optional work.
Possible degraded modes include:
- Hide optional recommendations
- Return fewer search facets
- Serve approved stale public content
- Delay analytics updates
- Reduce image or report detail
- Queue nonessential notifications
- Disable expensive sorting or export options temporarily
The degraded response should be explicit where users or callers need to know that information is delayed, reduced, or temporarily unavailable.
Do Not Degrade Correctness Silently
Some operations require authoritative correctness.
Examples include:
- Payment processing
- Inventory reservation
- Authorization changes
- Financial posting
- Certificate issuance
- Enrollment confirmation requiring a committed transaction
When safe completion is unavailable, reject, defer, or queue the operation according to its contract instead of returning a false success.
Caching during Overload
Caching can reduce repeated work against application and database resources.
Popular request
|
v
Cache lookup
|
+-- Hit:
| return response
|
+-- Miss:
request backend
populate cache
return result
Caching is useful only when the cached response is safe for the caller, tenant, authorization state, and required freshness.
Request Coalescing
Request coalescing allows concurrent requests for the same missing value to share one backend load.
1,000 simultaneous requests
for the same course
|
v
One backend retrieval
|
v
Result shared with
waiting requests
Without coalescing, a popular cache expiration can create a surge against the origin service.
Autoscaling and Overload Control
Autoscaling can add capacity, but it is not immediate.
Overload detected
|
v
Scaling decision
|
v
Provision instance
|
v
Start application
|
v
Pass readiness
|
v
Receive traffic
Rate limiting, admission control, bounded queues, and load shedding protect the service while additional capacity is starting.
Autoscaling also cannot fix:
- A hot database partition
- A serialized global lock
- An inefficient query
- A memory leak
- An exhausted external quota
- An unavailable dependency
HTTP Overload Responses
| Status | Possible Meaning | Important Note |
|---|---|---|
429 Too Many Requests |
The caller exceeded a rate or quota policy | Can include retry guidance when meaningful |
503 Service Unavailable |
The service cannot currently process the request safely | Clients should use bounded retry behaviour |
504 Gateway Timeout |
A gateway did not receive a timely dependency response | Automatic retry may be unsafe for non-idempotent requests |
The selected response should accurately represent the failure contract. Avoid returning a successful status for work that was not accepted or completed.
Conceptual Overload Policy
overloadControl:
admission:
maximumConcurrentRequests: approved-limit
maximumQueueDepth: approved-limit
rateLimits:
anonymous:
policy: approved-anonymous-policy
authenticated:
key: trusted-caller
policy: approved-caller-policy
tenant:
key: trusted-tenant
policy: approved-tenant-policy
priorities:
critical:
reservedCapacity: approved-reserve
normal:
admissionPolicy: standard
optional:
shedFirst: true
retries:
maximumAttempts: approved-limit
exponentialBackoff: true
jitter: true
requireIdempotencyForWrites: true
degradation:
disableOptionalRecommendations: true
deferAnalytics: true
allowApprovedStalePublicContent: true
observability:
metrics: enabled
structuredEvents: enabled
tracing: enabled
This is conceptual configuration. Limits and policies must come from load tests, service objectives, dependency capacity, security requirements, and the selected platform.
PHP Concurrency Guard
<?php
declare(strict_types=1);
final class ConcurrencyGuard
{
public function __construct(
private ConcurrencyStore $store,
private int $maximumConcurrentRequests
) {
}
public function execute(
string $limitKey,
callable $operation
): mixed {
$acquired =
$this->store->tryAcquire(
$limitKey,
$this->maximumConcurrentRequests
);
if (!$acquired) {
throw new RuntimeException(
'The service is currently at its safe concurrency limit.'
);
}
try {
return $operation();
} finally {
$this->store->release(
$limitKey
);
}
}
}
Production implementations require atomic acquisition, lease expiration, crash recovery, tenant isolation, metrics, and an approved failure policy.
PHP Overload Response
<?php
declare(strict_types=1);
function sendOverloadResponse(
int $retryAfterSeconds
): void {
http_response_code(
503
);
header(
'Content-Type: application/problem+json'
);
header(
'Cache-Control: no-store'
);
header(
'Retry-After: ' . $retryAfterSeconds
);
echo json_encode(
[
'type' => 'about:blank',
'title' => 'Service temporarily unavailable',
'status' => 503,
'detail' =>
'The service cannot safely accept additional work.'
],
JSON_THROW_ON_ERROR
);
}
A retry interval should be supplied only when the server has a reasonable basis for the value. Clients must still cap attempts and apply jitter.
Learning-platform Example
Incoming learning-platform traffic
|
v
API Gateway
|
+-- Per-caller rate limit
+-- Per-tenant rate limit
+-- Request-size limit
|
v
API Service
|
+-- Concurrency limit
+-- Deadline validation
+-- Optional feature shedding
|
v
Database and shared services
Suggested Priority Direction
| Operation | Overload Treatment | Reason |
|---|---|---|
| Authentication | Protect reserved capacity | Required for controlled access to the platform |
| Course access | Protect core read capacity and use safe caching | Primary user-facing learning function |
| Enrollment submission | Use admission control and idempotency | Requires correct transactional outcome |
| Search suggestions | Reduce or disable before core course access | Useful but optional enhancement |
| Recommendations | Omit during severe overload | Optional fan-out dependency |
| Analytics refresh | Queue or defer | Does not need to block the interactive request |
| Large report export | Apply strict concurrency and queue limits | Potentially expensive CPU, memory, and database work |
Observability
Useful overload-control metrics include:
- Incoming request rate
- Accepted request rate
- Rejected request rate
- Requests by rejection reason
- Active request count
- Queue depth
- Oldest queued-work age
- Request latency by priority
- Timeout count
- Retry rate
- Circuit-breaker state
- Rate-limit usage by tenant and route
- Load-shedding count
- Degraded-response count
- Database connection usage
- Thread or worker-pool usage
- Cache hit rate
- Autoscaling capacity and delay
Structured Rejection Event
{
"service": "course-api",
"operation": "generate-report",
"decision": "rejected",
"reason": "concurrency-limit",
"priority": "optional",
"policyVersion": "approved-policy-version"
}
Do not record access tokens, session identifiers, private request bodies, or other reusable credentials in overload-control logs.
Alert Conditions
Alert when:
- Rejection rate increases unexpectedly
- Critical traffic is being shed
- Queue depth or age continues increasing
- Concurrency remains at its limit
- Retry traffic exceeds its budget
- Timeouts increase while useful throughput falls
- All instances enter overload protection
- A tenant repeatedly consumes its complete allowance
- A circuit breaker remains open beyond its expected recovery period
- Autoscaling reaches maximum capacity
- Load shedding does not restore latency or throughput
- Database or dependency limits are approached
Troubleshooting Workflow
- Identify the saturated resource.
- Compare arrival rate with successful completion rate.
- Check active requests, queue depth, and oldest-work age.
- Check timeouts and retry traffic.
- Identify which tenant, route, key, or workload dominates demand.
- Check rate and concurrency-limit decisions.
- Check circuit-breaker states and dependency latency.
- Check whether optional fan-out is still active.
- Check load-balancer distribution and unhealthy instances.
- Check autoscaling capacity, startup delay, and maximum limits.
- Check database connections, locks, and hot partitions.
- Confirm that rejected requests are not retried immediately.
- Protect critical operations and reduce optional work.
- Verify recovery gradually before restoring full traffic.
Common Overload-control Mistakes
Accepting Unlimited Work
Queues, threads, memory, and connections eventually become exhausted.
Using an Unbounded Queue
The queue hides overload while waiting time and resource use grow without a controlled boundary.
Applying One Global Limit Only
One caller or tenant can consume the complete shared allowance.
Trusting a Client-supplied Priority
Every caller can claim the highest priority and defeat workload protection.
Retrying Overload Immediately
Controlled rejection becomes additional demand against the overloaded service.
Processing Expired Requests
Resources are spent producing responses that callers no longer need.
Relying Only on Autoscaling
Additional capacity arrives after a delay and can remain constrained by the same downstream bottleneck.
Shedding Work Randomly
Critical and optional operations lose capacity equally even when their business value differs.
Silently Degrading Correctness
The system reports success for work that was not completed correctly.
Ignoring Downstream Limits
The application accepts work that the database, cache, queue, or external service cannot safely process.
Failing Open on Limit-store Errors
A failed distributed limiter can permit unlimited traffic unless an explicit failure policy exists.
Testing Only Ordinary Load
The system remains untested against bursts, retries, hot tenants, dependency failures, and maximum-capacity conditions.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Normal traffic | Legitimate requests complete without unnecessary rejection |
| Short traffic burst | Bounded buffering absorbs the burst without instability |
| Sustained overload | Excess work is rejected or deferred while useful throughput remains stable |
| Per-tenant limit | One busy tenant does not consume every shared resource |
| Concurrency limit | Expensive operations never exceed their safe active count |
| Queue limit | Excess producer traffic receives the documented pressure response |
| Retry storm | Backoff, jitter, deadlines, and retry budgets limit amplification |
| Dependency slowdown | Circuit breaking and admission control preserve caller resources |
| Optional dependency failure | The core response remains available without optional fan-out |
| Maximum autoscaling capacity | Overload controls remain effective after scaling stops |
| Expired deadline | The request is rejected before expensive processing begins |
| Priority enforcement | Critical capacity cannot be claimed through an untrusted client value |
| Limiter-store failure | The documented fail-open or fail-closed policy is applied safely |
| Recovery | Limits relax gradually without causing a second overload wave |
Overload-control Best Practices
Recommended Practices
- Measure the safe capacity of each critical resource.
- Use admission control before expensive work begins.
- Apply both fleet-wide and caller-specific limits.
- Use trusted caller, tenant, route, and priority context.
- Limit concurrency for expensive and long-running operations.
- Use bounded queues with explicit full-queue behaviour.
- Propagate backpressure to upstream producers.
- Shed optional work before critical work.
- Reserve capacity for essential operations.
- Isolate interactive, batch, report, and media workloads.
- Use circuit breakers around failing dependencies.
- Reject requests unlikely to complete before their deadline.
- Use bounded retries, exponential backoff, and jitter.
- Protect state-changing retries with idempotency.
- Use caching and request coalescing for repeated reads.
- Do not silently degrade correctness-sensitive operations.
- Keep overload controls active while autoscaling adds capacity.
- Alert when maximum capacity or critical shedding is reached.
- Record rejection reasons and policy versions.
- Test overload, recovery, dependency failure, and retry amplification.
Practice Exercise
Design overload control for your online learning platform.
Requirements
- Measure the safe request capacity of one API instance.
- Set a fleet-wide concurrency boundary.
- Create separate anonymous, authenticated, and tenant rate limits.
- Reserve capacity for authentication and course access.
- Classify recommendations and analytics as optional work.
- Place report generation behind a bounded queue.
- Limit concurrent report workers.
- Propagate queue pressure to report producers.
- Use deadlines for interactive requests.
- Add circuit breakers for optional dependencies.
- Add idempotency to enrollment submissions.
- Apply bounded retries with backoff and jitter.
- Add approved stale caching for public course information.
- Add request coalescing for popular cache misses.
- Test sustained overload after maximum autoscaling capacity is reached.
- Verify that critical operations remain usable.
- Verify gradual recovery after demand decreases.
Overload-control Design Template
| Workload | Protected Resource | Control | Overload Outcome |
|---|---|---|---|
| Authentication | Identity service and database | Per-caller limit and reserved concurrency | Reject excess attempts without exhausting core capacity |
| Course reads | API, cache, and database | Caching, coalescing, and admission control | Serve safe cached data or a controlled unavailable response |
| Enrollment writes | Transactional database | Concurrency control and idempotency | Accept safely, defer, or reject without duplicate enrollment |
| Recommendations | Optional dependency capacity | Circuit breaker and load shedding | Return core course response without recommendations |
| Report exports | CPU, memory, database, and storage | Bounded queue and worker-concurrency limit | Queue within limits or reject excess requests |
| Video processing | Worker and media-processing capacity | Queue backpressure and processing limits | Delay work without exhausting interactive API resources |
Frequently Asked Questions
What is overload control?
Overload control is the collection of mechanisms that limit, delay, reject, prioritize, or degrade work when demand exceeds safe capacity.
Why not accept every request?
Accepting unlimited work can exhaust queues, memory, threads, connections, and dependencies, causing most or all requests to fail.
What is admission control?
Admission control determines whether enough safe capacity exists to accept new work.
What is rate limiting?
Rate limiting restricts how frequently a caller, tenant, application, or route can start requests.
What is backpressure?
Backpressure is a signal from a constrained consumer instructing upstream producers to slow, pause, or reduce work production.
What is load shedding?
Load shedding intentionally rejects or omits excess work so that the remaining useful work can complete successfully.
Why should queues be bounded?
A bound prevents waiting work from consuming unlimited memory, storage, and time, and forces an explicit overload decision.
What is graceful degradation?
Graceful degradation preserves essential functions while temporarily reducing optional features, detail, freshness, or background work.
Can autoscaling replace overload control?
No. New capacity takes time to become ready and can remain limited by the same database or downstream bottleneck.
Why are immediate retries dangerous?
Immediate retries increase traffic during the overload condition and can consume capacity needed by new requests.
Should every request have the same priority?
Not necessarily. Essential operations can receive protected capacity while optional work is delayed or shed, but priority must come from trusted policy.
How should overload recovery work?
Restore traffic and optional functions gradually while monitoring latency, queues, errors, retries, and dependency capacity to avoid a second overload wave.
Key Takeaway
Overload control keeps a system useful when demand exceeds safe processing capacity. Use admission control to decide whether new work can enter, rate limits to protect fairness and traffic boundaries, concurrency limits to restrict active expensive work, bounded queues to control waiting work, backpressure to slow producers, and load shedding to reject lower-value excess traffic. Protect critical operations with prioritization and workload isolation, stop repeated calls to failing dependencies with circuit breakers, and use deadlines so expired work does not consume capacity. Autoscaling can add resources, but overload controls must protect the system while capacity starts and after maximum scale is reached. Finally, prevent retry amplification, degrade only optional functionality, preserve correctness-sensitive operations, and test both overload and gradual recovery under realistic traffic and dependency failures.