backpressure
Backpressure
Learn how backpressure protects distributed systems when producers generate work faster than consumers can process it. Understand queue growth, Kafka consumer lag, bounded buffers, admission control, rate limiting, concurrency limits, pause and resume, prefetch, visibility timeout, load shedding, autoscaling, retry amplification, dead-letter queues, and overload recovery.
Introduction
Asynchronous messaging separates producers from consumers. A producer can publish work without waiting for the consumer to finish it.
Producer
|
v
Queue or Event Log
|
v
Consumer
|
v
Database or External Service
This separation improves resilience and allows a queue or log to absorb short traffic spikes. However, the messaging system does not create unlimited processing capacity.
When producers continuously generate messages faster than consumers complete them, the backlog grows.
Producer rate:
10,000 messages per second
Consumer completion rate:
6,000 messages per second
Backlog growth:
4,000 messages per second
Backpressure is the set of mechanisms used to slow, limit, defer, reject, or reshape incoming work when downstream capacity is unavailable.
Core idea: A queue absorbs temporary imbalance. Backpressure prevents that temporary imbalance from becoming an unlimited backlog, excessive latency, resource exhaustion, or system-wide failure.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Queues vs logs | Queues expose backlog, while logs commonly expose consumer lag. |
| 2 | Kafka, RabbitMQ, and Amazon SQS | Each platform offers different controls for consumption and buffering. |
| 3 | Consumer groups | Kafka processing capacity depends on partitions and active group members. |
| 4 | Retries | Retries add traffic during failures and can worsen overload. |
| 5 | Dead-letter queues | Poison messages must be removed from the normal processing path. |
| 6 | Rate limiting | Incoming work must sometimes be limited before it enters the queue. |
| 7 | Observability | Backlog age, lag, processing rate, and resource saturation must be measured. |
What Is Backpressure?
Backpressure is a flow-control response that occurs when a downstream component cannot safely accept or process additional work at the current rate.
Downstream capacity decreases
|
v
Backlog begins growing
|
v
System detects pressure
|
v
System slows, defers,
limits, rejects, or sheds work
|
v
Downstream component recovers
Backpressure can propagate toward the producer or be applied at an intermediate layer.
Buffering vs Backpressure
| Concept | Purpose | Limitation |
|---|---|---|
| Buffering | Stores work temporarily when production and consumption rates differ | A buffer eventually reaches its storage, latency, or retention boundary |
| Backpressure | Controls how much additional work can enter or progress through the system | Requires a documented response when capacity is unavailable |
| Load shedding | Rejects or drops selected work to preserve higher-priority processing | Some work is intentionally not accepted or completed |
| Autoscaling | Adds processing capacity | Cannot exceed partition, database, dependency, or cost constraints |
Buffer rule: A queue buys time. It does not remove the difference between production and consumption rates.
Backlog Growth
Let \(P\) represent the producer rate and \(C\) represent the successful consumer-completion rate.
A simplified backlog growth rate is:
\[ BacklogGrowthRate = P - C \]
The backlog grows when:
\[ P > C \]
The backlog remains approximately stable when:
\[ P = C \]
The backlog drains when:
\[ C > P \]
Backlog Drain Time
When consumer capacity exceeds the current producer rate, a simplified backlog-drain estimate is:
\[ EstimatedDrainTime = \frac{ CurrentBacklog }{ ConsumerRate - ProducerRate } \]
This is a planning approximation. Retries, poison messages, unequal message cost, partition imbalance, consumer restarts, and downstream failures can increase actual recovery time.
Backlog Count vs Message Age
Backlog count shows how many messages are waiting. Message age shows how long the oldest work has waited.
| Metric | Question Answered |
|---|---|
| Queue depth | How many messages are waiting? |
| Oldest-message age | How delayed is the oldest waiting operation? |
| Consumer lag | How far is a log consumer behind the latest partition position? |
| Processing rate | How quickly is useful work completing? |
| Arrival rate | How quickly is new work entering? |
| Retry rate | How much additional work is caused by failures? |
Internal reliability guidance explicitly recommends measuring queue message age because count alone does not show whether consumers are falling behind their processing objective.
Causes of Backpressure
| Cause | Example |
|---|---|
| Traffic spike | An examination opens and thousands of learners submit work simultaneously |
| Slow consumer | Every message performs an expensive database query |
| Downstream outage | The notification provider becomes unavailable |
| Database saturation | Consumers exhaust the database connection pool |
| Hot partition | Most Kafka events use the same partition key |
| Poison message | One invalid event repeatedly blocks ordered processing |
| Retry storm | Every failed operation retries immediately |
| Oversized batch | A consumer fetches more work than it can complete safely |
| Consumer deployment | Workers restart together and processing capacity temporarily decreases |
| External rate limit | A downstream API allows fewer calls than the queue receives |
Kafka Backpressure
Kafka producers append records to topic partitions. Consumers retrieve records and maintain offsets. When a consumer group processes records more slowly than producers append them, consumer lag grows.
Kafka Producers
|
v
Topic Partitions
|
v
Consumer Group
|
v
Consumer lag increases
A simplified partition-lag calculation is:
\[ ConsumerLag = LogEndOffset - ConsumerCommittedOffset \]
Kafka Controls
- Monitor lag for every partition
- Limit the number of records fetched per processing cycle
- Pause and resume partition fetching where appropriate
- Add consumers when useful unassigned partitions exist
- Increase partitions through a controlled design when additional parallelism is required
- Protect downstream databases with bounded concurrency
- Use retry and dead-letter topics for failing records
- Throttle producers when retained lag exceeds safe objectives
Kafka Pause and Resume
Consumer detects local saturation
|
v
Pause selected partitions
|
v
Complete bounded in-flight work
|
v
Downstream capacity recovers
|
v
Resume partition fetching
Pausing fetches does not stop producers from appending records. Lag continues growing until consumer processing catches up.
Hot Kafka Partition
Partition 0:
5,000 events per second
Partition 1:
200 events per second
Partition 2:
180 events per second
Partition 3:
210 events per second
Adding consumers cannot divide one partition among multiple active members of the same consumer group. The partition key or workload distribution must be corrected when one partition remains the bottleneck.
RabbitMQ Backpressure
RabbitMQ stores messages in queues and delivers them to consumers. Unacknowledged deliveries represent work already assigned to consumers but not yet completed.
Publisher
|
v
Exchange
|
v
Queue
|
v
Consumer
Pressure signals:
- Ready messages increase
- Unacknowledged messages increase
- Oldest-message age increases
RabbitMQ Controls
- Configure bounded consumer prefetch
- Limit consumer concurrency
- Use separate queues for independent workloads
- Use delayed retries instead of immediate requeue loops
- Route poison messages to a dead-letter destination
- Limit publisher input when backlog exceeds the safe boundary
- Scale consumers only within downstream capacity
RabbitMQ Prefetch
Prefetch bounds the number of unacknowledged deliveries assigned to a consumer according to the applicable RabbitMQ configuration.
Prefetch:
10
Consumer can hold:
Up to 10 unacknowledged
deliveries under the
configured scope.
A prefetch value that is too large can overload slow consumers and create uneven work distribution. A value that is too small can reduce throughput when processing is fast.
Amazon SQS Backpressure
SQS stores messages until consumers receive and delete them. Received messages remain hidden from other consumers during the visibility timeout.
Producer
|
v
SQS Queue
|
v
Consumers
Pressure signals:
- Available messages increase
- In-flight messages increase
- Oldest-message age increases
- Messages repeatedly become visible
SQS Controls
- Limit consumer concurrency
- Scale workers using backlog and message-age signals
- Set visibility timeout according to processing requirements
- Use bounded retries through receive-count policies
- Move poison messages to a DLQ
- Use separate queues for workloads with different priorities
- Limit message production when consumers or dependencies are saturated
Visibility Timeout and Pressure
Consumer receives message
|
v
Message becomes invisible
|
v
Processing takes longer
than visibility timeout
|
v
Message becomes available again
|
v
Another consumer receives it
|
v
Duplicate processing adds
more system pressure.
Configure visibility according to normal and maximum processing behaviour. Extend it only through a bounded and observable mechanism when work is still progressing.
Platform Comparison
| Area | Kafka | RabbitMQ | Amazon SQS |
|---|---|---|---|
| Primary pressure signal | Consumer lag and event age by partition | Ready and unacknowledged messages | Available, in-flight, and oldest-message age |
| Consumer flow control | Fetch sizing, pause/resume, group scaling | Prefetch and consumer concurrency | Polling and worker concurrency |
| Parallelism limit | Partitions within a consumer group | Queue, consumer, and downstream capacity | Worker and downstream capacity; FIFO also depends on message-group distribution |
| Failed-message isolation | Retry and dead-letter topics | Retry queues and dead-letter exchanges or queues | Redrive policy and DLQ |
| Main overload risk | Growing retained lag or hot partitions | Queue and unacknowledged-delivery growth | Backlog growth and repeated visibility expiry |
Admission Control
Admission control decides whether new work may enter the system.
New request
|
v
Check current capacity
|
+-- Capacity available:
| accept request
|
+-- Temporary pressure:
| defer or rate limit
|
+-- Critical saturation:
reject or shed
lower-priority work
Admission rules can use:
- Queue depth
- Message age
- Consumer lag
- Database saturation
- Dependency health
- Tenant quota
- Request priority
- Remaining storage or retention capacity
Rate Limiting
Rate limiting bounds how much work a producer, tenant, user, operation type, or consumer can generate during a defined interval.
Incoming messages
|
v
Check rate policy
|
+-- Within limit:
| accept
|
+-- Above limit:
delay, reject,
or route according
to business policy
Apply limits using trusted caller or tenant identity. Do not allow a client to select a higher processing priority through an untrusted request field.
Concurrency Limiting
Concurrency limiting bounds how many operations can execute simultaneously.
Queue contains:
100,000 messages
Safe database concurrency:
20 operations
Consumer pool:
Processes no more than
20 database operations
at the same time.
High worker count does not help when the database, connection pool, external API, or storage service supports only a smaller safe concurrency.
Bounded Buffers
A bounded buffer defines a maximum amount of waiting or in-flight work.
Buffer capacity:
1,000 tasks
Current tasks:
1,000
New task arrives
|
v
System must:
- block
- defer
- reject
- shed
- or route elsewhere
An unbounded in-memory queue can eventually consume the complete memory of a process. A durable queue can avoid process-memory exhaustion but can still exceed latency, retention, storage, or business-validity limits.
Load Shedding
Load shedding intentionally rejects, drops, expires, or deprioritizes selected work to preserve critical processing.
Possible candidates include:
- Optional analytics
- Duplicate refresh requests
- Expired notifications
- Outdated reports
- Low-priority background work
- Requests whose clients have already abandoned the result
Critical transactional work should not be silently dropped. Define explicit failure, rejection, or compensation behaviour.
Priority and Workload Classes
High priority:
Enrollment
Payment confirmation
Assessment submission
Medium priority:
Certificate generation
Notification delivery
Low priority:
Analytics backfill
Historical report generation
Separating workloads into queues, topics, partitions, or worker pools can stop a large low-priority backlog from delaying critical work.
Bulkheads
Bulkheads isolate resource pools so one workload cannot consume every worker, connection, thread, or queue slot.
Interactive Consumer Pool:
Reserved database connections
Batch Consumer Pool:
Separate bounded connections
Migration Consumer Pool:
Separate low-priority capacity
Isolation can be implemented through separate queues, consumer groups, connection pools, worker pools, topics, partitions, or deployment resources.
Fairness
One tenant or message type can consume the complete processing pool.
Tenant A:
90% of queued work
All other tenants:
10% of queued work
Fairness controls can include:
- Per-tenant concurrency limits
- Per-tenant rate limits
- Weighted scheduling
- Separate queues for large tenants
- Round-robin selection across tenant partitions
- Reserved capacity for critical workloads
Autoscaling
Consumers can scale according to backlog, message age, lag, or processing demand.
Backlog increases
|
v
Scaling policy adds consumers
|
v
Consumer completion rate increases
|
v
Backlog begins draining
Autoscaling works only while additional consumers can receive useful work and downstream systems have spare capacity.
Autoscaling Limits
- Kafka consumer parallelism can be limited by partition count.
- One hot partition cannot be divided automatically.
- RabbitMQ consumers can overload a shared database.
- SQS workers can exhaust external API quotas.
- Rapid scaling can create connection and cache-warming spikes.
- Scaling can increase cost without improving useful throughput.
Scaling rule: Scale consumers according to useful downstream capacity, not backlog alone.
Retries and Backpressure
Retries add work when the system is already failing.
Dependency slows down
|
v
Consumer operations time out
|
v
Every operation retries
|
v
Dependency receives more requests
|
v
Dependency slows further
Protect the system using:
- Failure classification
- Capped exponential backoff
- Jitter
- Maximum attempts
- Retry budgets
- Delayed retry queues or topics
- Circuit breakers
- Dead-letter handling
Circuit Breakers
A circuit breaker stops repeated calls to a dependency that is failing persistently.
Consumer calls dependency
|
v
Failure threshold reached
|
v
Circuit opens
|
v
New calls fail fast,
defer, or route elsewhere
|
v
Limited recovery probes
|
v
Normal traffic resumes gradually
A circuit breaker prevents consumers from spending all processing capacity waiting for the same failed dependency.
Dead-Letter Queues
Poison messages should leave the primary processing path after bounded attempts.
Primary queue or topic
|
v
Bounded retry
|
v
Terminal failure
|
v
DLQ
|
v
Investigation and
controlled redrive
A DLQ prevents one permanently invalid message from consuming retry capacity indefinitely. Monitor DLQ depth and oldest unresolved-message age.
Stale-message Handling
A message can become irrelevant while waiting in a long backlog.
Notification created:
Course begins at 10:00
Message processed:
Course already ended
Result:
Notification is no longer useful.
Every message class should define:
- Maximum useful age
- Expiration behaviour
- Whether stale work is dropped, recorded, or dead-lettered
- Whether a replacement current-state operation is required
Conceptual Backpressure Policy
backpressure:
detection:
metrics:
- producer-rate
- consumer-completion-rate
- queue-depth
- oldest-message-age
- consumer-lag
- retry-rate
- downstream-saturation
admission:
enabled: true
perTenantLimits: true
priorityAware: true
consumers:
concurrency: approved-safe-limit
batchSize: approved-bounded-size
autoscaling:
enabled: true
respectDownstreamCapacity: true
buffers:
bounded: true
staleMessagePolicy: workload-specific
retries:
delayed: true
exponentialBackoff: true
jitter: true
maximumAttempts: approved-limit
failureHandling:
deadLetterDestination: approved-dlq
poisonMessageIsolation: enabled
overload:
circuitBreaker: enabled
lowPriorityLoadShedding: workload-specific
observability:
backlogAge: enabled
lagByPartition: enabled
saturation: enabled
rejectedWork: enabled
This configuration is conceptual. Thresholds must be derived from measured processing cost, business latency objectives, retention, dependency capacity, and message-system behaviour.
Java Bounded Consumer Example
public final class BoundedMessageHandler {
private final Semaphore concurrencyLimit;
private final MessageProcessor processor;
public BoundedMessageHandler(
int maximumConcurrentOperations,
MessageProcessor processor
) {
this.concurrencyLimit =
new Semaphore(
maximumConcurrentOperations
);
this.processor = processor;
}
public void handle(
Message message
) throws Exception {
boolean capacityAvailable =
concurrencyLimit.tryAcquire();
if (!capacityAvailable) {
throw new TemporaryCapacityException(
"Consumer capacity is currently exhausted."
);
}
try {
processor.processIdempotently(
message
);
} finally {
concurrencyLimit.release();
}
}
}
This is an educational example. Production consumers require bounded retries, message-system acknowledgment handling, cancellation, graceful shutdown, observability, downstream timeouts, and protection against acknowledging a message before durable processing succeeds.
PHP Admission-control Example
<?php
declare(strict_types=1);
final class QueueAdmissionController
{
public function __construct(
private QueueHealth $queueHealth,
private CapacityPolicy $capacityPolicy
) {
}
public function canAccept(
int $tenantId,
string $workType
): bool {
$health =
$this->queueHealth
->current();
return $this->capacityPolicy
->allows(
tenantId:
$tenantId,
workType:
$workType,
queueDepth:
$health->getDepth(),
oldestMessageAge:
$health->getOldestMessageAge(),
consumerUtilization:
$health->getConsumerUtilization()
);
}
}
Admission decisions should use trusted tenant and workload identity. The application should return a documented response or defer the request rather than silently discarding important work.
Learning-platform Examples
| Workflow | Pressure Risk | Possible Control |
|---|---|---|
| Certificate generation | A course-completion event creates a large task burst | Bound worker concurrency and autoscale within database capacity |
| Enrollment email | The email provider applies a rate limit | Use a bounded delivery rate, delayed retries, and DLQ handling |
| Progress-event projection | Kafka consumer lag grows during peak activity | Monitor lag by partition and scale within partition and database limits |
| Assessment submissions | All learners submit around the deadline | Reserve capacity and prioritize durable submission acceptance |
| Report generation | Large reports consume the same resources as interactive requests | Use a separate queue, bounded consumers, and lower priority |
| Analytics backfill | Historical processing competes with live events | Use a separate consumer group or pool with limited capacity |
| Popular course notifications | One event creates millions of downstream commands | Batch, partition, rate limit, and isolate notification delivery |
Security and Tenant Isolation
- Apply rate and concurrency limits using trusted tenant identity.
- Do not let clients choose privileged processing priority.
- Protect administrative pause, resume, drain, and redrive operations.
- Use least-privilege access for queues, topics, and consumer groups.
- Preserve tenant authorization when messages are delayed or replayed.
- Protect backlog and lag metrics containing customer identifiers.
- Audit load shedding and rejected critical work.
- Encrypt buffered and dead-lettered messages according to policy.
Observability
Useful backpressure metrics include:
- Producer rate
- Consumer completion rate
- Queue depth
- Consumer lag by partition
- Oldest unprocessed-message age
- Ready and in-flight messages
- Consumer concurrency
- Batch size and processing duration
- Retry rate
- Dead-letter rate
- Rejected or shed work
- Database connection usage
- External API throttling
- Worker CPU and memory
- Estimated backlog drain time
Structured Pressure Event
{
"destination": "certificate-generation",
"pressureState": "admission-limited",
"queueDepth": "measured-value",
"oldestMessageAge": "measured-value",
"producerRate": "measured-value",
"consumerCompletionRate": "measured-value",
"constrainedDependency": "certificate-database",
"action": "consumer-concurrency-limited"
}
Avoid including message payloads, authentication credentials, or unnecessary personal information in pressure-control logs.
Alert Conditions
Alert when:
- Producer rate remains above consumer completion rate
- Queue depth continues growing
- Oldest-message age exceeds the processing objective
- Kafka lag increases on one or more partitions
- RabbitMQ ready or unacknowledged messages continue growing
- SQS in-flight messages remain near the safe processing boundary
- Messages repeatedly return after visibility timeout
- Retry traffic increases during dependency saturation
- Dead-letter volume grows
- Consumer scaling does not improve useful throughput
- Database or external-service saturation increases
- Load shedding affects critical work
Troubleshooting Workflow
- Identify the affected queue, topic, partition, or consumer group.
- Compare producer and consumer-completion rates.
- Inspect queue depth, consumer lag, and oldest-message age.
- Determine whether the pressure is global or limited to one partition or tenant.
- Check consumer CPU, memory, connections, and processing duration.
- Check downstream database and external-service saturation.
- Check retry traffic and poison-message loops.
- Check current consumer concurrency and batch size.
- Check Kafka partition count or RabbitMQ prefetch where applicable.
- Check SQS visibility behaviour and repeated receives where applicable.
- Limit intake or lower-priority work if useful capacity is threatened.
- Scale consumers only when partitions and downstream capacity permit it.
- Isolate permanently failing messages.
- Monitor recovery and estimated drain time.
Common Backpressure Mistakes
Treating the Queue as Unlimited Storage
Backlog continues growing until retention, storage, latency, or cost boundaries are exceeded.
Monitoring Count but Not Message Age
A small backlog can still contain work that has missed its business deadline.
Adding Consumers without Checking Downstream Capacity
More workers overload the database, connection pool, storage service, or external API.
Adding Kafka Consumers beyond Partition Count
Additional group members remain without active partition assignments.
Ignoring a Hot Partition
One consumer remains overloaded while other group members stay underused.
Using an Unbounded In-memory Queue
Waiting work consumes process memory until the application becomes unstable.
Retrying Immediately during Overload
Retries increase pressure while the dependency is least able to accept more work.
Using a Very Large RabbitMQ Prefetch
Slow consumers accumulate excessive unacknowledged work and distribute tasks unevenly.
Using a Short SQS Visibility Timeout
Messages reappear before normal processing finishes, creating concurrent duplicate work.
Failing to Isolate Poison Messages
Permanent failures consume processing capacity and can block ordered work.
Applying One Limit to Every Tenant
One large tenant can consume the complete shared processing pool.
Recovering at Full Speed
Releasing every delayed request simultaneously can overload the dependency again and restart the failure cycle.
Recommended Test Cases
| Test | Expected Evidence |
|---|---|
| Short producer spike | The buffer absorbs the spike and drains within the processing objective. |
| Sustained overload | Admission control prevents unlimited backlog growth. |
| Database slowdown | Bounded concurrency protects the database connection pool. |
| Kafka consumer lag | Lag is visible for every affected partition. |
| Kafka hot partition | Adding consumers does not hide the partition-key problem. |
| RabbitMQ slow consumer | Bounded prefetch prevents excessive unacknowledged deliveries. |
| SQS long-running task | Visibility remains valid while bounded processing continues. |
| Retry storm | Backoff, jitter, retry budgets, and circuit breaking limit amplification. |
| Poison message | The message reaches the controlled dead-letter path. |
| Low-priority overload | Critical work continues through isolated capacity. |
| Autoscaling | Consumer capacity increases without saturating downstream systems. |
| Recovery | Traffic and consumer concurrency increase gradually without a second overload. |
Backpressure Best Practices
Recommended Practices
- Measure producer and consumer-completion rates independently.
- Monitor queue depth, lag, and oldest-message age.
- Use bounded buffers rather than unlimited in-memory queues.
- Define the maximum useful age for each message class.
- Apply admission control before critical saturation.
- Use rate limits based on trusted tenant and workload identity.
- Limit consumer concurrency according to downstream capacity.
- Use separate resource pools for interactive and batch work.
- Use Kafka pause and resume for temporary local pressure where appropriate.
- Scale Kafka consumers only within useful partition parallelism.
- Tune RabbitMQ prefetch using measured consumer behaviour.
- Configure SQS visibility according to processing duration.
- Use delayed retries with exponential backoff and jitter.
- Use retry budgets and circuit breakers.
- Move poison messages to a DLQ after bounded attempts.
- Apply priority and fairness policies.
- Shed optional or stale work when business policy allows it.
- Scale consumers only when downstream capacity is available.
- Recover gradually after an overload.
- Test spikes, sustained overload, dependency failure, and recovery.
Practice Exercise
Design backpressure controls for your online learning platform.
Requirements
- Measure the producer rate for course-completion events.
- Measure certificate-consumer completion rate.
- Monitor queue depth and oldest-message age.
- Define the maximum useful delay for certificate generation.
- Apply a bounded consumer-concurrency limit.
- Protect the certificate database connection pool.
- Add delayed retries with backoff and jitter.
- Move poison messages to a DLQ.
- Create separate queues for interactive and batch reports.
- Apply per-tenant report limits.
- Simulate a short traffic spike.
- Simulate sustained overload.
- Scale consumers within downstream capacity.
- Reject or defer low-priority work when capacity is exhausted.
- Verify gradual recovery after the dependency becomes healthy.
Backpressure-design Template
| Decision | Selected Direction | Risk Controlled |
|---|---|---|
| Pressure signal | Backlog age, lag, completion rate, and saturation | Detects overload before complete failure |
| Buffer | Durable and bounded through policy | Prevents uncontrolled memory and storage growth |
| Admission | Tenant-aware rate and priority controls | Prevents one workload from consuming all capacity |
| Consumer concurrency | Limited by downstream capacity | Protects databases and external dependencies |
| Retries | Delayed, bounded, and randomized | Prevents retry amplification |
| Poison messages | Dead-letter after bounded attempts | Prevents permanent errors from blocking useful work |
| Scaling | Based on lag, age, partitions, and downstream headroom | Prevents ineffective or harmful scaling |
| Recovery | Gradual restart and traffic restoration | Prevents a second overload wave |
Frequently Asked Questions
What is backpressure?
Backpressure is flow control used when a downstream component cannot safely accept or process work at the current rate.
Why is a queue not enough?
A queue absorbs temporary imbalance, but a sustained producer-consumer difference creates a continuously growing backlog.
How is backpressure detected?
Compare arrival and completion rates, queue depth, consumer lag, oldest-message age, retry rate, and downstream saturation.
What happens when producers are faster than consumers?
Waiting work accumulates, processing latency grows, and messaging or downstream resources can eventually reach capacity.
How does Kafka represent backpressure?
Kafka consumers fall behind topic partitions, which appears as increasing partition and consumer-group lag.
How does RabbitMQ control consumer flow?
Consumer prefetch and concurrency can limit the number of unacknowledged deliveries assigned to workers.
How does SQS show processing pressure?
Available-message count, in-flight messages, message age, receive count, and repeated visibility expiry provide important signals.
Does adding consumers always solve backpressure?
No. Consumers remain limited by Kafka partitions, downstream capacity, database connections, external quotas, and hot keys.
What is load shedding?
Load shedding intentionally rejects or drops selected work to preserve critical processing during overload.
Why do retries worsen backpressure?
Retries generate additional calls while the dependency can already be slow or overloaded.
What is the most important queue metric?
Both backlog size and oldest-message age are important. Message age reveals whether processing is meeting its business latency objective.
What is a safe backpressure strategy?
Detect pressure early, bound intake and concurrency, prioritize important work, delay retries, isolate permanent failures, scale only within useful capacity, and recover gradually.
Key Takeaway
Backpressure protects a messaging system when producers create work faster than consumers and downstream dependencies can complete it. A queue or log can absorb temporary spikes, but sustained imbalance increases backlog, consumer lag, message age, storage use, and recovery time. Detect pressure using production rate, completion rate, queue depth, per-partition lag, message age, retries, and downstream saturation. Control intake through admission rules, rate limits, bounded buffers, per-tenant fairness, and priority policies. Control processing through bounded concurrency, Kafka pause and resume, RabbitMQ prefetch, SQS visibility configuration, worker isolation, and downstream connection budgets. Use delayed bounded retries, circuit breakers, and DLQs so failures do not amplify load indefinitely. Scale consumers only when useful partitions and downstream capacity are available. Finally, define stale-work and load-shedding policies and restore traffic gradually after recovery to avoid another overload wave.