latency vs throughput
Latency vs Throughput in System Design
Learn how latency measures the time required to complete an operation, while throughput measures how much work a system completes over time, and how load, concurrency, queues and bottlenecks connect the two.
Introduction
Latency and throughput are two fundamental performance measurements in system design. They are related, but they answer different questions.
- Latency asks how long one operation takes.
- Throughput asks how many operations are completed during a period.
A system can have high throughput while individual requests remain slow. A system can also process one request very quickly while supporting only a small number of requests per second.
Core distinction: Latency measures delay for an individual operation. Throughput measures completed work over time. Neither metric alone provides a complete description of system performance.
In the System Design curriculum, latency versus throughput is Topic 1.6 under System Design Process and Estimation. It follows high-level diagrams and precedes availability, reliability, QPS, storage and bandwidth estimation.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Functional requirements | Performance is measured for specific system operations. |
| 2 | Non-functional requirements | Latency and throughput targets are quality requirements. |
| 3 | Request flow | Every synchronous stage can contribute to end-to-end latency. |
| 4 | High-level diagrams | Components and dependencies help identify processing stages and bottlenecks. |
| 5 | Basic arithmetic | Rate, percentile and concurrency calculations use simple formulas. |
What Is Latency?
Latency is the elapsed time between the beginning of an operation and the observation of its result at a defined measurement boundary.
Examples include:
- Time from submitting a search to receiving results
- Time from sending an API request to receiving the response
- Time required to execute a database query
- Time from publishing a message until a consumer receives it
- Time from submitting a notification until delivery
Latency Units
| Unit | Symbol | Relationship |
|---|---|---|
| Second | s |
Base time unit |
| Millisecond | ms |
1 second = 1,000 milliseconds |
| Microsecond | µs |
1 millisecond = 1,000 microseconds |
| Nanosecond | ns |
1 microsecond = 1,000 nanoseconds |
Components of End-to-End Latency
End-to-end latency is normally composed of several delays across the request path.
\[ L_{total} = L_{client} + L_{network} + L_{queue} + L_{application} + L_{dependencies} + L_{response} \]
| Component | Description |
|---|---|
| Client latency | Time spent preparing the request or rendering the result |
| Network latency | Time spent transferring data between network locations |
| Queueing latency | Time waiting for processing capacity or a shared resource |
| Application latency | Time spent executing application logic |
| Dependency latency | Time spent waiting for databases, caches or external services |
| Serialization latency | Time spent encoding or decoding data |
| Response latency | Time required to return and display the result |
Sequential Dependencies
When dependencies are called sequentially, their delays contribute one after another to the critical path.
\[ L_{sequential} = L_1 + L_2 + L_3 + \cdots + L_n \]
Client
|
| 20 ms
v
Application
|
| 30 ms
v
Database
|
| 40 ms
v
External Service
Simplified total = 20 + 30 + 40 = 90 ms
Parallel Dependencies
Independent calls can sometimes run concurrently. The dependency portion of latency is then influenced by the slowest required branch rather than by adding every branch duration.
\[ L_{parallel} \approx \max(L_1, L_2, \ldots, L_n) + L_{coordination} \]
+--> Service A: 30 ms --+
Application -+--> Service B: 50 ms --+--> Combine results
+--> Service C: 40 ms --+
Dependency path is influenced by the 50 ms branch,
plus coordination and response assembly.
Parallel calls can reduce elapsed time, but they can increase downstream load and introduce partial-failure handling.
What Is Throughput?
Throughput is the amount of successfully completed work produced by a system during a defined period.
\[ Throughput = \frac{Completed\ Operations}{Time} \]
Throughput can be expressed as:
- Requests per second
- Transactions per second
- Messages processed per second
- Jobs completed per minute
- Records processed per hour
- Bytes transferred per second
Measurement rule: State what counts as completed work. Incoming requests, accepted requests and successfully completed requests can represent different measurements.
Latency vs Throughput
| Comparison Area | Latency | Throughput |
|---|---|---|
| Meaning | Time required for one operation | Amount of work completed over time |
| Primary question | How long does it take? | How much can be completed? |
| Typical unit | Milliseconds or seconds | Requests, transactions or bytes per second |
| User impact | Responsiveness of an individual interaction | Ability to serve the total workload |
| Common visualization | Distribution or percentile chart | Rate over time |
| Typical problem | Slow requests | Insufficient processing capacity |
| Common remedy | Reduce critical-path work and waiting | Remove bottlenecks or add safe parallel capacity |
Highway Analogy
A highway provides a useful conceptual analogy:
- Latency is how long one vehicle takes to complete the journey.
- Throughput is how many vehicles complete the journey per hour.
- Capacity or bandwidth is the maximum potential volume supported by the road.
A wide road can have high potential capacity but poor actual throughput during congestion. One vehicle can also complete a journey quickly while the road carries very few vehicles overall.
Average Latency Is Not Enough
Request latency is a distribution. Different requests can take different amounts of time due to queues, cache misses, storage delays, retries, garbage collection, network variation or resource contention.
Latency Percentiles
| Metric | Meaning |
|---|---|
| p50 | Half of measured operations completed at or below this value |
| p90 | Ninety percent completed at or below this value |
| p95 | Ninety-five percent completed at or below this value |
| p99 | Ninety-nine percent completed at or below this value |
| Maximum | The slowest measured operation in the sample |
Percentiles reveal slow requests that an average can hide. The measurement report should also define the sampling period, workload and observation boundary.
Average Can Hide Tail Latency
Nine requests complete in 10 ms.
One request completes in 910 ms.
Average latency:
(9 x 10 + 910) / 10 = 100 ms
The average does not show that most requests were fast
while one request was much slower.
Concurrency
Concurrency is the amount of work in progress at the same time. Concurrency can increase throughput by allowing independent operations to make progress simultaneously.
However, increasing concurrency beyond resource limits can increase queueing, contention, timeouts and failures.
| Low Concurrency | Controlled Concurrency | Excessive Concurrency |
|---|---|---|
| Resources can remain underused | Independent work uses available capacity | Queues and contention increase |
| Throughput can remain low | Throughput improves | Latency and failure rate can rise |
| Simple operation | Requires limits and monitoring | Can overload downstream services |
Little’s Law
For a stable system, Little’s Law relates the average number of operations in the system, average completion rate and average time spent in the system.
\[ L = \lambda W \]
Where:
- \(L\) is the average work in progress
- \(\lambda\) is the average throughput
- \(W\) is the average time in the system
Example
If a system completes 200 requests per second and average end-to-end latency is 0.25 seconds:
\[ L = 200 \times 0.25 = 50 \]
The system has approximately 50 requests in progress on average under those stated conditions.
Important: Little’s Law is a relationship among average values in a stable system. It does not by itself determine the required worker, thread or connection count.
Queueing and Saturation
A request can spend significant time waiting before processing begins. When work arrives faster than the limiting stage can complete it, a queue grows.
\[ Queue\ Growth\ Rate = Arrival\ Rate - Completion\ Rate \]
Arrival rate: 1,200 requests/second
Completion rate: 1,000 requests/second
Queue growth:
1,200 - 1,000 = 200 requests/second
If the difference continues, waiting time, timeouts and memory consumption can increase until requests are rejected or the system fails.
Typical Saturation Pattern
Low load:
Low queueing
Stable latency
Available spare capacity
Rising load:
Higher throughput
Increasing resource utilization
Near saturation:
Queueing rises
Tail latency increases
Little spare capacity
Overload:
Arrival exceeds sustainable completion
Timeouts and failures increase
Throughput can stop improving or decline
The Bottleneck Determines Throughput
A multi-stage system can process work only as fast as its limiting stage over a sustained period.
Gateway capacity: 10,000 requests/second
Application capacity: 5,000 requests/second
Database capacity: 2,000 requests/second
External provider: 1,000 requests/second
The external provider is the limiting stage for a flow
that requires the provider for every completed request.
Adding capacity to a non-limiting stage does not necessarily improve end-to-end throughput.
Can Throughput Increase Without Lower Latency?
Yes. Batching and concurrency can increase completed work without reducing the time required for an individual request. They can sometimes increase individual latency.
Batching Example
Without batching:
Process one event per storage operation.
With batching:
Collect several events.
Write the batch in one storage operation.
Possible effect:
Higher throughput
Fewer storage operations
Additional waiting while the batch fills
Batching is appropriate when throughput gains justify the additional delay and the workload permits grouped processing.
Strategies to Reduce Latency
| Strategy | Primary Benefit | Trade-off or Risk |
|---|---|---|
| Cache eligible reads | Avoid repeated slower data access | Staleness and invalidation complexity |
| Reduce network distance | Lower propagation delay | Regional replication and consistency complexity |
| Remove unnecessary calls | Shorten the critical path | May require redesign or additional local state |
| Run independent calls in parallel | Reduce sequential waiting | Higher downstream concurrency and partial failures |
| Optimize queries and indexes | Reduce storage-access time | Additional storage and write cost |
| Move optional work asynchronously | Shorten user-facing processing | Eventual completion and status tracking |
| Use bounded payloads | Reduce transmission and processing work | Pagination or additional requests may be needed |
| Set deadlines and timeouts | Prevent indefinite waiting | Timeouts must reflect realistic dependency behaviour |
Strategies to Increase Throughput
| Strategy | Primary Benefit | Trade-off or Risk |
|---|---|---|
| Increase safe concurrency | Process more independent work simultaneously | Contention and dependency pressure |
| Scale stateless workers | Add processing capacity | Shared stores can become bottlenecks |
| Batch compatible work | Reduce per-operation overhead | Can increase waiting latency |
| Partition data or workload | Distribute work across independent capacity | Routing, skew and cross-partition operations |
| Use queues | Buffer bursts and protect workers | Backlog, delayed completion and duplicate handling |
| Optimize the limiting stage | Increase end-to-end sustainable capacity | The bottleneck can move elsewhere |
| Reduce repeated work | Free capacity for new operations | Requires caching, deduplication or reuse logic |
| Apply admission control | Protect accepted work during overload | Some incoming requests are delayed or rejected |
Example: URL Shortener
A URL shortener contains at least two performance-sensitive flows with different characteristics.
Create-link Flow
Client
|
v
Gateway
|
v
Link Service
|
+--> Validate destination
+--> Generate short code
+--> Write mapping
|
v
Link Store
|
v
Response
Important measurements include:
- End-to-end creation latency
- Creation requests completed per second
- Storage-write latency
- Duplicate or conflict rate
- Error rate under peak load
Redirect Flow
Visitor
|
v
Edge / Gateway
|
v
Redirect Service
|
+--> Cache hit ----------> Redirect response
|
+--> Cache miss
|
v
Link Store
|
v
Populate cache
|
v
Redirect response
Important measurements include:
- End-to-end redirect latency
- Redirects completed per second
- Cache-hit and cache-miss latency
- Cache-hit ratio
- Link-store query rate
- Tail latency under peak load
- Timeout and error rate
Performance Interpretation
A high cache-hit ratio may reduce average database workload and improve redirect latency. However, the design must still define cache-miss behaviour, staleness, invalidation and cache-failure handling.
Analytics processing can be removed from the redirect critical path when synchronous analytics completion is not required. This can improve user-facing latency, but introduces asynchronous-delivery, backlog and duplicate-processing considerations.
Writing Performance Requirements
The URL shortener must be extremely fast and handle many users.
Under:
<defined workload, data set and environment>
The operation:
<specific request or background job>
Shall achieve:
<latency percentile, throughput and error-rate targets>
Measured from:
<defined entry and completion boundaries>
For:
<defined test duration>
Numeric targets should come from stakeholder requirements, existing measurements, approved forecasts or documented assumptions.
Performance Testing
| Test Type | Purpose |
|---|---|
| Baseline test | Measure behaviour under a small controlled workload |
| Load test | Verify expected normal and peak workload |
| Stress test | Identify behaviour beyond expected capacity |
| Spike test | Evaluate sudden workload increases |
| Soak test | Evaluate behaviour under sustained workload |
| Scalability test | Measure behaviour as resources and workload change |
| Dependency-failure test | Measure timeout, retry and degraded-mode behaviour |
Test Configuration Template
Test name:
<Descriptive name>
Operation:
<Endpoint, query, message or job>
Environment:
<Infrastructure and configuration>
Data set:
<Size and distribution>
Workload:
<Arrival rate, concurrency and request mix>
Duration:
<Warm-up and measurement periods>
Measurements:
- Successful throughput
- p50 latency
- p95 latency
- p99 latency
- Error rate
- Timeout rate
- Queue depth
- Resource utilization
Pass criteria:
<Approved measurable targets>
Monitor Latency and Throughput Together
Throughput should be interpreted together with latency, failures, queue depth and resource saturation.
| Observed Pattern | Possible Interpretation |
|---|---|
| Throughput increases and latency remains stable | The system may still have spare capacity |
| Throughput increases and latency rises gradually | Queueing or contention may be increasing |
| Throughput stops increasing while latency rises | A limiting stage may be saturated |
| Throughput declines and errors increase | The system may be overloaded or experiencing cascading failure |
| Average latency is stable but p99 rises | A subset of requests may be affected by a slow path |
| Queue depth grows continuously | Arrival rate may exceed sustainable completion rate |
Performance Review Checklist
Review Before Capacity Decisions
- The measured operation is clearly defined.
- The latency boundary is documented.
- Successful completion is defined for throughput.
- Average and peak workloads are distinguished.
- Request mix and data distribution are documented.
- Latency percentiles are reported.
- Error and timeout rates are reported.
- Cache-hit and cache-miss paths are measured separately.
- Queueing time is distinguished from processing time.
- Dependency latency is measured.
- The critical path is identified.
- The limiting stage is identified through evidence.
- Concurrency and connection limits are documented.
- Queue depth and backlog age are monitored.
- Performance targets are connected to requirements.
- Testing uses a representative environment and data set.
- Optimization results are compared with a baseline.
- Cost and correctness effects are documented.
Common Mistakes
Confusing Latency with Throughput
Latency measures elapsed time for an operation. Throughput measures completed work per unit of time.
Reporting Only Average Latency
Report percentiles so slower requests are not hidden by the average.
Counting Incoming Requests as Throughput
Arrival rate and successfully completed throughput are different measurements.
Testing Without a Defined Workload
Performance results require request rate, concurrency, request mix, data distribution and test duration.
Increasing Concurrency Without Limits
Excessive concurrency can overload connection pools, databases, networks and external dependencies.
Scaling the Wrong Component
Increasing capacity in a non-limiting stage does not necessarily improve end-to-end throughput.
Ignoring Queueing Time
Processing can be fast while requests spend most of their time waiting for capacity.
Assuming Batching Improves Everything
Batching can improve throughput while increasing the latency of individual items.
Ignoring Errors During Performance Tests
High reported throughput is not useful when a significant portion of requests fail or time out.
Optimizing Without a Baseline
Record the original measurement, make one controlled change and compare the same workload afterward.
Practice Exercise
Analyze the latency and throughput of a notification service that accepts requests through an API and delivers them asynchronously.
Flow
Client
|
v
Notification API
|
v
Notification Store
|
v
Delivery Queue
|
v
Channel Worker
|
v
External Provider
|
v
Delivery Status Store
Tasks
- Define API acceptance latency.
- Define end-to-end delivery latency.
- Define accepted-request throughput.
- Define successful-delivery throughput.
- Identify the likely limiting stages.
- Determine which queue measurements are required.
- Explain how batching could affect throughput and latency.
- Explain how provider rate limits affect sustainable throughput.
- Create a load-test configuration.
- Define the measurements required to identify saturation.
Model Analysis
| Measurement | Boundary |
|---|---|
| Acceptance latency | Client submission to accepted API response |
| Queue waiting time | Message publication to worker processing start |
| Provider latency | Provider request to provider response |
| End-to-end delivery latency | Accepted request to final delivery status |
| Acceptance throughput | Requests durably accepted per second |
| Delivery throughput | Notifications successfully completed per second |
Frequently Asked Questions
What is latency?
Latency is the elapsed time required for an operation to move from a defined start point to a defined completion point.
What is throughput?
Throughput is the amount of completed work produced during a defined period.
Can a system have high throughput and high latency?
Yes. A system can process many operations concurrently while each operation still spends substantial time waiting or processing.
Does lower latency always increase throughput?
Not always. Reducing work in a limiting stage can improve both, but some optimizations primarily affect one metric.
What is tail latency?
Tail latency describes the slower portion of the latency distribution, commonly examined through higher percentiles such as p95 or p99.
Why does latency increase near capacity?
Requests increasingly wait for workers, connections, storage, CPU or other shared resources as utilization and contention rise.
Is throughput the same as bandwidth?
No. Bandwidth describes potential transfer capacity in a communication context. Throughput describes the actual completed work or transferred volume achieved over time.
How does caching affect performance?
Eligible cache hits can reduce access latency and load on the authoritative store. The design must still handle misses, staleness, invalidation and cache failure.
What should be optimized first?
Measure the complete flow, identify the limiting stage or largest critical path contribution and optimize according to the approved requirements.
What comes after latency and throughput?
The next step is to study availability and reliability, then estimate QPS, storage and bandwidth for the expected workload.
Key Takeaway
Latency measures how long an operation takes, while throughput measures how much work the system completes over time. Measure latency as a distribution, define exactly what throughput counts and observe both metrics together with errors, queue depth and resource saturation. Concurrency, caching, batching and scaling can improve performance, but each introduces trade-offs. Optimize the measured bottleneck according to explicit workload and quality requirements rather than relying on isolated averages or assumptions.