Table of Contents

    latency vs throughput

    SYSTEM DESIGN FOUNDATIONS

    Latency vs Throughput in System Design

    Learn how latency measures the time required to complete an operation, while throughput measures how much work a system completes over time, and how load, concurrency, queues and bottlenecks connect the two.

    Introduction

    Latency and throughput are two fundamental performance measurements in system design. They are related, but they answer different questions.

    • Latency asks how long one operation takes.
    • Throughput asks how many operations are completed during a period.

    A system can have high throughput while individual requests remain slow. A system can also process one request very quickly while supporting only a small number of requests per second.

    Core distinction: Latency measures delay for an individual operation. Throughput measures completed work over time. Neither metric alone provides a complete description of system performance.

    In the System Design curriculum, latency versus throughput is Topic 1.6 under System Design Process and Estimation. It follows high-level diagrams and precedes availability, reliability, QPS, storage and bandwidth estimation.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Functional requirements Performance is measured for specific system operations.
    2 Non-functional requirements Latency and throughput targets are quality requirements.
    3 Request flow Every synchronous stage can contribute to end-to-end latency.
    4 High-level diagrams Components and dependencies help identify processing stages and bottlenecks.
    5 Basic arithmetic Rate, percentile and concurrency calculations use simple formulas.

    What Is Latency?

    Latency is the elapsed time between the beginning of an operation and the observation of its result at a defined measurement boundary.

    Examples include:

    • Time from submitting a search to receiving results
    • Time from sending an API request to receiving the response
    • Time required to execute a database query
    • Time from publishing a message until a consumer receives it
    • Time from submitting a notification until delivery

    Latency Units

    Unit Symbol Relationship
    Second s Base time unit
    Millisecond ms 1 second = 1,000 milliseconds
    Microsecond µs 1 millisecond = 1,000 microseconds
    Nanosecond ns 1 microsecond = 1,000 nanoseconds

    Components of End-to-End Latency

    End-to-end latency is normally composed of several delays across the request path.

    \[ L_{total} = L_{client} + L_{network} + L_{queue} + L_{application} + L_{dependencies} + L_{response} \]

    Component Description
    Client latency Time spent preparing the request or rendering the result
    Network latency Time spent transferring data between network locations
    Queueing latency Time waiting for processing capacity or a shared resource
    Application latency Time spent executing application logic
    Dependency latency Time spent waiting for databases, caches or external services
    Serialization latency Time spent encoding or decoding data
    Response latency Time required to return and display the result

    Sequential Dependencies

    When dependencies are called sequentially, their delays contribute one after another to the critical path.

    \[ L_{sequential} = L_1 + L_2 + L_3 + \cdots + L_n \]

    Client
      |
      | 20 ms
      v
    Application
      |
      | 30 ms
      v
    Database
      |
      | 40 ms
      v
    External Service
    
    Simplified total = 20 + 30 + 40 = 90 ms

    Parallel Dependencies

    Independent calls can sometimes run concurrently. The dependency portion of latency is then influenced by the slowest required branch rather than by adding every branch duration.

    \[ L_{parallel} \approx \max(L_1, L_2, \ldots, L_n) + L_{coordination} \]

                 +--> Service A: 30 ms --+
    Application -+--> Service B: 50 ms --+--> Combine results
                 +--> Service C: 40 ms --+
    
    Dependency path is influenced by the 50 ms branch,
    plus coordination and response assembly.

    Parallel calls can reduce elapsed time, but they can increase downstream load and introduce partial-failure handling.

    What Is Throughput?

    Throughput is the amount of successfully completed work produced by a system during a defined period.

    \[ Throughput = \frac{Completed\ Operations}{Time} \]

    Throughput can be expressed as:

    • Requests per second
    • Transactions per second
    • Messages processed per second
    • Jobs completed per minute
    • Records processed per hour
    • Bytes transferred per second

    Measurement rule: State what counts as completed work. Incoming requests, accepted requests and successfully completed requests can represent different measurements.

    Latency vs Throughput

    Comparison Area Latency Throughput
    Meaning Time required for one operation Amount of work completed over time
    Primary question How long does it take? How much can be completed?
    Typical unit Milliseconds or seconds Requests, transactions or bytes per second
    User impact Responsiveness of an individual interaction Ability to serve the total workload
    Common visualization Distribution or percentile chart Rate over time
    Typical problem Slow requests Insufficient processing capacity
    Common remedy Reduce critical-path work and waiting Remove bottlenecks or add safe parallel capacity

    Highway Analogy

    A highway provides a useful conceptual analogy:

    • Latency is how long one vehicle takes to complete the journey.
    • Throughput is how many vehicles complete the journey per hour.
    • Capacity or bandwidth is the maximum potential volume supported by the road.

    A wide road can have high potential capacity but poor actual throughput during congestion. One vehicle can also complete a journey quickly while the road carries very few vehicles overall.

    Average Latency Is Not Enough

    Request latency is a distribution. Different requests can take different amounts of time due to queues, cache misses, storage delays, retries, garbage collection, network variation or resource contention.

    Latency Percentiles

    Metric Meaning
    p50 Half of measured operations completed at or below this value
    p90 Ninety percent completed at or below this value
    p95 Ninety-five percent completed at or below this value
    p99 Ninety-nine percent completed at or below this value
    Maximum The slowest measured operation in the sample

    Percentiles reveal slow requests that an average can hide. The measurement report should also define the sampling period, workload and observation boundary.

    Average Can Hide Tail Latency

    Nine requests complete in 10 ms.
    One request completes in 910 ms.
    
    Average latency:
    (9 x 10 + 910) / 10 = 100 ms
    
    The average does not show that most requests were fast
    while one request was much slower.

    Concurrency

    Concurrency is the amount of work in progress at the same time. Concurrency can increase throughput by allowing independent operations to make progress simultaneously.

    However, increasing concurrency beyond resource limits can increase queueing, contention, timeouts and failures.

    Low Concurrency Controlled Concurrency Excessive Concurrency
    Resources can remain underused Independent work uses available capacity Queues and contention increase
    Throughput can remain low Throughput improves Latency and failure rate can rise
    Simple operation Requires limits and monitoring Can overload downstream services

    Little’s Law

    For a stable system, Little’s Law relates the average number of operations in the system, average completion rate and average time spent in the system.

    \[ L = \lambda W \]

    Where:

    • \(L\) is the average work in progress
    • \(\lambda\) is the average throughput
    • \(W\) is the average time in the system

    Example

    If a system completes 200 requests per second and average end-to-end latency is 0.25 seconds:

    \[ L = 200 \times 0.25 = 50 \]

    The system has approximately 50 requests in progress on average under those stated conditions.

    Important: Little’s Law is a relationship among average values in a stable system. It does not by itself determine the required worker, thread or connection count.

    Queueing and Saturation

    A request can spend significant time waiting before processing begins. When work arrives faster than the limiting stage can complete it, a queue grows.

    \[ Queue\ Growth\ Rate = Arrival\ Rate - Completion\ Rate \]

    Arrival rate:    1,200 requests/second
    Completion rate: 1,000 requests/second
    
    Queue growth:
    1,200 - 1,000 = 200 requests/second

    If the difference continues, waiting time, timeouts and memory consumption can increase until requests are rejected or the system fails.

    Typical Saturation Pattern

    Low load:
    Low queueing
    Stable latency
    Available spare capacity
    
    Rising load:
    Higher throughput
    Increasing resource utilization
    
    Near saturation:
    Queueing rises
    Tail latency increases
    Little spare capacity
    
    Overload:
    Arrival exceeds sustainable completion
    Timeouts and failures increase
    Throughput can stop improving or decline

    The Bottleneck Determines Throughput

    A multi-stage system can process work only as fast as its limiting stage over a sustained period.

    Gateway capacity:       10,000 requests/second
    Application capacity:    5,000 requests/second
    Database capacity:       2,000 requests/second
    External provider:       1,000 requests/second
    
    The external provider is the limiting stage for a flow
    that requires the provider for every completed request.

    Adding capacity to a non-limiting stage does not necessarily improve end-to-end throughput.

    Can Throughput Increase Without Lower Latency?

    Yes. Batching and concurrency can increase completed work without reducing the time required for an individual request. They can sometimes increase individual latency.

    Batching Example

    Without batching:
    Process one event per storage operation.
    
    With batching:
    Collect several events.
    Write the batch in one storage operation.
    
    Possible effect:
    Higher throughput
    Fewer storage operations
    Additional waiting while the batch fills

    Batching is appropriate when throughput gains justify the additional delay and the workload permits grouped processing.

    Strategies to Reduce Latency

    Strategy Primary Benefit Trade-off or Risk
    Cache eligible reads Avoid repeated slower data access Staleness and invalidation complexity
    Reduce network distance Lower propagation delay Regional replication and consistency complexity
    Remove unnecessary calls Shorten the critical path May require redesign or additional local state
    Run independent calls in parallel Reduce sequential waiting Higher downstream concurrency and partial failures
    Optimize queries and indexes Reduce storage-access time Additional storage and write cost
    Move optional work asynchronously Shorten user-facing processing Eventual completion and status tracking
    Use bounded payloads Reduce transmission and processing work Pagination or additional requests may be needed
    Set deadlines and timeouts Prevent indefinite waiting Timeouts must reflect realistic dependency behaviour

    Strategies to Increase Throughput

    Strategy Primary Benefit Trade-off or Risk
    Increase safe concurrency Process more independent work simultaneously Contention and dependency pressure
    Scale stateless workers Add processing capacity Shared stores can become bottlenecks
    Batch compatible work Reduce per-operation overhead Can increase waiting latency
    Partition data or workload Distribute work across independent capacity Routing, skew and cross-partition operations
    Use queues Buffer bursts and protect workers Backlog, delayed completion and duplicate handling
    Optimize the limiting stage Increase end-to-end sustainable capacity The bottleneck can move elsewhere
    Reduce repeated work Free capacity for new operations Requires caching, deduplication or reuse logic
    Apply admission control Protect accepted work during overload Some incoming requests are delayed or rejected

    Example: URL Shortener

    A URL shortener contains at least two performance-sensitive flows with different characteristics.

    Create-link Flow

    Client
      |
      v
    Gateway
      |
      v
    Link Service
      |
      +--> Validate destination
      +--> Generate short code
      +--> Write mapping
      |
      v
    Link Store
      |
      v
    Response

    Important measurements include:

    • End-to-end creation latency
    • Creation requests completed per second
    • Storage-write latency
    • Duplicate or conflict rate
    • Error rate under peak load

    Redirect Flow

    Visitor
      |
      v
    Edge / Gateway
      |
      v
    Redirect Service
      |
      +--> Cache hit ----------> Redirect response
      |
      +--> Cache miss
               |
               v
           Link Store
               |
               v
           Populate cache
               |
               v
           Redirect response

    Important measurements include:

    • End-to-end redirect latency
    • Redirects completed per second
    • Cache-hit and cache-miss latency
    • Cache-hit ratio
    • Link-store query rate
    • Tail latency under peak load
    • Timeout and error rate

    Performance Interpretation

    A high cache-hit ratio may reduce average database workload and improve redirect latency. However, the design must still define cache-miss behaviour, staleness, invalidation and cache-failure handling.

    Analytics processing can be removed from the redirect critical path when synchronous analytics completion is not required. This can improve user-facing latency, but introduces asynchronous-delivery, backlog and duplicate-processing considerations.

    Writing Performance Requirements

    Unmeasurable requirement
    The URL shortener must be extremely fast and handle many users.
    Measurable requirement structure
    Under:
    <defined workload, data set and environment>
    
    The operation:
    <specific request or background job>
    
    Shall achieve:
    <latency percentile, throughput and error-rate targets>
    
    Measured from:
    <defined entry and completion boundaries>
    
    For:
    <defined test duration>

    Numeric targets should come from stakeholder requirements, existing measurements, approved forecasts or documented assumptions.

    Performance Testing

    Test Type Purpose
    Baseline test Measure behaviour under a small controlled workload
    Load test Verify expected normal and peak workload
    Stress test Identify behaviour beyond expected capacity
    Spike test Evaluate sudden workload increases
    Soak test Evaluate behaviour under sustained workload
    Scalability test Measure behaviour as resources and workload change
    Dependency-failure test Measure timeout, retry and degraded-mode behaviour

    Test Configuration Template

    Test name:
    <Descriptive name>
    
    Operation:
    <Endpoint, query, message or job>
    
    Environment:
    <Infrastructure and configuration>
    
    Data set:
    <Size and distribution>
    
    Workload:
    <Arrival rate, concurrency and request mix>
    
    Duration:
    <Warm-up and measurement periods>
    
    Measurements:
    - Successful throughput
    - p50 latency
    - p95 latency
    - p99 latency
    - Error rate
    - Timeout rate
    - Queue depth
    - Resource utilization
    
    Pass criteria:
    <Approved measurable targets>

    Monitor Latency and Throughput Together

    Throughput should be interpreted together with latency, failures, queue depth and resource saturation.

    Observed Pattern Possible Interpretation
    Throughput increases and latency remains stable The system may still have spare capacity
    Throughput increases and latency rises gradually Queueing or contention may be increasing
    Throughput stops increasing while latency rises A limiting stage may be saturated
    Throughput declines and errors increase The system may be overloaded or experiencing cascading failure
    Average latency is stable but p99 rises A subset of requests may be affected by a slow path
    Queue depth grows continuously Arrival rate may exceed sustainable completion rate

    Performance Review Checklist

    Review Before Capacity Decisions

    • The measured operation is clearly defined.
    • The latency boundary is documented.
    • Successful completion is defined for throughput.
    • Average and peak workloads are distinguished.
    • Request mix and data distribution are documented.
    • Latency percentiles are reported.
    • Error and timeout rates are reported.
    • Cache-hit and cache-miss paths are measured separately.
    • Queueing time is distinguished from processing time.
    • Dependency latency is measured.
    • The critical path is identified.
    • The limiting stage is identified through evidence.
    • Concurrency and connection limits are documented.
    • Queue depth and backlog age are monitored.
    • Performance targets are connected to requirements.
    • Testing uses a representative environment and data set.
    • Optimization results are compared with a baseline.
    • Cost and correctness effects are documented.

    Common Mistakes

    1

    Confusing Latency with Throughput

    Latency measures elapsed time for an operation. Throughput measures completed work per unit of time.

    2

    Reporting Only Average Latency

    Report percentiles so slower requests are not hidden by the average.

    3

    Counting Incoming Requests as Throughput

    Arrival rate and successfully completed throughput are different measurements.

    4

    Testing Without a Defined Workload

    Performance results require request rate, concurrency, request mix, data distribution and test duration.

    5

    Increasing Concurrency Without Limits

    Excessive concurrency can overload connection pools, databases, networks and external dependencies.

    6

    Scaling the Wrong Component

    Increasing capacity in a non-limiting stage does not necessarily improve end-to-end throughput.

    7

    Ignoring Queueing Time

    Processing can be fast while requests spend most of their time waiting for capacity.

    8

    Assuming Batching Improves Everything

    Batching can improve throughput while increasing the latency of individual items.

    9

    Ignoring Errors During Performance Tests

    High reported throughput is not useful when a significant portion of requests fail or time out.

    10

    Optimizing Without a Baseline

    Record the original measurement, make one controlled change and compare the same workload afterward.

    Practice Exercise

    Analyze the latency and throughput of a notification service that accepts requests through an API and delivers them asynchronously.

    Flow

    Client
      |
      v
    Notification API
      |
      v
    Notification Store
      |
      v
    Delivery Queue
      |
      v
    Channel Worker
      |
      v
    External Provider
      |
      v
    Delivery Status Store

    Tasks

    1. Define API acceptance latency.
    2. Define end-to-end delivery latency.
    3. Define accepted-request throughput.
    4. Define successful-delivery throughput.
    5. Identify the likely limiting stages.
    6. Determine which queue measurements are required.
    7. Explain how batching could affect throughput and latency.
    8. Explain how provider rate limits affect sustainable throughput.
    9. Create a load-test configuration.
    10. Define the measurements required to identify saturation.

    Model Analysis

    Measurement Boundary
    Acceptance latency Client submission to accepted API response
    Queue waiting time Message publication to worker processing start
    Provider latency Provider request to provider response
    End-to-end delivery latency Accepted request to final delivery status
    Acceptance throughput Requests durably accepted per second
    Delivery throughput Notifications successfully completed per second

    Frequently Asked Questions

    1

    What is latency?

    Latency is the elapsed time required for an operation to move from a defined start point to a defined completion point.

    2

    What is throughput?

    Throughput is the amount of completed work produced during a defined period.

    3

    Can a system have high throughput and high latency?

    Yes. A system can process many operations concurrently while each operation still spends substantial time waiting or processing.

    4

    Does lower latency always increase throughput?

    Not always. Reducing work in a limiting stage can improve both, but some optimizations primarily affect one metric.

    5

    What is tail latency?

    Tail latency describes the slower portion of the latency distribution, commonly examined through higher percentiles such as p95 or p99.

    6

    Why does latency increase near capacity?

    Requests increasingly wait for workers, connections, storage, CPU or other shared resources as utilization and contention rise.

    7

    Is throughput the same as bandwidth?

    No. Bandwidth describes potential transfer capacity in a communication context. Throughput describes the actual completed work or transferred volume achieved over time.

    8

    How does caching affect performance?

    Eligible cache hits can reduce access latency and load on the authoritative store. The design must still handle misses, staleness, invalidation and cache failure.

    9

    What should be optimized first?

    Measure the complete flow, identify the limiting stage or largest critical path contribution and optimize according to the approved requirements.

    10

    What comes after latency and throughput?

    The next step is to study availability and reliability, then estimate QPS, storage and bandwidth for the expected workload.

    Key Takeaway

    Latency measures how long an operation takes, while throughput measures how much work the system completes over time. Measure latency as a distribution, define exactly what throughput counts and observe both metrics together with errors, queue depth and resource saturation. Concurrency, caching, batching and scaling can improve performance, but each introduces trade-offs. Optimize the measured bottleneck according to explicit workload and quality requirements rather than relying on isolated averages or assumptions.