Table of Contents

    QPS, storage and bandwidth estimation

    SYSTEM DESIGN FOUNDATIONS

    QPS, Storage and Bandwidth Estimation

    Learn how to convert users, actions, payload sizes, retention periods and traffic patterns into approximate request rates, storage requirements and network bandwidth for early system design decisions.

    Introduction

    Capacity estimation converts product expectations into approximate numbers that can guide architecture decisions.

    Before selecting databases, caches, queues, partitions or server counts, a system designer should estimate:

    • How many requests arrive during an average second
    • How many requests arrive during peak periods
    • How much traffic consists of reads and writes
    • How much new data is created
    • How long the data must be retained
    • How much storage is required after replication and overhead
    • How much data enters and leaves the system
    • How much spare capacity is required for growth and failure

    Capacity estimation is often called back-of-the-envelope estimation because the initial goal is not perfect precision. The goal is to create an understandable, internally consistent model that exposes the likely scale and bottlenecks.

    Core idea: Every estimate must remain connected to its assumptions. A calculated number without its workload, unit, retention and peak assumptions can be misleading.

    In the System Design curriculum, QPS, storage and bandwidth estimation is Topic 1.8 and completes the System Design Process and Estimation module. The module uses a URL shortener as the practical exercise, progressing from requirements through request flows, bottlenecks and capacity estimates.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Requirements discovery User counts, retention and usage patterns begin as product requirements or assumptions.
    2 Functional requirements Each operation can have a different request rate and payload size.
    3 Non-functional requirements Peak load, growth, recovery and latency targets affect provisioned capacity.
    4 Request flow One user request can generate several internal requests or messages.
    5 Latency and throughput QPS describes a rate and must be interpreted alongside latency and saturation.
    6 Availability and reliability Capacity must account for failures, maintenance and remaining healthy resources.
    7 Basic arithmetic and unit conversion Estimation requires rate, size, time and replication calculations.

    Capacity Estimation Workflow

    Estimation Flow
    users → actions → requests → QPS → bytes → storage and bandwidth → headroom
    1. Define the operations being estimated.
    2. List the workload assumptions and their sources.
    3. Calculate daily request volume.
    4. Convert daily volume to average QPS.
    5. Apply a justified peak factor.
    6. Separate read and write traffic.
    7. Calculate logical data created per period.
    8. Apply retention, replication and storage overhead.
    9. Calculate ingress and egress bandwidth separately.
    10. Add growth and failure headroom without double-counting.
    11. Identify the assumptions that most affect the result.
    12. Validate estimates through measurements and load testing.

    Start with an Assumption Sheet

    Estimates should begin with explicit inputs rather than hidden mental assumptions.

    Input Example Value Source or Status
    Daily active users 10,000,000 Example assumption
    Links created per active user per day 0.1 Example assumption
    Redirects per active user per day 20 Example assumption
    Peak-to-average traffic factor 5 Example assumption
    Average mapping-record size 500 bytes Example assumption
    Average redirect response size 600 bytes Example assumption
    Retention period 5 years Example assumption
    Replication factor 3 copies Example assumption
    Annual growth Not yet confirmed Open question

    Important: The sample values in this article are synthetic teaching assumptions. A real design must replace them with approved forecasts, current metrics, policy requirements or clearly documented assumptions.

    What Is QPS?

    QPS stands for queries per second. In system design, the term is frequently used as a general estimate of requests or operations processed each second.

    \[ Average\ QPS = \frac{Requests\ During\ Period} {Seconds\ During\ Period} \]

    One day contains:

    \[ 24 \times 60 \times 60 = 86{,}400\ seconds \]

    Therefore:

    \[ Average\ Daily\ QPS = \frac{Requests\ Per\ Day}{86{,}400} \]

    Estimate Requests per Day

    When usage is expressed through daily active users:

    \[ Requests\ Per\ Day = Daily\ Active\ Users \times Requests\ Per\ User\ Per\ Day \]

    Example

    Assume:

    • 10 million daily active users
    • 20 redirect requests per user per day

    \[ Redirects\ Per\ Day = 10{,}000{,}000 \times 20 \]

    \[ Redirects\ Per\ Day = 200{,}000{,}000 \]

    Calculate Average QPS

    \[ Average\ Redirect\ QPS = \frac{200{,}000{,}000}{86{,}400} \]

    \[ Average\ Redirect\ QPS \approx 2{,}315 \]

    This means the synthetic workload averages approximately 2,315 redirects per second when distributed uniformly across the entire day.

    Actual traffic is rarely uniform. The average is a useful baseline, not a sufficient capacity target.

    Calculate Peak QPS

    An initial peak estimate can apply a peak-to-average factor:

    \[ Peak\ QPS = Average\ QPS \times Peak\ Factor \]

    Using an illustrative peak factor of five:

    \[ Peak\ Redirect\ QPS = 2{,}315 \times 5 \]

    \[ Peak\ Redirect\ QPS \approx 11{,}575 \]

    The peak factor should be based on historical traffic, campaign behaviour, geographic usage patterns, business events or an approved planning assumption.

    Incomplete estimate
    The system receives 2,315 QPS.
    Reproducible estimate
    Assuming 10 million daily active users, 20 redirects per active user per day,
    86,400 seconds per day and a peak factor of 5:
    
    Average redirect QPS ≈ 2,315
    Estimated peak redirect QPS ≈ 11,575

    Separate Read and Write QPS

    Reads and writes often create different load patterns and architecture requirements.

    Estimate them separately:

    \[ Read\ QPS = Total\ QPS \times Read\ Fraction \]

    \[ Write\ QPS = Total\ QPS \times Write\ Fraction \]

    For a URL shortener, link creation is a write operation while redirection is primarily a read operation.

    Link Creation Example

    Assume each of 10 million daily active users creates an average of 0.1 links per day:

    \[ New\ Links\ Per\ Day = 10{,}000{,}000 \times 0.1 = 1{,}000{,}000 \]

    \[ Average\ Creation\ QPS = \frac{1{,}000{,}000}{86{,}400} \approx 11.57 \]

    With the same illustrative peak factor of five:

    \[ Peak\ Creation\ QPS \approx 57.87 \]

    The synthetic example is strongly read-heavy:

    \[ Read\ To\ Write\ Ratio = \frac{200{,}000{,}000} {1{,}000{,}000} = 200:1 \]

    This ratio helps explain why the redirect path may require different scaling and caching decisions from the link-creation path.

    External QPS vs Internal QPS

    One public request can generate multiple internal operations.

    One redirect request
            |
            +--> One cache lookup
            |
            +--> On cache miss, one database lookup
            |
            +--> One analytics event
            |
            +--> Selected observability operations

    Therefore, external request QPS should not be assumed to equal:

    • Database query QPS
    • Cache-operation QPS
    • Message-publication QPS
    • Worker-processing QPS
    • External-provider QPS

    Cache-adjusted Database QPS

    A simplified database-read estimate is:

    \[ Database\ Read\ QPS = Read\ QPS \times (1 - Cache\ Hit\ Ratio) \]

    If peak redirect QPS is 11,575 and the observed cache-hit ratio is 90%:

    \[ Database\ Read\ QPS = 11{,}575 \times 0.10 \]

    \[ Database\ Read\ QPS \approx 1{,}158 \]

    This simplified estimate assumes every cache miss leads to one database read. Real request flows can include retries, duplicate misses, negative caching and additional lookups.

    Storage Estimation

    Storage estimation begins with the amount of logical data created and then applies retention, replication, indexing, metadata and operational overhead.

    \[ Logical\ Storage = New\ Records \times Average\ Record\ Size \times Retention \]

    A broader physical-storage estimate is:

    \[ Physical\ Storage = Logical\ Storage \times Replication\ Factor + Indexes + Metadata + Operational\ Overhead \]

    Estimate the Record Size

    A URL-mapping record may contain:

    Field Illustrative Size
    Short code 16 bytes
    Destination URL 300 bytes
    Owner identifier 16 bytes
    Creation and expiration values 16 bytes
    Status and configuration 8 bytes
    Record-format and row overhead 144 bytes
    Total illustrative record size 500 bytes

    These are synthetic assumptions for calculation practice. Production record size should be measured from the selected encoding and storage engine.

    Daily Storage Growth

    With one million new links per day and an average logical record size of 500 bytes:

    \[ Daily\ Logical\ Storage = 1{,}000{,}000 \times 500\ bytes \]

    \[ Daily\ Logical\ Storage = 500{,}000{,}000\ bytes \]

    Using decimal units:

    \[ Daily\ Logical\ Storage \approx 500\ MB \]

    Annual Storage

    \[ Annual\ Logical\ Storage = 500\ MB \times 365 \]

    \[ Annual\ Logical\ Storage \approx 182.5\ GB \]

    Retained Storage

    For five years of retention:

    \[ Five\ Year\ Logical\ Storage = 182.5\ GB \times 5 \]

    \[ Five\ Year\ Logical\ Storage \approx 912.5\ GB \]

    Apply Replication

    If the approved storage design keeps three complete copies:

    \[ Replicated\ Record\ Storage = 912.5\ GB \times 3 \]

    \[ Replicated\ Record\ Storage \approx 2.7375\ TB \]

    This number covers only the synthetic mapping-record estimate. It does not yet include indexes, backups, logs, temporary files or reserved free capacity.

    Storage Components Often Missed

    Storage Component Why It Matters
    Primary records Stores the application data itself
    Indexes Accelerate lookups but consume additional storage
    Metadata Stores object, row, partition or file-management information
    Replication Keeps additional copies for the selected availability or durability design
    Backups Support restoration after data loss or corruption
    Logs and journals Support recovery, replication or auditing
    Temporary storage Supports compaction, sorting, imports, exports or migrations
    Deleted but retained data May remain until cleanup or retention expires
    Reserved free capacity Supports growth, rebalancing and operational safety

    Logical vs Physical Storage

    Measurement Meaning
    Logical storage Size of application data before storage-system copies and overhead
    Replicated storage Logical data multiplied by the number of stored copies
    Physical storage Actual consumed capacity after encoding, indexes, metadata and system overhead
    Provisioned storage Capacity allocated, including required free space and headroom

    Compression and Deduplication

    Compression and deduplication can reduce physical storage, but their effect must be measured for representative data.

    If a measured compression ratio is defined as:

    \[ Compression\ Ratio = \frac{Uncompressed\ Size}{Compressed\ Size} \]

    Then:

    \[ Compressed\ Size = \frac{Uncompressed\ Size}{Compression\ Ratio} \]

    Do not apply an assumed compression benefit to already compressed images, audio or video without evidence.

    Bandwidth Estimation

    Bandwidth estimation calculates the rate at which data must be transmitted.

    \[ Bandwidth = Operations\ Per\ Second \times Bytes\ Per\ Operation \]

    Calculate the two directions independently:

    • Ingress is data entering the measured boundary.
    • Egress is data leaving the measured boundary.

    Ingress Bandwidth

    \[ Ingress\ Bandwidth = Write\ QPS \times Average\ Request\ Size \]

    Suppose peak link-creation QPS is approximately 58 and the average creation request is 500 bytes:

    \[ Peak\ Ingress = 58 \times 500 \]

    \[ Peak\ Ingress = 29{,}000\ bytes/second \]

    \[ Peak\ Ingress \approx 29\ KB/s \]

    Protocol headers, encryption overhead, retries and observability traffic are not included in this simplified application-payload estimate.

    Egress Bandwidth

    \[ Egress\ Bandwidth = Read\ QPS \times Average\ Response\ Size \]

    Suppose peak redirect QPS is approximately 11,575 and the average redirect response is 600 bytes:

    \[ Peak\ Egress = 11{,}575 \times 600 \]

    \[ Peak\ Egress = 6{,}945{,}000\ bytes/second \]

    \[ Peak\ Egress \approx 6.945\ MB/s \]

    Convert Bytes per Second to Bits per Second

    Storage and payload sizes are commonly expressed in bytes, while network capacity is frequently expressed in bits per second.

    \[ Bits\ Per\ Second = Bytes\ Per\ Second \times 8 \]

    \[ 6.945\ MB/s \times 8 = 55.56\ Mb/s \]

    Therefore, the illustrative peak application-response traffic is approximately 55.56 megabits per second before protocol, retry and safety headroom.

    Unit rule: Always label bytes and bits explicitly. Confusing MB/s with Mb/s creates an eightfold error.

    Network Traffic Is More Than Payload Size

    A fuller estimate can include:

    • Request and response headers
    • Transport and encryption overhead
    • Retries and retransmissions
    • Replication traffic
    • Service-to-service calls
    • Queue and stream traffic
    • Database protocol traffic
    • Logs, metrics and traces
    • Backup and migration traffic
    • Cross-region transfer

    Keep user-facing ingress and egress separate from internal network traffic so the estimate remains understandable.

    Large-object Example

    Bandwidth becomes especially important for images, audio, video and file downloads.

    Suppose a system serves 2,000 objects per second and the average object is 2 MB:

    \[ Egress = 2{,}000 \times 2\ MB \]

    \[ Egress = 4{,}000\ MB/s = 4\ GB/s \]

    \[ Network\ Rate = 4\ GB/s \times 8 = 32\ Gb/s \]

    This type of result can motivate investigation of object storage, content delivery, regional distribution and caching, subject to the system's requirements and constraints.

    Estimate Cache Capacity

    A cache-capacity estimate requires the number of hot entries and the average memory consumed per entry.

    \[ Cache\ Memory = Hot\ Entries \times Memory\ Per\ Entry \]

    Suppose:

    • Five million mappings are considered hot
    • Each cached mapping consumes an estimated 700 bytes including cache overhead

    \[ Cache\ Memory = 5{,}000{,}000 \times 700 \]

    \[ Cache\ Memory = 3.5\ GB \]

    Actual cache memory should be measured because key representation, object headers, allocator behaviour, expiration metadata and fragmentation can significantly affect per-entry memory.

    Estimate Instance Count Only After Measurement

    QPS alone does not determine how many application instances are required. The supported QPS per instance depends on:

    • Request complexity
    • CPU usage
    • Memory usage
    • Payload size
    • Database and dependency latency
    • Connection limits
    • Runtime configuration
    • Latency and error-rate objectives

    After load testing establishes sustainable per-instance throughput:

    \[ Required\ Instances = \left\lceil \frac{Peak\ QPS} {Sustainable\ QPS\ Per\ Instance} \right\rceil \]

    If measured sustainable throughput is 1,000 QPS per instance:

    \[ Base\ Instance\ Count = \left\lceil \frac{11{,}575}{1{,}000} \right\rceil = 12 \]

    This is only the base count before accounting for target utilization, maintenance, deployment and failure headroom.

    Headroom

    Headroom is spare capacity reserved for variability, growth, maintenance and failure.

    One target-utilization approach is:

    \[ Provisioned\ Capacity = \frac{Required\ Peak\ Capacity} {Target\ Utilization} \]

    If required peak capacity is 11,575 QPS and target utilization is 70%:

    \[ Provisioned\ Capacity = \frac{11{,}575}{0.70} \]

    \[ Provisioned\ Capacity \approx 16{,}536\ QPS \]

    If one tested instance sustains 1,000 QPS:

    \[ Provisioned\ Instances = \left\lceil \frac{16{,}536}{1{,}000} \right\rceil = 17 \]

    Avoid Double-counting

    Peak factor, growth allowance, failure reserve and utilization target are different concepts. Define each one and avoid applying two factors for the same uncertainty.

    Factor Purpose
    Peak factor Represents nonuniform traffic within the planning period
    Growth factor Represents expected future workload increase
    Failure reserve Preserves capacity when components are unavailable
    Target utilization Limits normal usage to preserve responsiveness and flexibility
    Safety margin Covers explicitly identified model uncertainty

    Capacity During Failure

    Availability requirements can require the remaining healthy components to serve the workload after a failure.

    For \(N\) identical instances with one-instance failure tolerance:

    \[ Remaining\ Capacity = (N - 1) \times Sustainable\ Capacity\ Per\ Instance \]

    Verify:

    \[ Remaining\ Capacity \ge Required\ Failure\ Workload \]

    Merely running multiple instances does not prove that adequate capacity remains after failure.

    Estimate Queue Backlog

    If asynchronous work arrives faster than workers complete it:

    \[ Backlog\ Growth\ Per\ Second = Arrival\ Rate - Completion\ Rate \]

    Suppose notification work arrives at 5,000 messages per second while workers complete 4,200:

    \[ Backlog\ Growth = 5{,}000 - 4{,}200 = 800\ messages/second \]

    After ten minutes:

    \[ Backlog = 800 \times 600 = 480{,}000\ messages \]

    Drain Time

    If arrivals later return to 3,000 messages per second while workers still complete 4,200:

    \[ Spare\ Processing\ Rate = 4{,}200 - 3{,}000 = 1{,}200\ messages/second \]

    \[ Drain\ Time = \frac{480{,}000}{1{,}200} = 400\ seconds \]

    \[ Drain\ Time \approx 6.67\ minutes \]

    The estimate assumes rates remain stable and ignores retries, failures and priority classes.

    Model Growth Over Time

    If workload grows at an expected annual rate \(g\):

    \[ Future\ Load = Current\ Load \times (1 + g)^n \]

    Where \(n\) is the number of years.

    For example, with a current peak of 10,000 QPS and an illustrative annual growth assumption of 30%:

    \[ Year\ 1 = 10{,}000 \times 1.30 = 13{,}000 \]

    \[ Year\ 2 = 10{,}000 \times 1.30^2 = 16{,}900 \]

    Growth assumptions should be reviewed periodically because actual usage can differ significantly from an early forecast.

    Sensitivity Analysis

    Sensitivity analysis identifies which assumptions most strongly influence the design.

    Changed Assumption Primary Effect
    Users double Request rate, storage growth and bandwidth can approximately double
    Peak factor increases Required short-period processing capacity increases
    Retention changes from one year to five Retained logical storage increases approximately fivefold for stable daily writes
    Average object size doubles Storage and relevant bandwidth approximately double
    Cache-hit ratio decreases Authoritative-store read load increases
    Replication factor increases Physical storage and replication traffic increase

    When the design changes dramatically after a small assumption change, that assumption deserves validation or a flexible architecture.

    Complete URL Shortener Estimate

    Synthetic Inputs

    Input Illustrative Value
    Daily active users 10 million
    Redirects per user per day 20
    Links created per user per day 0.1
    Peak factor 5
    Average mapping size 500 bytes
    Average redirect response 600 bytes
    Retention 5 years
    Replication factor 3

    Calculated Outputs

    Output Illustrative Estimate
    Redirect requests per day 200 million
    Average redirect QPS Approximately 2,315
    Peak redirect QPS Approximately 11,575
    New links per day 1 million
    Average creation QPS Approximately 11.57
    Peak creation QPS Approximately 57.87
    Daily logical mapping storage Approximately 500 MB
    Five-year logical mapping storage Approximately 912.5 GB
    Three-copy mapping storage Approximately 2.7375 TB before indexes and overhead
    Peak redirect-response egress Approximately 6.945 MB/s or 55.56 Mb/s before overhead

    Architecture Implications

    1. The illustrative workload is heavily read-oriented.
    2. The redirect path needs substantially more request capacity than the creation path.
    3. Caching can materially reduce authoritative-store read traffic when locality exists.
    4. The required mapping storage is manageable in the synthetic example, but indexes and operational copies still matter.
    5. Redirect bandwidth is larger than creation-request ingress because redirects dominate the workload.
    6. Analytics events can create a separate high-volume storage and processing requirement.
    7. Instance count still requires measured sustainable throughput per instance.
    8. Failure and growth capacity must be added according to approved objectives.

    Estimation Worksheet Template

    Section Input or Formula Unit
    Users Daily active users Users/day
    Behaviour Actions per user per day Actions/user/day
    Daily volume Users × actions per user Requests/day
    Average QPS Requests per day ÷ 86,400 Requests/second
    Peak QPS Average QPS × peak factor Requests/second
    Daily storage Writes per day × bytes per write Bytes/day
    Retained storage Daily storage × retention Bytes
    Replicated storage Retained storage × replication factor Bytes
    Ingress Write QPS × average request bytes Bytes/second
    Egress Read QPS × average response bytes Bytes/second
    Network rate Bytes per second × 8 Bits/second
    Base instances Peak QPS ÷ tested QPS per instance Instances
    Provisioned capacity Required capacity ÷ target utilization QPS or instances

    Common Estimation Mistakes

    1

    Hiding Assumptions

    Every result should show the inputs and formula that produced it.

    2

    Using Average QPS as the Capacity Target

    Estimate peak traffic separately using evidence or a documented planning assumption.

    3

    Mixing Read and Write Traffic

    Reads and writes can have different payloads, processing costs, storage effects and scaling requirements.

    4

    Confusing External QPS with Database QPS

    One public request can produce several internal operations, while caches can prevent some database operations.

    5

    Ignoring Retention

    Daily storage growth must be multiplied by the required retention period.

    6

    Ignoring Replication and Indexes

    Logical application data can be significantly smaller than consumed physical storage.

    7

    Confusing Bytes and Bits

    Network capacity expressed in bits per second requires multiplying bytes per second by eight.

    8

    Ignoring Protocol and Internal Traffic

    Payload-only estimates exclude headers, retries, replication, observability and service-to-service communication.

    9

    Calculating Server Count from QPS Alone

    Measure sustainable per-instance performance using representative load, data and dependencies.

    10

    Double-counting Safety Factors

    Document whether peak, growth, failure reserve and utilization factors cover different risks.

    11

    Using False Precision

    Early estimates are approximate. Preserve meaningful units and rounding rather than implying unsupported accuracy.

    12

    Failing to Revisit the Model

    Replace assumptions with observed measurements as the product, architecture and workload become clearer.

    Capacity Review Checklist

    Review Before Component Sizing

    • Every input has a source or is marked as an assumption.
    • The estimated operations are named explicitly.
    • Daily volume is calculated before QPS.
    • Average and peak QPS are shown separately.
    • Read and write QPS are calculated separately.
    • External and internal request rates are distinguished.
    • Cache-hit and cache-miss paths are considered.
    • Average record and payload sizes are documented.
    • Daily storage growth is calculated.
    • Retention is applied.
    • Replication is applied separately.
    • Indexes, metadata, logs and backups are considered.
    • Ingress and egress are calculated separately.
    • Bytes and bits use clear units.
    • Peak network traffic is estimated.
    • Growth and failure capacity are considered.
    • Safety factors are not double-counted.
    • Instance capacity is based on a load test.
    • Sensitivity analysis identifies the most influential assumptions.
    • The estimate is connected to architecture decisions.

    Practice Exercise

    Estimate capacity for a notification platform using the following synthetic assumptions:

    Input Value
    Daily active applications 500,000
    Notifications per application per day 40
    Peak factor 4
    Average accepted notification record 1 KB
    Average provider request 2 KB
    Average provider response 500 bytes
    Status retention 90 days
    Replication factor 3

    Tasks

    1. Calculate notifications per day.
    2. Calculate average acceptance QPS.
    3. Calculate estimated peak QPS.
    4. Calculate daily logical status storage.
    5. Calculate 90-day logical status storage.
    6. Apply the replication factor.
    7. Calculate peak provider-request bandwidth.
    8. Calculate peak provider-response bandwidth.
    9. Identify storage components missing from the basic calculation.
    10. Explain how provider rate limits could become the bottleneck.

    Model Solution

    \[ Notifications\ Per\ Day = 500{,}000 \times 40 = 20{,}000{,}000 \]

    \[ Average\ QPS = \frac{20{,}000{,}000}{86{,}400} \approx 231.48 \]

    \[ Peak\ QPS = 231.48 \times 4 \approx 925.93 \]

    Using decimal units and 1 KB as 1,000 bytes:

    \[ Daily\ Logical\ Storage = 20{,}000{,}000 \times 1{,}000 \]

    \[ Daily\ Logical\ Storage = 20\ GB \]

    \[ 90\ Day\ Logical\ Storage = 20\ GB \times 90 = 1.8\ TB \]

    \[ Replicated\ Storage = 1.8\ TB \times 3 = 5.4\ TB \]

    Peak provider-request traffic:

    \[ 925.93 \times 2{,}000 \approx 1.852\ MB/s \]

    \[ 1.852\ MB/s \times 8 \approx 14.81\ Mb/s \]

    Peak provider-response traffic:

    \[ 925.93 \times 500 \approx 0.463\ MB/s \]

    \[ 0.463\ MB/s \times 8 \approx 3.70\ Mb/s \]

    The calculation does not yet include indexes, event queues, retry traffic, message metadata, audit records, backups, logs, protocol overhead or free operating capacity.

    Frequently Asked Questions

    1

    What is QPS?

    QPS is the number of queries, requests or operations processed during one second, provided that the measured operation is clearly defined.

    2

    How is average QPS calculated?

    Divide the total request count during a period by the number of seconds in that period.

    3

    Why is peak QPS different from average QPS?

    User traffic is usually uneven. Peak QPS represents concentrated periods when request rates exceed the daily average.

    4

    How should a peak factor be selected?

    Prefer historical measurements, expected events and approved forecasts. When evidence is unavailable, record the selected value as an assumption and test the sensitivity of the design.

    5

    Why calculate read and write QPS separately?

    Reads and writes often use different code paths, storage operations, payload sizes, caching strategies and consistency guarantees.

    6

    Is database QPS equal to API QPS?

    Not necessarily. One API request can produce several database operations, while cache hits can prevent database reads.

    7

    Why apply a replication factor to storage?

    When the storage design keeps several full copies, physical capacity must account for those copies. The exact factor depends on the approved durability and availability design.

    8

    What is the difference between ingress and egress?

    Ingress is data entering the measured system boundary. Egress is data leaving that boundary.

    9

    Can QPS determine the server count?

    Not by itself. Representative load testing must establish sustainable per-instance performance under the required latency and error targets.

    10

    How accurate should an early estimate be?

    An early estimate should be correct enough in scale and internally consistent enough to guide design discussion. Replace assumptions with measurements as more evidence becomes available.

    11

    What should happen after capacity estimation?

    Use the estimates to identify likely bottlenecks, compare architecture options and plan focused load tests. Continue updating the model with observed traffic, payload and storage measurements.

    Key Takeaway

    Capacity estimation turns user behaviour into architecture inputs. Calculate daily requests, average and peak QPS, and separate read and write traffic. Estimate logical storage from write volume, record size and retention, then account for replication, indexes, metadata and operational overhead. Calculate ingress and egress independently, keep bytes and bits explicit, and add justified growth and failure headroom. Preserve every assumption beside the result and validate machine-level capacity through representative load testing.