QPS, storage and bandwidth estimation
QPS, Storage and Bandwidth Estimation
Learn how to convert users, actions, payload sizes, retention periods and traffic patterns into approximate request rates, storage requirements and network bandwidth for early system design decisions.
Introduction
Capacity estimation converts product expectations into approximate numbers that can guide architecture decisions.
Before selecting databases, caches, queues, partitions or server counts, a system designer should estimate:
- How many requests arrive during an average second
- How many requests arrive during peak periods
- How much traffic consists of reads and writes
- How much new data is created
- How long the data must be retained
- How much storage is required after replication and overhead
- How much data enters and leaves the system
- How much spare capacity is required for growth and failure
Capacity estimation is often called back-of-the-envelope estimation because the initial goal is not perfect precision. The goal is to create an understandable, internally consistent model that exposes the likely scale and bottlenecks.
Core idea: Every estimate must remain connected to its assumptions. A calculated number without its workload, unit, retention and peak assumptions can be misleading.
In the System Design curriculum, QPS, storage and bandwidth estimation is Topic 1.8 and completes the System Design Process and Estimation module. The module uses a URL shortener as the practical exercise, progressing from requirements through request flows, bottlenecks and capacity estimates.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Requirements discovery | User counts, retention and usage patterns begin as product requirements or assumptions. |
| 2 | Functional requirements | Each operation can have a different request rate and payload size. |
| 3 | Non-functional requirements | Peak load, growth, recovery and latency targets affect provisioned capacity. |
| 4 | Request flow | One user request can generate several internal requests or messages. |
| 5 | Latency and throughput | QPS describes a rate and must be interpreted alongside latency and saturation. |
| 6 | Availability and reliability | Capacity must account for failures, maintenance and remaining healthy resources. |
| 7 | Basic arithmetic and unit conversion | Estimation requires rate, size, time and replication calculations. |
Capacity Estimation Workflow
- Define the operations being estimated.
- List the workload assumptions and their sources.
- Calculate daily request volume.
- Convert daily volume to average QPS.
- Apply a justified peak factor.
- Separate read and write traffic.
- Calculate logical data created per period.
- Apply retention, replication and storage overhead.
- Calculate ingress and egress bandwidth separately.
- Add growth and failure headroom without double-counting.
- Identify the assumptions that most affect the result.
- Validate estimates through measurements and load testing.
Start with an Assumption Sheet
Estimates should begin with explicit inputs rather than hidden mental assumptions.
| Input | Example Value | Source or Status |
|---|---|---|
| Daily active users | 10,000,000 | Example assumption |
| Links created per active user per day | 0.1 | Example assumption |
| Redirects per active user per day | 20 | Example assumption |
| Peak-to-average traffic factor | 5 | Example assumption |
| Average mapping-record size | 500 bytes | Example assumption |
| Average redirect response size | 600 bytes | Example assumption |
| Retention period | 5 years | Example assumption |
| Replication factor | 3 copies | Example assumption |
| Annual growth | Not yet confirmed | Open question |
Important: The sample values in this article are synthetic teaching assumptions. A real design must replace them with approved forecasts, current metrics, policy requirements or clearly documented assumptions.
What Is QPS?
QPS stands for queries per second. In system design, the term is frequently used as a general estimate of requests or operations processed each second.
\[ Average\ QPS = \frac{Requests\ During\ Period} {Seconds\ During\ Period} \]
One day contains:
\[ 24 \times 60 \times 60 = 86{,}400\ seconds \]
Therefore:
\[ Average\ Daily\ QPS = \frac{Requests\ Per\ Day}{86{,}400} \]
Estimate Requests per Day
When usage is expressed through daily active users:
\[ Requests\ Per\ Day = Daily\ Active\ Users \times Requests\ Per\ User\ Per\ Day \]
Example
Assume:
- 10 million daily active users
- 20 redirect requests per user per day
\[ Redirects\ Per\ Day = 10{,}000{,}000 \times 20 \]
\[ Redirects\ Per\ Day = 200{,}000{,}000 \]
Calculate Average QPS
\[ Average\ Redirect\ QPS = \frac{200{,}000{,}000}{86{,}400} \]
\[ Average\ Redirect\ QPS \approx 2{,}315 \]
This means the synthetic workload averages approximately 2,315 redirects per second when distributed uniformly across the entire day.
Actual traffic is rarely uniform. The average is a useful baseline, not a sufficient capacity target.
Calculate Peak QPS
An initial peak estimate can apply a peak-to-average factor:
\[ Peak\ QPS = Average\ QPS \times Peak\ Factor \]
Using an illustrative peak factor of five:
\[ Peak\ Redirect\ QPS = 2{,}315 \times 5 \]
\[ Peak\ Redirect\ QPS \approx 11{,}575 \]
The peak factor should be based on historical traffic, campaign behaviour, geographic usage patterns, business events or an approved planning assumption.
The system receives 2,315 QPS.
Assuming 10 million daily active users, 20 redirects per active user per day,
86,400 seconds per day and a peak factor of 5:
Average redirect QPS ≈ 2,315
Estimated peak redirect QPS ≈ 11,575
Separate Read and Write QPS
Reads and writes often create different load patterns and architecture requirements.
Estimate them separately:
\[ Read\ QPS = Total\ QPS \times Read\ Fraction \]
\[ Write\ QPS = Total\ QPS \times Write\ Fraction \]
For a URL shortener, link creation is a write operation while redirection is primarily a read operation.
Link Creation Example
Assume each of 10 million daily active users creates an average of 0.1 links per day:
\[ New\ Links\ Per\ Day = 10{,}000{,}000 \times 0.1 = 1{,}000{,}000 \]
\[ Average\ Creation\ QPS = \frac{1{,}000{,}000}{86{,}400} \approx 11.57 \]
With the same illustrative peak factor of five:
\[ Peak\ Creation\ QPS \approx 57.87 \]
The synthetic example is strongly read-heavy:
\[ Read\ To\ Write\ Ratio = \frac{200{,}000{,}000} {1{,}000{,}000} = 200:1 \]
This ratio helps explain why the redirect path may require different scaling and caching decisions from the link-creation path.
External QPS vs Internal QPS
One public request can generate multiple internal operations.
One redirect request
|
+--> One cache lookup
|
+--> On cache miss, one database lookup
|
+--> One analytics event
|
+--> Selected observability operations
Therefore, external request QPS should not be assumed to equal:
- Database query QPS
- Cache-operation QPS
- Message-publication QPS
- Worker-processing QPS
- External-provider QPS
Cache-adjusted Database QPS
A simplified database-read estimate is:
\[ Database\ Read\ QPS = Read\ QPS \times (1 - Cache\ Hit\ Ratio) \]
If peak redirect QPS is 11,575 and the observed cache-hit ratio is 90%:
\[ Database\ Read\ QPS = 11{,}575 \times 0.10 \]
\[ Database\ Read\ QPS \approx 1{,}158 \]
This simplified estimate assumes every cache miss leads to one database read. Real request flows can include retries, duplicate misses, negative caching and additional lookups.
Storage Estimation
Storage estimation begins with the amount of logical data created and then applies retention, replication, indexing, metadata and operational overhead.
\[ Logical\ Storage = New\ Records \times Average\ Record\ Size \times Retention \]
A broader physical-storage estimate is:
\[ Physical\ Storage = Logical\ Storage \times Replication\ Factor + Indexes + Metadata + Operational\ Overhead \]
Estimate the Record Size
A URL-mapping record may contain:
| Field | Illustrative Size |
|---|---|
| Short code | 16 bytes |
| Destination URL | 300 bytes |
| Owner identifier | 16 bytes |
| Creation and expiration values | 16 bytes |
| Status and configuration | 8 bytes |
| Record-format and row overhead | 144 bytes |
| Total illustrative record size | 500 bytes |
These are synthetic assumptions for calculation practice. Production record size should be measured from the selected encoding and storage engine.
Daily Storage Growth
With one million new links per day and an average logical record size of 500 bytes:
\[ Daily\ Logical\ Storage = 1{,}000{,}000 \times 500\ bytes \]
\[ Daily\ Logical\ Storage = 500{,}000{,}000\ bytes \]
Using decimal units:
\[ Daily\ Logical\ Storage \approx 500\ MB \]
Annual Storage
\[ Annual\ Logical\ Storage = 500\ MB \times 365 \]
\[ Annual\ Logical\ Storage \approx 182.5\ GB \]
Retained Storage
For five years of retention:
\[ Five\ Year\ Logical\ Storage = 182.5\ GB \times 5 \]
\[ Five\ Year\ Logical\ Storage \approx 912.5\ GB \]
Apply Replication
If the approved storage design keeps three complete copies:
\[ Replicated\ Record\ Storage = 912.5\ GB \times 3 \]
\[ Replicated\ Record\ Storage \approx 2.7375\ TB \]
This number covers only the synthetic mapping-record estimate. It does not yet include indexes, backups, logs, temporary files or reserved free capacity.
Storage Components Often Missed
| Storage Component | Why It Matters |
|---|---|
| Primary records | Stores the application data itself |
| Indexes | Accelerate lookups but consume additional storage |
| Metadata | Stores object, row, partition or file-management information |
| Replication | Keeps additional copies for the selected availability or durability design |
| Backups | Support restoration after data loss or corruption |
| Logs and journals | Support recovery, replication or auditing |
| Temporary storage | Supports compaction, sorting, imports, exports or migrations |
| Deleted but retained data | May remain until cleanup or retention expires |
| Reserved free capacity | Supports growth, rebalancing and operational safety |
Logical vs Physical Storage
| Measurement | Meaning |
|---|---|
| Logical storage | Size of application data before storage-system copies and overhead |
| Replicated storage | Logical data multiplied by the number of stored copies |
| Physical storage | Actual consumed capacity after encoding, indexes, metadata and system overhead |
| Provisioned storage | Capacity allocated, including required free space and headroom |
Compression and Deduplication
Compression and deduplication can reduce physical storage, but their effect must be measured for representative data.
If a measured compression ratio is defined as:
\[ Compression\ Ratio = \frac{Uncompressed\ Size}{Compressed\ Size} \]
Then:
\[ Compressed\ Size = \frac{Uncompressed\ Size}{Compression\ Ratio} \]
Do not apply an assumed compression benefit to already compressed images, audio or video without evidence.
Bandwidth Estimation
Bandwidth estimation calculates the rate at which data must be transmitted.
\[ Bandwidth = Operations\ Per\ Second \times Bytes\ Per\ Operation \]
Calculate the two directions independently:
- Ingress is data entering the measured boundary.
- Egress is data leaving the measured boundary.
Ingress Bandwidth
\[ Ingress\ Bandwidth = Write\ QPS \times Average\ Request\ Size \]
Suppose peak link-creation QPS is approximately 58 and the average creation request is 500 bytes:
\[ Peak\ Ingress = 58 \times 500 \]
\[ Peak\ Ingress = 29{,}000\ bytes/second \]
\[ Peak\ Ingress \approx 29\ KB/s \]
Protocol headers, encryption overhead, retries and observability traffic are not included in this simplified application-payload estimate.
Egress Bandwidth
\[ Egress\ Bandwidth = Read\ QPS \times Average\ Response\ Size \]
Suppose peak redirect QPS is approximately 11,575 and the average redirect response is 600 bytes:
\[ Peak\ Egress = 11{,}575 \times 600 \]
\[ Peak\ Egress = 6{,}945{,}000\ bytes/second \]
\[ Peak\ Egress \approx 6.945\ MB/s \]
Convert Bytes per Second to Bits per Second
Storage and payload sizes are commonly expressed in bytes, while network capacity is frequently expressed in bits per second.
\[ Bits\ Per\ Second = Bytes\ Per\ Second \times 8 \]
\[ 6.945\ MB/s \times 8 = 55.56\ Mb/s \]
Therefore, the illustrative peak application-response traffic is approximately 55.56 megabits per second before protocol, retry and safety headroom.
Unit rule: Always label bytes and bits explicitly. Confusing MB/s with Mb/s creates an eightfold error.
Network Traffic Is More Than Payload Size
A fuller estimate can include:
- Request and response headers
- Transport and encryption overhead
- Retries and retransmissions
- Replication traffic
- Service-to-service calls
- Queue and stream traffic
- Database protocol traffic
- Logs, metrics and traces
- Backup and migration traffic
- Cross-region transfer
Keep user-facing ingress and egress separate from internal network traffic so the estimate remains understandable.
Large-object Example
Bandwidth becomes especially important for images, audio, video and file downloads.
Suppose a system serves 2,000 objects per second and the average object is 2 MB:
\[ Egress = 2{,}000 \times 2\ MB \]
\[ Egress = 4{,}000\ MB/s = 4\ GB/s \]
\[ Network\ Rate = 4\ GB/s \times 8 = 32\ Gb/s \]
This type of result can motivate investigation of object storage, content delivery, regional distribution and caching, subject to the system's requirements and constraints.
Estimate Cache Capacity
A cache-capacity estimate requires the number of hot entries and the average memory consumed per entry.
\[ Cache\ Memory = Hot\ Entries \times Memory\ Per\ Entry \]
Suppose:
- Five million mappings are considered hot
- Each cached mapping consumes an estimated 700 bytes including cache overhead
\[ Cache\ Memory = 5{,}000{,}000 \times 700 \]
\[ Cache\ Memory = 3.5\ GB \]
Actual cache memory should be measured because key representation, object headers, allocator behaviour, expiration metadata and fragmentation can significantly affect per-entry memory.
Estimate Instance Count Only After Measurement
QPS alone does not determine how many application instances are required. The supported QPS per instance depends on:
- Request complexity
- CPU usage
- Memory usage
- Payload size
- Database and dependency latency
- Connection limits
- Runtime configuration
- Latency and error-rate objectives
After load testing establishes sustainable per-instance throughput:
\[ Required\ Instances = \left\lceil \frac{Peak\ QPS} {Sustainable\ QPS\ Per\ Instance} \right\rceil \]
If measured sustainable throughput is 1,000 QPS per instance:
\[ Base\ Instance\ Count = \left\lceil \frac{11{,}575}{1{,}000} \right\rceil = 12 \]
This is only the base count before accounting for target utilization, maintenance, deployment and failure headroom.
Headroom
Headroom is spare capacity reserved for variability, growth, maintenance and failure.
One target-utilization approach is:
\[ Provisioned\ Capacity = \frac{Required\ Peak\ Capacity} {Target\ Utilization} \]
If required peak capacity is 11,575 QPS and target utilization is 70%:
\[ Provisioned\ Capacity = \frac{11{,}575}{0.70} \]
\[ Provisioned\ Capacity \approx 16{,}536\ QPS \]
If one tested instance sustains 1,000 QPS:
\[ Provisioned\ Instances = \left\lceil \frac{16{,}536}{1{,}000} \right\rceil = 17 \]
Avoid Double-counting
Peak factor, growth allowance, failure reserve and utilization target are different concepts. Define each one and avoid applying two factors for the same uncertainty.
| Factor | Purpose |
|---|---|
| Peak factor | Represents nonuniform traffic within the planning period |
| Growth factor | Represents expected future workload increase |
| Failure reserve | Preserves capacity when components are unavailable |
| Target utilization | Limits normal usage to preserve responsiveness and flexibility |
| Safety margin | Covers explicitly identified model uncertainty |
Capacity During Failure
Availability requirements can require the remaining healthy components to serve the workload after a failure.
For \(N\) identical instances with one-instance failure tolerance:
\[ Remaining\ Capacity = (N - 1) \times Sustainable\ Capacity\ Per\ Instance \]
Verify:
\[ Remaining\ Capacity \ge Required\ Failure\ Workload \]
Merely running multiple instances does not prove that adequate capacity remains after failure.
Estimate Queue Backlog
If asynchronous work arrives faster than workers complete it:
\[ Backlog\ Growth\ Per\ Second = Arrival\ Rate - Completion\ Rate \]
Suppose notification work arrives at 5,000 messages per second while workers complete 4,200:
\[ Backlog\ Growth = 5{,}000 - 4{,}200 = 800\ messages/second \]
After ten minutes:
\[ Backlog = 800 \times 600 = 480{,}000\ messages \]
Drain Time
If arrivals later return to 3,000 messages per second while workers still complete 4,200:
\[ Spare\ Processing\ Rate = 4{,}200 - 3{,}000 = 1{,}200\ messages/second \]
\[ Drain\ Time = \frac{480{,}000}{1{,}200} = 400\ seconds \]
\[ Drain\ Time \approx 6.67\ minutes \]
The estimate assumes rates remain stable and ignores retries, failures and priority classes.
Model Growth Over Time
If workload grows at an expected annual rate \(g\):
\[ Future\ Load = Current\ Load \times (1 + g)^n \]
Where \(n\) is the number of years.
For example, with a current peak of 10,000 QPS and an illustrative annual growth assumption of 30%:
\[ Year\ 1 = 10{,}000 \times 1.30 = 13{,}000 \]
\[ Year\ 2 = 10{,}000 \times 1.30^2 = 16{,}900 \]
Growth assumptions should be reviewed periodically because actual usage can differ significantly from an early forecast.
Sensitivity Analysis
Sensitivity analysis identifies which assumptions most strongly influence the design.
| Changed Assumption | Primary Effect |
|---|---|
| Users double | Request rate, storage growth and bandwidth can approximately double |
| Peak factor increases | Required short-period processing capacity increases |
| Retention changes from one year to five | Retained logical storage increases approximately fivefold for stable daily writes |
| Average object size doubles | Storage and relevant bandwidth approximately double |
| Cache-hit ratio decreases | Authoritative-store read load increases |
| Replication factor increases | Physical storage and replication traffic increase |
When the design changes dramatically after a small assumption change, that assumption deserves validation or a flexible architecture.
Complete URL Shortener Estimate
Synthetic Inputs
| Input | Illustrative Value |
|---|---|
| Daily active users | 10 million |
| Redirects per user per day | 20 |
| Links created per user per day | 0.1 |
| Peak factor | 5 |
| Average mapping size | 500 bytes |
| Average redirect response | 600 bytes |
| Retention | 5 years |
| Replication factor | 3 |
Calculated Outputs
| Output | Illustrative Estimate |
|---|---|
| Redirect requests per day | 200 million |
| Average redirect QPS | Approximately 2,315 |
| Peak redirect QPS | Approximately 11,575 |
| New links per day | 1 million |
| Average creation QPS | Approximately 11.57 |
| Peak creation QPS | Approximately 57.87 |
| Daily logical mapping storage | Approximately 500 MB |
| Five-year logical mapping storage | Approximately 912.5 GB |
| Three-copy mapping storage | Approximately 2.7375 TB before indexes and overhead |
| Peak redirect-response egress | Approximately 6.945 MB/s or 55.56 Mb/s before overhead |
Architecture Implications
- The illustrative workload is heavily read-oriented.
- The redirect path needs substantially more request capacity than the creation path.
- Caching can materially reduce authoritative-store read traffic when locality exists.
- The required mapping storage is manageable in the synthetic example, but indexes and operational copies still matter.
- Redirect bandwidth is larger than creation-request ingress because redirects dominate the workload.
- Analytics events can create a separate high-volume storage and processing requirement.
- Instance count still requires measured sustainable throughput per instance.
- Failure and growth capacity must be added according to approved objectives.
Estimation Worksheet Template
| Section | Input or Formula | Unit |
|---|---|---|
| Users | Daily active users | Users/day |
| Behaviour | Actions per user per day | Actions/user/day |
| Daily volume | Users × actions per user | Requests/day |
| Average QPS | Requests per day ÷ 86,400 | Requests/second |
| Peak QPS | Average QPS × peak factor | Requests/second |
| Daily storage | Writes per day × bytes per write | Bytes/day |
| Retained storage | Daily storage × retention | Bytes |
| Replicated storage | Retained storage × replication factor | Bytes |
| Ingress | Write QPS × average request bytes | Bytes/second |
| Egress | Read QPS × average response bytes | Bytes/second |
| Network rate | Bytes per second × 8 | Bits/second |
| Base instances | Peak QPS ÷ tested QPS per instance | Instances |
| Provisioned capacity | Required capacity ÷ target utilization | QPS or instances |
Common Estimation Mistakes
Hiding Assumptions
Every result should show the inputs and formula that produced it.
Using Average QPS as the Capacity Target
Estimate peak traffic separately using evidence or a documented planning assumption.
Mixing Read and Write Traffic
Reads and writes can have different payloads, processing costs, storage effects and scaling requirements.
Confusing External QPS with Database QPS
One public request can produce several internal operations, while caches can prevent some database operations.
Ignoring Retention
Daily storage growth must be multiplied by the required retention period.
Ignoring Replication and Indexes
Logical application data can be significantly smaller than consumed physical storage.
Confusing Bytes and Bits
Network capacity expressed in bits per second requires multiplying bytes per second by eight.
Ignoring Protocol and Internal Traffic
Payload-only estimates exclude headers, retries, replication, observability and service-to-service communication.
Calculating Server Count from QPS Alone
Measure sustainable per-instance performance using representative load, data and dependencies.
Double-counting Safety Factors
Document whether peak, growth, failure reserve and utilization factors cover different risks.
Using False Precision
Early estimates are approximate. Preserve meaningful units and rounding rather than implying unsupported accuracy.
Failing to Revisit the Model
Replace assumptions with observed measurements as the product, architecture and workload become clearer.
Capacity Review Checklist
Review Before Component Sizing
- Every input has a source or is marked as an assumption.
- The estimated operations are named explicitly.
- Daily volume is calculated before QPS.
- Average and peak QPS are shown separately.
- Read and write QPS are calculated separately.
- External and internal request rates are distinguished.
- Cache-hit and cache-miss paths are considered.
- Average record and payload sizes are documented.
- Daily storage growth is calculated.
- Retention is applied.
- Replication is applied separately.
- Indexes, metadata, logs and backups are considered.
- Ingress and egress are calculated separately.
- Bytes and bits use clear units.
- Peak network traffic is estimated.
- Growth and failure capacity are considered.
- Safety factors are not double-counted.
- Instance capacity is based on a load test.
- Sensitivity analysis identifies the most influential assumptions.
- The estimate is connected to architecture decisions.
Practice Exercise
Estimate capacity for a notification platform using the following synthetic assumptions:
| Input | Value |
|---|---|
| Daily active applications | 500,000 |
| Notifications per application per day | 40 |
| Peak factor | 4 |
| Average accepted notification record | 1 KB |
| Average provider request | 2 KB |
| Average provider response | 500 bytes |
| Status retention | 90 days |
| Replication factor | 3 |
Tasks
- Calculate notifications per day.
- Calculate average acceptance QPS.
- Calculate estimated peak QPS.
- Calculate daily logical status storage.
- Calculate 90-day logical status storage.
- Apply the replication factor.
- Calculate peak provider-request bandwidth.
- Calculate peak provider-response bandwidth.
- Identify storage components missing from the basic calculation.
- Explain how provider rate limits could become the bottleneck.
Model Solution
\[ Notifications\ Per\ Day = 500{,}000 \times 40 = 20{,}000{,}000 \]
\[ Average\ QPS = \frac{20{,}000{,}000}{86{,}400} \approx 231.48 \]
\[ Peak\ QPS = 231.48 \times 4 \approx 925.93 \]
Using decimal units and 1 KB as 1,000 bytes:
\[ Daily\ Logical\ Storage = 20{,}000{,}000 \times 1{,}000 \]
\[ Daily\ Logical\ Storage = 20\ GB \]
\[ 90\ Day\ Logical\ Storage = 20\ GB \times 90 = 1.8\ TB \]
\[ Replicated\ Storage = 1.8\ TB \times 3 = 5.4\ TB \]
Peak provider-request traffic:
\[ 925.93 \times 2{,}000 \approx 1.852\ MB/s \]
\[ 1.852\ MB/s \times 8 \approx 14.81\ Mb/s \]
Peak provider-response traffic:
\[ 925.93 \times 500 \approx 0.463\ MB/s \]
\[ 0.463\ MB/s \times 8 \approx 3.70\ Mb/s \]
The calculation does not yet include indexes, event queues, retry traffic, message metadata, audit records, backups, logs, protocol overhead or free operating capacity.
Frequently Asked Questions
What is QPS?
QPS is the number of queries, requests or operations processed during one second, provided that the measured operation is clearly defined.
How is average QPS calculated?
Divide the total request count during a period by the number of seconds in that period.
Why is peak QPS different from average QPS?
User traffic is usually uneven. Peak QPS represents concentrated periods when request rates exceed the daily average.
How should a peak factor be selected?
Prefer historical measurements, expected events and approved forecasts. When evidence is unavailable, record the selected value as an assumption and test the sensitivity of the design.
Why calculate read and write QPS separately?
Reads and writes often use different code paths, storage operations, payload sizes, caching strategies and consistency guarantees.
Is database QPS equal to API QPS?
Not necessarily. One API request can produce several database operations, while cache hits can prevent database reads.
Why apply a replication factor to storage?
When the storage design keeps several full copies, physical capacity must account for those copies. The exact factor depends on the approved durability and availability design.
What is the difference between ingress and egress?
Ingress is data entering the measured system boundary. Egress is data leaving that boundary.
Can QPS determine the server count?
Not by itself. Representative load testing must establish sustainable per-instance performance under the required latency and error targets.
How accurate should an early estimate be?
An early estimate should be correct enough in scale and internally consistent enough to guide design discussion. Replace assumptions with measurements as more evidence becomes available.
What should happen after capacity estimation?
Use the estimates to identify likely bottlenecks, compare architecture options and plan focused load tests. Continue updating the model with observed traffic, payload and storage measurements.
Key Takeaway
Capacity estimation turns user behaviour into architecture inputs. Calculate daily requests, average and peak QPS, and separate read and write traffic. Estimate logical storage from write volume, record size and retention, then account for replication, indexes, metadata and operational overhead. Calculate ingress and egress independently, keep bytes and bits explicit, and add justified growth and failure headroom. Preserve every assumption beside the result and validate machine-level capacity through representative load testing.