CPU, memory, disk and network costs
CPU, Memory, Disk and Network Costs
Learn how processor time, memory access, storage I/O and network communication contribute different costs to application latency, throughput, capacity and reliability.
Introduction
Every software system eventually depends on four foundational resources:
- CPU
- Memory
- Disk or persistent storage
- Network
Cloud services, containers, virtual machines, databases and distributed platforms provide useful abstractions, but the underlying work still consumes processor cycles, memory capacity, storage operations and network bandwidth.
These resources have different performance characteristics. An operation that accesses CPU cache can be much faster than one that waits for main memory. A memory access can be considerably faster than a storage request, while a remote network call can include transmission, queueing, processing and geographical delay.
Core idea: Performance problems are rarely solved by looking at one utilization percentage. Identify where work is executing, where it is waiting, which resource is saturated and whether the workload is CPU-bound, memory-bound, disk-bound or network-bound.
In your System Design.xlsx, this is Topic 2.1 under Computer Systems, Linux and Concurrency. The module’s learning objective is to explain local resource bottlenecks and recognize concurrency hazards before distributing a system. Its practical lab profiles a small service using Linux tools before analyzing concurrency defects.
Prerequisites
| # | Prerequisite | Why It Is Needed |
|---|---|---|
| 1 | Latency and throughput | Resource costs influence request completion time and sustainable processing rate. |
| 2 | Basic operating-system knowledge | The operating system schedules CPU work and manages memory, files and network sockets. |
| 3 | Processes and threads | Applications execute through schedulable threads within processes. |
| 4 | Request-flow analysis | Each processing stage can consume or wait for a different resource. |
| 5 | Basic Linux commands | System tools expose resource usage, pressure, queues and errors. |
The Resource Hierarchy
A simplified computing hierarchy moves from resources close to the CPU toward slower and more distant resources:
This is a conceptual hierarchy, not a universal benchmark. Real values vary by processor, workload, hardware, operating system, data-access pattern, payload size and network location.
| Resource | Primary Strength | Typical Limitation |
|---|---|---|
| CPU | Rapid instruction execution and computation | Finite cores, instruction throughput and scheduling capacity |
| Memory | Fast access to active data | Finite capacity, memory bandwidth and allocation pressure |
| Disk or storage | Durable, comparatively large data capacity | Higher latency, finite IOPS and queueing |
| Network | Communication across process and machine boundaries | Latency, bandwidth, congestion and remote failure |
Performance rule: Memorized latency numbers become outdated and are not substitutes for measurement. Use the hierarchy to guide reasoning, then benchmark the actual target environment.
Latency Scale
Resource operations occur at different orders of magnitude. The following unit relationships help compare them:
\[ 1\ second = 1{,}000\ milliseconds \]
\[ 1\ millisecond = 1{,}000\ microseconds \]
\[ 1\ microsecond = 1{,}000\ nanoseconds \]
A delay measured in milliseconds represents millions of nanoseconds. Therefore, a thread waiting for storage or a remote response can potentially miss a large number of CPU execution opportunities.
Part 1: CPU Costs
The CPU executes instructions, performs arithmetic, evaluates branches and coordinates data movement.
Common CPU-consuming activities include:
- Parsing and validating requests
- Encryption and decryption
- Compression and decompression
- Serialization and deserialization
- Sorting, searching and aggregation
- Image, audio or video processing
- Regular-expression evaluation
- Garbage collection or memory-management work
- Operating-system and network-protocol processing
CPU Execution Costs
CPU cost is affected by more than the number of source-code statements. Important factors include:
- Instruction count
- Instruction complexity
- Branch predictability
- Cache locality
- Vectorization
- Lock contention
- Context switching
- System-call frequency
- Runtime and compiler optimization
Branch Prediction
Modern processors attempt to predict the path of conditional branches. A misprediction can discard speculative work and restart execution along the correct path.
for (size_t index = 0;
index < count;
++index) {
if (values[index] >
threshold) {
process_value(
values[index]
);
}
}
Performance depends partly on the data distribution and predictability of the condition. Avoid rewriting clear code solely to remove branches unless profiling shows that the branch is a material bottleneck.
CPU Cache Locality
CPU caches keep recently accessed data closer to the processor. Code that accesses nearby memory locations can often benefit from spatial and temporal locality.
| Locality Type | Meaning |
|---|---|
| Temporal locality | Recently accessed data is accessed again soon |
| Spatial locality | Data near a recently accessed location is accessed soon |
long long sum = 0;
for (size_t index = 0;
index < count;
++index) {
sum += values[index];
}
long long sum = 0;
for (size_t index = 0;
index < count;
++index) {
sum += values[
random_indexes[index]
];
}
The second pattern is not automatically incorrect. Some algorithms require unpredictable access. The important point is that memory-access pattern can influence CPU performance.
CPU Utilization vs CPU Pressure
High CPU utilization means the processor is busy. It does not automatically mean the system has a performance defect.
CPU pressure is more likely when:
- Runnable work waits for processor time
- Request latency increases as load rises
- Throughput no longer increases
- Context switching becomes excessive
- CPU throttling occurs
- Important workloads cannot meet deadlines
CPU monitoring should consider how processor time is spent, scheduling queues and whether useful work is delayed, rather than interpreting one utilization percentage in isolation.
Context-switching Cost
A context switch occurs when the processor changes from one executing thread or process to another. The operating system must preserve and restore execution state, and the new workload can have a different cache footprint.
Context switching increases when:
- Too many runnable threads compete for CPU cores
- Threads frequently block and wake
- Locks create contention
- Very small tasks are scheduled independently
- Interrupt activity is high
Concurrency rule: More threads do not guarantee more throughput. Beyond the useful concurrency level, scheduling and contention overhead can reduce performance.
Part 2: Memory Costs
Main memory stores active program code, objects, buffers, stacks, caches and operating-system data.
Memory cost includes:
- Allocation and deallocation
- Reading and writing memory
- Cache misses
- Page faults
- Copying large buffers
- Memory reclamation
- Swapping or paging to storage
- Garbage collection in managed runtimes
Capacity Is Not the Only Memory Metric
High memory usage alone does not always indicate a problem. Operating systems can use available memory for filesystem caching and release that memory when applications need it.
More useful warning signals can include:
- Allocation failures
- Increasing swap activity
- Frequent page reclamation
- Major page faults
- Out-of-memory termination
- Repeated cache eviction
- Growing process resident memory without stabilization
- Latency increasing during memory pressure
Page Faults
Virtual memory allows a process to use a logical address space. A page fault occurs when the referenced virtual-memory page requires operating-system handling.
| Page-fault Type | General Meaning |
|---|---|
| Minor page fault | The page can be resolved without reading its content from persistent storage |
| Major page fault | Resolving the page requires slower storage-related work |
Frequent major page faults can cause large application-latency increases because the process waits for storage-backed data.
Memory Copy Cost
Copying data consumes CPU time and memory bandwidth. An architecture that repeatedly copies large payloads between buffers, processes or protocol layers can spend substantial capacity moving data rather than transforming it.
Network buffer
|
| Copy
v
Application buffer
|
| Copy
v
Serialization buffer
|
| Copy
v
Storage or network output
Possible improvements, when supported by evidence, include:
- Streaming instead of buffering the complete payload
- Reusing bounded buffers
- Avoiding unnecessary transformations
- Using references or views when lifetime rules permit
- Applying zero-copy facilities supported by the platform
Data Structure Overhead
Logical data size can be smaller than actual memory usage because data structures can contain:
- Object headers
- Pointers
- Alignment padding
- Unused collection capacity
- Hash-table buckets
- Allocator metadata
- Fragmentation
For example, storing millions of small records as separately allocated objects can use significantly more memory than the sum of the visible field sizes.
Part 3: Disk and Storage Costs
Persistent storage provides durability beyond process and machine memory. Storage operations generally have higher latency than CPU-cache or main memory access.
Storage performance depends on:
- Storage technology
- Random vs sequential access
- Read vs write workload
- Operation size
- Queue depth
- Filesystem behaviour
- Durability and synchronization settings
- Compaction or background maintenance
- Contention from other workloads
Sequential vs Random I/O
| Access Pattern | Description | Typical Use |
|---|---|---|
| Sequential I/O | Reads or writes nearby storage locations in order | Logs, scans, backups and large-file transfer |
| Random I/O | Accesses scattered locations | Point lookups and database-page access |
Sequential access often provides higher transfer efficiency because it reduces seek, command and coordination overhead. The actual difference depends on the storage technology and workload.
IOPS, Throughput and Latency
| Metric | Meaning |
|---|---|
| IOPS | Input/output operations completed per second |
| Storage throughput | Bytes transferred per second |
| I/O latency | Time required to complete one storage operation |
| Queue depth | Outstanding or waiting storage operations |
| Utilization | How busy the storage resource is during the measurement period |
Small random operations can be limited by IOPS, while large sequential transfers can be limited by bytes per second.
Simplified Throughput Relationship
\[ Storage\ Throughput \approx IOPS \times Average\ Operation\ Size \]
If a device completes 10,000 operations per second and each operation moves 4 KB:
\[ Throughput \approx 10{,}000 \times 4\ KB = 40{,}000\ KB/s \]
\[ Throughput \approx 40\ MB/s \]
This is an illustrative relationship. Real throughput can be affected by queueing, protocol overhead, caching and device behaviour.
Storage Queueing
Storage latency can increase before the device reaches full capacity. When new operations arrive faster than storage can complete them, queue depth and waiting time increase.
Application requests
|
v
Storage queue
|
| Wait for device capacity
v
Storage device
|
v
Completed I/O
Disk monitoring should therefore include latency, queue depth, completed operations and error rates rather than only free capacity.
Durability Has a Cost
A write acknowledged only after durable persistence can require more work than a write acknowledged after being placed in memory.
Faster but weaker acknowledgement:
Request -> Memory buffer -> Acknowledge -> Persist later
Stronger durability:
Request -> Durable log or storage -> Confirm persistence -> Acknowledge
The correct policy depends on the data-loss requirement. A financial update may require a stronger durability guarantee than a reconstructable analytics event.
Storage Capacity vs Storage Performance
A storage device can have sufficient free space but still become a performance bottleneck.
Separate these questions:
- Is there enough capacity to store the data?
- Can the system complete the required IOPS?
- Can the system deliver the required transfer throughput?
- Does latency satisfy the application deadline?
- How does performance change under queueing?
- Can storage recover from device or node failure?
Part 4: Network Costs
Network communication allows processes and machines to exchange data. A network request can include:
- Name resolution
- Connection establishment
- Transport and encryption handshakes
- Serialization
- Transmission
- Routing
- Queueing
- Remote processing
- Response transfer
- Deserialization
Network Latency
Network latency depends on more than bandwidth. It can include:
| Delay | Description |
|---|---|
| Propagation delay | Time for a signal to travel through the physical medium |
| Transmission delay | Time required to place all payload bits onto the link |
| Queueing delay | Time waiting in network-device or system buffers |
| Processing delay | Time for endpoints and network devices to process the traffic |
| Remote-service delay | Time spent by the destination application and its dependencies |
A high-bandwidth link can still have significant round-trip latency, especially over long geographical distances or during congestion.
Bandwidth and Transfer Time
A simplified payload transmission time is:
\[ Transmission\ Time = \frac{Payload\ Size\ In\ Bits} {Link\ Rate\ In\ Bits\ Per\ Second} \]
Example
To transmit a 10 MB payload over an effective 100 Mb/s link, using decimal units:
\[ Payload = 10\ MB \times 8 = 80\ Mb \]
\[ Transmission\ Time = \frac{80\ Mb} {100\ Mb/s} = 0.8\ seconds \]
This simplified calculation excludes connection setup, protocol overhead, congestion, retries and remote processing.
Round Trips
A request that requires several sequential network round trips accumulates latency.
Client
|
| Round trip 1: authenticate
v
Identity Service
|
| Round trip 2: business request
v
Application Service
|
| Round trip 3: data request
v
Remote Database
Reducing unnecessary round trips can improve latency, but combining operations must preserve correctness, authorization and maintainability.
Serialization Cost
Data must often be transformed between in-memory objects and a transport representation.
Serialization cost depends on:
- Payload size
- Encoding format
- Schema complexity
- Compression
- Allocation and copying
- Validation
- Encryption
Smaller payloads can reduce network transfer work, but highly compressed formats can consume additional CPU.
Remote Calls Can Fail Independently
A local function call normally shares the current process. A remote call adds independent failure possibilities:
- Name resolution failure
- Connection refusal
- Connection reset
- Timeout
- Rate limiting
- Partial data transfer
- Remote overload
- Remote application failure
- Uncertain result after a lost response
Distributed-call rule: Treat a network call as a slower, failure-prone operation with explicit deadlines, bounded retries and a defined duplicate-handling policy.
Comparing Resource Costs
| Operation | Main Cost | Design Concern |
|---|---|---|
| Arithmetic on cached data | CPU execution | Instruction count and branch behaviour |
| Accessing a large in-memory structure | Memory latency and bandwidth | Locality, allocation and pressure |
| Random persistent lookup | Storage latency and IOPS | Indexing, queue depth and cache misses |
| Large sequential file transfer | Storage and network throughput | Buffering and transfer size |
| Remote API call | Network plus remote processing | Timeout, retry and dependency availability |
| Compression before transfer | CPU exchanged for fewer bytes | Payload size, CPU cost and latency target |
| Application cache | Memory exchanged for less storage or network access | Capacity, eviction and staleness |
CPU-bound vs I/O-bound Workloads
| Workload Type | Primary Behaviour | Examples |
|---|---|---|
| CPU-bound | Spends most active time performing computation | Compression, encryption, transcoding and complex calculations |
| Memory-bound | Limited by memory latency, bandwidth or capacity | Large graph traversal and in-memory analytics |
| Disk-bound | Limited by persistent-storage latency, IOPS or throughput | Large scans, random database reads and backup operations |
| Network-bound | Limited by transfer rate, round trips or remote responses | File transfer, remote APIs and distributed queries |
| Mixed | Different stages are limited by different resources | Most production request flows |
Identify the Bottleneck
The bottleneck is the limiting resource or stage that prevents the system from increasing useful throughput or meeting latency objectives.
A structured investigation asks:
- Which operation is slow or capacity-limited?
- What workload reproduces the behaviour?
- Where is time spent?
- Which resource is saturated or queued?
- Is the problem local or caused by a dependency?
- Does the issue affect every request or only a subset?
- Which change is expected to relieve the limiting stage?
- Did the same workload improve after the change?
Symptoms by Resource
| Resource | Possible Symptoms | Measurements to Inspect |
|---|---|---|
| CPU | Runnable queues, increasing request latency and throughput flattening | CPU time, load, run queue, per-process usage and context switches |
| Memory | Reclamation, swap activity, allocation failures and unstable latency | Available memory, resident memory, faults, swap and pressure |
| Disk | Increasing I/O latency, queue depth and I/O wait | Read/write latency, IOPS, throughput, queue depth and errors |
| Network | Higher round-trip time, retransmissions, connection delay and bandwidth saturation | Traffic rate, errors, drops, retransmissions and dependency latency |
Linux Investigation Tools
The following tools help collect evidence. Availability and output fields can vary by Linux distribution and installed packages.
| Tool | Primary Use |
|---|---|
top |
Interactive process and system activity |
vmstat |
CPU, runnable work, memory, paging and system activity |
iostat |
CPU and storage-device activity |
free |
Memory and swap summary |
pidstat |
Per-process CPU and selected I/O activity |
lsof |
Open files and sockets associated with processes |
ss |
Socket and network-connection information |
sar |
Historical and interval-based system activity where configured |
perf |
CPU profiling and hardware or software event analysis |
General System Activity
top
vmstat 1
Memory Summary
free -h
Storage Activity
iostat -xz 1
Per-process Activity
pidstat 1
pidstat -d 1
Socket Information
ss -s
ss -tupn
Open Files and Sockets
lsof -p PROCESS_ID
CPU Profiling
perf stat ./application
perf record -g ./application
perf report
Operational caution: Some profiling and tracing tools introduce overhead or require elevated permissions. Follow the approved process before using them in production.
Example: Diagnose a Slow API
Request Flow
Client
|
v
API Service
|
+--> Parse and validate request
|
+--> Query database
|
+--> Call external service
|
+--> Serialize response
|
v
Client
Investigation
| Observation | Possible Interpretation |
|---|---|
| CPU is saturated and runnable work is growing | Application computation or excessive concurrency may limit progress |
| CPU is low but storage latency is high | The request can be waiting for disk or database I/O |
| Memory pressure and swap activity increase | The working set may exceed practical memory capacity |
| External-call latency dominates traces | The remote dependency may be the critical-path bottleneck |
| Network retransmissions increase | Packet loss or network congestion can contribute to delay |
| Database connection queue grows | The connection pool or database can be saturated |
The correct optimization depends on the measured bottleneck. Adding CPU does not solve a saturated database, and increasing database capacity does not solve unnecessary payload transfer.
Example: Database Query Cost
Application
|
| Send query
v
Database connection pool
|
| Wait for connection if pool is exhausted
v
Database
|
+--> Parse and plan
+--> Read index pages
+--> Read table pages
+--> Filter and sort
+--> Return rows
|
v
Application
|
+--> Deserialize
+--> Build response
Potential costs include:
- Network round trip
- Connection-pool waiting
- CPU used for planning and execution
- Memory used for sorting or hashing
- Storage reads after cache misses
- Lock waiting
- Response serialization and transfer
Caching as a Resource Trade-off
Caching typically spends memory to reduce repeated computation, disk access or network calls.
Without cache:
Request -> Application -> Database or remote service
With cache:
Request -> Application -> Memory cache
|
+--> Miss -> Database or remote service
Caching can improve performance when:
- Requests reuse the same data
- The cached result remains valid long enough
- The authoritative operation is comparatively expensive
- The cache-hit rate is sufficient
Caching introduces costs involving:
- Memory capacity
- Invalidation
- Staleness
- Eviction
- Cache-miss amplification
- Warm-up after restart
Compression as a CPU-Network Trade-off
Compression exchanges CPU work for fewer transferred or stored bytes.
Original payload
|
| CPU compression
v
Smaller payload
|
| Network or storage transfer
v
Receiver
|
| CPU decompression
v
Original data
Compression is more attractive when:
- The payload compresses significantly
- Network or storage is the bottleneck
- CPU headroom is available
- The additional latency remains acceptable
Compression can be unattractive for tiny payloads, already compressed data or CPU-saturated services.
Queues as a Time-Capacity Trade-off
A queue can buffer bursts and separate request acceptance from background processing.
Producer
|
| Accept work
v
Queue
|
| Wait for worker capacity
v
Worker
|
v
Storage or external provider
A queue does not remove work. It moves waiting time and provides controlled buffering.
Monitor:
- Arrival rate
- Completion rate
- Queue depth
- Age of the oldest item
- Retry volume
- Failure or dead-letter volume
Scaling by Resource Type
| Bottleneck | Possible Responses | Important Verification |
|---|---|---|
| CPU | Optimize hot code, reduce repeated work, add cores or partition computation | Profile shows the changed code was limiting performance |
| Memory capacity | Reduce working set, bound caches, improve representation or add memory | Pressure, faults and latency improve |
| Memory bandwidth | Improve locality, reduce copies or partition processing | Measured memory stalls or bandwidth pressure decline |
| Storage IOPS | Improve indexes, cache reads, batch writes or partition workload | I/O queue and latency improve under the same workload |
| Storage throughput | Use larger sequential operations, compression or parallel storage paths | Transfer rate improves without violating latency or durability |
| Network latency | Reduce round trips, colocate services or use regional endpoints | End-to-end traces show lower critical-path delay |
| Network bandwidth | Reduce payload, compress, cache closer or increase link capacity | Transfer rate and application latency improve |
Common Performance Mistakes
Assuming High CPU Usage Is Always Bad
High utilization can be healthy when throughput and latency targets are satisfied and sufficient headroom remains.
Looking Only at Used Memory
Inspect pressure, allocation failures, reclaim, faults and swap activity, not only the amount shown as used.
Monitoring Only Free Disk Space
Storage can become latency- or IOPS-limited while substantial free capacity remains.
Confusing Bandwidth with Latency
Bandwidth describes transfer capacity. Latency describes delay. A high-bandwidth connection can still have a long round trip.
Adding Threads Without Measuring Contention
More threads can increase scheduling overhead, lock contention and downstream pressure.
Adding Memory to a CPU-bound Workload
Additional memory does not automatically improve a workload limited by computation.
Optimizing Disk When Data Already Comes from Memory
Confirm the active path and cache behaviour before changing storage.
Ignoring Payload Size
Large payloads consume network, memory, serialization and sometimes storage capacity.
Using Historical Latency Numbers as Current Guarantees
Hardware and environments differ. Benchmark the actual deployment target.
Optimizing Without Reproducing the Workload
Establish a baseline, change one factor and test the same workload again.
Recommended Performance Tests
| Test | Purpose |
|---|---|
| CPU-intensive benchmark | Measure compute capacity and scaling across cores |
| Sequential memory scan | Observe locality and memory-transfer behaviour |
| Random memory access | Observe the cost of poor locality |
| Sequential storage transfer | Measure large-block storage throughput |
| Random storage operations | Measure IOPS and latency under queueing |
| Local network test | Measure communication within the deployment environment |
| Remote dependency test | Measure end-to-end latency and failure behaviour |
| Mixed service workload | Observe interaction among CPU, memory, disk and network |
| Sustained-load test | Detect leaks, thermal effects, queue growth and background maintenance |
Resource Review Checklist
Performance Investigation Checklist
- The slow or capacity-limited operation is clearly identified.
- The test workload is documented.
- Average and peak load are distinguished.
- CPU utilization is interpreted with runnable queues and latency.
- User and kernel CPU time are distinguished where useful.
- Context-switch activity is inspected.
- Memory pressure is inspected instead of only memory usage.
- Page faults and swap activity are reviewed.
- Process resident memory is monitored over time.
- Storage latency, IOPS and throughput are measured separately.
- Random and sequential storage workloads are distinguished.
- Storage queue depth is inspected.
- Network latency and bandwidth are measured separately.
- Retransmissions, drops and errors are reviewed.
- Payload sizes are documented.
- Remote-dependency time is visible in traces.
- The limiting stage is supported by evidence.
- One controlled optimization is applied at a time.
- The same workload is rerun after optimization.
- Correctness, reliability and cost are checked after performance changes.
Practice Exercise
Profile a small Linux service that reads a file, transforms records and sends results to a remote endpoint.
Workload
Input file
|
v
Read records
|
v
Parse and transform
|
v
Serialize batches
|
v
Send to remote endpoint
|
v
Record completion
Tasks
- Measure total execution time.
- Measure CPU utilization during parsing.
- Inspect memory growth and page faults.
- Measure storage read throughput and latency.
- Measure outbound network throughput.
- Measure remote-call latency.
- Determine whether the workload is CPU-, memory-, disk- or network-bound.
- Change the batch size and rerun the same workload.
- Compare throughput, latency and memory usage.
- Document the bottleneck and the evidence supporting the conclusion.
Result Template
Workload:
<Input size, number of records and request destination>
Environment:
<CPU, memory, storage, operating system and network location>
Baseline:
- Total duration:
- Records per second:
- CPU utilization:
- Maximum resident memory:
- Major page faults:
- Storage read throughput:
- Storage latency:
- Network throughput:
- Remote-call latency:
Observed bottleneck:
<CPU, memory, disk, network or mixed>
Evidence:
<Measurements supporting the conclusion>
Change:
<One optimization or configuration change>
Result:
<Comparison using the same workload>
Trade-offs:
<Effects on correctness, memory, cost or reliability>
Frequently Asked Questions
Why is CPU cache faster than main memory?
Processor caches are smaller and located closer to CPU execution units, allowing lower-latency access than larger main memory.
Is high CPU utilization always a bottleneck?
No. CPU becomes a practical bottleneck when processing demand delays work, prevents throughput growth or violates latency and headroom objectives.
Why can high memory usage be normal?
Operating systems use otherwise available memory for useful filesystem caching. Memory pressure, allocation failure and paging behaviour are more informative than one used-memory value.
What is a major page fault?
It is a page fault whose resolution requires slower storage-related work rather than only updating in-memory mappings.
What is the difference between storage IOPS and throughput?
IOPS measures completed operations per second. Throughput measures bytes transferred per second.
Can a disk have free capacity and still be slow?
Yes. Storage can be limited by latency, IOPS, transfer throughput or queue depth long before free space is exhausted.
What is the difference between network latency and bandwidth?
Latency measures communication delay. Bandwidth measures potential data transfer rate.
Why are remote calls more expensive than function calls?
Remote calls add serialization, network transfer, remote scheduling, remote processing and independent failure possibilities.
Does caching always improve performance?
No. Caching helps when reuse is sufficient and freshness rules permit it. Cache misses, invalidation and memory overhead can reduce the benefit.
How should a bottleneck be identified?
Reproduce the workload, measure the complete flow, locate waiting and saturation, make one targeted change and rerun the same test.
Should latency numbers be memorized?
Understanding relative orders of magnitude is useful, but actual designs should use benchmark and production measurements from the target environment.
What comes after CPU, memory, disk and network costs?
The next topic is processes versus threads, followed by synchronization, race conditions, deadlocks, files, sockets and Linux diagnostic tools.
Key Takeaway
CPU, memory, disk and network resources have different costs and failure characteristics. CPU performance depends on computation, locality, branches and scheduling. Memory performance depends on locality, capacity, bandwidth and pressure. Storage performance depends on latency, IOPS, throughput and queueing. Network performance depends on round trips, transfer rate, payload size and remote-service behaviour. Measure the complete workload, identify where work waits, optimize the actual bottleneck and verify the result using the same representative test.