Table of Contents

    CPU, memory, disk and network costs

    COMPUTER SYSTEMS, LINUX AND CONCURRENCY

    CPU, Memory, Disk and Network Costs

    Learn how processor time, memory access, storage I/O and network communication contribute different costs to application latency, throughput, capacity and reliability.

    Introduction

    Every software system eventually depends on four foundational resources:

    • CPU
    • Memory
    • Disk or persistent storage
    • Network

    Cloud services, containers, virtual machines, databases and distributed platforms provide useful abstractions, but the underlying work still consumes processor cycles, memory capacity, storage operations and network bandwidth.

    These resources have different performance characteristics. An operation that accesses CPU cache can be much faster than one that waits for main memory. A memory access can be considerably faster than a storage request, while a remote network call can include transmission, queueing, processing and geographical delay.

    Core idea: Performance problems are rarely solved by looking at one utilization percentage. Identify where work is executing, where it is waiting, which resource is saturated and whether the workload is CPU-bound, memory-bound, disk-bound or network-bound.

    In your System Design.xlsx, this is Topic 2.1 under Computer Systems, Linux and Concurrency. The module’s learning objective is to explain local resource bottlenecks and recognize concurrency hazards before distributing a system. Its practical lab profiles a small service using Linux tools before analyzing concurrency defects.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Latency and throughput Resource costs influence request completion time and sustainable processing rate.
    2 Basic operating-system knowledge The operating system schedules CPU work and manages memory, files and network sockets.
    3 Processes and threads Applications execute through schedulable threads within processes.
    4 Request-flow analysis Each processing stage can consume or wait for a different resource.
    5 Basic Linux commands System tools expose resource usage, pressure, queues and errors.

    The Resource Hierarchy

    A simplified computing hierarchy moves from resources close to the CPU toward slower and more distant resources:

    Simplified Resource Hierarchy
    CPU registers → CPU caches → main memory → local storage → network service

    This is a conceptual hierarchy, not a universal benchmark. Real values vary by processor, workload, hardware, operating system, data-access pattern, payload size and network location.

    Resource Primary Strength Typical Limitation
    CPU Rapid instruction execution and computation Finite cores, instruction throughput and scheduling capacity
    Memory Fast access to active data Finite capacity, memory bandwidth and allocation pressure
    Disk or storage Durable, comparatively large data capacity Higher latency, finite IOPS and queueing
    Network Communication across process and machine boundaries Latency, bandwidth, congestion and remote failure

    Performance rule: Memorized latency numbers become outdated and are not substitutes for measurement. Use the hierarchy to guide reasoning, then benchmark the actual target environment.

    Latency Scale

    Resource operations occur at different orders of magnitude. The following unit relationships help compare them:

    \[ 1\ second = 1{,}000\ milliseconds \]

    \[ 1\ millisecond = 1{,}000\ microseconds \]

    \[ 1\ microsecond = 1{,}000\ nanoseconds \]

    A delay measured in milliseconds represents millions of nanoseconds. Therefore, a thread waiting for storage or a remote response can potentially miss a large number of CPU execution opportunities.

    Part 1: CPU Costs

    The CPU executes instructions, performs arithmetic, evaluates branches and coordinates data movement.

    Common CPU-consuming activities include:

    • Parsing and validating requests
    • Encryption and decryption
    • Compression and decompression
    • Serialization and deserialization
    • Sorting, searching and aggregation
    • Image, audio or video processing
    • Regular-expression evaluation
    • Garbage collection or memory-management work
    • Operating-system and network-protocol processing

    CPU Execution Costs

    CPU cost is affected by more than the number of source-code statements. Important factors include:

    • Instruction count
    • Instruction complexity
    • Branch predictability
    • Cache locality
    • Vectorization
    • Lock contention
    • Context switching
    • System-call frequency
    • Runtime and compiler optimization

    Branch Prediction

    Modern processors attempt to predict the path of conditional branches. A misprediction can discard speculative work and restart execution along the correct path.

    for (size_t index = 0;
         index < count;
         ++index) {
    
        if (values[index] >
            threshold) {
    
            process_value(
                values[index]
            );
        }
    }

    Performance depends partly on the data distribution and predictability of the condition. Avoid rewriting clear code solely to remove branches unless profiling shows that the branch is a material bottleneck.

    CPU Cache Locality

    CPU caches keep recently accessed data closer to the processor. Code that accesses nearby memory locations can often benefit from spatial and temporal locality.

    Locality Type Meaning
    Temporal locality Recently accessed data is accessed again soon
    Spatial locality Data near a recently accessed location is accessed soon
    Sequential access often has useful locality
    long long sum = 0;
    
    for (size_t index = 0;
         index < count;
         ++index) {
    
        sum += values[index];
    }
    Random access can create more cache misses
    long long sum = 0;
    
    for (size_t index = 0;
         index < count;
         ++index) {
    
        sum += values[
            random_indexes[index]
        ];
    }

    The second pattern is not automatically incorrect. Some algorithms require unpredictable access. The important point is that memory-access pattern can influence CPU performance.

    CPU Utilization vs CPU Pressure

    High CPU utilization means the processor is busy. It does not automatically mean the system has a performance defect.

    CPU pressure is more likely when:

    • Runnable work waits for processor time
    • Request latency increases as load rises
    • Throughput no longer increases
    • Context switching becomes excessive
    • CPU throttling occurs
    • Important workloads cannot meet deadlines

    CPU monitoring should consider how processor time is spent, scheduling queues and whether useful work is delayed, rather than interpreting one utilization percentage in isolation.

    Context-switching Cost

    A context switch occurs when the processor changes from one executing thread or process to another. The operating system must preserve and restore execution state, and the new workload can have a different cache footprint.

    Context switching increases when:

    • Too many runnable threads compete for CPU cores
    • Threads frequently block and wake
    • Locks create contention
    • Very small tasks are scheduled independently
    • Interrupt activity is high

    Concurrency rule: More threads do not guarantee more throughput. Beyond the useful concurrency level, scheduling and contention overhead can reduce performance.

    Part 2: Memory Costs

    Main memory stores active program code, objects, buffers, stacks, caches and operating-system data.

    Memory cost includes:

    • Allocation and deallocation
    • Reading and writing memory
    • Cache misses
    • Page faults
    • Copying large buffers
    • Memory reclamation
    • Swapping or paging to storage
    • Garbage collection in managed runtimes

    Capacity Is Not the Only Memory Metric

    High memory usage alone does not always indicate a problem. Operating systems can use available memory for filesystem caching and release that memory when applications need it.

    More useful warning signals can include:

    • Allocation failures
    • Increasing swap activity
    • Frequent page reclamation
    • Major page faults
    • Out-of-memory termination
    • Repeated cache eviction
    • Growing process resident memory without stabilization
    • Latency increasing during memory pressure

    Page Faults

    Virtual memory allows a process to use a logical address space. A page fault occurs when the referenced virtual-memory page requires operating-system handling.

    Page-fault Type General Meaning
    Minor page fault The page can be resolved without reading its content from persistent storage
    Major page fault Resolving the page requires slower storage-related work

    Frequent major page faults can cause large application-latency increases because the process waits for storage-backed data.

    Memory Copy Cost

    Copying data consumes CPU time and memory bandwidth. An architecture that repeatedly copies large payloads between buffers, processes or protocol layers can spend substantial capacity moving data rather than transforming it.

    Network buffer
          |
          | Copy
          v
    Application buffer
          |
          | Copy
          v
    Serialization buffer
          |
          | Copy
          v
    Storage or network output

    Possible improvements, when supported by evidence, include:

    • Streaming instead of buffering the complete payload
    • Reusing bounded buffers
    • Avoiding unnecessary transformations
    • Using references or views when lifetime rules permit
    • Applying zero-copy facilities supported by the platform

    Data Structure Overhead

    Logical data size can be smaller than actual memory usage because data structures can contain:

    • Object headers
    • Pointers
    • Alignment padding
    • Unused collection capacity
    • Hash-table buckets
    • Allocator metadata
    • Fragmentation

    For example, storing millions of small records as separately allocated objects can use significantly more memory than the sum of the visible field sizes.

    Part 3: Disk and Storage Costs

    Persistent storage provides durability beyond process and machine memory. Storage operations generally have higher latency than CPU-cache or main memory access.

    Storage performance depends on:

    • Storage technology
    • Random vs sequential access
    • Read vs write workload
    • Operation size
    • Queue depth
    • Filesystem behaviour
    • Durability and synchronization settings
    • Compaction or background maintenance
    • Contention from other workloads

    Sequential vs Random I/O

    Access Pattern Description Typical Use
    Sequential I/O Reads or writes nearby storage locations in order Logs, scans, backups and large-file transfer
    Random I/O Accesses scattered locations Point lookups and database-page access

    Sequential access often provides higher transfer efficiency because it reduces seek, command and coordination overhead. The actual difference depends on the storage technology and workload.

    IOPS, Throughput and Latency

    Metric Meaning
    IOPS Input/output operations completed per second
    Storage throughput Bytes transferred per second
    I/O latency Time required to complete one storage operation
    Queue depth Outstanding or waiting storage operations
    Utilization How busy the storage resource is during the measurement period

    Small random operations can be limited by IOPS, while large sequential transfers can be limited by bytes per second.

    Simplified Throughput Relationship

    \[ Storage\ Throughput \approx IOPS \times Average\ Operation\ Size \]

    If a device completes 10,000 operations per second and each operation moves 4 KB:

    \[ Throughput \approx 10{,}000 \times 4\ KB = 40{,}000\ KB/s \]

    \[ Throughput \approx 40\ MB/s \]

    This is an illustrative relationship. Real throughput can be affected by queueing, protocol overhead, caching and device behaviour.

    Storage Queueing

    Storage latency can increase before the device reaches full capacity. When new operations arrive faster than storage can complete them, queue depth and waiting time increase.

    Application requests
            |
            v
    Storage queue
            |
            | Wait for device capacity
            v
    Storage device
            |
            v
    Completed I/O

    Disk monitoring should therefore include latency, queue depth, completed operations and error rates rather than only free capacity.

    Durability Has a Cost

    A write acknowledged only after durable persistence can require more work than a write acknowledged after being placed in memory.

    Faster but weaker acknowledgement:
    Request -> Memory buffer -> Acknowledge -> Persist later
    
    Stronger durability:
    Request -> Durable log or storage -> Confirm persistence -> Acknowledge

    The correct policy depends on the data-loss requirement. A financial update may require a stronger durability guarantee than a reconstructable analytics event.

    Storage Capacity vs Storage Performance

    A storage device can have sufficient free space but still become a performance bottleneck.

    Separate these questions:

    • Is there enough capacity to store the data?
    • Can the system complete the required IOPS?
    • Can the system deliver the required transfer throughput?
    • Does latency satisfy the application deadline?
    • How does performance change under queueing?
    • Can storage recover from device or node failure?

    Part 4: Network Costs

    Network communication allows processes and machines to exchange data. A network request can include:

    • Name resolution
    • Connection establishment
    • Transport and encryption handshakes
    • Serialization
    • Transmission
    • Routing
    • Queueing
    • Remote processing
    • Response transfer
    • Deserialization

    Network Latency

    Network latency depends on more than bandwidth. It can include:

    Delay Description
    Propagation delay Time for a signal to travel through the physical medium
    Transmission delay Time required to place all payload bits onto the link
    Queueing delay Time waiting in network-device or system buffers
    Processing delay Time for endpoints and network devices to process the traffic
    Remote-service delay Time spent by the destination application and its dependencies

    A high-bandwidth link can still have significant round-trip latency, especially over long geographical distances or during congestion.

    Bandwidth and Transfer Time

    A simplified payload transmission time is:

    \[ Transmission\ Time = \frac{Payload\ Size\ In\ Bits} {Link\ Rate\ In\ Bits\ Per\ Second} \]

    Example

    To transmit a 10 MB payload over an effective 100 Mb/s link, using decimal units:

    \[ Payload = 10\ MB \times 8 = 80\ Mb \]

    \[ Transmission\ Time = \frac{80\ Mb} {100\ Mb/s} = 0.8\ seconds \]

    This simplified calculation excludes connection setup, protocol overhead, congestion, retries and remote processing.

    Round Trips

    A request that requires several sequential network round trips accumulates latency.

    Client
      |
      | Round trip 1: authenticate
      v
    Identity Service
      |
      | Round trip 2: business request
      v
    Application Service
      |
      | Round trip 3: data request
      v
    Remote Database

    Reducing unnecessary round trips can improve latency, but combining operations must preserve correctness, authorization and maintainability.

    Serialization Cost

    Data must often be transformed between in-memory objects and a transport representation.

    Serialization cost depends on:

    • Payload size
    • Encoding format
    • Schema complexity
    • Compression
    • Allocation and copying
    • Validation
    • Encryption

    Smaller payloads can reduce network transfer work, but highly compressed formats can consume additional CPU.

    Remote Calls Can Fail Independently

    A local function call normally shares the current process. A remote call adds independent failure possibilities:

    • Name resolution failure
    • Connection refusal
    • Connection reset
    • Timeout
    • Rate limiting
    • Partial data transfer
    • Remote overload
    • Remote application failure
    • Uncertain result after a lost response

    Distributed-call rule: Treat a network call as a slower, failure-prone operation with explicit deadlines, bounded retries and a defined duplicate-handling policy.

    Comparing Resource Costs

    Operation Main Cost Design Concern
    Arithmetic on cached data CPU execution Instruction count and branch behaviour
    Accessing a large in-memory structure Memory latency and bandwidth Locality, allocation and pressure
    Random persistent lookup Storage latency and IOPS Indexing, queue depth and cache misses
    Large sequential file transfer Storage and network throughput Buffering and transfer size
    Remote API call Network plus remote processing Timeout, retry and dependency availability
    Compression before transfer CPU exchanged for fewer bytes Payload size, CPU cost and latency target
    Application cache Memory exchanged for less storage or network access Capacity, eviction and staleness

    CPU-bound vs I/O-bound Workloads

    Workload Type Primary Behaviour Examples
    CPU-bound Spends most active time performing computation Compression, encryption, transcoding and complex calculations
    Memory-bound Limited by memory latency, bandwidth or capacity Large graph traversal and in-memory analytics
    Disk-bound Limited by persistent-storage latency, IOPS or throughput Large scans, random database reads and backup operations
    Network-bound Limited by transfer rate, round trips or remote responses File transfer, remote APIs and distributed queries
    Mixed Different stages are limited by different resources Most production request flows

    Identify the Bottleneck

    The bottleneck is the limiting resource or stage that prevents the system from increasing useful throughput or meeting latency objectives.

    Performance Investigation
    measure → localize waiting → identify saturation → change one factor → remeasure

    A structured investigation asks:

    1. Which operation is slow or capacity-limited?
    2. What workload reproduces the behaviour?
    3. Where is time spent?
    4. Which resource is saturated or queued?
    5. Is the problem local or caused by a dependency?
    6. Does the issue affect every request or only a subset?
    7. Which change is expected to relieve the limiting stage?
    8. Did the same workload improve after the change?

    Symptoms by Resource

    Resource Possible Symptoms Measurements to Inspect
    CPU Runnable queues, increasing request latency and throughput flattening CPU time, load, run queue, per-process usage and context switches
    Memory Reclamation, swap activity, allocation failures and unstable latency Available memory, resident memory, faults, swap and pressure
    Disk Increasing I/O latency, queue depth and I/O wait Read/write latency, IOPS, throughput, queue depth and errors
    Network Higher round-trip time, retransmissions, connection delay and bandwidth saturation Traffic rate, errors, drops, retransmissions and dependency latency

    Linux Investigation Tools

    The following tools help collect evidence. Availability and output fields can vary by Linux distribution and installed packages.

    Tool Primary Use
    top Interactive process and system activity
    vmstat CPU, runnable work, memory, paging and system activity
    iostat CPU and storage-device activity
    free Memory and swap summary
    pidstat Per-process CPU and selected I/O activity
    lsof Open files and sockets associated with processes
    ss Socket and network-connection information
    sar Historical and interval-based system activity where configured
    perf CPU profiling and hardware or software event analysis

    General System Activity

    top
    vmstat 1

    Memory Summary

    free -h

    Storage Activity

    iostat -xz 1

    Per-process Activity

    pidstat 1
    pidstat -d 1

    Socket Information

    ss -s
    ss -tupn

    Open Files and Sockets

    lsof -p PROCESS_ID

    CPU Profiling

    perf stat ./application
    perf record -g ./application
    perf report

    Operational caution: Some profiling and tracing tools introduce overhead or require elevated permissions. Follow the approved process before using them in production.

    Example: Diagnose a Slow API

    Request Flow

    Client
      |
      v
    API Service
      |
      +--> Parse and validate request
      |
      +--> Query database
      |
      +--> Call external service
      |
      +--> Serialize response
      |
      v
    Client

    Investigation

    Observation Possible Interpretation
    CPU is saturated and runnable work is growing Application computation or excessive concurrency may limit progress
    CPU is low but storage latency is high The request can be waiting for disk or database I/O
    Memory pressure and swap activity increase The working set may exceed practical memory capacity
    External-call latency dominates traces The remote dependency may be the critical-path bottleneck
    Network retransmissions increase Packet loss or network congestion can contribute to delay
    Database connection queue grows The connection pool or database can be saturated

    The correct optimization depends on the measured bottleneck. Adding CPU does not solve a saturated database, and increasing database capacity does not solve unnecessary payload transfer.

    Example: Database Query Cost

    Application
        |
        | Send query
        v
    Database connection pool
        |
        | Wait for connection if pool is exhausted
        v
    Database
        |
        +--> Parse and plan
        +--> Read index pages
        +--> Read table pages
        +--> Filter and sort
        +--> Return rows
        |
        v
    Application
        |
        +--> Deserialize
        +--> Build response

    Potential costs include:

    • Network round trip
    • Connection-pool waiting
    • CPU used for planning and execution
    • Memory used for sorting or hashing
    • Storage reads after cache misses
    • Lock waiting
    • Response serialization and transfer

    Caching as a Resource Trade-off

    Caching typically spends memory to reduce repeated computation, disk access or network calls.

    Without cache:
    Request -> Application -> Database or remote service
    
    With cache:
    Request -> Application -> Memory cache
                             |
                             +--> Miss -> Database or remote service

    Caching can improve performance when:

    • Requests reuse the same data
    • The cached result remains valid long enough
    • The authoritative operation is comparatively expensive
    • The cache-hit rate is sufficient

    Caching introduces costs involving:

    • Memory capacity
    • Invalidation
    • Staleness
    • Eviction
    • Cache-miss amplification
    • Warm-up after restart

    Compression as a CPU-Network Trade-off

    Compression exchanges CPU work for fewer transferred or stored bytes.

    Original payload
          |
          | CPU compression
          v
    Smaller payload
          |
          | Network or storage transfer
          v
    Receiver
          |
          | CPU decompression
          v
    Original data

    Compression is more attractive when:

    • The payload compresses significantly
    • Network or storage is the bottleneck
    • CPU headroom is available
    • The additional latency remains acceptable

    Compression can be unattractive for tiny payloads, already compressed data or CPU-saturated services.

    Queues as a Time-Capacity Trade-off

    A queue can buffer bursts and separate request acceptance from background processing.

    Producer
       |
       | Accept work
       v
    Queue
       |
       | Wait for worker capacity
       v
    Worker
       |
       v
    Storage or external provider

    A queue does not remove work. It moves waiting time and provides controlled buffering.

    Monitor:

    • Arrival rate
    • Completion rate
    • Queue depth
    • Age of the oldest item
    • Retry volume
    • Failure or dead-letter volume

    Scaling by Resource Type

    Bottleneck Possible Responses Important Verification
    CPU Optimize hot code, reduce repeated work, add cores or partition computation Profile shows the changed code was limiting performance
    Memory capacity Reduce working set, bound caches, improve representation or add memory Pressure, faults and latency improve
    Memory bandwidth Improve locality, reduce copies or partition processing Measured memory stalls or bandwidth pressure decline
    Storage IOPS Improve indexes, cache reads, batch writes or partition workload I/O queue and latency improve under the same workload
    Storage throughput Use larger sequential operations, compression or parallel storage paths Transfer rate improves without violating latency or durability
    Network latency Reduce round trips, colocate services or use regional endpoints End-to-end traces show lower critical-path delay
    Network bandwidth Reduce payload, compress, cache closer or increase link capacity Transfer rate and application latency improve

    Common Performance Mistakes

    1

    Assuming High CPU Usage Is Always Bad

    High utilization can be healthy when throughput and latency targets are satisfied and sufficient headroom remains.

    2

    Looking Only at Used Memory

    Inspect pressure, allocation failures, reclaim, faults and swap activity, not only the amount shown as used.

    3

    Monitoring Only Free Disk Space

    Storage can become latency- or IOPS-limited while substantial free capacity remains.

    4

    Confusing Bandwidth with Latency

    Bandwidth describes transfer capacity. Latency describes delay. A high-bandwidth connection can still have a long round trip.

    5

    Adding Threads Without Measuring Contention

    More threads can increase scheduling overhead, lock contention and downstream pressure.

    6

    Adding Memory to a CPU-bound Workload

    Additional memory does not automatically improve a workload limited by computation.

    7

    Optimizing Disk When Data Already Comes from Memory

    Confirm the active path and cache behaviour before changing storage.

    8

    Ignoring Payload Size

    Large payloads consume network, memory, serialization and sometimes storage capacity.

    9

    Using Historical Latency Numbers as Current Guarantees

    Hardware and environments differ. Benchmark the actual deployment target.

    10

    Optimizing Without Reproducing the Workload

    Establish a baseline, change one factor and test the same workload again.

    Recommended Performance Tests

    Test Purpose
    CPU-intensive benchmark Measure compute capacity and scaling across cores
    Sequential memory scan Observe locality and memory-transfer behaviour
    Random memory access Observe the cost of poor locality
    Sequential storage transfer Measure large-block storage throughput
    Random storage operations Measure IOPS and latency under queueing
    Local network test Measure communication within the deployment environment
    Remote dependency test Measure end-to-end latency and failure behaviour
    Mixed service workload Observe interaction among CPU, memory, disk and network
    Sustained-load test Detect leaks, thermal effects, queue growth and background maintenance

    Resource Review Checklist

    Performance Investigation Checklist

    • The slow or capacity-limited operation is clearly identified.
    • The test workload is documented.
    • Average and peak load are distinguished.
    • CPU utilization is interpreted with runnable queues and latency.
    • User and kernel CPU time are distinguished where useful.
    • Context-switch activity is inspected.
    • Memory pressure is inspected instead of only memory usage.
    • Page faults and swap activity are reviewed.
    • Process resident memory is monitored over time.
    • Storage latency, IOPS and throughput are measured separately.
    • Random and sequential storage workloads are distinguished.
    • Storage queue depth is inspected.
    • Network latency and bandwidth are measured separately.
    • Retransmissions, drops and errors are reviewed.
    • Payload sizes are documented.
    • Remote-dependency time is visible in traces.
    • The limiting stage is supported by evidence.
    • One controlled optimization is applied at a time.
    • The same workload is rerun after optimization.
    • Correctness, reliability and cost are checked after performance changes.

    Practice Exercise

    Profile a small Linux service that reads a file, transforms records and sends results to a remote endpoint.

    Workload

    Input file
        |
        v
    Read records
        |
        v
    Parse and transform
        |
        v
    Serialize batches
        |
        v
    Send to remote endpoint
        |
        v
    Record completion

    Tasks

    1. Measure total execution time.
    2. Measure CPU utilization during parsing.
    3. Inspect memory growth and page faults.
    4. Measure storage read throughput and latency.
    5. Measure outbound network throughput.
    6. Measure remote-call latency.
    7. Determine whether the workload is CPU-, memory-, disk- or network-bound.
    8. Change the batch size and rerun the same workload.
    9. Compare throughput, latency and memory usage.
    10. Document the bottleneck and the evidence supporting the conclusion.

    Result Template

    Workload:
    <Input size, number of records and request destination>
    
    Environment:
    <CPU, memory, storage, operating system and network location>
    
    Baseline:
    - Total duration:
    - Records per second:
    - CPU utilization:
    - Maximum resident memory:
    - Major page faults:
    - Storage read throughput:
    - Storage latency:
    - Network throughput:
    - Remote-call latency:
    
    Observed bottleneck:
    <CPU, memory, disk, network or mixed>
    
    Evidence:
    <Measurements supporting the conclusion>
    
    Change:
    <One optimization or configuration change>
    
    Result:
    <Comparison using the same workload>
    
    Trade-offs:
    <Effects on correctness, memory, cost or reliability>

    Frequently Asked Questions

    1

    Why is CPU cache faster than main memory?

    Processor caches are smaller and located closer to CPU execution units, allowing lower-latency access than larger main memory.

    2

    Is high CPU utilization always a bottleneck?

    No. CPU becomes a practical bottleneck when processing demand delays work, prevents throughput growth or violates latency and headroom objectives.

    3

    Why can high memory usage be normal?

    Operating systems use otherwise available memory for useful filesystem caching. Memory pressure, allocation failure and paging behaviour are more informative than one used-memory value.

    4

    What is a major page fault?

    It is a page fault whose resolution requires slower storage-related work rather than only updating in-memory mappings.

    5

    What is the difference between storage IOPS and throughput?

    IOPS measures completed operations per second. Throughput measures bytes transferred per second.

    6

    Can a disk have free capacity and still be slow?

    Yes. Storage can be limited by latency, IOPS, transfer throughput or queue depth long before free space is exhausted.

    7

    What is the difference between network latency and bandwidth?

    Latency measures communication delay. Bandwidth measures potential data transfer rate.

    8

    Why are remote calls more expensive than function calls?

    Remote calls add serialization, network transfer, remote scheduling, remote processing and independent failure possibilities.

    9

    Does caching always improve performance?

    No. Caching helps when reuse is sufficient and freshness rules permit it. Cache misses, invalidation and memory overhead can reduce the benefit.

    10

    How should a bottleneck be identified?

    Reproduce the workload, measure the complete flow, locate waiting and saturation, make one targeted change and rerun the same test.

    11

    Should latency numbers be memorized?

    Understanding relative orders of magnitude is useful, but actual designs should use benchmark and production measurements from the target environment.

    12

    What comes after CPU, memory, disk and network costs?

    The next topic is processes versus threads, followed by synchronization, race conditions, deadlocks, files, sockets and Linux diagnostic tools.

    Key Takeaway

    CPU, memory, disk and network resources have different costs and failure characteristics. CPU performance depends on computation, locality, branches and scheduling. Memory performance depends on locality, capacity, bandwidth and pressure. Storage performance depends on latency, IOPS, throughput and queueing. Network performance depends on round trips, transfer rate, payload size and remote-service behaviour. Measure the complete workload, identify where work waits, optimize the actual bottleneck and verify the result using the same representative test.