Table of Contents

    Linux process and network tools

    COMPUTER SYSTEMS, LINUX AND CONCURRENCY

    Linux Process and Network Tools

    Learn how to inspect Linux processes, threads, CPU usage, memory pressure, storage activity, open files, sockets, network interfaces, routes and system calls through a structured troubleshooting workflow.

    Introduction

    Linux provides command-line tools for observing running processes, resource usage, open files, network connections and operating-system activity.

    These tools help answer practical troubleshooting questions such as:

    • Which process is consuming CPU?
    • Which process is using the most memory?
    • Are runnable tasks waiting for processor time?
    • Is the system reclaiming or swapping memory?
    • Is storage latency increasing?
    • Which process is listening on a port?
    • Which files and sockets belong to a process?
    • Which network route will Linux use?
    • Are connections accumulating in an unexpected state?
    • Which system call is blocking an application?

    Core idea: Do not diagnose a system from one utilization percentage. Start with the user-visible symptom, inspect the complete request path and correlate process, CPU, memory, storage and network evidence.

    In the System Design curriculum, Linux Process and Network Tools is Topic 2.7 and completes the Computer Systems, Linux and Concurrency module. The practical objective is to profile a small service using tools such as top, vmstat, iostat and lsof.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 CPU, memory, disk and network costs Linux tools expose measurements for these underlying resources.
    2 Processes versus threads Many commands report process IDs, thread IDs and scheduling states.
    3 Files and sockets Descriptors, ports and connection states appear throughout Linux diagnostics.
    4 Synchronization and deadlocks Blocked threads and excessive context switching can indicate concurrency problems.
    5 Basic shell commands Diagnostic commands are commonly filtered, combined and redirected through a shell.
    6 Appropriate system permissions Some process, socket and tracing information requires elevated access.

    Structured Troubleshooting Workflow

    Investigation Flow
    define symptom → identify process → inspect resources → inspect files and sockets → trace activity → test hypothesis
    1. Define the user-visible symptom.
    2. Record when the problem occurs.
    3. Identify the affected service and process.
    4. Inspect overall CPU, memory, storage and network activity.
    5. Inspect the affected process and its threads.
    6. Inspect open files, sockets and connection states.
    7. Capture interval-based evidence while the problem is active.
    8. Form one specific hypothesis.
    9. Make one controlled change.
    10. Repeat the same workload and compare the measurements.

    Operational caution: Some commands expose sensitive arguments, paths, addresses or process information. Tracing and profiling can also add overhead. Use approved permissions and production procedures.

    Understanding Process Identifiers

    Linux assigns identifiers to processes and threads. Common terms include:

    Term Meaning
    PID Process identifier
    PPID Parent process identifier
    TID Thread identifier
    UID User identity associated with the process
    GID Group identity associated with the process
    PGID Process-group identifier
    SID Session identifier

    ps: Process Snapshot

    The ps command displays a snapshot of process information.

    List Processes

    ps -ef

    List Processes for the Current User

    ps -u "$USER"

    Inspect One Process

    ps -p PROCESS_ID -o pid,ppid,user,state,etime,%cpu,%mem,rss,vsz,cmd

    Sort by CPU Usage

    ps -eo pid,ppid,user,state,%cpu,%mem,etime,cmd --sort=-%cpu

    Sort by Memory Usage

    ps -eo pid,ppid,user,state,%cpu,%mem,rss,vsz,etime,cmd --sort=-rss

    Display Threads

    ps -T -p PROCESS_ID

    Common ps Fields

    Field Meaning
    PID Process identifier
    PPID Parent process identifier
    STAT or S Process state and selected attributes
    %CPU Reported processor usage according to the command's calculation
    %MEM Reported proportion of physical memory
    RSS Resident memory currently present in physical memory
    VSZ Virtual address-space size
    ETIME Elapsed time since process start
    CMD Command or command line

    Common Process States

    Process-state letters help identify whether work is running, waiting or stopped. Exact reporting can vary by command and kernel.

    State General Meaning
    R Running or runnable
    S Interruptible sleep or waiting
    D Uninterruptible sleep, commonly while waiting in the kernel
    T Stopped or traced
    Z Zombie process whose termination status has not yet been collected
    I Idle kernel thread where reported

    Interpretation rule: One process-state snapshot does not establish a root cause. Observe the state repeatedly and correlate it with I/O, CPU, memory and application evidence.

    pstree: Process Hierarchy

    The pstree command displays parent-child process relationships.

    pstree -p

    Inspect One Process Tree

    pstree -p PROCESS_ID

    This is useful when investigating:

    • Worker processes created by a parent service
    • Unexpected child processes
    • Shell and script process relationships
    • Supervised application processes
    • Processes left after an incomplete shutdown

    top: Interactive System Activity

    The top command provides an interactive view of system and process activity.

    top

    Monitor One Process

    top -p PROCESS_ID

    Display Threads

    top -H -p PROCESS_ID

    Important areas include:

    • Load average
    • Task counts and states
    • CPU-time categories
    • Memory and swap summaries
    • Per-process CPU and memory activity
    • Process runtime and command

    Useful Interactive Keys

    Key Purpose
    P Sort by CPU usage
    M Sort by memory usage
    H Toggle thread display
    1 Toggle per-CPU display where supported
    k Prompt to send a signal to a process
    q Quit

    Load Average

    Linux commonly reports load averages over several intervals.

    uptime

    The values represent average demand involving runnable tasks and selected uninterruptible tasks during the reported intervals.

    Load average should be interpreted with:

    • Available CPU count
    • Runnable queue length
    • Uninterruptible waiting
    • CPU utilization
    • Storage latency
    • Application throughput and latency
    Incomplete conclusion
    Load average is high, so the CPU is definitely the bottleneck.
    Evidence-based interpretation
    Load average is elevated.
    
    Next checks:
    - Runnable task count
    - CPU idle and saturation
    - Uninterruptible task count
    - Storage latency and queue depth
    - Per-process and per-thread activity

    vmstat: CPU, Memory and System Activity

    The vmstat command reports interval-based virtual-memory, process, CPU and system activity.

    vmstat 1

    The first row can represent statistics since boot, while later rows represent the requested reporting interval. Verify the behaviour of the installed implementation.

    Important vmstat Fields

    Field General Meaning
    r Runnable tasks waiting for or using CPU
    b Tasks blocked in uninterruptible waits
    swpd Amount of swap space in use
    free Reported unused memory
    si Swap input activity
    so Swap output activity
    bi Blocks received from block devices
    bo Blocks sent to block devices
    in Interrupt activity
    cs Context switches
    us User-space CPU time
    sy Kernel-space CPU time
    id Idle CPU time
    wa Reported CPU time associated with I/O waiting

    Example Investigation

    Observation:
    - Runnable count remains high
    - CPU idle remains low
    - Application latency increases
    - Throughput no longer increases
    
    Possible hypothesis:
    CPU scheduling capacity is saturated.
    
    
    Observation:
    - Blocked count increases
    - Storage latency increases
    - CPU remains partly idle
    
    Possible hypothesis:
    Tasks are waiting for storage rather than CPU execution.

    free: Memory Summary

    free -h

    The command summarizes physical memory and swap. The human-readable option displays sizes using scaled units.

    Important memory concepts include:

    • Total memory
    • Used memory
    • Free memory
    • Shared memory
    • Buffers and cache
    • Available memory
    • Swap usage

    Memory rule: Low completely free memory is not sufficient evidence of a memory problem. Linux can use available memory for useful caching. Inspect available capacity, sustained reclaim, swap activity, faults and application latency.

    iostat: CPU and Storage Activity

    The iostat command reports CPU and block-device statistics. It is commonly provided by the sysstat package.

    iostat -xz 1

    Extended output can include request rate, transferred data, queueing, latency and utilization measurements. Exact field names depend on the installed version.

    Common Storage Measurements

    Measurement What It Helps Explain
    Read requests per second Read-operation rate
    Write requests per second Write-operation rate
    Read throughput Amount of data read per second
    Write throughput Amount of data written per second
    Average request size Typical transfer size per operation
    Queue size Outstanding storage work
    Read and write latency Time required to complete storage operations
    Device utilization How busy the device remained during the sample

    Device utilization must be interpreted together with latency, queue depth, workload type and the underlying storage architecture.

    pidstat: Per-process Statistics

    The pidstat command reports activity for individual processes and, with selected options, threads.

    Per-process CPU

    pidstat -p PROCESS_ID 1

    Per-process I/O

    pidstat -d -p PROCESS_ID 1

    Context Switching

    pidstat -w -p PROCESS_ID 1

    Thread-level Reporting

    pidstat -t -p PROCESS_ID 1

    pidstat is useful for correlating system-wide pressure with one application or selected application threads.

    lsof: Open Files and Sockets

    The lsof command lists open files. In Linux, the output can include regular files, directories, devices, pipes and sockets.

    Inspect One Process

    lsof -p PROCESS_ID

    Find Processes Using a File

    lsof /path/to/file

    List Network Sockets

    lsof -i

    Inspect One TCP Port

    lsof -iTCP:8080

    Show Listening TCP Sockets

    lsof -iTCP -sTCP:LISTEN

    lsof can help investigate:

    • Descriptor leaks
    • Unexpectedly retained files
    • Which process owns a listening port
    • Deleted files that remain open
    • Open pipes and sockets
    • Unexpected network destinations

    The /proc Filesystem

    Linux exposes process and kernel information through the /proc virtual filesystem.

    Process Status

    cat /proc/PROCESS_ID/status

    Process Command Line

    tr '\0' ' ' < /proc/PROCESS_ID/cmdline
    printf '\n'

    Open Descriptors

    ls -l /proc/PROCESS_ID/fd

    Memory Maps

    cat /proc/PROCESS_ID/maps

    Memory Summary

    cat /proc/PROCESS_ID/smaps_rollup

    Process Tasks

    ls /proc/PROCESS_ID/task

    Process I/O Counters

    cat /proc/PROCESS_ID/io

    Access to information about other processes can be restricted through permissions and system configuration.

    kill: Sending Signals

    Despite its name, kill sends a signal to a process or process group. The effect depends on the signal and the receiving process.

    List Signal Names

    kill -l

    Request Termination

    kill -TERM PROCESS_ID

    Force Termination

    kill -KILL PROCESS_ID

    Signal rule: Prefer a graceful termination request when the application can perform controlled cleanup. Forced termination does not allow the process to handle the signal or complete normal cleanup.

    Zombie Processes

    A zombie is a terminated process whose parent has not yet collected its termination status.

    Child process terminates
            |
            v
    Kernel retains termination status
            |
            v
    Parent has not called wait()
            |
            v
    Zombie entry remains

    A zombie does not continue executing application code, but its process-table entry remains until collected by the appropriate parent or reparenting mechanism.

    Find Zombie States

    ps -eo pid,ppid,state,cmd | awk '$3 == "Z"'

    Investigation should identify the parent process and determine why child termination status is not being collected.

    nice and renice

    Niceness influences CPU scheduling priority for applicable normal scheduling policies.

    Start with an Adjusted Niceness

    nice -n 10 ./application

    Adjust a Running Process

    renice 10 -p PROCESS_ID

    Niceness does not directly limit memory, disk, network or total CPU usage. Permissions and allowed adjustment direction depend on the environment.

    ip: Interfaces, Addresses and Routes

    The ip command displays and manages network interfaces, addresses, routes and related networking configuration.

    Show Interfaces

    ip link show

    Show IP Addresses

    ip address show

    Show Routes

    ip route show

    Determine Route to a Destination

    ip route get 203.0.113.10

    The output can help answer:

    • Which interface will be used?
    • Which source address is selected?
    • Which gateway is used?
    • Is a route available?
    • Is traffic using an unexpected path?

    ss: Socket Statistics

    The ss command displays socket information and is useful for investigating listening ports, established connections and socket states.

    Socket Summary

    ss -s

    Listening TCP Sockets

    ss -ltn

    Listening TCP Sockets with Processes

    ss -ltnp

    Established TCP Connections

    ss -tn state established

    UDP Sockets

    ss -lunp

    Filter by Port

    ss -ltnp 'sport = :8080'

    Common ss Options

    Option Purpose
    -l Show listening sockets
    -t Show TCP sockets
    -u Show UDP sockets
    -n Avoid name and service resolution
    -p Show process information where permissions allow
    -a Show listening and non-listening sockets

    Common TCP States

    State General Meaning
    LISTEN The local endpoint is accepting connection requests
    ESTABLISHED A connection has been established
    SYN-SENT A connection request was sent and completion is pending
    SYN-RECV A connection request was received and handshake completion is pending
    FIN-WAIT-1 Local closure has started and acknowledgement is pending
    FIN-WAIT-2 Local closure was acknowledged and remote closure is pending
    CLOSE-WAIT The remote side closed and the local application has not completed closure
    LAST-ACK Local closure is waiting for final acknowledgement
    TIME-WAIT The endpoint remains temporarily after active closure according to TCP handling

    State rule: A TCP state is not automatically an error. Investigate whether the count, duration and application behaviour are consistent with the expected workload and connection lifecycle.

    ping: Reachability and Round-trip Observation

    The ping command sends supported diagnostic requests and reports responses and timing.

    ping -c 4 203.0.113.10

    It can help observe:

    • Whether the target responds to the selected diagnostic traffic
    • Approximate round-trip timing
    • Packet-loss observations during the sample

    Failure to receive a response does not prove that the host or application is unavailable. Intermediate devices or host policies can filter the diagnostic protocol while application traffic remains permitted.

    traceroute and tracepath

    Route-tracing tools attempt to reveal intermediate network hops toward a destination.

    traceroute 203.0.113.10
    tracepath 203.0.113.10

    These tools can help investigate:

    • Unexpected routing paths
    • Where responses stop appearing
    • Changes in path between environments
    • Potential path maximum transmission unit observations

    Missing hop responses do not necessarily mean traffic stops at that hop. Network devices can handle or filter diagnostic responses differently from normal application traffic.

    dig: DNS Investigation

    The dig command queries Domain Name System information.

    Query an Address Record

    dig example.com A

    Query an IPv6 Address Record

    dig example.com AAAA

    Query Mail-exchange Records

    dig example.com MX

    Query a Specific Resolver

    dig @RESOLVER_ADDRESS example.com A

    Trace Delegation

    dig +trace example.com

    Review:

    • Response status
    • Returned records
    • Time to live
    • Authoritative information
    • Resolver used
    • Query duration

    curl: Application-level Network Testing

    The curl command transfers data using supported protocols and is especially useful for testing HTTP endpoints.

    Fetch Headers and Body

    curl -i https://example.com/

    Verbose Request

    curl -v https://example.com/

    Display Timing Measurements

    curl -o /dev/null -sS \
      -w 'dns=%{time_namelookup}\nconnect=%{time_connect}\ntls=%{time_appconnect}\nfirst_byte=%{time_starttransfer}\ntotal=%{time_total}\n' \
      https://example.com/

    Test a Specific Address

    curl --resolve example.com:443:203.0.113.10 https://example.com/

    Replace example names and addresses with approved test targets. Avoid placing credentials or sensitive values directly in shell history.

    nc: Basic TCP and UDP Testing

    The nc command, commonly called netcat, can create basic client or listening connections. Available options vary among implementations.

    Test TCP Connectivity

    nc -vz 203.0.113.10 443

    Start a Test Listener

    nc -l 8080

    Connect to the Test Listener

    nc 127.0.0.1 8080

    Security caution: Open test listeners only in approved environments and bind them according to the intended network scope.

    strace: System-call Tracing

    The strace command traces system calls and related results for supported Linux processes.

    Trace a New Process

    strace ./application

    Attach to a Running Process

    strace -p PROCESS_ID

    Follow Child Processes

    strace -f ./application

    Trace File and Network Calls

    strace -f -e trace=file,network ./application

    Include Timing

    strace -f -tt -T ./application

    System-call tracing can help identify:

    • Missing files
    • Permission failures
    • Repeated connection attempts
    • Blocked reads or receives
    • Unexpected process creation
    • Frequent small I/O calls
    • System-call errors

    Tracing can affect application timing and produce sensitive output. Capture only the information needed for the approved investigation.

    perf: CPU Profiling

    The perf tool can collect CPU performance measurements and execution profiles where supported and permitted.

    Measure a Command

    perf stat ./application

    Record a Profile

    perf record -g ./application

    Review the Recorded Profile

    perf report

    Profiling helps identify where CPU time is spent. It does not by itself explain waiting in a database, remote service or storage system.

    sar: Historical and Interval Statistics

    The sar command reports system activity and can access historical data when system-activity collection is configured.

    CPU Activity

    sar -u 1

    Memory Activity

    sar -r 1

    Device Activity

    sar -d 1

    Network-interface Activity

    sar -n DEV 1

    Historical availability depends on whether data collection was enabled before the incident.

    systemctl and journalctl

    On systems using systemd, systemctl and journalctl help inspect service state and logs.

    Service Status

    systemctl status SERVICE_NAME

    Service Logs

    journalctl -u SERVICE_NAME

    Recent Service Logs

    journalctl -u SERVICE_NAME --since "30 minutes ago"

    Follow New Log Entries

    journalctl -u SERVICE_NAME -f

    Service and journal availability depend on the init and logging systems used by the Linux environment.

    Scenario 1: High CPU Usage

    Investigation Flow

    User reports slow responses
            |
            v
    Check overall CPU and load
            |
            v
    Identify high-CPU process
            |
            v
    Inspect per-thread CPU
            |
            v
    Capture CPU profile
            |
            v
    Locate hot function or loop
            |
            v
    Apply one controlled correction
            |
            v
    Repeat the same workload

    Commands

    uptime
    top
    ps -eo pid,ppid,%cpu,%mem,etime,cmd --sort=-%cpu
    top -H -p PROCESS_ID
    pidstat -t -p PROCESS_ID 1
    perf stat -p PROCESS_ID

    Investigate:

    • Whether throughput is still increasing
    • Whether runnable work is accumulating
    • Whether one thread dominates CPU usage
    • Whether context switching is excessive
    • Whether CPU time is in application or kernel work
    • Whether throttling or limits apply

    Scenario 2: Memory Growth

    Commands

    free -h
    vmstat 1
    ps -p PROCESS_ID -o pid,rss,vsz,%mem,etime,cmd
    cat /proc/PROCESS_ID/status
    cat /proc/PROCESS_ID/smaps_rollup

    Investigate:

    • Whether resident memory grows continuously
    • Whether memory stabilizes after the workload stops
    • Whether swap input and output increase
    • Whether page faults increase
    • Whether caches are bounded
    • Whether child processes or threads are also growing
    • Whether the process was terminated for memory exhaustion

    Scenario 3: Storage Bottleneck

    Commands

    vmstat 1
    iostat -xz 1
    pidstat -d -p PROCESS_ID 1
    lsof -p PROCESS_ID
    cat /proc/PROCESS_ID/io

    Investigate:

    • Read and write request rates
    • Transfer throughput
    • Storage-operation latency
    • Queue depth
    • Processes producing the workload
    • Filesystem capacity
    • Unexpected repeated small writes
    • Log, backup, compaction or batch activity

    Scenario 4: Service Is Not Reachable

    Investigation Flow

    Confirm process and service status
            |
            v
    Confirm listening socket
            |
            v
    Confirm local address and route
            |
            v
    Test local connection
            |
            v
    Test remote path
            |
            v
    Test application protocol
            |
            v
    Inspect logs and packet path

    Commands

    systemctl status SERVICE_NAME
    ss -ltnp
    ip address show
    ip route show
    curl -v http://127.0.0.1:PORT/
    nc -vz SERVER_ADDRESS PORT
    journalctl -u SERVICE_NAME

    Investigate:

    • Whether the application process is running
    • Whether the expected port is listening
    • Whether the listener is bound to the expected address
    • Whether local requests succeed
    • Whether routing reaches the target
    • Whether a firewall or policy blocks the traffic
    • Whether the protocol and TLS expectations match
    • Whether the application rejects the request

    Scenario 5: Connection Accumulation

    Commands

    ss -s
    ss -tan
    ss -tan state established
    ss -tan state close-wait
    lsof -iTCP -sTCP:ESTABLISHED
    ls /proc/PROCESS_ID/fd | wc -l

    Investigate:

    • Which connection states are increasing
    • Whether the application closes connections correctly
    • Whether clients are slow
    • Whether connection deadlines exist
    • Whether descriptor limits are being approached
    • Whether a connection pool returns resources after errors
    • Whether retries create excessive connections

    Scenario 6: Descriptor Leak

    A descriptor leak occurs when an application opens or accepts resources but fails to release them after use.

    Commands

    ls /proc/PROCESS_ID/fd | wc -l
    lsof -p PROCESS_ID
    watch -n 2 'ls /proc/PROCESS_ID/fd | wc -l'

    Possible symptoms include:

    • Steadily increasing descriptor count
    • Failure to open files
    • Failure to accept new connections
    • Old sockets remaining open
    • Deleted files continuing to consume storage

    Scenario 7: Process Is Running but Not Progressing

    Commands

    ps -T -p PROCESS_ID
    top -H -p PROCESS_ID
    pidstat -w -t -p PROCESS_ID 1
    strace -f -p PROCESS_ID
    gdb -p PROCESS_ID

    Investigate whether threads are:

    • Waiting for mutexes
    • Waiting for condition variables
    • Blocked in file or socket I/O
    • Waiting for child processes
    • Sleeping on timers
    • Repeatedly executing without completing work

    A debugger thread dump can help distinguish deadlock, long I/O waiting, livelock and normal waiting.

    Diagnostic Capture Template

    Incident:
    <Short description>
    
    Observed symptom:
    <Latency, failure, timeout, resource growth or unavailability>
    
    Affected service:
    <Service and process identifier>
    
    Observation time:
    <Time and timezone>
    
    Workload:
    <Request rate, users, batch or test>
    
    Process evidence:
    - Process state:
    - CPU:
    - Memory:
    - Threads:
    - Context switching:
    
    Storage evidence:
    - Read/write rate:
    - Latency:
    - Queue depth:
    - Errors:
    
    Network evidence:
    - Listening address and port:
    - Connection states:
    - Route:
    - Application request result:
    - Errors or retransmissions:
    
    Open resources:
    - Descriptor count:
    - Important files:
    - Important sockets:
    
    Hypothesis:
    <Specific testable explanation>
    
    Controlled change:
    <One change applied>
    
    Result:
    <Comparison using the same workload>

    Common Diagnostic Mistakes

    1

    Using Only One Snapshot

    Many problems require interval-based observations to reveal trends, queueing and sustained pressure.

    2

    Assuming High CPU Usage Is Automatically Bad

    High CPU usage can be productive when latency, throughput and headroom remain within requirements.

    3

    Assuming Low Free Memory Means a Leak

    Linux can use memory for caches. Inspect available memory, process growth, faults and swap activity.

    4

    Looking Only at Disk Space

    A storage resource can have free capacity while being limited by latency, IOPS, throughput or queue depth.

    5

    Assuming ping Proves Application Health

    A host can respond to network diagnostics while the application remains unavailable or incorrect.

    6

    Assuming a Listening Port Proves Readiness

    A socket can be listening while downstream dependencies or application functions are failing.

    7

    Killing the Process Before Capturing Evidence

    Restarting can remove thread, descriptor and connection-state evidence needed for root-cause analysis.

    8

    Using Forced Termination First

    Forced termination prevents controlled cleanup and can complicate recovery from partially completed work.

    9

    Ignoring Permissions

    Missing process or socket details can reflect access restrictions rather than absence of activity.

    10

    Running Heavy Tracing During Peak Production Load

    Profiling and tracing can add overhead or expose sensitive information. Follow the approved operational process.

    11

    Changing Several Things at Once

    Multiple simultaneous changes make it difficult to identify which change affected the result.

    12

    Ignoring Application-level Metrics

    System metrics must be correlated with request rate, latency, errors, queue delay and business-operation outcomes.

    Practical Profiling Lab

    Profile a small service under a repeatable workload using Linux process and network tools.

    Lab Tasks

    1. Start the service and record its process ID.
    2. Record the service command, parent process and elapsed runtime.
    3. Capture a baseline with no client workload.
    4. Apply a controlled request workload.
    5. Observe overall CPU, memory and load with top and vmstat.
    6. Observe per-process CPU and I/O with pidstat.
    7. Observe storage behaviour with iostat.
    8. List the service's open descriptors with lsof.
    9. Confirm its listening socket with ss.
    10. Test its application endpoint with curl.
    11. Inspect system calls for one controlled request using strace.
    12. Identify the strongest bottleneck hypothesis.
    13. Apply one controlled change.
    14. Repeat the same workload and compare the metrics.

    Lab Evidence Table

    Area Baseline Under Load After Change
    Request throughput Record measurement Record measurement Record measurement
    p95 latency Record measurement Record measurement Record measurement
    CPU usage Record measurement Record measurement Record measurement
    Resident memory Record measurement Record measurement Record measurement
    Runnable tasks Record measurement Record measurement Record measurement
    Context switches Record measurement Record measurement Record measurement
    Storage latency Record measurement Record measurement Record measurement
    Open descriptors Record measurement Record measurement Record measurement
    Active connections Record measurement Record measurement Record measurement
    Error rate Record measurement Record measurement Record measurement

    Diagnostic Best Practices

    Recommended Practices

    • Begin with the user-visible symptom.
    • Record exact observation times and timezones.
    • Capture interval-based measurements rather than one snapshot.
    • Identify the exact process and threads involved.
    • Correlate CPU, memory, storage and network evidence.
    • Distinguish resource usage from resource pressure.
    • Use numeric output when name resolution can delay commands.
    • Check process and service boundaries separately.
    • Confirm both listening sockets and application-level responses.
    • Inspect open descriptors when leaks or retained resources are suspected.
    • Capture thread states before restarting a stalled process when permitted.
    • Use tracing and profiling only when lighter evidence is insufficient.
    • Protect credentials, arguments, paths and payload data.
    • Save the commands and relevant outputs used in the investigation.
    • State one testable hypothesis at a time.
    • Apply one controlled change and repeat the same workload.
    • Verify correctness and reliability after performance changes.
    • Document environment, permissions and tool versions where relevant.

    Frequently Asked Questions

    1

    What does ps show?

    ps displays a snapshot of process information, including identifiers, states, resource measurements and commands according to the selected options.

    2

    What is the difference between ps and top?

    ps provides a process snapshot. top provides an interactive view that refreshes system and process activity.

    3

    What does vmstat help diagnose?

    vmstat reports interval-based information about runnable and blocked tasks, memory, swapping, system activity and CPU-time categories.

    4

    What does iostat measure?

    iostat reports CPU and storage-device measurements, including operation rates, transfer volume, latency, queueing and device activity where supported.

    5

    What is lsof used for?

    lsof lists open files and can show regular files, directories, pipes, devices and network sockets associated with processes.

    6

    What is ss used for?

    ss displays socket information, including listening sockets, active connections and protocol states.

    7

    Does a listening port prove the service is healthy?

    No. It proves that a socket is listening. The application can still fail authentication, processing or dependency operations.

    8

    Does high load average always mean high CPU usage?

    No. Load can include runnable demand and selected uninterruptible waiting. Correlate it with CPU, task states and I/O measurements.

    9

    What is strace used for?

    strace observes system-call activity and can help identify file, network, process and blocking behaviour.

    10

    What is perf used for?

    perf can collect processor-related measurements and profiles to show where CPU execution time is spent.

    11

    Why should diagnostics be captured before restarting?

    Restarting can remove thread states, open-resource evidence, queue conditions and connection states needed to explain the original failure.

    12

    What comes after Linux process and network tools?

    The next module is Networking and Web Request Lifecycle, beginning with TCP/IP and UDP.

    Key Takeaway

    Linux process and network tools convert operating-system behaviour into evidence. Use ps, top, vmstat, free, iostat and pidstat to inspect processes and resources. Use lsof, /proc, ip, ss, dig and curl to inspect files, sockets, routes, DNS and application requests. Use strace and perf when deeper tracing or profiling is justified. Always correlate measurements with the user-visible symptom, capture evidence before changing the system and verify one hypothesis at a time.