Table of Contents

    race conditions

    COMPUTER SYSTEMS, LINUX AND CONCURRENCY

    Race Conditions

    Learn how uncontrolled timing between concurrent operations can produce lost updates, duplicate actions, inconsistent state, unsafe object access and unpredictable program behaviour.

    Introduction

    Concurrent programs allow multiple threads, processes, tasks or requests to make progress during overlapping periods. This can improve responsiveness and resource utilization, but it also introduces execution-order uncertainty.

    A race condition occurs when program correctness depends on the relative timing or ordering of concurrent events, and one or more possible orderings produce an incorrect result.

    Race conditions can affect:

    • Shared counters
    • Account balances
    • Inventory quantities
    • Queues and collections
    • Files and shared records
    • Object lifetimes
    • Permission checks
    • Cache entries
    • Database updates
    • Distributed requests and retries

    Core idea: A race condition is not simply “two threads running at the same time.” It is a correctness defect in which an uncontrolled ordering can change the outcome of the program.

    In your System Design curriculum, Race Conditions is Topic 2.4 under Computer Systems, Linux and Concurrency. The module introduces synchronization, race conditions and deadlocks before moving to files, sockets and Linux diagnostic tools. The practical lab includes identifying and correcting a race condition.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Processes versus threads Threads within one process commonly share memory and resources.
    2 Synchronization Mutexes, atomics and ownership rules can prevent selected races.
    3 Shared mutable state Many race conditions involve concurrently accessed changing data.
    4 Object lifetime Concurrent deallocation and access can create dangling-pointer defects.
    5 Basic testing and debugging Race conditions can appear only under selected timing conditions.

    What Is a Race Condition?

    A race condition is a semantic correctness problem caused by an uncontrolled ordering or timing dependency between concurrent operations.

    Race Condition
    concurrent operations + uncontrolled ordering + order-sensitive logic → incorrect possible outcome

    A race condition commonly requires:

    • Two or more overlapping execution flows
    • Shared state, a shared resource or an ordering dependency
    • Missing, incomplete or incorrect coordination
    • At least one harmful execution ordering

    Thread scheduling is nondeterministic from the application's perspective. A program must therefore remain correct under every execution ordering permitted by its synchronization design.

    Race Condition vs Data Race

    The terms race condition and data race are related but should not always be treated as identical.

    Concept Description
    Race condition A correctness defect caused by an undesirable timing or ordering of concurrent events
    Data race Conflicting unsynchronized memory accesses from different threads, where at least one access modifies the memory location

    Many race conditions are caused by data races, but ordering defects can exist even when individual memory accesses are protected.

    Synchronized Access but Incorrect Ordering

    pthread_mutex_lock(
        &state_mutex
    );
    
    system_state =
        STATE_READY;
    
    pthread_mutex_unlock(
        &state_mutex
    );
    pthread_mutex_lock(
        &state_mutex
    );
    
    system_state =
        STATE_STOPPED;
    
    pthread_mutex_unlock(
        &state_mutex
    );

    The mutex prevents the writes from occurring simultaneously. However, if application correctness requires STATE_STOPPED to be the final state, an uncontrolled order can still create a race condition even though the individual accesses use a lock.

    Important: Absence of a detected data race does not automatically prove that the program's higher-level ordering is correct.

    Why Race Conditions Are Nondeterministic

    Operating systems can pause and resume threads at many points. Execution order can change because of:

    • Scheduler decisions
    • Available CPU cores
    • Interrupts
    • Storage and network timing
    • Lock contention
    • Memory allocation
    • Compiler optimization
    • Runtime behaviour
    • Diagnostic logging
    • Differences in machine load
    Execution 1:
    
    Thread A: Read
    Thread A: Modify
    Thread A: Write
    Thread B: Read
    Thread B: Modify
    Thread B: Write
    
    Result appears correct.
    
    
    Execution 2:
    
    Thread A: Read
    Thread B: Read
    Thread A: Modify
    Thread B: Modify
    Thread A: Write
    Thread B: Write
    
    One update is lost.

    This timing dependency explains why a race condition can disappear during debugging and reappear under production load.

    Read-Modify-Write Race

    A read-modify-write operation reads a value, calculates a new value and writes the result.

    ++shared_counter;

    The statement can conceptually involve:

    1. Read shared_counter
    2. Add one
    3. Write the new value

    Lost-update Interleaving

    Initial counter: 100
    
    Thread A reads: 100
    Thread B reads: 100
    
    Thread A calculates: 101
    Thread B calculates: 101
    
    Thread A writes: 101
    Thread B writes: 101
    
    Expected after two increments: 102
    Observed result: 101

    Unsafe POSIX Thread Program

    The following program intentionally contains a data race for educational analysis. It must not be used as a correct shared-counter implementation.

    #define _POSIX_C_SOURCE 200809L
    
    #include <pthread.h>
    #include <stdio.h>
    #include <stdlib.h>
    
    enum
    {
        THREAD_COUNT = 4,
        INCREMENTS_PER_THREAD = 100000
    };
    
    static long long shared_counter = 0;
    
    static void *increment_counter(
        void *argument)
    {
        (void)argument;
    
        for (int index = 0;
             index < INCREMENTS_PER_THREAD;
             ++index) {
    
            ++shared_counter;
        }
    
        return NULL;
    }
    
    int main(void)
    {
        pthread_t threads[
            THREAD_COUNT
        ];
    
        int created = 0;
    
        for (int index = 0;
             index < THREAD_COUNT;
             ++index) {
    
            if (pthread_create(
                    &threads[index],
                    NULL,
                    increment_counter,
                    NULL) != 0) {
    
                break;
            }
    
            ++created;
        }
    
        for (int index = 0;
             index < created;
             ++index) {
    
            pthread_join(
                threads[index],
                NULL
            );
        }
    
        printf(
            "Expected: %lld\n",
            (long long)THREAD_COUNT *
            INCREMENTS_PER_THREAD
        );
    
        printf(
            "Observed: %lld\n",
            shared_counter
        );
    
        return
            created == THREAD_COUNT
                ? EXIT_SUCCESS
                : EXIT_FAILURE;
    }

    Why the Program Is Invalid

    • All worker threads access the same object.
    • Every worker modifies that object.
    • No synchronization orders the conflicting accesses.
    • The increment is not guaranteed to be an indivisible operation.
    • In C, the conflicting unsynchronized accesses create undefined behaviour.

    The program is not merely allowed to print a smaller number. Once undefined behaviour occurs, the language does not provide a reliable result contract.

    Correct with a Mutex

    static long long shared_counter = 0;
    
    static pthread_mutex_t counter_mutex =
        PTHREAD_MUTEX_INITIALIZER;
    
    static void *increment_counter(
        void *argument)
    {
        (void)argument;
    
        for (int index = 0;
             index < INCREMENTS_PER_THREAD;
             ++index) {
    
            if (pthread_mutex_lock(
                    &counter_mutex) != 0) {
    
                return NULL;
            }
    
            ++shared_counter;
    
            if (pthread_mutex_unlock(
                    &counter_mutex) != 0) {
    
                return NULL;
            }
        }
    
        return NULL;
    }

    Every participating access to the counter must follow the same synchronization policy.

    Correct with an Atomic Counter

    An independent statistical counter can be implemented with a C atomic object.

    #include <stdatomic.h>
    
    static atomic_llong shared_counter = 0;
    
    static void *increment_counter(
        void *argument)
    {
        (void)argument;
    
        for (int index = 0;
             index < INCREMENTS_PER_THREAD;
             ++index) {
    
            atomic_fetch_add_explicit(
                &shared_counter,
                1,
                memory_order_relaxed
            );
        }
    
        return NULL;
    }

    Relaxed ordering is suitable here only because the counter is independent and is not used to publish or coordinate other shared state.

    Atomic limitation: Making individual fields atomic does not automatically preserve a business invariant involving multiple fields or multiple operations.

    Check-Then-Act Race

    A check-then-act race occurs when a thread checks a condition and later performs an action based on that condition, but shared state can change between the check and the action.

    Inventory Example

    if (inventory_count > 0) {
        --inventory_count;
    
        create_order();
    }

    Harmful Interleaving

    Initial inventory: 1
    
    Thread A checks inventory > 0: true
    Thread B checks inventory > 0: true
    
    Thread A decreases inventory
    Thread B decreases inventory
    
    Two orders can be accepted for one item.

    The check and update must be treated as one coordinated operation.

    Protected Check and Update

    static int reserve_item(
        struct Inventory *inventory)
    {
        if (inventory == NULL) {
            return 0;
        }
    
        if (pthread_mutex_lock(
                &inventory->mutex) != 0) {
    
            return 0;
        }
    
        int reserved = 0;
    
        if (inventory->available > 0) {
            --inventory->available;
            ++inventory->reserved;
    
            reserved = 1;
        }
    
        pthread_mutex_unlock(
            &inventory->mutex
        );
    
        return reserved;
    }

    The mutex protects the relationship between available and reserved quantities throughout the complete decision and state change.

    Check-Then-Insert Race

    Checking that a record does not exist and then inserting it as two separate unprotected operations can create duplicates.

    Thread A:
    Check whether username exists -> No
    
    Thread B:
    Check whether username exists -> No
    
    Thread A:
    Insert username
    
    Thread B:
    Insert same username

    Possible protections include:

    • A mutex around the complete in-memory operation
    • A database uniqueness constraint
    • An atomic insert-if-absent operation
    • A compare-and-exchange loop for a suitable structure
    • Single ownership of the affected key or partition

    Database rule: Application-level checking alone is not a reliable replacement for a database-enforced uniqueness constraint when concurrent database writers are possible.

    Time-of-Check to Time-of-Use Race

    A time-of-check to time-of-use race, often abbreviated as TOCTOU, occurs when a program checks a resource and uses it later, while another actor can change the resource between those operations.

    Check:
    Path refers to an acceptable file.
    
    Race window:
    Another process replaces or redirects the path.
    
    Use:
    Program opens or modifies a different object.

    File Example

    Separate existence check and open
    if (access(
            filename,
            F_OK) == 0) {
    
        FILE *file =
            fopen(
                filename,
                "r"
            );
    }

    The file associated with the pathname can change between access() and fopen(). Security-sensitive file operations should use platform interfaces and designs that avoid relying on a separately checked pathname.

    Object-lifetime Race

    One thread can free or invalidate an object while another thread still holds a pointer to it.

    Thread A:
    Read pointer to shared object
    
    Thread B:
    Remove object from collection
    Free object
    
    Thread A:
    Dereference old pointer
    
    Result:
    Use after free

    Unsafe Example

    struct Session *session =
        find_session(
            session_id
        );
    
    /*
     * Another thread can remove
     * and free the session here.
     */
    
    process_session(
        session
    );

    Possible designs include:

    • Holding the protecting lock while acquiring valid ownership
    • Reference counting with correct atomic and lifetime rules
    • Copying immutable data while protected
    • Deferred reclamation
    • Single-owner state
    • Message passing instead of shared pointers

    Collection Race

    A collection can become invalid when one thread changes its storage while another thread iterates through it.

    Thread A:
    Begin iterating through dynamic array
    
    Thread B:
    Insert item
    Array reallocates
    Old storage is released
    
    Thread A:
    Continue using old array pointer

    Protecting only the collection's item count is insufficient. The synchronization design must cover the collection storage, size, capacity, indexes and object lifetimes as one coherent invariant.

    Ordering Race

    An ordering race occurs when operations are individually protected but can occur in an invalid sequence.

    Initialization Example

    Required order:
    
    1. Initialize configuration
    2. Publish ready state
    3. Start processing requests
    
    
    Unsafe order:
    
    Request worker reads ready state
    Request worker accesses incomplete configuration
    Initialization finishes afterward

    Correctness may require a condition variable, initialization barrier, one-time initialization primitive or another explicit ordering mechanism.

    Balance-update Race

    Consider two concurrent withdrawals from the same account.

    Initial balance: 100
    
    Withdrawal A: 80
    Withdrawal B: 40
    
    Thread A reads balance: 100
    Thread B reads balance: 100
    
    Thread A verifies sufficient funds
    Thread B verifies sufficient funds
    
    Thread A writes balance: 20
    Thread B writes balance: 60
    
    Both withdrawals were accepted,
    although their combined value exceeded the balance.

    Protecting the final write alone does not solve the problem. Balance validation and modification must occur in one atomic business operation.

    Duplicate-processing Race

    Two workers can receive or discover the same task and both process it.

    Worker A checks task status: pending
    Worker B checks task status: pending
    
    Worker A performs external action
    Worker B performs external action
    
    Worker A marks completed
    Worker B marks completed

    The system can produce duplicate notifications, orders, payments or file transformations.

    Possible protections include:

    • Atomic task claiming
    • Leases with expiry and ownership
    • Idempotency keys
    • Unique operation identifiers
    • Transactional state transitions
    • Idempotent downstream operations

    Race Conditions Beyond Threads

    Race conditions also occur across processes, services and distributed requests.

    Environment Example Race
    Multiple processes Two processes update the same file without locking
    Database clients Two requests read the same version and overwrite each other's update
    Message workers Two consumers process the same logical operation
    Cache and database A stale cache value overwrites a newer database value
    Deployment Two administrators apply conflicting configuration changes
    Retries A timed-out request is repeated after the first attempt already committed

    A local mutex protects only participating threads in the same process. It does not coordinate independent processes or services unless the selected synchronization mechanism explicitly spans those boundaries.

    File-update Race

    Several processes can corrupt or lose updates when they independently rewrite the same file.

    Process A reads version 1
    Process B reads version 1
    
    Process A creates version 2A
    Process B creates version 2B
    
    Process A replaces file with version 2A
    Process B replaces file with version 2B
    
    Process A's update is lost.

    Temporary-file replacement prevents some partial-write failures, but it does not automatically serialize several writers. Multi-writer designs require an appropriate locking, versioning, transaction or ownership mechanism.

    Prevention Strategies

    Strategy Best Used For Important Limitation
    Mutex Protecting a compound in-process invariant Can create contention or deadlock
    Atomic operation Selected counters, flags and lock-free state transitions Does not automatically protect multi-field invariants
    Condition variable Waiting for a synchronized predicate The predicate still requires correct protection
    Single ownership Avoiding concurrent mutation The owner can become a bottleneck
    Immutable data Sharing completed snapshots safely Updates require replacement or versioning
    Message passing Coordinating independent workers or processes Delivery, ordering and duplicate rules are required
    Database transaction Coordinating related persistent changes Isolation level and contention must be understood
    Unique constraint Preventing duplicate persistent identifiers The application must handle constraint failure
    Optimistic concurrency Detecting updates based on an old version Conflicts require retry or rejection
    Idempotency key Protecting retryable business operations Key scope, retention and response reuse must be defined

    Optimistic Concurrency Control

    Optimistic concurrency allows operations to proceed without holding an exclusive lock for the entire workflow. The update succeeds only if the state has not changed since it was read.

    Read record:
    balance = 100
    version = 7
    
    Calculate:
    new balance = 80
    
    Conditional update:
    Update balance to 80
    only where version = 7
    
    If one row is updated:
    Success
    
    If zero rows are updated:
    Another operation changed the record
    Retry or reject according to policy

    SQL-style Example

    UPDATE accounts
    SET
        balance = :new_balance,
        version = version + 1
    WHERE
        account_id = :account_id
        AND version = :expected_version;

    The application must verify the affected-row count and handle a conflict explicitly.

    Single-owner Design

    A strong method for avoiding races is to allow only one execution context to modify a given piece of state.

    Worker A ----+
                 |
    Worker B ----+--> Command Queue --> State Owner
                 |                         |
    Worker C ----+                         v
                                      Mutable State

    The state owner processes commands sequentially. Other threads communicate through messages rather than modifying the state directly.

    This design can simplify correctness, but it requires:

    • A bounded queue
    • Defined command ordering
    • Failure handling
    • Backpressure
    • State-owner recovery
    • Capacity analysis

    Immutable Data

    Immutable objects do not change after construction. Several threads can read a correctly published immutable object without coordinating modifications to that object.

    Build new immutable configuration
            |
            v
    Publish completed configuration
            |
            +--> Reader A
            +--> Reader B
            +--> Reader C

    Publication and object lifetime still require correct coordination. An object being conceptually immutable does not make an unsafe pointer publication valid.

    How to Find Race Conditions

    Race-condition investigation combines design review, stress testing, tracing and specialized tools.

    Code-review Questions

    • Which objects can be accessed by several threads?
    • Which shared objects can change?
    • What synchronization protects each shared invariant?
    • Do all readers and writers follow the same policy?
    • Can state change between a check and a later action?
    • Can an object be freed while another thread holds a pointer?
    • Can two workers claim the same logical task?
    • Can a retry repeat an already completed business operation?
    • Are external libraries and functions thread-safe?
    • Are callbacks executed concurrently?

    Stress Techniques

    • Run the test repeatedly
    • Increase the number of worker threads
    • Use barriers to start workers simultaneously
    • Add controlled delays around suspicious operations
    • Vary CPU load
    • Use several processor cores
    • Repeat with different optimization levels
    • Test start-up and shutdown concurrently
    • Force retries and timeouts
    • Check invariants after every run

    Testing limitation: Repeated success does not prove that a concurrent program is race-free. Testing observes only the execution orderings that occurred during those runs.

    ThreadSanitizer

    Supported compilers can instrument a program to detect many executed thread-level data races.

    Build

    cc -std=c17 -Wall -Wextra -Wpedantic -g -O1 -fsanitize=thread -pthread race_example.c -o race_example

    Run

    ./race_example

    Toolchain and platform support vary. ThreadSanitizer can add substantial execution and memory overhead, so it is intended for diagnostic builds rather than normal production deployment.

    A clean detector run does not prove the absence of all race conditions. It may not observe unexecuted paths, and higher-level ordering races can remain even when individual memory accesses are synchronized.

    Linux Observation Commands

    Show Threads

    ps -T -p PROCESS_ID

    Observe Per-thread Activity

    top -H -p PROCESS_ID

    Observe Context Switching

    pidstat -w -p PROCESS_ID 1

    Record CPU Profile

    perf record -g ./application
    perf report

    These commands can help explain concurrency and performance behaviour, but ordinary process-monitoring tools do not prove the absence of a data race.

    Debugging Strategy

    Race Investigation
    reproduce → identify shared state → capture ordering → define invariant → apply synchronization → stress test
    1. Identify the incorrect observable result.
    2. Reduce the problem to the smallest reproducible test.
    3. List every thread, process or request involved.
    4. Identify the shared state or ordering dependency.
    5. Document the required invariant.
    6. Write one harmful interleaving that violates the invariant.
    7. Apply an appropriate ownership, transaction or synchronization strategy.
    8. Retest under varied timing and load.
    9. Add a regression test for the corrected behaviour.
    10. Measure contention introduced by the correction.

    Race-condition Review Checklist

    Review Shared and Concurrent Operations

    • Shared mutable objects are identified.
    • Each shared object has a documented owner or synchronization policy.
    • Every participating read follows the policy.
    • Every participating write follows the policy.
    • Read-modify-write operations are protected as one operation.
    • Check-then-act operations are protected as one decision.
    • Persistent uniqueness is enforced at the appropriate storage boundary.
    • Object lifetime is protected during concurrent access.
    • Collection iteration and modification are coordinated.
    • Initialization is completed before dependent use.
    • Shutdown cannot invalidate objects still in use.
    • Retries cannot create unintended duplicate effects.
    • Task claiming is atomic or conflict-aware.
    • Atomic operations use an appropriate memory ordering.
    • Multi-field invariants are not incorrectly split across independent atomics.
    • Process-local locks are not assumed to coordinate remote writers.
    • File writes account for multiple writers.
    • Thread-safety of called libraries is verified.
    • Concurrency tests check final invariants.
    • Supported race-detection tools are included in diagnostic testing.

    Common Mistakes

    1

    Assuming a Simple Statement Is Atomic

    Assignment, increment, collection updates and compound expressions can involve several machine-level operations.

    2

    Protecting Writes but Not Reads

    Unsynchronized reads can conflict with writes and can observe invalid or inconsistent state.

    3

    Locking Only the Final Update

    The condition check and dependent update must often be protected as one logical operation.

    4

    Using Different Locks for the Same Invariant

    Mutual exclusion works only when participating accesses use a compatible, shared synchronization policy.

    5

    Assuming volatile Prevents Race Conditions

    In C, volatile does not provide mutual exclusion or make ordinary shared updates atomic.

    6

    Making Each Field Atomic

    Independent atomic fields do not automatically preserve relationships across several fields.

    7

    Relying on Sleep to Enforce Order

    A delay does not establish a correctness guarantee. Machine load and scheduling can change the relative timing.

    8

    Adding Logging and Assuming the Defect Is Fixed

    Logging can change execution timing and temporarily hide a race without correcting the underlying ordering problem.

    9

    Ignoring Object Lifetime

    Synchronizing field updates is insufficient if an object can be destroyed while another thread still uses it.

    10

    Assuming a Local Mutex Protects Distributed State

    A process-local mutex cannot coordinate independent service instances or machines.

    11

    Testing Only Once

    One successful execution represents only one possible scheduling history.

    12

    Fixing the Race but Creating Excessive Contention

    Establish correctness first, then measure whether the synchronization strategy creates a material throughput or latency problem.

    Recommended Test Cases

    Test Expected Evidence
    Single-thread baseline The underlying algorithm produces the correct result
    Simultaneous start Workers begin from a barrier to increase overlapping execution
    Repeated concurrent execution The invariant remains valid across many runs
    High worker count Shared state remains correct under increased contention
    Check-then-act conflict Only one operation succeeds when one resource remains
    Duplicate task delivery The business effect follows the idempotency policy
    Concurrent deletion and access No object is freed while a valid user still owns it
    Concurrent shutdown No worker accesses destroyed synchronization or state objects
    ThreadSanitizer run Executed data-race paths are reported for investigation
    Contention benchmark The correction meets the required latency and throughput targets

    Practice Exercise

    Analyze and correct a race condition in an inventory reservation system.

    Defective Program Logic

    struct Inventory
    {
        int available;
        int reserved;
    };
    
    static int reserve_item(
        struct Inventory *inventory)
    {
        if (inventory->available <= 0) {
            return 0;
        }
    
        --inventory->available;
        ++inventory->reserved;
    
        return 1;
    }

    Tasks

    1. Identify the shared mutable fields.
    2. Write a harmful interleaving involving two threads.
    3. Define the inventory invariant.
    4. Add a mutex to the inventory structure.
    5. Protect validation and modification as one critical section.
    6. Ensure every return path releases the mutex correctly.
    7. Create several worker threads that attempt reservations.
    8. Verify that successful reservations do not exceed initial availability.
    9. Run the program repeatedly.
    10. Run a supported ThreadSanitizer build.

    Model Correction

    struct Inventory
    {
        int available;
        int reserved;
        pthread_mutex_t mutex;
    };
    
    static int reserve_item(
        struct Inventory *inventory)
    {
        if (inventory == NULL) {
            return 0;
        }
    
        if (pthread_mutex_lock(
                &inventory->mutex) != 0) {
    
            return 0;
        }
    
        int succeeded = 0;
    
        if (inventory->available > 0) {
            --inventory->available;
            ++inventory->reserved;
    
            succeeded = 1;
        }
    
        if (pthread_mutex_unlock(
                &inventory->mutex) != 0) {
    
            return 0;
        }
    
        return succeeded;
    }

    Inventory Invariant

    available >= 0
    reserved >= 0
    
    available + reserved = initial inventory
    when no cancellation, return or restocking operation occurs

    The test should validate the invariant after all worker threads have completed.

    Frequently Asked Questions

    1

    What is a race condition?

    A race condition is a correctness defect in which an uncontrolled timing or ordering of concurrent events can produce an incorrect result.

    2

    What is a data race?

    A data race involves conflicting unsynchronized accesses to the same memory location from different threads, with at least one modifying access.

    3

    Are race conditions and data races identical?

    No. A data race concerns conflicting memory accesses. A race condition is a broader ordering or timing defect that can occur even when individual accesses use synchronization.

    4

    Why is incrementing a shared integer unsafe?

    Increment commonly requires reading the value, calculating a new value and writing it. Concurrent increments can overlap and lose updates.

    5

    Does volatile make a shared variable thread-safe?

    No. In C, volatile does not provide mutual exclusion, atomicity or a general inter-thread synchronization guarantee.

    6

    Can a mutex prevent every race condition?

    No. A mutex prevents selected overlapping accesses only when all participating operations use it correctly. Higher-level ordering and lifetime races can still remain.

    7

    Why are race conditions difficult to reproduce?

    The harmful ordering can depend on scheduling, load, hardware, I/O timing and other conditions that change between executions.

    8

    Can race conditions occur with processes?

    Yes. Processes can race through shared memory, files, databases, devices or other shared resources.

    9

    Can race conditions occur in distributed systems?

    Yes. Concurrent requests, retries, stale reads and independent workers can create order-dependent business defects across service boundaries.

    10

    Does ThreadSanitizer find every race condition?

    No. It can detect many executed data races, but it does not explore every possible path and does not prove higher-level ordering correctness.

    11

    What is the best way to prevent race conditions?

    Minimize shared mutable state, define ownership and invariants, and apply an appropriate synchronization, transaction, uniqueness or idempotency mechanism at the boundary where concurrency occurs.

    12

    What comes after race conditions?

    The next topic is deadlocks, including deadlock conditions, lock ordering, prevention, detection and recovery.

    Key Takeaway

    Race conditions occur when program correctness depends on an uncontrolled ordering of concurrent operations. Common forms include lost updates, check-then-act defects, duplicate processing, unsafe object lifetimes and persistent-data conflicts. Protect complete invariants rather than individual statements, minimize shared mutable state, use database and distributed protections at the correct boundary, and test under varied timing with supported race-detection tools. Correctness must hold for every permitted execution ordering, not only the ordering observed during one successful run.