Table of Contents

    Leader-follower, multi-leader and leaderless replication

    DISTRIBUTED DATA & REPLICATION

    Leader-Follower, Multi-Leader, and Leaderless Replication

    Learn how distributed databases maintain copies of data across multiple nodes, how the three major replication topologies accept and propagate writes, and how replication lag, synchronous and asynchronous replication, failover, split brain, write conflicts, quorums, read repair, anti-entropy, consistency, availability, latency, and operational complexity influence the correct design.

    Introduction

    A database running on one machine creates a major failure boundary. If that machine becomes unavailable, the application can lose access to its data. One machine can also limit read capacity and force geographically distant users to access data through one location.

    Replication maintains copies of data on multiple machines, commonly called replicas or nodes.

    Application
        |
        v
    Database cluster
        |
        +-- Replica A
        +-- Replica B
        +-- Replica C

    Replication can support:

    • Higher availability during node failure
    • Recovery from selected hardware or server failures
    • Read scalability across several replicas
    • Geographically closer data access
    • Continued operation during selected network problems

    Replication also introduces a difficult question: what happens when copies of the same data temporarily disagree?

    Core idea: Replication creates several copies of data. The replication topology determines which nodes can accept writes, how writes reach other replicas, how conflicting versions are detected, and what consistency and availability behaviour clients observe.

    Prerequisites

    # Prerequisite Why It Is Needed
    1 Database transactions A replicated write begins with a database operation whose durability and visibility must be understood.
    2 Consistency models Replicas can temporarily return different versions of the same data.
    3 Network partitions Nodes can remain operational while being unable to communicate with one another.
    4 Logical clocks and versioning Multi-writer systems need a way to identify ordering and concurrent updates.
    5 Quorums Leaderless systems commonly read from and write to several replicas.
    6 Failover and health checks A failed leader can require safe promotion of another replica.
    7 Idempotency Retries after uncertain outcomes must not create unintended duplicate operations.

    What Is Data Replication?

    Data replication is the process of distributing copies of data across multiple machines in one location or across different locations.

    Original write
          |
          v
    Replica A
          |
          +-- Copy to Replica B
          +-- Copy to Replica C

    Replication is different from backup. A replica is normally part of the operational system and can serve reads or participate in failover. A backup preserves restorable data for recovery from corruption, deletion, or another failure scenario.

    Replication Design Flow
    identify write authority → choose propagation timing → define read behaviour → handle failure → resolve disagreement

    Three Main Replication Topologies

    Topology Who Accepts Writes? Central Challenge
    Leader-follower One leader accepts writes Leader availability, failover, and replication lag
    Multi-leader Several leaders accept writes Concurrent-write conflicts and cross-leader coordination
    Leaderless Clients or coordinators write to several replicas Quorum behaviour, stale replicas, version reconciliation, and repair

    Leader-Follower Replication

    Leader-follower replication is also called single-leader, primary-replica, or active-passive replication.

    One node acts as the leader. Clients send writes to the leader, which writes the data locally and propagates changes to follower replicas.

    Writes
      |
      v
    +-----------+
    |  Leader   |
    +-----------+
       |     |
       |     +------------------+
       |                        |
       v                        v
    +------------+         +------------+
    | Follower A |         | Follower B |
    +------------+         +------------+
       |                        |
       v                        v
     Reads                    Reads

    Write Path

    1. Client sends write to leader.
    
    2. Leader validates and applies the write.
    
    3. Leader records the change in its replication log.
    
    4. Change is sent to followers.
    
    5. Followers apply the change locally.
    
    6. Write becomes visible according to
       the configured replication policy.

    Read Path

    Reads can be routed to:

    • The leader when the freshest committed data is required
    • A follower when some replication delay is acceptable
    • A selected follower that satisfies a freshness requirement

    Leader-Follower Advantages

    • One node provides a clear write authority
    • Writes have one primary ordering point
    • Followers can distribute read traffic
    • Conflict handling is simpler because followers do not independently accept ordinary writes
    • Backup, reporting, and analytical reads can use selected followers
    • A follower can be promoted after leader failure

    Leader-Follower Limitations

    • The leader can become a write bottleneck
    • Leader failure requires detection and failover
    • Followers can return stale data
    • Asynchronous replication can lose recent acknowledged writes after leader failure
    • Failover can briefly interrupt write availability
    • An unsafe promotion can create split brain or divergent histories

    Replication Log

    A leader commonly records data changes in an ordered log and sends those changes to followers.

    Leader replication log:
    
    Position 1001: Update Course 42
    Position 1002: Create Enrollment 981
    Position 1003: Update Progress 72%
    Position 1004: Revoke Session 123

    Each follower tracks how far it has applied the leader's change stream.

    Synchronous vs Asynchronous Replication

    Synchronous Replication

    Client
      |
      v
    Leader applies write
      |
      v
    Follower receives or applies write
      |
      v
    Follower acknowledges
      |
      v
    Leader confirms success to client

    Synchronous replication can strengthen durability and freshness guarantees, but follower or network delay contributes to write latency. Depending on the required acknowledgment policy, an unavailable follower can also affect write availability.

    Asynchronous Replication

    Client
      |
      v
    Leader applies write
      |
      v
    Leader confirms success
      |
      v
    Followers receive change later

    Asynchronous replication reduces the follower acknowledgment from the client write path, but creates replication lag. If the leader fails before a recent change reaches another replica, that acknowledged change can be unavailable after failover.

    Timing Comparison

    Area Synchronous Asynchronous
    Write acknowledgment Waits for the configured replica acknowledgment Can complete before followers catch up
    Write latency Includes replica and network coordination Usually shorter from the client's perspective
    Follower failure Can affect writes depending on the policy Leader can continue while followers recover
    Recent-write loss after leader failure Reduced according to the acknowledgment guarantee Possible when acknowledged writes have not replicated
    Follower freshness Stronger under the configured acknowledgment scope Followers can lag

    Replication Lag

    Replication lag is the difference between the leader's latest committed state and the state applied by a follower.

    Leader:
    
    Latest log position = 10,500
    
    
    Follower A:
    
    Applied position = 10,500
    
    
    Follower B:
    
    Applied position = 10,420
    
    
    Follower B lag:
    
    80 log positions

    Lag can be caused by network delay, follower overload, storage latency, expensive change application, or temporary disconnection.

    Read-After-Write Problem

    1. User updates profile on leader.
    
    2. Leader confirms the write.
    
    3. User refreshes immediately.
    
    4. Read is routed to a lagging follower.
    
    5. Follower returns the old profile.

    Possible design directions include:

    • Route the user's read to the leader after that user's write
    • Track a required replication position and choose a sufficiently current replica
    • Wait until the chosen replica has applied the required version
    • Use a datastore-provided read consistency option where available

    Monotonic-read Problem

    Read 1 from Follower A:
    
    Course progress = 72%
    
    
    Read 2 from lagging Follower B:
    
    Course progress = 65%
    
    
    User observes data moving backward.

    A system requiring monotonic reads must prevent a client from moving from a newer observed version to an older one.

    Consistent-prefix Problem

    Write 1:
    
    Create course discussion question
    
    
    Write 2:
    
    Create reply to that question
    
    
    Lagging or differently ordered view:
    
    Reply appears before the question

    Consistent-prefix guarantees preserve a causally sensible ordering for related operations.

    Leader Failure and Failover

    Leader stops responding
          |
          v
    Failure detector evaluates health
          |
          v
    Select eligible follower
          |
          v
    Fence old leader
          |
          v
    Promote follower
          |
          v
    Redirect writes
          |
          v
    Rebuild replication topology

    Failover must avoid allowing the old leader and new leader to accept independent writes simultaneously.

    Fencing

    Fencing prevents an old or isolated leader from continuing to make writes after another node has been promoted.

    Old leader
        |
        | loses coordination access
        v
    New leader elected
    
    
    Fencing mechanism:
    
    Old leader's write authority
    is invalidated before the new
    leader serves writes.

    Failover rule: Detecting that a leader appears unavailable is not enough. The system must also ensure that the previous leader cannot continue accepting valid writes after promotion.

    Split Brain

    Split brain occurs when more than one node believes it has exclusive write authority for the same data.

    Network partition
          |
          +----------------------+
          |                      |
          v                      v
    
    Old Leader             Promoted Leader
    accepts writes         accepts writes
    
          |                      |
          v                      v
    
    History A              History B

    When communication returns, the two histories can contain conflicting writes. Safe leader election, quorum participation, leases, epochs, and fencing are common mechanisms used to prevent or control this condition.

    Multi-Leader Replication

    Multi-leader replication allows more than one leader to accept writes. Each leader replicates its accepted changes to the other leaders and their followers.

    Data Center A                    Data Center B
    
    +-----------+                    +-----------+
    | Leader A  | <--- replication ---> | Leader B  |
    +-----------+                    +-----------+
       |    |                           |    |
       v    v                           v    v
    Follower A1                     Follower B1
    Follower A2                     Follower B2

    Multi-leader replication is also called multi-primary or active-active replication in some systems.

    Typical Motivation

    • Accept writes in several geographic regions
    • Reduce write latency for geographically distributed users
    • Continue selected local writes during cross-region disconnection
    • Support disconnected or intermittently connected operation

    Multi-Leader Advantages

    • Several locations can accept writes
    • Users can write to a geographically closer leader
    • A remote-region communication failure does not always stop local writes
    • Write traffic can be distributed among leaders when data ownership permits it

    Multi-Leader Limitations

    • Concurrent writes can conflict
    • Conflict resolution becomes part of application correctness
    • Cross-region replication delay can expose stale data
    • Uniqueness constraints are difficult across disconnected leaders
    • Automatic conflict resolution can discard intended changes
    • Schema changes require compatibility among independently writing regions

    Concurrent-write Conflict

    Initial course title:
    
    "System Design Basics"
    
    
    Leader A updates title to:
    
    "Practical System Design"
    
    
    Leader B concurrently updates title to:
    
    "System Design Fundamentals"
    
    
    Replication resumes:
    
    Both leaders receive
    different valid updates.

    The system must detect or resolve the conflicting updates instead of pretending that both can remain the single current title.

    Conflict-resolution Strategies

    Strategy General Behaviour Risk
    Last-write-wins Selects one write according to a timestamp or ordering rule A valid concurrent update can be silently discarded
    Deterministic leader priority Chooses the update from the preferred leader Other accepted writes can be lost
    Field-level merge Combines changes to independent fields Unsafe when business fields are related
    Application-specific merge Uses domain rules to reconcile the versions More complex and must handle every valid conflict
    Preserve siblings Stores concurrent versions for later resolution Clients or operators must resolve the conflict
    Conflict-free data type Uses operations designed to merge deterministically Only suitable for specific data models and semantics

    Conflict rule: Conflict resolution is a business-data decision, not only a database setting. A technically deterministic winner can still produce an incorrect business result.

    Uniqueness across Multiple Leaders

    Leader A creates:
    
    username = learner42
    
    
    Leader B concurrently creates:
    
    username = learner42
    
    
    Each local leader sees
    the value as available.

    Possible design directions include:

    • Assigning non-overlapping identifier ranges or namespaces
    • Routing one entity or partition to one write leader
    • Using globally unique generated identifiers
    • Coordinating uniqueness through a strongly consistent service
    • Detecting and resolving duplicates after replication

    The correct direction depends on whether uniqueness is a convenience or a strict business invariant.

    Avoid Multi-Leader for Unsafe Invariants

    Multi-leader replication is risky when disconnected leaders must enforce one global invariant.

    Examples include:

    • One remaining inventory unit
    • A unique financial reference
    • A balance that must never become negative
    • One active ownership assignment
    • A globally unique username without coordination

    Partition ownership or a single authoritative write path can be safer for these operations.

    Leaderless Replication

    In leaderless replication, no single replica acts as the permanent write leader for the replicated item. A client or coordinating node sends reads and writes to several replicas.

                     Client or Coordinator
                         /     |     \
                        /      |      \
                       v       v       v
    
                  Replica A Replica B Replica C

    The coordinating component collects responses and determines whether the required read or write threshold has been achieved.

    Leaderless Write Path

    Client sends write
    to Replica A, B, and C
          |
          +-- Replica A acknowledges
          +-- Replica B acknowledges
          +-- Replica C unavailable
          |
          v
    Write succeeds if the configured
    write threshold is satisfied.

    Leaderless Read Path

    Client sends read
    to Replica A, B, and C
          |
          +-- Replica A returns Version 5
          +-- Replica B returns Version 5
          +-- Replica C returns Version 4
          |
          v
    Coordinator selects or merges
    the correct version
          |
          v
    Outdated replica can be repaired.

    Quorum Parameters

    Let:

    • \(N\) be the number of replicas
    • \(W\) be the number of write acknowledgments required
    • \(R\) be the number of replicas read

    A commonly discussed quorum intersection condition is:

    \[ W + R > N \]

    The intention is for the read and write sets to overlap, increasing the chance that a read observes a replica containing the latest successfully acknowledged write.

    Example

    Replication factor:
    
    N = 3
    
    
    Required write acknowledgments:
    
    W = 2
    
    
    Required read responses:
    
    R = 2
    
    
    Intersection:
    
    W + R = 4
    
    4 > 3

    Quorum arithmetic alone does not provide every consistency guarantee. Concurrent writes, sloppy quorums, delayed writes, clock assumptions, failed repairs, and implementation details can still affect observed behaviour.

    Quorum Trade-offs

    Configuration Direction Potential Benefit Potential Cost
    Larger \(W\) More replicas confirm the write Higher write latency and lower write availability during failures
    Smaller \(W\) Faster and more available writes More replicas can remain stale after acknowledgment
    Larger \(R\) Reads consult more replicas Higher read latency and reduced read availability
    Smaller \(R\) Faster and more available reads Greater risk of reading an old replica

    Concurrent Versions

    Two leaderless replicas can accept concurrent writes during a network problem.

    Initial value:
    
    Version 4
    
    
    Client A writes Value X
    to Replicas A and B
    
    
    Client B concurrently writes Value Y
    to Replicas B and C
    
    
    Result:
    
    More than one version can exist.

    Version vectors, logical clocks, causal metadata, or datastore-specific versioning can help identify whether one version supersedes another or the versions are concurrent.

    Read Repair

    Read repair updates stale replicas while serving a read.

    Read three replicas
          |
          v
    Detect stale Replica C
          |
          v
    Return selected current value
          |
          v
    Send repair to Replica C

    Read repair is most effective for frequently read data. Rarely read data can remain inconsistent unless another repair process exists.

    Anti-Entropy Repair

    Anti-entropy compares replicas in the background and repairs differences.

    Replica comparison
          |
          v
    Detect missing or divergent ranges
          |
          v
    Transfer required versions
          |
          v
    Replicas converge

    Background repair is important because one replica can miss writes while unavailable and not receive reads that would otherwise trigger read repair.

    Hinted Handoff

    When an intended replica is unavailable, some leaderless systems can temporarily store a write hint elsewhere and deliver it when the target returns.

    Replica C unavailable
          |
          v
    Temporary node stores
    a hint for Replica C
          |
          v
    Replica C recovers
          |
          v
    Hinted write is delivered

    This can improve write availability, but the unavailable replica remains stale until handoff or repair completes.

    Leaderless Advantages

    • No permanent leader failover is required for ordinary writes
    • Several replicas can accept a write
    • Read and write consistency can be tuned through supported policies
    • Selected node failures can be tolerated while thresholds remain achievable
    • Requests can be coordinated near the client where the platform supports it

    Leaderless Limitations

    • Replicas can return different versions
    • Concurrent writes require detection and reconciliation
    • Repair processes are essential
    • Quorum settings can be misunderstood
    • Stale reads remain possible under selected configurations
    • Application semantics can become more complex
    • Operations spanning several keys can require additional coordination

    Complete Comparison

    Area Leader-Follower Multi-Leader Leaderless
    Write authority One leader Several leaders Several replicas through a coordinator
    Ordinary write conflicts Lower conflict surface Concurrent leader writes can conflict Concurrent replica versions can conflict
    Write ordering Centralized at the leader Separate order at each leader before reconciliation Version and causality metadata are needed
    Read scaling Followers can serve reads Leaders and followers can serve reads Reads can query several replicas
    Leader failover Required Local leaders remain, but leader and region failure still require handling No permanent leader promotion for ordinary writes
    Replication lag Followers can lag Leaders and followers can lag across locations Replicas can contain older versions
    Conflict resolution Usually simpler Central design requirement Central design requirement
    Operational complexity Moderate High High
    Common fit Transactional systems with one write authority Geographically distributed or disconnected writers Availability-focused key-value and partitioned workloads

    Choosing a Topology

    Choose Leader-Follower When

    • One ordered write path is acceptable
    • Transactional consistency is more important than accepting writes everywhere
    • Read scaling is required
    • Conflict avoidance is preferred
    • Leader failover can meet the recovery objective

    Consider Multi-Leader When

    • Several locations must accept writes independently
    • Cross-region write latency through one leader is unacceptable
    • Disconnected operation is required
    • Conflicting writes can be detected and resolved safely
    • Global invariants can be partitioned or coordinated separately

    Consider Leaderless When

    • High write availability is a major requirement
    • The data model supports version reconciliation
    • Quorum-based reads and writes fit the workload
    • Read repair and background repair can be operated reliably
    • Selected stale or concurrent versions can be handled explicitly

    Selection rule: Do not select replication topology only from database popularity. Begin with write ownership, conflict semantics, read freshness, partition behaviour, failover objectives, geographic latency, and the invariants the application must preserve.

    Conceptual Replication Policy

    replication:
      topology: leader-follower
      replicationFactor: approved-replica-count
    
      writes:
        authority: elected-leader
        acknowledgment: approved-durability-policy
    
      reads:
        stronglyFresh:
          target: leader
    
        staleTolerant:
          target: eligible-follower
          maximumLag: approved-lag-boundary
    
      failover:
        automatic: true
        fencing: required
        minimumElectionQuorum: approved-policy
    
      monitoring:
        replicationLag: enabled
        replicaHealth: enabled
        failoverEvents: enabled
        divergentHistoryDetection: enabled

    This configuration is conceptual. Product-specific settings, guarantees, and terminology must be verified in the selected database documentation.

    Learning-Platform Examples

    Data Possible Direction Reasoning
    Enrollment and payment state Leader-based transactional authority Requires controlled ordering, uniqueness, and transactional invariants
    Public course catalog reads Leader with read replicas or projections High read scale can tolerate a documented freshness delay
    Offline learner notes Multi-writer synchronization with explicit conflict handling Separate devices can edit while disconnected
    High-volume activity events Partitioned or leaderless-style storage where supported Workload can prioritize write availability and later convergence
    Certificate issuance Single authoritative transactional path Duplicate issuance and conflicting status require strict control

    These are architecture directions, not universal product prescriptions. Actual selection requires workload measurement and database-specific guarantees.

    Data Loss and Durability Questions

    For every topology, document:

    • When the client receives write success
    • How many replicas have persisted the write at that moment
    • Whether the write can be lost after acknowledged success
    • Whether acknowledged writes can be rolled back after failover
    • How divergent versions are retained or discarded
    • How node replacement restores the replication factor

    Replication Is Not Backup

    Accidental deletion
          |
          v
    Deletion replicated rapidly
          |
          v
    Every replica removes the record

    Replication protects against selected node failures. It can also replicate accidental deletion, corruption, or application mistakes. Maintain separate backup, retention, point-in-time recovery, and restore testing according to recovery requirements.

    Schema Evolution

    Replicas can run different software or schema versions during rolling deployment.

    Plan for:

    • Backward-compatible replication records
    • Readers that tolerate newly added fields
    • Writers that do not break older replicas
    • Controlled removal of obsolete fields
    • Failover to a replica running a compatible version
    • Conflict resolution across mixed versions

    Observability

    Useful replication metrics include:

    • Replica health
    • Replication lag by replica
    • Log position or version difference
    • Replication throughput
    • Pending replication bytes or operations
    • Leader-election count
    • Failover duration
    • Write acknowledgment latency
    • Follower read latency
    • Conflict count
    • Conflict-resolution outcome
    • Quorum success and failure rate
    • Read-repair count
    • Anti-entropy repair backlog
    • Unavailable or recovering replicas

    Alert Conditions

    Alert when:

    • Replication lag exceeds the approved freshness boundary
    • The replication factor falls below the safe requirement
    • No write leader is available
    • More than one leader claims exclusive authority unexpectedly
    • Failover exceeds its recovery objective
    • A follower cannot apply replication records
    • Conflict rate increases unexpectedly
    • Read repair or anti-entropy backlog grows
    • Quorum thresholds cannot be achieved
    • Replica storage approaches capacity
    • Acknowledged writes are missing after failover

    Troubleshooting Workflow

    1. Identify the replication topology.
    2. Identify which nodes are allowed to accept writes.
    3. Locate the authoritative or newest known version.
    4. Check leader, follower, or replica health.
    5. Check replication position and lag.
    6. Check recent network partitions and failovers.
    7. Check whether synchronous acknowledgments were achieved.
    8. Check for split brain or concurrent versions.
    9. Inspect quorum response counts.
    10. Inspect read-repair and anti-entropy status.
    11. Check schema and software compatibility.
    12. Compare client-observed version metadata.
    13. Rebuild or repair replicas through the approved procedure.

    Common Replication Mistakes

    1

    Assuming Every Replica Is Always Current

    Asynchronous propagation allows followers and replicas to return older versions.

    2

    Reading from a Follower after Writing to the Leader

    A user can observe an older value and believe the write failed.

    3

    Failing Over without Fencing

    The old and new leaders can accept divergent writes.

    4

    Using Multi-Leader without Conflict Semantics

    Concurrent accepted writes reach other leaders without a safe resolution rule.

    5

    Using Last-Write-Wins for Critical Data

    One valid business update can be discarded because of an ordering rule.

    6

    Enforcing Global Uniqueness Independently at Each Leader

    Disconnected leaders can accept duplicate values.

    7

    Assuming Quorum Arithmetic Solves Every Consistency Problem

    Concurrent versions, clock assumptions, delayed writes, and repair failures still require design.

    8

    Ignoring Background Repair

    Rarely read replicas can remain stale when no repair process reconciles them.

    9

    Treating Replication as Backup

    Accidental deletion or corruption can be copied to every replica.

    10

    Promoting the Most Available Follower Instead of the Safest Follower

    A lagging follower can omit recent acknowledged changes after promotion.

    11

    Ignoring Replica Capacity

    Followers can fall behind because of CPU, storage, network, or query load.

    12

    Skipping Partition and Failover Tests

    The most important conflict, durability, and availability behaviours remain unverified.

    Recommended Test Cases

    Test Expected Evidence
    Normal replication Every healthy replica converges to the expected version
    Follower lag Freshness-sensitive reads avoid an excessively stale follower
    Read after write The writer observes the committed update according to the contract
    Monotonic reads A client does not move from a newer observed version to an older one
    Leader failure An eligible follower is promoted and the old leader is fenced
    Leader recovery The old leader returns as a follower without accepting divergent writes
    Network partition Write availability and consistency follow the documented policy
    Concurrent multi-leader write The conflict is detected and resolved according to domain rules
    Duplicate unique value The design prevents or safely resolves a cross-leader duplicate
    Leaderless quorum write The write succeeds only after the required acknowledgments
    Leaderless stale replica The read selects the appropriate version and repairs the stale copy
    Replica outage and recovery Hints or repair restore the recovered replica
    Schema rollout Mixed-version replicas exchange compatible changes
    Backup restoration Data can be recovered independently of live replication

    Replication Best Practices

    Recommended Practices

    • Choose the topology from write ownership and consistency requirements.
    • Document exactly when a write is acknowledged.
    • Define how much replication lag each read path can tolerate.
    • Route freshness-sensitive reads appropriately.
    • Monitor lag for every replica.
    • Use safe leader election and fencing.
    • Test leader failure and old-leader recovery.
    • Define multi-leader conflict semantics before deployment.
    • Avoid last-write-wins for critical data unless loss is acceptable.
    • Partition strict invariants under one authoritative writer where practical.
    • Choose quorum parameters from measured requirements.
    • Retain version metadata needed to detect concurrent writes.
    • Use read repair and background anti-entropy where required.
    • Rebuild failed replicas through a controlled procedure.
    • Keep replication traffic isolated from uncontrolled query load.
    • Use compatible schema and software rollouts.
    • Maintain backups independently of replication.
    • Test partitions, lag, failover, conflicts, and recovery.
    • Record failover, conflict, and repair events.
    • Verify product-specific guarantees rather than assuming them.

    Practice Exercise

    Select replication strategies for your online learning platform.

    Requirements

    1. Replicate course, enrollment, progress, and activity data.
    2. Identify which data requires one authoritative writer.
    3. Define which reads can tolerate replication lag.
    4. Define read-after-write behaviour for profile and progress updates.
    5. Choose synchronous or asynchronous acknowledgment requirements.
    6. Design leader detection, promotion, and fencing.
    7. Define behaviour during a network partition.
    8. Model an offline notes feature with concurrent edits.
    9. Define conflict detection and user-visible resolution.
    10. Model activity events using a quorum-based store.
    11. Select \(N\), \(W\), and \(R\) from the intended guarantees.
    12. Design read repair and background repair.
    13. Add replication-lag and conflict monitoring.
    14. Create a backup and restore plan independent of replication.
    15. Test leader failure and replica recovery.

    Frequently Asked Questions

    1

    What is replication?

    Replication maintains copies of data on multiple machines to support availability, read scaling, fault tolerance, or geographic proximity.

    2

    What is leader-follower replication?

    One leader accepts writes and propagates its changes to followers, which can serve selected reads and participate in failover.

    3

    What is multi-leader replication?

    Several leaders accept writes and exchange their changes, requiring explicit handling of concurrent conflicting updates.

    4

    What is leaderless replication?

    Clients or coordinators send reads and writes to several replicas without using one permanent write leader.

    5

    What is replication lag?

    Replication lag is the difference between the latest accepted data and the state applied by another replica.

    6

    What is split brain?

    Split brain occurs when multiple nodes incorrectly believe they have exclusive write authority and accept divergent writes.

    7

    What is a quorum?

    A quorum is a required number of replica responses used to determine whether a read or write operation succeeds.

    8

    What does \(W + R > N\) mean?

    It defines an intended overlap between write and read replica sets, but it does not by itself solve every stale-read or concurrent-write problem.

    9

    What is read repair?

    Read repair detects an outdated replica during a read and updates that replica with the selected current version.

    10

    What is anti-entropy?

    Anti-entropy is a background process that compares replicas and repairs missing or divergent data.

    11

    Does replication replace backups?

    No. Replication can copy accidental deletion or corruption to every replica. Independent backups and tested restoration remain necessary.

    12

    Which replication model should I choose?

    Choose based on write authority, read freshness, conflict semantics, geographic latency, failure behaviour, availability requirements, and the business invariants that must remain correct.

    Key Takeaway

    Leader-follower replication sends all writes through one leader and propagates changes to followers. It simplifies write ordering and conflict avoidance but introduces leader failover and follower lag. Multi-leader replication allows several locations to accept writes, improving local write availability and latency while making concurrent conflicts, uniqueness, and global invariants substantially harder. Leaderless replication sends reads and writes to several replicas and uses quorum, versioning, read repair, and anti-entropy to converge data without a permanent leader. No topology is universally best. Select one from the workload's write authority, consistency, availability, latency, conflict, and failure-management requirements. Monitor replication lag and repair, fence old leaders during failover, test network partitions and concurrent writes, and maintain independent backups because replication alone does not protect against every form of data loss.