conflict resolution - Distributed Systems Foundations
Conflict Resolution
Understand what happens after concurrent writes are detected, why last-writer-wins silently destroys data, and how to choose a merge strategy that reflects what your domain actually means.
Prerequisites
Recommended Knowledge
- Logical and vector clocks
- Happens-before and concurrency detection
- Eventual and causal consistency
- Replication topologies and partitions
- Version vectors and siblings
- Ordering guarantees and their limits
- Idempotency and duplicate handling
- Basic set and counter semantics
Why Conflicts Exist
A conflict occurs when two writes to the same data are made independently, neither aware of the other. There is no correct ordering between them, because causally they are unordered.
| Origin | Mechanism | Frequency |
|---|---|---|
| Network partition | Both sides accept writes independently | Rare but produces many conflicts at once |
| Multi-leader replication | Each region writes locally | Continuous background occurrence |
| Offline clients | Local edits sync later | Common in mobile applications |
| Concurrent users | Two people edit the same record | Proportional to collaboration |
| Leaderless quorum | Writes accepted without coordination | Depends on quorum configuration |
| Retry after timeout | Original and retry both applied | Whenever retries exist |
Simple Analogy
Two people edit printed copies of the same document in separate rooms. When the copies are compared, neither is wrong. Deciding which changes survive requires knowing what the document is for.
The Last-Writer-Wins Trap
The most common strategy compares timestamps and keeps the later write. It is trivial to implement, requires no metadata, and is the default in many systems. It is also the strategy most likely to destroy data.
| Problem | Consequence |
|---|---|
| Clock skew | The genuinely newer write may lose |
| Silent discard | No party learns their change vanished |
| Timestamp collision | Arbitrary tie-break decides the outcome |
| Whole-record replacement | Unrelated fields are overwritten |
| Unbounded data loss | Loss scales with conflict volume |
The Strategy Landscape
| Strategy | Data Loss | Complexity | Suitable For |
|---|---|---|---|
| Last writer wins | Guaranteed | Minimal | Disposable latest-value data |
| First writer wins | Guaranteed | Minimal | Claim and reservation semantics |
| Field-level merge | Only on same field | Moderate | Records with independent fields |
| Semantic merge | None if modelled well | High | Domain-rich entities |
| Conflict-free types | None by construction | Moderate, constrained | Counters, sets, collaborative text |
| Preserve siblings | None, deferred | Moderate | Documents needing judgment |
| Reject and retry | None, pushed to client | Low | Interactive single-writer flows |
| Manual resolution | None | Operationally heavy | Rare high-value conflicts |
Field-Level Merge
Rather than treating a record as one indivisible value, track versions per field. Concurrent writes to different fields then merge cleanly, and only same-field conflicts require a decision.
{
"entityId": "profile-4821",
"fields": {
"displayName": {
"value": "R. Ansari",
"version": { "replica-east": 4 }
},
"email": {
"value": "user@example.com",
"version": { "replica-west": 3 }
},
"timezone": {
"value": "Asia/Kolkata",
"version": { "replica-east": 2 }
}
}
}
function mergeByField(siblingA, siblingB) {
const merged = {};
const unresolved = [];
const fields = new Set([
...Object.keys(siblingA.fields),
...Object.keys(siblingB.fields)
]);
for (const field of fields) {
const a = siblingA.fields[field];
const b = siblingB.fields[field];
if (!a) { merged[field] = b; continue; }
if (!b) { merged[field] = a; continue; }
const relation = compareVectors(a.version, b.version);
if (relation === "a-after-b") merged[field] = a;
else if (relation === "b-after-a") merged[field] = b;
else if (relation === "equal") merged[field] = a;
else {
unresolved.push({ field, candidates: [a, b] });
merged[field] = a;
}
}
return { merged, unresolved };
}
Semantic Merge
The strongest approach applies domain knowledge. What two concurrent changes should produce together is a business question, and only the application can answer it.
| Scenario | Concurrent Writes | Correct Merge |
|---|---|---|
| Shopping cart | Different items added | Union both additions |
| Item quantity | Set to 3 and set to 5 | Take maximum, or prompt |
| Order status | Cancelled and shipped | Domain precedence, likely shipped |
| Permission set | Grant and revoke | Revoke wins on safety grounds |
| Account balance | Two independent debits | Apply both as deltas |
| Tag list | Different tags added | Union both sets |
| Document title | Two different titles | Preserve both, ask the user |
const MERGE_POLICY = {
"cart.items": { rule: "union", onTie: "merge-quantities" },
"post.tags": { rule: "union", onTie: "deduplicate" },
"order.status": { rule: "precedence", order: ["delivered", "shipped", "packed", "cancelled", "confirmed", "created"] },
"account.entries": { rule: "append-all", onTie: "deduplicate-by-id" },
"permissions.roles": { rule: "safest-wins", bias: "restrictive" },
"document.title": { rule: "preserve-both", onTie: "require-user-choice" }
};
function semanticMerge(path, candidates) {
const policy = MERGE_POLICY[path];
if (!policy) {
return {
resolved: false,
reason: `no merge policy declared for ${path}`,
siblings: candidates
};
}
switch (policy.rule) {
case "union":
return { resolved: true, value: unionValues(candidates) };
case "precedence":
return {
resolved: true,
value: candidates
.slice()
.sort((a, b) =>
policy.order.indexOf(a.value) -
policy.order.indexOf(b.value)
)[0].value
};
case "safest-wins":
return { resolved: true, value: mostRestrictive(candidates) };
case "append-all":
return { resolved: true, value: dedupeById(candidates) };
default:
return { resolved: false, siblings: candidates };
}
}
Conflict-Free Replicated Data Types
CRDTs sidestep resolution entirely by constraining data to shapes whose merge function is mathematically guaranteed to converge, regardless of order or repetition.
| Required Property | Meaning | Why It Matters |
|---|---|---|
| Commutative | Order of merging is irrelevant | Replicas may receive updates in any sequence |
| Associative | Grouping of merges is irrelevant | Partial merges compose correctly |
| Idempotent | Merging twice changes nothing | Duplicate delivery is harmless |
Common CRDT Shapes
| Type | Supports | Limitation |
|---|---|---|
| Grow-only counter | Increment only | Cannot decrement |
| Positive-negative counter | Increment and decrement | Two vectors of state |
| Grow-only set | Add only | Cannot remove |
| Two-phase set | Add and remove | Removed items cannot return |
| Observed-remove set | Add, remove, re-add | Metadata per element |
| Last-writer-wins register | Single value | Still discards a write |
| Sequence type | Ordered collaborative text | Substantial metadata overhead |
class PNCounter {
constructor(nodeId) {
this.nodeId = nodeId;
this.increments = {};
this.decrements = {};
}
increment(amount = 1) {
this.increments[this.nodeId] =
(this.increments[this.nodeId] ?? 0) + amount;
}
decrement(amount = 1) {
this.decrements[this.nodeId] =
(this.decrements[this.nodeId] ?? 0) + amount;
}
value() {
const sum = obj =>
Object.values(obj).reduce((a, b) => a + b, 0);
return sum(this.increments) - sum(this.decrements);
}
merge(other) {
for (const side of ["increments", "decrements"]) {
for (const [node, count] of Object.entries(other[side])) {
this[side][node] = Math.max(
this[side][node] ?? 0,
count
);
}
}
}
}
CRDT Strengths
- Convergence is guaranteed, not hoped for
- No coordination on the write path
- Fully available during partitions
- Duplicate delivery is harmless
- No application merge logic required
CRDT Limitations
- Data must fit a supported shape
- Metadata can exceed the payload
- Cannot enforce global invariants
- Tombstones accumulate over time
- Converged does not mean intended
Deletion and Tombstones
Deletion is the hardest case in any convergent system, because the absence of data is indistinguishable from never having received it.
{
"setId": "cart:user-4821",
"elements": [
{ "value": "sku-501", "addedBy": "replica-east", "tag": "e-14" },
{ "value": "sku-733", "addedBy": "replica-west", "tag": "w-09" }
],
"tombstones": [
{
"value": "sku-910",
"removedTags": ["e-11"],
"removedAt": "2026-09-24T05:42:18Z",
"expiresAt": "2026-10-24T05:42:18Z"
}
]
}
| Concern | Problem | Mitigation |
|---|---|---|
| Unbounded growth | Tombstones accumulate indefinitely | Expire after a retention window |
| Premature expiry | A lagging replica resurrects the item | Retention exceeding maximum lag |
| Privacy obligation | Tombstone retains the deleted value | Store a hash rather than the content |
| Add after remove | Re-adding conflicts with the tombstone | Tag-based observed-remove semantics |
| Storage overhead | Deleted data still occupies space | Compaction during garbage collection |
Preserving Siblings
Where no automatic rule is defensible, the honest approach is to return both versions and let the application or user decide.
async function readWithConflictHandling(key, resolver) {
const result = await datastore.read(key);
if (result.siblings.length === 1) {
return { value: result.siblings[0].value, conflicted: false };
}
metrics.increment("conflict.detected", {
key: key,
siblingCount: result.siblings.length
});
const resolution = resolver(result.siblings);
if (!resolution.resolved) {
return {
conflicted: true,
siblings: result.siblings,
requiresUserChoice: true
};
}
await datastore.write(key, resolution.value, {
context: result.causalContext
});
return { value: resolution.value, conflicted: true, autoResolved: true };
}
Avoiding Conflicts Entirely
The most effective strategy is arranging for conflicts not to arise. Several design choices achieve this.
Single Writer Per Key
Route all writes for a given entity to one replica. Concurrency across entities remains, but no two writers touch the same key.
Append-Only Modelling
Record events rather than mutating state. Two concurrent appends are both retained, and the current state is derived from the sequence.
Partition Ownership
Assign each data range an owning region. Writes elsewhere are forwarded rather than applied locally.
Narrower Granularity
Split large documents into independently written parts, so concurrent edits touch different records entirely.
Commutative Operations
Express changes as deltas rather than absolute values. "Add two" and "add three" compose; "set to five" and "set to six" conflict.
| Conflicting Formulation | Conflict-Free Reformulation |
|---|---|
| Set quantity to 5 | Adjust quantity by +2 |
| Replace the tag list | Add tag, remove tag |
| Overwrite the balance | Append a ledger entry |
| Set status to cancelled | Record a cancellation event |
| Update the whole profile | Update individual fields |
Choosing a Strategy
| Question | If Yes |
|---|---|
| Can conflicts be prevented by routing? | Use single-writer ownership |
| Can changes be expressed as deltas? | Use commutative operations or CRDTs |
| Do writes touch independent fields? | Use field-level merge |
| Does the domain define a correct outcome? | Use semantic merge with a declared policy |
| Would a user know which is right? | Preserve siblings and prompt |
| Is losing a write genuinely harmless? | Last writer wins is acceptable |
| Is the conflict rare and high value? | Queue for manual resolution |
A Worked Example
| Data | Strategy | Reasoning |
|---|---|---|
| Cart contents | Observed-remove set | Additions must never be lost, removals must stick |
| Item quantity | Delta counter | Adjustments compose naturally |
| Account balance | Append-only ledger | Every entry preserved, balance derived |
| Order status | Precedence rule | Domain defines which state dominates |
| User profile | Field-level merge | Fields are independent |
| Document body | Sequence CRDT or siblings | Collaborative editing needs true merge |
| Inventory count | Single-writer ownership | Invariant requires coordination |
| Session last-seen | Last writer wins | Only the newest value has meaning |
Monitoring
Signals Worth Tracking
- Conflict rate per key space
- Sibling count distribution
- Conflicts auto-resolved versus escalated
- Writes discarded by last-writer-wins
- Keys exceeding a sibling threshold
- Tombstone count and storage share
- Items resurrected after deletion
- Time from conflict to resolution
- Manual resolution queue depth
- CRDT metadata size relative to payload
Testing Merge Logic
describe("merge function properties", function () {
it("is commutative", function () {
const a = makeReplicaState("A");
const b = makeReplicaState("B");
expect(merge(a, b)).toEqual(merge(b, a));
});
it("is associative", function () {
const [a, b, c] = ["A", "B", "C"].map(makeReplicaState);
expect(merge(merge(a, b), c))
.toEqual(merge(a, merge(b, c)));
});
it("is idempotent", function () {
const a = makeReplicaState("A");
expect(merge(a, a)).toEqual(a);
});
it("does not resurrect deleted elements", function () {
const a = addThenRemove("sku-910");
const b = stillContaining("sku-910");
expect(merge(a, b).elements)
.not.toContain("sku-910");
});
});
Common Design Mistakes
Weak Design
- Accepting last-writer-wins as the default
- Resolving by timestamp despite clock skew
- Discarding siblings without inspection
- Merging sets without tombstones
- Expiring tombstones faster than replica lag
- Treating whole records as one value
- Assuming CRDTs enforce invariants
- Never measuring conflict rate
- Testing merges only on two inputs
Strong Design
- Declares a merge policy per field
- Detects concurrency before resolving
- Models changes as intent, not final state
- Uses tombstones with safe retention
- Splits records into independent fields
- Coordinates where invariants demand it
- Escalates ambiguous conflicts to users
- Monitors conflicts and discarded writes
- Property-tests merge functions
System Design Interview Discussion
| Question | What Your Answer Should Cover |
|---|---|
| How are conflicts detected? | Version vectors, not timestamps |
| Why avoid last-writer-wins? | Silent loss and clock skew |
| What merge rule applies here? | Domain-specific reasoning per field |
| How is deletion handled? | Tombstones and resurrection risk |
| When would you use a CRDT? | Data fitting a convergent shape |
| What do CRDTs not solve? | Global invariants and semantic correctness |
| Can conflicts be avoided? | Single-writer routing and delta modelling |
| How is this verified? | Property tests for merge algebra |
Design Checklist
Production Checklist
- Detect concurrency with version vectors
- Declare a merge policy for every replicated field
- Avoid last-writer-wins on meaningful data
- Model changes as deltas where possible
- Split records into independently versioned fields
- Use tombstones for deletion
- Set tombstone retention above maximum lag
- Store hashes in tombstones where privacy applies
- Cap sibling count and alert on breaches
- Escalate undecidable conflicts to users
- Coordinate where invariants span replicas
- Property-test commutativity, associativity, idempotence
- Test three-way merges, not only pairs
- Monitor conflict rate and discarded writes
- Test concurrent writes under injected partitions
Knowledge Check
Why is last-writer-wins dangerous?
Both writes represent real intent, so one is discarded silently. Clock skew can also cause the genuinely newer write to lose.
What three properties must a CRDT merge satisfy?
Commutativity, associativity, and idempotence, which together guarantee convergence regardless of delivery order or duplication.
Why are tombstones necessary?
Absence is indistinguishable from never having received the value, so merging would resurrect deleted items without an explicit deletion marker.
What do CRDTs not guarantee?
Correctness against business invariants. A counter converges reliably but may settle on a value that violates a rule such as a non-negative balance.
How does data modelling reduce conflicts?
Expressing changes as intent rather than final state makes concurrent updates compose instead of compete, removing most conflicts before they occur.
Summary
Conflicts arise when writes are made independently with no causal relationship between them. Neither is wrong, so resolution is a question about what two intentions should mean together, which is a domain decision rather than a storage one.
Last-writer-wins is the default in many systems and guarantees data loss. It discards a valid write silently, and clock skew can mean the discarded one was actually newer. It is acceptable only where the value is a disposable snapshot.
Field-level merge reduces conflicts by narrowing granularity. Semantic merge applies domain rules and produces the best outcomes where the correct answer is definable. CRDTs guarantee convergence mathematically for data fitting supported shapes, but cannot enforce invariants spanning replicas.
Deletion requires tombstones, whose retention must exceed maximum replica lag to prevent resurrection. Most importantly, many conflicts are created by the data model itself: expressing changes as deltas and events rather than as absolute final state removes them before any resolution is required.
Key Takeaway
Prevent conflicts through modelling, resolve the rest with domain knowledge. Express changes as intent rather than final state, declare a merge policy for every replicated field, use tombstones with retention above your worst lag, escalate genuinely ambiguous cases to users, and never accept last-writer-wins for data someone cares about.