Construct lost-update, concurrent-write, duplicate-delivery, and ambiguous-retry histories, then identify which conflicts are detectable only when AtlasMart retains versions or operation identity.
Lost Updates, Concurrent Writes, Duplicate Events, and Retry-Induced Conflicts
Conflict resolution begins with evidence. This lesson separates four mechanisms that are often mislabeled as one consistency problem, then makes silent overwrite and retry ambiguity observable.
Distinguish lost updates, truly concurrent writes, duplicate delivery, and ambiguous retry as different conflict mechanisms.
Use versions or causal metadata to identify which conflicting histories are detectable and which become invisible after overwrite.
Explain why a timeout does not reveal whether a remote side effect committed.
Connect conflict evidence to safer conditional writes, deduplication, reconciliation, and later resolution strategies.
1. One word, four different failure mechanisms
AtlasMart receives complaints that “concurrency duplicated a charge and also lost a profile edit.” Those symptoms sound related, but the mechanism matters. A lost update occurs when one writer overwrites another writer's change after both operated from an older state. Concurrent writes are updates for which neither is causally known to follow the other. A duplicate delivery is the same logical event or command processed more than once. An ambiguous timeout occurs when the caller does not know whether the server committed an operation before the response was lost.
Calling all four “eventual consistency” hides the evidence needed for repair. The first modeling job is therefore to preserve enough identity and version information to tell histories apart.
2. Lost update: the overwrite can erase the evidence
Suppose customer profile version 7 contains email
a@atlas.test and phone 111. Client A
reads version 7 and changes the email. Client B also reads
version 7 and changes the phone. If each client sends the entire
record and the store accepts both writes unconditionally, B's
later arrival can restore the old email. The final record looks
internally valid, but A's accepted intent has disappeared.
| History step | State/version observed | Action | Evidence |
|---|---|---|---|
| t0 | v7: email=a, phone=111 | A and B both read | Two writers share the same base |
| t1 | client-local v7 | A writes email=new | Without a conditional, store may accept |
| t2 | client-local v7 | B writes phone=222 as whole record | May overwrite A silently |
| t3 | final v8-ish record | email is old, phone is new | No conflict is reconstructable if base/version metadata was discarded |
A compare-and-set (CAS), entity tag (ETag), or explicit expected-version condition converts the silent overwrite into a visible rejection: B can no longer claim that its version-7 mutation applies to version 8. Detection is not resolution; the application must still re-read and decide how to merge or ask a user.
3. Causality metadata can distinguish descendant from concurrent
A scalar version generated by one authority is enough to reject stale writes at that authority. In a disconnected multi-writer system, a vector clock/version vector can record progress per writer. If version X dominates Y component-wise, X causally descends from Y. If neither dominates, the versions are concurrent and discarding either requires a resolution policy.
Dynamo famously used versioning and application-assisted reconciliation so a read could return concurrent siblings instead of silently inventing a winner. The important idea is not that every modern system should expose Dynamo's exact vector-clock API; it is that causal metadata preserves information that a bare value does not.
4. Duplicate delivery and ambiguous retries
At-least-once messaging and client retries are normal
reliability mechanisms. They become correctness bugs when a
logical operation is not identifiable. If loyalty event
evt-9 is delivered twice and the consumer
increments points twice, the final state contains no clue that
both increments came from one event. Likewise, if payment
request pay-44 commits but the response is lost, a
retry without an idempotency identity can capture twice.
A transport timeout proves only that the caller failed to receive a timely success response. It does not prove that the server rolled back, failed to persist, or never invoked a downstream side effect.
5. AtlasMart lab: make each conflict observable
Python 3.13+ standard library only. The generated lab was verified with Python 3.13.5. No database server, Docker, cloud account, paid feature, credential, firewall change, clock manipulation, process killing, or destructive failure injection is required. All conflicts, partitions, retries, and failures are deterministic in-memory simulations.
The lab first performs a whole-record lost update, then repeats the history with a version check. It also compares two vector versions, processes a duplicate loyalty event, and simulates an ambiguous payment response followed by an unsafe retry.
from copy import deepcopy
print("1) LOST UPDATE WITHOUT VERSION CHECK")
record = {"email": "a@atlas.test", "phone": "111", "version": 7}
a = deepcopy(record)
b = deepcopy(record)
a["email"] = "new@atlas.test"
b["phone"] = "222"
# Naive whole-record writes: B writes after A and silently restores A's old email.
record = {**a, "version": 8}
record = {**b, "version": 8}
print("final naive record:", record)
print("email update lost:", record["email"] != "new@atlas.test")
print("\n2) VERSION CHECK MAKES THE RACE DETECTABLE")
record = {"email": "a@atlas.test", "phone": "111", "version": 7}
def cas_write(expected, patch):
global record
if record["version"] != expected:
return False
record = {**record, **patch, "version": expected + 1}
return True
print("A succeeds:", cas_write(7, {"email": "new@atlas.test"}), record)
print("B stale write rejected:", not cas_write(7, {"phone": "222"}), record)
print("\n3) CAUSAL METADATA CAN MARK TRUE CONCURRENCY")
def dominates(x, y):
keys = set(x) | set(y)
ge = all(x.get(k, 0) >= y.get(k, 0) for k in keys)
gt = any(x.get(k, 0) > y.get(k, 0) for k in keys)
return ge and gt
west = {"W": 3, "E": 1}
east = {"W": 2, "E": 2}
concurrent = not dominates(west, east) and not dominates(east, west)
print("west clock:", west, "east clock:", east, "concurrent:", concurrent)
print("\n4) DUPLICATE EVENT DELIVERY")
points = 100
for event_id in ["evt-9", "evt-9"]:
points += 10
print("naive points after duplicate:", points)
points = 100
seen = set()
for event_id in ["evt-9", "evt-9"]:
if event_id not in seen:
points += 10
seen.add(event_id)
print("deduplicated points:", points, "seen:", sorted(seen))
print("\n5) AMBIGUOUS TIMEOUT + RETRY")
charges = []
def charge(request_id):
charges.append(request_id)
return "captured"
charge("pay-44")
print("server committed pay-44, response is lost")
charge("pay-44")
print("naive retry side effects:", charges, "count:", len(charges))
print("lesson: without a version/event/request identity, some conflicts are invisible after overwrite")
The naive whole-record write loses the email edit. CAS turns the second stale writer into a detectable rejection. The vector clocks are concurrent. The duplicate event produces 120 points without deduplication but 110 with an event-ID set. The payment retry creates two side effects. These observations prove the mechanics of the deterministic histories, not any specific database's API semantics.
6. Deliberately wrong approach: “just timestamp every write”
A timestamp can impose a deterministic winner, but that does not recover causality, business intent, or duplicate identity. Two disconnected clients can have skewed clocks; two operations can be semantically mergeable; and two retries can carry the same timestamp or different timestamps while still representing one logical request. Choosing one by clock order turns a detectable conflict into silent information loss.
The safer diagnostic sequence is: identify the logical command/event, capture the base version or causal context, record durable operation IDs for retry-prone side effects, and preserve concurrent versions long enough for the chosen resolver to act.
7. Production judgment and bridge
Use optimistic version checks when conflicts are uncommon and a single authoritative update path can reject stale bases cheaply. Preserve causal metadata when disconnected/multi-writer operation is a requirement and sibling detection matters. Deduplicate retryable business operations where duplicate effects violate invariants. Observe conflict rate, conditional-write failures, duplicate event IDs, retry counts, ambiguous timeouts, reconciliation backlog, and user-visible conflict frequency. Security also matters: request IDs are not authorization tokens, and tenant scoping must prevent one tenant's dedup key or version from colliding with another's.
The next lesson studies Last-Write-Wins (LWW): a tempting policy that converges easily by selecting one winner, but can silently discard exactly the information this lesson worked to preserve.
Check your understanding
- Why can a lost update be impossible to diagnose after the fact?
- What does a conditional expected-version write add?
- What does a timeout prove about a remote side effect?
- Why are duplicate delivery and concurrent writes different?
- What can vector/version metadata reveal that a bare timestamp winner cannot?
Review the answers
1. If the store accepted unconditional whole-record overwrites and retained no base/version history, the overwritten intent may leave no evidence in the final value.
2. It converts a stale overwrite into an explicit conflict/rejection, preserving the fact that the writer acted on an old base.
3. Only that the caller did not receive a timely definitive response; the side effect may have committed, failed, or still be in progress.
4. A duplicate is the same logical operation repeated; concurrent writes are distinct operations without a causal order. They need different metadata and repair logic.
5. Whether one version descends from another or whether the writes are concurrent, preserving information for reconciliation.
References
Foundational claims use primary research/specifications where practical. Product documentation is used only as a current implementation example and is not required for the mandatory labs.
- DeCandia et al. — Dynamo — Primary system paper for versioning, concurrent siblings, and application-assisted reconciliation.
- Saito & Shapiro — Optimistic Replication — Survey of optimistic replicated systems, conflicts, and reconciliation.
- RFC 9110 — HTTP Semantics — Standards-track definition of idempotent HTTP method semantics and retry motivation.
- Apache Cassandra 5.0.9 download — Current implementation/version context only; mandatory labs remain vendor-neutral.