Chapter 06 · Quorums, Consistency Levels, Read Repair, and Anti-Entropy
Measure Staleness and Correctness Instead of Treating Consistency Labels as Guarantees
Measure version/time staleness and invariant outcomes under controlled propagation delay instead of treating consistency labels as empirical guarantees.
Learning outcomes
AtlasMart’s architecture review contains labels—“eventual,” “quorum,” “strong”—but no evidence about what users actually observe. This lesson builds a small history recorder that measures visibility delay, stale-read frequency, and invariant outcomes under controlled propagation delays. The goal is not to benchmark a database; it is to define what production experiments must record.
Define version staleness, time staleness/visibility delay, stale-read window, invariant violation, and client-visible history.
Design an experiment that records write completion separately from per-replica visibility time.
Compute stale-read observations instead of inferring freshness from a configured consistency label.
Inject controlled delay/failure and distinguish performance symptoms from correctness violations.
Connect measured histories to SLOs, alerting, repair cadence, and the invariant decision notebook from Chapters 01–03.
A consistency label is a contract claim. A test history is evidence about one implementation/topology/workload under stated conditions. Neither replaces the other: documentation says what should happen; measured histories help verify that your deployment and client usage actually satisfy the required property.
1. Measure freshness in dimensions the business understands
Version staleness asks how many logical versions behind a read is. Time staleness asks how long after a write completed before a replica/read path exposed it. A stale-read window is the interval in which a client can still observe an older version under the chosen routing/consistency policy. These measures are more actionable than “eventual.”
For a search index, AtlasMart may accept 5 seconds of catalog lag. For a profile avatar, 2 seconds may be fine if the session gets read-your-writes. For inventory reservation, even a 20 ms stale read can be unacceptable if that read authorizes a second sale of the last item. Correctness depends on the invariant, not the absolute lag alone.
2. A useful history has write completion and visibility timestamps
Record at least: operation/request ID, key, logical version, client invocation/completion time, coordinator, acknowledgement set, each replica’s visible version/time, read target(s), returned version, error/timeout, and any repair/hint activity. For cross-region systems also capture region/zone and topology epoch.
Do not compare clocks from different hosts naïvely. Chapter 04 explained clock uncertainty. Production measurements need synchronized/monotonic tracing strategies or server-side event correlation that make the timing error explicit. A toy simulator can use one deterministic clock because all events share one process.
3. Correctness experiments need invariants, not only latency percentiles
A latency dashboard can look excellent while the system returns stale stock and oversells. Add invariant probes: after each history, test conditions such as inventory never negative, one idempotency key maps to one effect, order state transitions are monotonic, and deleted/consent data does not reappear.
Likewise, a stale read is not automatically a correctness violation. Search/catalog derived views may intentionally lag. The experiment should classify each anomaly as tolerated, SLO-breaching, or invariant-breaking.
4. Deliberately wrong approach — benchmark averages and declare consistency “good”
AtlasMart runs 1,000 reads against the nearest replica, reports a 4 ms average, and calls the database “strong enough.” The test never records which write version each read returned, never injects propagation delay, and never checks inventory invariants. It has measured speed while remaining blind to correctness.
The repair is to preserve versions in the workload, record the full client-visible history, inject controlled delay/replica unavailability, calculate stale windows and tail distributions, and run invariant checks. Real production benchmarking must also avoid coordinated omission and include warmup, realistic key skew, durability settings, and failure-state capacity—topics expanded in Chapter 23.
5. AtlasMart lab — build a staleness history
from dataclasses import dataclass
@dataclass(frozen=True)
class Write:
version: int
complete_ms: int
visible_ms: dict
writes = [
Write(1, 0, {"A":0, "B":20, "C":80}),
Write(2, 100, {"A":100, "B":140, "C":260}),
Write(3, 300, {"A":300, "B":330, "C":500}),
]
reads = [
(30, "A"), (30, "C"), (160, "B"), (160, "C"),
(350, "A"), (350, "C"), (520, "C")
]
def latest_completed(t):
return max((w.version for w in writes if w.complete_ms <= t), default=0)
def visible_version(node, t):
visible = [w.version for w in writes if w.visible_ms[node] <= t]
return max(visible, default=0)
stale = []
print("READ HISTORY")
for t,node in reads:
expected = latest_completed(t)
observed = visible_version(node,t)
is_stale = observed < expected
stale.append(is_stale)
print(f"t={t:>3}ms node={node} observed=v{observed} latest_completed=v{expected} stale={is_stale}")
print("\nVISIBILITY WINDOWS")
windows=[]
for w in writes:
for node,vis in w.visible_ms.items():
windows.append(vis-w.complete_ms)
print(f"v{w.version} node={node} staleness_window_ms={vis-w.complete_ms}")
print("max_visibility_delay_ms", max(windows))
print("stale_read_fraction", sum(stale), "/", len(stale))
print("\nINVARIANT PROBE")
stock_authoritative = 0
stale_stock_replica = 1
print("authoritative stock", stock_authoritative, "stale replica stock", stale_stock_replica)
print("wrong local decision: accept purchase?", stale_stock_replica > 0)
print("correctness test: inventory invariant requires authoritative/serialized reservation, not a marketing label")
print("times are controlled experiment inputs, not production latency measurements")
Verification checklist
- Each read prints both the observed version and the latest completed version at that simulated time.
- The C replica exhibits larger visibility delays than A/B.
- The experiment reports a maximum visibility delay and stale-read fraction.
- The stock probe demonstrates that even a small stale window can break an inventory decision.
- All times are explicitly controlled inputs rather than claims about a real product.
Check your understanding
- What is time staleness?
- Why is stale-read frequency alone insufficient?
- Why record logical versions?
- What does a successful experiment prove?
- How should results feed operations?
Review the answers
1. The elapsed time between a write completing under the stated contract and a particular read path/replica making that version visible.
2. Different stale reads have different business consequences; the experiment must also evaluate invariant violations and tolerated degraded modes.
3. They let the test determine whether a read is behind without trusting wall-clock ordering alone.
4. Only that the tested implementation/topology/client policy under the stated workload/fault schedule produced the recorded history; it does not prove all possible executions.
5. Convert them into freshness/error-budget SLOs, alerts, repair-age targets, routing rules, retry policy, and invariant-focused game-day tests.
6. Chapter synthesis — from algebra to evidence
Chapter 06 began with a mathematical overlap condition and ended with client-visible histories. In between, tunable thresholds moved latency/availability exposure, repair mechanisms restored missed state, and edge cases showed why timeout outcomes can be ambiguous. The recurring engineering rule is to state assumptions, observe actual replica/version state, and test the business invariant.
Keep the AtlasMart decision notebook updated with
N/R/W or analogous settings, strict/sloppy
placement, version/conflict rule, repair mechanism/cadence,
maximum measured staleness, tolerated anomalies, timeout/retry
behavior, durability domains, and rollback plan. Chapter 07 now
changes the question from “how many replicas?” to “which
partition owns the key?”—partitioning and sharding for capacity,
throughput, locality, and failure isolation.
7. Production judgment
Production evidence should combine latency/error/saturation with freshness, divergence, repair age, and invariant checks. Alert on trends, not one magic threshold copied from another system. Protect tracing and repair metadata because keys, tenant IDs, and values can be sensitive. When experimenting with faults, use synthetic/test data and isolated replicas unless a formally reviewed game day defines blast radius and rollback.
Authoritative references
- Dynamo: Amazon’s Highly Available Key-value Store — primary quorum/repair background and motivation for explicit tradeoffs
- Apache Cassandra 5.0 — Repair — current operational repair behavior and validation options
- Apache Cassandra 5.0 — cqlsh consistency levels — implementation example for per-operation consistency selection
- Bailis et al. — Probabilistically Bounded Staleness — research example of reasoning about partial quorums using measured/probabilistic staleness rather than labels