Chapter 06 · Quorums, Consistency Levels, Read Repair, and Anti-Entropy

Tunable Consistency: Per-Operation Choices and Their Failure/Latency Consequences

Compare per-operation response thresholds by latency, availability, freshness, and failure behavior without pretending named consistency levels have universal semantics.

Beginner → Advanced85–105 minutestunable-consistency simulationVendor-neutral · Python 3.13.5 simulatorFree/local · no database or cloud requiredLast reviewed: August 2026

Learning outcomes

AtlasMart does not have one universal read/write path. Product browse traffic may prefer low latency; checkout inventory needs a much stronger freshness/correctness contract. Tunable consistency means the client or operation can request different acknowledgement/response thresholds. That flexibility is useful only when the team understands the failure and latency consequences of each choice.

01

Distinguish consistency-level thresholds from universal product semantics and define local versus cross-region scope explicitly.

02

Explain why waiting for more replicas usually increases evidence/freshness while consuming availability and tail-latency budget.

03

Compare ONE-like, QUORUM-like, ALL-like, LOCAL_ONE-like, and LOCAL_QUORUM-like behavior without treating those names as portable guarantees.

04

Show why a successful low-threshold read can legally return stale state under an eventually convergent design.

05

Choose per-operation levels from business invariants and SLOs rather than one cluster-wide superstition.

Vocabulary discipline

Cassandra uses names such as ONE, QUORUM, ALL, LOCAL_ONE, and LOCAL_QUORUM. Other systems use different names, placement rules, read reconciliation, and acknowledgement semantics. In this lesson, “ONE-like” means a conceptual threshold of one qualifying response; it is not a claim that every database named ONE behaves identically.

1. A consistency level is an operation contract

For an operation over replicated state, the selected level decides how much evidence the coordinator must collect before completing. On writes, that may be acknowledgements from one, a local majority, a global majority, or all eligible replicas. On reads, it may be one local copy, multiple digests/values, or another product-specific protocol. The selected level therefore changes both the success condition and the latency path.

A low threshold generally gives the coordinator more ways to succeed when replicas are slow or unavailable. But if replicas diverge, a low-threshold read is more likely to land on a stale copy. A high threshold asks for more corroboration, but can turn one slow/unreachable replica into a tail-latency spike or an unavailable operation.

2. Latency follows the slowest required responder

Suppose A responds in 3 ms, B in 8 ms, and a remote C in 72 ms. A ONE-like operation can complete after A. A two-response quorum-like operation completes after B if A and B are eligible. An ALL-like operation cannot complete before C. These are controlled numbers for reasoning, not benchmark predictions.

The same pattern appears in healthy steady state as PACELC reasoning from Chapter 03: even without a network partition, stronger coordination can add network and queueing delay. Tail latency matters more than averages because a quorum waits for enough individual tails to arrive.

Conceptual level Healthy-path benefit Failure/latency cost Typical misuse
ONE-like low response threshold higher stale-read exposure assuming “success” means latest
QUORUM-like more overlap/evidence needs a majority-like set; waits for more tails calling it linearizable without checking assumptions
ALL-like all eligible replicas participate one unavailable/slow replica can fail or delay operation using it as a blanket durability substitute
LOCAL_* like avoid remote-region latency remote replicas can lag; region loss semantics matter confusing local quorum with global convergence

3. Read and write levels form a policy pair

Choosing a read level without the write level is incomplete. If writes acknowledge locally on one replica and reads also accept one arbitrary replica, stale reads are unsurprising. If both operations use intersecting strict quorums with compatible version semantics, freshness improves. If a business rule cannot tolerate an oversell, the team should not “tune” away the serialization/ownership mechanism protecting inventory merely to lower read latency.

Also separate availability from success rate during a particular outage. An ALL-like operation can be perfectly reliable in healthy conditions yet intentionally fail during one-replica loss. That can be the correct choice for a strong invariant if the business prefers an explicit failure to an ambiguous result.

4. Deliberately wrong approach — one global level for every AtlasMart request

Setting every operation to ALL-like makes browse and profile reads inherit the slowest replica and fail during routine maintenance. Setting everything to ONE-like makes checkout depend on whichever replica happens to answer first. Both choices ignore the application’s different invariants.

The repair is an operation matrix. For each path, record authoritative owner, tolerated staleness, write conflict semantics, failure mode, required response set, latency SLO, and fallback behavior. Review that matrix when topology, replication factor, or region placement changes.

5. AtlasMart lab — latency, stale state, and unavailability by threshold

python · tunable_consistency.py
from dataclasses import dataclass

@dataclass(frozen=True)
class Replica:
    name: str
    region: str
    latency_ms: int
    version: int
    reachable: bool = True

replicas = [
    Replica("A", "east", 3, 12),
    Replica("B", "east", 8, 12),
    Replica("C", "west", 72, 11),
]

levels = {
    "ONE-like": 1,
    "QUORUM-like": 2,
    "ALL-like": 3,
}

for name, needed in levels.items():
    reachable = sorted([r for r in replicas if r.reachable], key=lambda r: r.latency_ms)
    if len(reachable) < needed:
        print(name, "UNAVAILABLE")
        continue
    chosen = reachable[:needed]
    latency = max(r.latency_ms for r in chosen)
    observed = max(chosen, key=lambda r: r.version).version
    print(name, "acks/responses=", [r.name for r in chosen],
          "latency_ms=", latency, "observed_version=", observed)

print("\nFAILURE: B becomes unreachable")
replicas2 = [Replica("A", "east", 3, 12), Replica("B", "east", 8, 12, False), Replica("C", "west", 72, 11)]
for name, needed in levels.items():
    reachable = sorted([r for r in replicas2 if r.reachable], key=lambda r: r.latency_ms)
    status = "SUCCESS" if len(reachable) >= needed else "UNAVAILABLE"
    print(name, status, "reachable=", [r.name for r in reachable])

print("\nLOCAL-ONE-LIKE READ")
local = min([r for r in replicas if r.region == "west"], key=lambda r: r.latency_ms)
print("nearest local response", local.name, "version", local.version, "latency_ms", local.latency_ms)
print("fast and available does not mean freshest")
print("latencies are deterministic teaching inputs, not database benchmark results")

Verification checklist

  • ONE-like completes using the fastest eligible responder.
  • QUORUM-like waits for two responders and returns the highest version among those sampled by the toy reconciler.
  • ALL-like includes the remote 72 ms path and becomes unavailable when one replica is unreachable.
  • A west-local read can be fast yet return version 11 while east holds version 12.
  • The displayed latency values are labeled model inputs, not product benchmark results.

Check your understanding

  1. Why can increasing a consistency threshold increase tail latency?
  2. Why is ONE-like success not proof of freshness?
  3. Why is ALL-like not automatically the best production setting?
  4. What is the value of LOCAL_QUORUM-like behavior?
  5. How should AtlasMart choose levels?
Review the answers

1. The coordinator must wait for more qualifying responses, so additional network, queue, disk, and replica tails become part of the completion path.

2. One responding replica can be behind the latest completed write if propagation or repair has not reached it.

3. It maximizes participation but can make routine replica loss or slowness fail/delay operations and may still not solve unrelated concurrency or durability problems.

4. It can keep coordination within one region/data center to reduce WAN latency while still requiring multiple local replicas; global freshness/DR semantics remain separate.

5. Per operation, from invariant strength, stale-read tolerance, latency SLO, failure-domain requirements, and the implementation’s documented semantics.

6. Production judgment

Observe per-level latency histograms, timeouts, unavailable errors, replica lag/divergence, cross-region traffic, and retry volume. Capacity plans must include failure mode: a local quorum that is comfortable with three healthy replicas can saturate the remaining two during maintenance. Security and tenant authorization must remain identical regardless of consistency level; “eventual” is never permission to expose unauthorized state.

The next lesson explains how replicas that missed writes are discovered and repaired rather than merely tolerated.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.