Chapter 03 · CAP, PACELC, and Consistency/Availability Tradeoffs
PACELC: Latency vs Consistency Tradeoffs Even When the Network Is Healthy
Extend failure reasoning into the healthy-network case: coordination and geographic distance can trade lower latency for stronger freshness/order.
Learning outcomes
AtlasMart operates in several regions. There is no outage, yet a design question remains: should a shopper in Europe read the nearby replica in roughly one regional hop, or wait for coordination with a distant authority to obtain the freshest possible value? CAP does not answer this healthy-network question. PACELC is a useful design lens because it asks about both partition behavior and the else-case when communication works.
Explain PACELC as “if Partition, trade Availability/Consistency; Else, trade Latency/Consistency” without treating it as a rigid product taxonomy.
Distinguish network distance, coordination, and replication lag as separate contributors to read behavior.
Compare local, leader-coordinated, and multi-replica confirmation paths using deterministic latency/freshness inputs.
Explain why an AP/CP label says little about steady-state tail latency.
Choose healthy-operation paths from freshness/invariant requirements rather than from a slogan.
1. CAP talks about a fault; users experience latency every day
A globally distributed service spends most of its life outside a total partition. Reads and writes still cross machines, racks, zones, or regions. Waiting for more remote participants can improve freshness or ordering confidence, but it increases the response’s dependency on network delay and the slowest required participant. Serving a nearby replica can reduce latency while accepting some staleness.
Daniel Abadi’s PACELC formulation captures this: if there is a Partition (P), how does the system trade Availability (A) and Consistency (C); Else (E), how does it trade Latency (L) and Consistency (C)? It is a lens, not a promise that every operation in a named product fits one immutable pair of letters.
2. Three healthy read paths
| Path | Typical mechanism | Latency dependency | Freshness/order consequence |
|---|---|---|---|
| Local replica | read nearest copy without remote coordination | local processing + local network | may observe replication lag |
| Leader/authority read | route to current authoritative owner | client→leader distance + leader work | can support stronger real-time semantics if the protocol guarantees them |
| Multi-replica confirmation | wait for a required set of replicas | order statistic of required responses | can reduce stale-read risk under assumptions; not automatically linearizable |
These are abstractions. Real products differ in quorum algorithms, leader leases, follower-read modes, consistency tokens, timestamp semantics, and caches. The lesson’s goal is to identify where latency comes from and what guarantee the extra communication actually buys.
3. Tail latency matters more than a single average
If a request must wait for two of three replicas, it depends on the second-fastest response for that request. If it must wait for all replicas, one slow path becomes the tail. A local read can avoid that coordination but may return an older version. Production decisions therefore need latency distributions and freshness distributions, not one “fast” or “consistent” adjective.
The numbers below are deterministic inputs chosen for reasoning. They are not measurements and must never be copied into a capacity plan. A real benchmark must disclose regions/topology, payloads, client concurrency, consistency/durability settings, warmup, caches, network conditions, and percentile methodology.
4. Deliberately wrong approach — infer latency from “CP”
Calling a system “CP” does not tell you whether reads are served by the leader, by followers with leases, by local replicas under bounded-staleness rules, or by a quorum. Two deployments of the same software can have radically different normal-operation latency because of region placement and configuration. Similarly, calling a system “AP” does not mean every read is local or every write avoids coordination.
The repair is to draw the actual request path. Mark the client region, coordinator, replicas, required acknowledgements, and freshness source. Then measure the resulting percentiles under representative load and failure.
5. AtlasMart lab — model latency and freshness together
from statistics import median
# Deterministic teaching data, NOT measured production latency.
paths = {
"local eventual read": [4, 5, 4, 6, 5],
"leader-coordinated read": [74, 82, 79, 77, 86],
"two-replica confirmation": [39, 47, 42, 55, 44],
}
freshness_ms = {
"local eventual read": 900,
"leader-coordinated read": 0,
"two-replica confirmation": 80,
}
for name, samples in paths.items():
print(f"{name:27} median={median(samples):>5.1f} ms | modeled freshness lag <= {freshness_ms[name]} ms")
print("\nThese numbers are inputs to a model, not a benchmark.")
print("PACELC asks what the design does during P, and what it trades in E (no partition).")
Expected output
local eventual read median= 5.0 ms | modeled freshness lag <= 900 ms
leader-coordinated read median= 79.0 ms | modeled freshness lag <= 0 ms
two-replica confirmation median= 44.0 ms | modeled freshness lag <= 80 ms
These numbers are inputs to a model, not a benchmark.
PACELC asks what the design does during P, and what it trades in E (no partition).
The model intentionally makes the local read fastest and stalest, the leader path freshest and slowest, and a two-replica path intermediate. Those outcomes are inputs, not universal truths. Change the values and the result changes—which is exactly why topology and workload evidence matter.
Check your understanding
- What does the E in PACELC represent?
- Why does a CAP classification not determine steady-state latency?
- What hidden variable often makes globally coordinated reads slower than local reads?
- Does waiting for multiple replicas automatically prove linearizability?
- What should a real benchmark report in addition to latency percentiles?
Review the answers
The Else case: normal operation when the relevant network communication is available.
CAP’s core impossibility concerns partitioned operation; normal paths depend on routing, topology, coordination, and configuration.
Network distance plus the need to wait for remote participants; tail latency can also be dominated by the slowest required response.
No. Quorum overlap can be useful under specific assumptions, but linearizability also depends on write/read protocol, concurrency, failures, and version/order rules.
Freshness/staleness, consistency and durability settings, dataset and key distribution, concurrency, topology, warmup/cache state, errors, and failure conditions.
6. Production judgment
For AtlasMart, catalog pages may prefer low local latency with a freshness SLO, while inventory reservation may accept extra coordination to preserve scarcity. Payment status might route to an authoritative region after a write while recommendation reads remain local. Document these decisions per operation and expose freshness/consistency metadata where useful rather than hiding all reads behind one repository method.
The next lesson gives names to the client-visible guarantees behind these choices: linearizable/strong, eventual, bounded-staleness, session, causal, and tunable consistency. Those models should be selected by the anomalies they permit or forbid, not by marketing tiers.
Authoritative references
- Daniel Abadi — Consistency Tradeoffs in Modern Distributed Database System Design: CAP is Only Part of the Story — source of the PACELC framing
- Gilbert & Lynch — CAP theorem formalization — the partition-side result PACELC complements
- Amazon Dynamo paper — real system discussion of latency/availability-driven tradeoffs and conflict handling