Chapter 05 · Replication Patterns: Leaders, Multi-Leader, and Leaderless Systems
Single-Leader Replication: Logs, Followers, Failover, and Read Scaling
Trace an AtlasMart order write through a single leader, replication log, follower durability/apply, stale follower reads, and safe failover.
Learning outcomes
AtlasMart has reached the point where a single database process is no longer the only copy of important state. The replication decision now determines who may accept a write, when the client can be told “success,” which copies may serve reads, and what can disappear during failover. This lesson starts with the most common ownership model: one writable leader and one or more followers.
Explain leader, follower, replication log, durable acknowledgement, apply/replay, replica lag, failover, and promotion before using those terms operationally.
Trace one AtlasMart order transition from client request to leader log, follower durability, follower apply, and client acknowledgement.
Distinguish a stale follower read from a lost acknowledged write and connect each symptom to a different control.
Explain why promotion safety depends on the acknowledgement/commit rule rather than on the mere existence of replicas.
Use observable log positions and replica state to decide whether a candidate follower is safe to promote.
Logical versions and conditional updates answer “which state am I changing?” Replication adds a different question: “which machines have durably received that state, and which one owns future writes?” A version token does not make a lagging follower current, and a replica count does not define a commit rule.
1. Single-leader replication gives write ownership a single address
In a single-leader topology, one replica owns ordinary writes for a replication group. Followers receive an ordered stream of changes from that leader. The exact log may be a write-ahead log (WAL), operation log, logical change stream, or another database-specific representation; the architectural point is that followers learn an ordered sequence and advance through it.
The leader is not “the only durable server.” It is the current write owner. A follower can be durable yet behind. A follower can have received log bytes but not yet applied them to a queryable state. A replica can also be fully caught up and still not be eligible for promotion because membership, term/epoch, fencing, or topology policy says otherwise.
| State | Meaning | Evidence |
|---|---|---|
| leader accepted | current write owner validated/accepted the operation | leader log position, request trace |
| leader durable | leader persisted the relevant log/state according to its durability contract | WAL/fsync acknowledgement, commit LSN/offset |
| follower received | replication transport delivered the record | receive offset |
| follower durable | follower persisted it | flush/durable offset |
| follower applied | readable state reflects it | replay/apply offset, visible version |
2. One AtlasMart order write, end to end
Suppose order order-841 moves from
AUTHORIZED to PAID. The client sends
the command to the current leader. The leader validates the
version/invariant, appends the new state to durable local
history, streams it to followers, waits for whatever
acknowledgements the configured policy requires, and then
replies. Followers later apply the log record to their local
queryable state.
If AtlasMart allows reads from followers, the client may observe
a time window where the leader has acknowledged
PAID but one follower still returns
AUTHORIZED. That is replica lag.
It is a freshness problem, not proof that the acknowledged write
was lost.
PostgreSQL 18 streaming replication sends WAL records from primary to standby and is asynchronous by default. Its synchronous replication options can make commits wait for standby feedback. MongoDB replica sets similarly distinguish primary ownership, secondary replication, read concern, and write concern. Their exact semantics differ; the mechanism-first lesson is to inspect the acknowledgement and read contracts rather than infer them from the word “replica.”
3. Failover is a state transfer decision, not merely “pick a live follower”
When the leader is unreachable, a control plane or election mechanism may choose a new leader. A safe promotion process needs to know which candidate has sufficiently complete authoritative history, and it needs a way to prevent the old leader from continuing to accept writes after a newer owner exists. Later chapters cover terms, epochs, leases, consensus, and fencing in depth.
A lagging follower can be a dangerous promotion target if the old acknowledgement policy allowed the client to receive success before that follower received the write. If the old leader is permanently lost, the new leader cannot invent a record it never received. This is why “three replicas” is not, by itself, a durability statement.
4. Read scaling buys capacity by exposing freshness choices
Followers are attractive read targets because they can spread query load. That benefit comes with a contract question: may the read be stale? A product page description may tolerate a short delay; a just-completed payment confirmation often cannot. Routing every read to “nearest replica” without session or freshness rules can violate read-your-writes expectations even though replication is healthy.
Production routing should therefore identify which reads may use followers, what maximum lag is acceptable, how lag is observed, and what the fallback is when followers exceed that threshold. “Read scaling” is an application contract, not a free consequence of adding replicas.
5. Deliberately wrong approach — acknowledge after leader-local durability and promote any live follower
This combination creates a concrete data-loss window. The client receives success after the leader persists locally. Before replication reaches a follower, the leader is destroyed. An arbitrary follower is promoted. The new leader lacks the acknowledged update, so AtlasMart has converted a successful payment/order transition into missing state.
The repair is not “always wait for every replica.” The repair is to define the failure domains you must survive and make the acknowledgement rule match that durability requirement. Then promote only candidates consistent with that rule and use ownership fencing so the old leader cannot reappear as a second writer.
6. AtlasMart lab — observe leader WAL, follower lag, and failover safety
This deterministic simulator models durable log receipt separately from apply. It also contrasts a safer leader-plus-one-follower acknowledgement with a leader-only acknowledgement. It does not model a specific database election algorithm.
from dataclasses import dataclass, field
@dataclass
class Replica:
name: str
log: list[str] = field(default_factory=list)
applied: list[str] = field(default_factory=list)
alive: bool = True
def persist(self, entry: str):
if self.alive:
self.log.append(entry)
def apply_all(self):
if self.alive:
self.applied = list(self.log)
leader = Replica("L-eu-a")
f1 = Replica("F-eu-b")
f2 = Replica("F-eu-c")
entry = "v41 order-841 status=PAID"
print("WRITE PATH")
leader.persist(entry)
print("t1 leader WAL:", leader.log)
f1.persist(entry)
print("t2 follower F1 durable:", f1.log)
print("t3 client ACK policy=leader+one-follower -> ACK")
# F2 is lagging: it has not received the entry yet.
print("t4 follower F2 lagging:", f2.log)
leader.apply_all(); f1.apply_all(); f2.apply_all()
print("\nREAD STATE before catch-up")
for r in (leader, f1, f2):
print(r.name, "applied", r.applied)
print("\nFAILOVER")
leader.alive = False
promoted = f1
print("leader failed; promote", promoted.name)
print("promoted log contains acknowledged entry?", entry in promoted.log)
print("F2 stale read contains entry?", entry in f2.applied)
print("\nCATCH-UP")
f2.persist(entry); f2.apply_all()
print("F2 after catch-up:", f2.applied)
print("\nUNSAFE COUNTERFACTUAL")
leader_only = Replica("L2")
standby = Replica("F2")
leader_only.persist("v42 order-842 status=PAID")
print("client ACK after leader-only durability")
leader_only.alive = False
print("leader dies before replication; promotable standby log:", standby.log)
print("acknowledged write survives?", bool(standby.log))
Verification checklist
-
The acknowledged
v41record exists on follower F1 before the simulated leader failure. - F2 can return stale state before catch-up without implying that F1 lost the record.
- Promoting F1 preserves the acknowledged record under the simulator's stated policy.
- The leader-only counterfactual demonstrates an acknowledged record disappearing when the sole durable copy is lost.
- You can state exactly what evidence would be required before promoting a real follower.
Check your understanding
- Why does three replicas not automatically mean an acknowledged write survives one failure?
- What is the difference between follower receive/flush and apply?
- Why can follower reads be stale even when replication is working correctly?
- What makes failover unsafe besides missing data?
- When is follower read scaling appropriate?
Review the answers
1. Because survival depends on which replicas durably received the write before acknowledgement and which replica is promoted after failure. Replica count alone does not define that set.
2. Receive/flush means replication data arrived and was persisted; apply means local queryable state has replayed it. A follower may be durable but not yet current for reads.
3. Asynchronous transport and apply create lag. The follower may simply not have reached the leader's latest acknowledged position yet.
4. Two owners can accept writes if the old leader is not fenced or does not learn that a newer term/epoch exists.
5. When the application explicitly tolerates the measured freshness window or uses a stronger routing/session rule for reads that require recent writes.
7. Production judgment and next bridge
Single-leader replication is often attractive when AtlasMart needs simple write ordering, centralized conflict avoidance, mature failover tooling, and many read replicas. Its costs are leader write concentration, failover complexity, cross-region write latency when the leader is remote, and explicit freshness management for followers.
Monitor replication receive/flush/apply positions, lag distributions, follower eligibility, failed elections/promotions, read routing, and the number of durable copies per failure domain. Test leader loss before and after each acknowledgement stage. Chapter 05 Lesson 2 now isolates that acknowledgement choice and asks how much latency AtlasMart should pay for durability across zones or regions.
Authoritative references
- PostgreSQL 18 — High Availability, Load Balancing, and Replication — current official documentation for primary/standby replication patterns and tradeoffs
- PostgreSQL 18 — Log-Shipping Standby Servers — streaming replication, asynchronous default, synchronous replication, failover, and standby behavior
- MongoDB — Replication — official replica-set behavior and read/write semantics; used only as a concrete mapping
- MongoDB — Write Concern for Replica Sets — acknowledgement-count semantics and rollback risk under weaker write concern