Chapter 05 · Replication Patterns: Leaders, Multi-Leader, and Leaderless Systems
Choose a Replication Pattern from Write Locality, Failure Tolerance, and Conflict Semantics
Choose replication patterns for AtlasMart orders, profiles, and catalog projections from invariants, write locality, failure tolerance, conflict semantics, and operator skill.
Learning outcomes
Replication is not a ranking of technologies. The correct topology depends on where AtlasMart writes originate, which anomalies the business can tolerate, which failure domains must survive, how quickly failover must complete, and how much conflict/repair complexity the team can operate. This lesson turns the previous four mechanisms into a reusable decision record.
Compare single-leader, multi-leader, and leaderless patterns using workload evidence rather than product branding.
Use write locality, staleness tolerance, failover target, conflict semantics, cross-region latency, durability, and operator skill as explicit decision variables.
Defend a replication choice for orders and a different choice for mergeable customer-profile/catalog read-model data.
Name the failure mode that would invalidate each proposed topology.
Record observability, testing, migration, rollback, and recovery requirements alongside the topology choice.
1. Start with write ownership and invariants
Ask first: who is allowed to create authoritative history for this object? Orders/payments often benefit from one write owner per aggregate or shard because state transitions must not diverge. User profile preferences may be mergeable across regions. A catalog search/read model may be derived and rebuildable, allowing different freshness and replication choices than the authoritative price/inventory system.
This is why “AtlasMart uses multi-leader” is too coarse. The architecture can use different replication models for different bounded data responsibilities while still keeping ownership explicit.
2. Decision table
| Dimension | Single leader | Multi-leader | Leaderless/Dynamo-style |
|---|---|---|---|
| write locality | best near leader; remote writes pay routing distance | region-local writes possible | coordinator can write to replica set from many entry points |
| conflict surface | low for writes through current leader | explicit concurrent-write conflicts | concurrent versions/siblings possible depending on semantics |
| failover | promotion/election plus fencing | regional ownership and reconciliation | coordinator/replica availability and quorum/repair behavior |
| read freshness | leader fresh; followers may lag | region-local copies can diverge | depends on read set/consistency/version resolution |
| operator complexity | election, lag, promotion, read routing | plus conflict/loop/schema coordination | plus quorum, hints, repair, version divergence |
3. AtlasMart choice: orders
For orders and payment-state transitions, choose a single leader per order shard with acknowledgement only after a durable copy exists in an independent zone. The reason is not that single-leader is universally “strong.” It is that a single write owner simplifies the ordering of non-mergeable state transitions and reduces conflict surface. The exact durability rule still has to survive the failure domains in the order service's SLO.
Unacceptable failure mode: AtlasMart
acknowledges PAID and later promotes a replica that
lacks that transition, or two active owners both commit
incompatible transitions. That failure would force a redesign of
acknowledgement, promotion, or fencing.
4. AtlasMart choice: customer profile and catalog read model
For selected customer-profile preference fields, multi-leader replication can be justified because users write from multiple regions and independent fields can have an explicit merge policy. Security-sensitive consent/identity fields are excluded unless their conflict semantics are separately proven. For a catalog search/read model, a leaderless/read-optimized replicated projection can tolerate short staleness because the authoritative catalog source remains elsewhere and the projection can be rebuilt.
Unacceptable failure modes: the profile resolver silently discards a non-mergeable security/consent update; or the catalog projection becomes the accidental system of record and cannot be reconstructed after divergence.
5. Operational skill is part of the database requirement
A topology that looks elegant on a whiteboard can be a poor production choice if the team cannot observe and repair it. Single-leader requires election/promotion/fencing competence. Multi-leader requires conflict metrics, reconciliation tooling, and schema compatibility across leaders. Leaderless requires repair scheduling, hint monitoring, consistency-level reasoning, and version/conflict diagnosis.
Include runbooks, restore tests, topology-change procedures, alert ownership, and game-day scenarios in the decision. “Managed service” can shift some operational work to a provider, but it does not remove application invariants or the need to understand service semantics.
6. Deliberately wrong approach — choose one replication pattern for every AtlasMart store
A platform team mandates multi-leader everywhere to achieve “global availability.” Orders now need conflict resolution for states that should never have had concurrent owners; derived catalog data pays conflict-management cost it did not need; operational complexity rises while correctness becomes less obvious.
The repair is workload-scoped selection. Preserve a small number of well-understood patterns, but permit different replication contracts where invariants and locality justify them. Document data ownership and synchronization so polyglot replication does not become ambiguous authority.
7. AtlasMart lab — machine-readable decision record
from dataclasses import dataclass
@dataclass(frozen=True)
class Workload:
name: str
write_locality: str
conflict_tolerance: str
stale_read_tolerance: str
failover_target: str
cross_region_latency_sensitivity: str
operator_skill: str
workloads = [
Workload("orders", "single owning region per order", "very low", "very low", "automatic within a zone set", "medium", "strong"),
Workload("customer-profile", "active users write in multiple regions", "mergeable for selected fields", "seconds acceptable for some fields", "keep local writes during regional isolation", "high", "strong"),
Workload("catalog-read-model", "few authoritative writers; global reads", "derived view may rebuild", "seconds acceptable", "serve stale but valid view", "high", "moderate"),
]
choices = {
"orders": "single leader + synchronous durable copy across independent zones",
"customer-profile": "multi-leader only for explicitly mergeable fields with version/conflict metadata",
"catalog-read-model": "leaderless/read-optimized replicated derived view; source of truth remains separately owned",
}
for w in workloads:
print("\n", w.name.upper())
print("requirements:", w)
print("choice:", choices[w.name])
print("\nUNACCEPTABLE FAILURE MODES")
print("orders: acknowledged payment/order transition disappears or two active owners both commit incompatible state")
print("customer-profile: conflict resolver silently discards a non-mergeable security or consent field")
print("catalog-read-model: derived replicas become unrecoverable or are mistaken for the authoritative price/inventory source")
Verification checklist
- Orders and profile/catalog data do not receive the same replication choice automatically.
- Each choice cites a workload property, not a vendor name.
- The order choice names acknowledged-data loss/split ownership as unacceptable.
- The profile choice limits mergeability to fields with proven conflict semantics.
- The catalog read model remains explicitly derived/rebuildable rather than silently becoming authoritative.
Check your understanding
- What is the first replication-design question?
- Why can different AtlasMart data sets use different replication patterns?
- Why is multi-leader not automatically best for global applications?
- What makes a leaderless catalog projection safer than a leaderless order ledger in this design?
- What operational evidence belongs in the decision record?
Review the answers
1. Who owns authoritative writes for the object/aggregate, and which invariants must that ownership protect.
2. Their write locality, conflict semantics, freshness tolerance, recovery requirements, and operational costs differ.
3. It improves some local-write availability/latency paths but creates concurrent-write conflict and reconciliation requirements that may be unacceptable for strong invariants.
4. The catalog projection is derived, may tolerate bounded staleness, and can be rebuilt; order/payment history is authoritative and contains non-mergeable transitions.
5. Lag/conflict/hint/repair metrics as relevant, acknowledgement/failure-domain assumptions, failover tests, restore/rebuild paths, topology-change runbooks, and rollback/migration criteria.
8. Chapter summary and bridge to quorums
Chapter 05 compared three write-ownership families. Single-leader centralizes ordinary write order and makes follower lag/failover safety explicit. Multi-leader moves writes near users but requires conflict and loop semantics. Leaderless/Dynamo-style replication distributes coordination across a replica set and depends heavily on acknowledgement/read sets, versions, hints, and repair.
Chapter 06 focuses on that last set of mechanics: replication
factor N, write requirement W, read
requirement R, overlap assumptions, tunable
consistency, read repair, Merkle-tree/anti-entropy reasoning,
and the edge cases where a simple quorum formula is not enough.
Authoritative references
- Dynamo: Amazon’s Highly Available Key-value Store — primary leaderless replication reference
- PostgreSQL 18 — High Availability and Replication — current single-primary/standby reference and sync/async tradeoff example
- Apache CouchDB 3.5 — Replication and conflicts — current conflict-preserving decentralized replication example
- Apache Cassandra 5.0 — Hints — current leaderless repair/hinted-handoff implementation example