Separate consensus-backed cluster metadata from application-data replication, then prove that a database can have perfectly current ownership metadata and still serve a stale user record.

Distinguish Consensus for Cluster Metadata from Replication of Application Data

The presence of Raft or Paxos somewhere in a database does not define every user-operation guarantee. Trace the actual write/read path, acknowledgement policy, replication lag, and read mode.

Advanced115–150 minutesControl-plane vs data-plane labPython 3.13+ · standard libraryVendor-neutral · free/local mandatory pathLast reviewed: August 2026
01

Trace a system where consensus governs shard metadata while application records use a separate replication protocol.

02

Diagnose the false inference that “the database uses Raft/Paxos” means every user read/write is linearizable.

03

Use replica versions/lag and client-visible histories to state the actual guarantee of an application-data path.

04

Document when per-shard data consensus is used instead and how that architectural choice changes write/read semantics.

1. The misleading architecture diagram

An AtlasMart architecture slide says “Raft-based database.” Engineers then assume a read from any replica after a successful write must return the new value. That conclusion is not justified by the word Raft. The first question is: what state does the consensus group actually replicate?

In this lesson's example, consensus commits only the mapping catalog-17 → leader A, epoch 7. Application data under that shard uses leader/follower replication with a configurable acknowledgement boundary. Leader A can acknowledge version 10 while followers B and C still hold version 9. A follower-local read can therefore be stale even though every control-plane node agrees perfectly on leader A.

2. Trace guarantees by operation path

For any database claim, draw the concrete path: client → router → metadata lookup → data owner → replication acknowledgements → client response → later read target. Ask which steps use consensus, which use quorum overlap, which are asynchronous, and which state is durable before acknowledgement. A single product can expose multiple consistency modes.

Path component Possible mechanism Guarantee question
Membership/shard map Raft/Paxos consensus Is there one authoritative current topology/epoch?
User write Leader local + async followers When is the write acknowledged and what data-loss window remains?
User write Per-shard Raft/Paxos Is the command committed by that shard consensus group before acknowledgement?
Follower read Local replica state Can it be stale relative to a completed leader write?
Linearizable read Leader/read-index/quorum protocol depending system How does the implementation exclude stale leadership/state?

3. AtlasMart lab: metadata is current while data is stale

Mandatory lab environment

Python 3.13+ standard library only. The generated lab was verified with Python 3.13.5. No database server, Docker, cloud account, paid feature, credential, firewall change, clock manipulation, process killing, or destructive failure injection is required. All failures and partitions are deterministic in-memory simulations.

The control-plane mapping is already committed at metadata revision 104. The data plane then contains version 10 only on leader A while B and C remain version 9. Reading B demonstrates the exact false inference this lesson is designed to prevent.

python · AtlasMart deterministic simulation
metadata = {
    "epoch": 7,
    "shard": "catalog-17",
    "leader": "A",
    "replicas": ["A", "B", "C"],
    "metadata_commit": 104,
}
data = {
    "A": {"version": 10, "price": 90},
    "B": {"version": 9, "price": 100},
    "C": {"version": 9, "price": 100},
}

print("CONSENSUS-COMMITTED CONTROL PLANE")
print(metadata)
print("all clients can agree that A owns catalog-17 at epoch 7")

print("\nAPPLICATION DATA USES DIFFERENT REPLICATION SEMANTICS")
print("A accepted version 10 and acknowledged after its configured local durability boundary")
print("replication to B/C is still pending")
print("data replicas:", data)

print("\nCLIENT READS FOLLOWER B")
read = data["B"]
print("B returned:", read)
print("stale relative to leader A:", read["version"] < data["A"]["version"])
print("metadata consensus did NOT make this follower read linearizable")

print("\nREPLICATION CATCH-UP")
for replica in ["B", "C"]:
    data[replica] = dict(data["A"])
print("data replicas:", data)
print("converged:", len({(v['version'], v['price']) for v in data.values()}) == 1)

print("\nARCHITECTURE NOTE")
print("a different system may run Raft/Paxos per data shard and consensus-commit user writes")
print("therefore infer guarantees from the user-operation path, not from the mere presence of Raft/Paxos somewhere")
Expected evidence

The client can know with consensus-backed certainty that A owns catalog-17 at epoch 7 and still receive stale application data from B. After asynchronous catch-up, replicas converge. This is a teaching architecture, not a claim about every database. Some systems deliberately run a Raft/Paxos group per data shard, in which case the user write path itself may be consensus-committed.

4. Deliberately wrong approach: market the strongest subsystem guarantee as the whole-system guarantee

A system may have linearizable metadata, strongly consistent schema changes, or consensus-based leader election while offering eventual/tunable consistency for ordinary records. Documentation and architecture reviews must state guarantees per operation, not per brand. Likewise, a “quorum write” is not automatically a consensus decision; Chapters 5–6 showed that replica quorums and consensus quorums solve related but different problems.

The safe correction is to publish a guarantee matrix: operation, routing path, acknowledgement set, durability boundary, read modes, failover behavior, stale-read allowance, maximum tolerated lag if bounded, and recovery/reconciliation path.

5. What if data shards also use consensus?

Many production systems do use Raft/Paxos-family protocols for each replicated data range/shard. Then a successful write can mean the shard's consensus log committed the command before acknowledgement, and linearizable reads may be available through leader/read-index/lease mechanisms depending on the implementation. But the performance envelope changes: every write normally pays quorum/leader coordination within that shard, hot shards can bottleneck, and cross-shard transactions require additional coordination.

Therefore the right statement is conditional: this user operation is consensus-committed by group G under policy P, not the database uses Raft.

6. Observability proves which path happened

Expose metadata revision/epoch, shard ID, data leader, write acknowledgement policy, commit/applied index where applicable, replica version/lag, read consistency mode, and client request ID in traces. During incident review, these fields answer whether the client hit stale routing, a lagging follower, a lost leadership term, or a genuinely committed data-log entry. “Consensus healthy” is not enough if the data plane is 30 seconds behind.

For backups and disaster recovery, separately test restoration of consensus metadata and application data. A perfectly restored control plane pointing at incomplete data is not a successful restore.

7. Security and operational boundaries

Consensus metadata is often extremely sensitive: membership, credentials/leases, encryption configuration, network endpoints, and tenant routing can create broad blast radius. Authenticate members, encrypt peer/client traffic, restrict administrative APIs, patch security issues, and audit changes. etcd's July 2026 patch release is a useful reminder that authorization bugs can exist even in a mature consensus-backed store; consensus safety does not imply access-control correctness.

8. Chapter synthesis and bridge to conflict resolution

Consensus gives a group a safe authoritative decision sequence under stated fault/timing assumptions. Raft expresses that through leaders, terms, and a replicated log; Paxos through ballots, promises, acceptors, and chosen values. Linearizability turns that ordered state into a real-time client contract when the operation path implements it. Fencing makes ownership robust against paused stale actors. Finally, architecture determines whether consensus protects only control-plane metadata or every application-data write.

Chapter 18 now assumes that distinction and studies what happens when application writes are allowed to occur concurrently without one total consensus order: lost updates, duplicate events, last-write-wins, application merges, conflict-free replicated data types (CRDTs), and idempotency.

Check your understanding

  1. Why can a follower read be stale even when shard ownership metadata is consensus-committed?
  2. What should an architecture guarantee statement name?
  3. Is a replication quorum automatically a consensus protocol?
  4. How does per-shard consensus change the picture?
  5. What telemetry distinguishes control-plane health from data-plane freshness?
Review the answers

1. The metadata and application data follow different replication paths; consensus on ownership does not replicate the record itself.

2. The specific operation/path, acknowledgement and durability policy, read mode, failure assumptions, and whether the data itself is consensus-committed.

3. No. Quorum replication can provide overlap/freshness properties without the ballot/term/log rules needed to choose one ordered value.

4. User writes for that shard may themselves be committed through Raft/Paxos, at the cost of quorum coordination and shard-level hotspots.

5. Metadata epoch/revision plus data leader, commit/applied indexes where relevant, replica versions/lag, acknowledgement policy, and read consistency mode.

References

Foundational claims use primary research where practical. Product documentation is used only as a current implementation example and is not required for the mandatory labs.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.