Turn concerns and preferences into an operation-level decision record tied to invariants, SLOs, and regional failure assumptions.
Choose Concerns per Operation from Invariants, Latency Budgets, and Regional Failure Requirements
Benchmark the same order-version invariant under several policies, record tail latency and visibility, and justify every override.
Learning objectives
Translate business invariants into per-operation write concern, read concern, read preference, and session choices.
Run one AtlasMart order invariant through several policies and record acknowledgement, source member, visibility, and latency.
Distinguish durability, freshness, linearizability, causal ordering, and availability instead of collapsing them into “consistency”.
Use latency/error distributions and failure requirements to reject settings that are correct but operationally unacceptable.
Document rollback and migration paths before changing application-wide defaults.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where the driver is used. The mandatory
topology is a disposable three-member replica set on one Docker
host, with loopback-published diagnostic ports 27112–27114.
Members communicate over a lesson-specific Docker bridge network
by container DNS name. Authentication and TLS are disabled only
for this isolated local learning topology. Feature Compatibility
Version (FCV) is observed and never changed. Unless the
experiment states otherwise, reads target the primary and writes
use an explicitly named concern rather than assuming a global
default. Run only one Chapter 15 topology at a time and budget
roughly 3–4 GB of free RAM plus disk headroom. Local member tags
such as east/west demonstrate routing
semantics only; they do not simulate real WAN latency or failure
domains. Atlas, Search, Vector Search, KMS, and Enterprise
Advanced are not mandatory. The decision lab tags members for
routing and uses one short delayed secondary so policies can be
compared. Real regional latency and correlated failures must be
tested in a production-like multi-zone environment; this
single-host lab cannot reproduce them. Product commands were not
executed in this generation environment because Docker, mongod,
mongosh, and PyMongo are unavailable here; expected invariants
are documentation-derived and measured values must be recorded
on the learner’s machine.
1. There is no single correct “MongoDB consistency setting”
AtlasMart has different invariants. Checkout needs a durable stock reservation. A customer opening the just-confirmed order page needs read-your-writes. A recommendation widget can tolerate stale secondary data. An administrator deciding whether a unique payment has already settled might require a primary-only real-time read. Applying one strongest configuration globally can inflate latency and reduce availability without protecting any additional invariant.
| Operation | Invariant / tolerance | Candidate policy |
|---|---|---|
| Reserve inventory | Accepted reservation should survive ordinary replica failover. |
w:"majority" with bounded timeout;
idempotent operation ID.
|
| Read own confirmation from secondary | Same workflow must not regress after its write. | Causal session + majority read/write concerns; suitable secondary read preference. |
| Recommendation browse | Minutes of staleness acceptable. | Secondary/nearest with explicit staleness policy and ordinary read concern chosen from rollback tolerance. |
| Real-time single-document guard | Must reflect majority writes completed before the read starts. |
Primary + linearizable +
maxTimeMS.
|
| Historical point-in-time analysis | Needs internally consistent recent snapshot, not latest value. |
snapshot where supported; monitor
snapshot-history limits.
|
docker rm -f atlasmart-mongo-ch15-l5-a atlasmart-mongo-ch15-l5-b atlasmart-mongo-ch15-l5-c 2>/dev/null || truedocker network rm atlasmart-ch15-l5-net 2>/dev/null || truedocker volume rm atlasmart-mongo-ch15-l5-a-data atlasmart-mongo-ch15-l5-b-data atlasmart-mongo-ch15-l5-c-data 2>/dev/null || truedocker network create atlasmart-ch15-l5-netdocker run -d --name atlasmart-mongo-ch15-l5-a --network atlasmart-ch15-l5-net -p 127.0.0.1:27112:27017 -v atlasmart-mongo-ch15-l5-a-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs15-l5 --bind_ip_alldocker run -d --name atlasmart-mongo-ch15-l5-b --network atlasmart-ch15-l5-net -p 127.0.0.1:27113:27017 -v atlasmart-mongo-ch15-l5-b-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs15-l5 --bind_ip_alldocker run -d --name atlasmart-mongo-ch15-l5-c --network atlasmart-ch15-l5-net -p 127.0.0.1:27114:27017 -v atlasmart-mongo-ch15-l5-c-data:/data/db mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs15-l5 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27112/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27112/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-rs15-l5",members:[ {_id:0,host:"atlasmart-mongo-ch15-l5-a:27017"}, {_id:1,host:"atlasmart-mongo-ch15-l5-b:27017"}, {_id:2,host:"atlasmart-mongo-ch15-l5-c:27017"}]})'until mongosh "mongodb://127.0.0.1:27112/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donefor i in $(seq 1 60); do STATES=$(mongosh "mongodb://127.0.0.1:27112/admin?directConnection=true" --quiet --eval 'const s=rs.status(); print(s.members.map(m=>m.stateStr).sort().join(","))' || true) [ "$STATES" = "PRIMARY,SECONDARY,SECONDARY" ] && break sleep 1donemongosh "mongodb://127.0.0.1:27112/admin?directConnection=true" --quiet --eval 'printjson({server:db.version(),hello:db.hello(),fcv:db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion,status:rs.status().members.map(m=>({name:m.name,stateStr:m.stateStr}))})'
2. Prepare routing metadata and one controlled lagging secondary
let cfg=rs.conf();const a=cfg.members.find(m=>m.host.startsWith("atlasmart-mongo-ch15-l5-a:"));const b=cfg.members.find(m=>m.host.startsWith("atlasmart-mongo-ch15-l5-b:"));const c=cfg.members.find(m=>m.host.startsWith("atlasmart-mongo-ch15-l5-c:"));a.tags={region:"east",role:"operational"};b.tags={region:"east",role:"read"};c.tags={region:"west",role:"read"};c.priority=0; c.votes=0; c.secondaryDelaySecs=15;rs.reconfig(cfg);const app=db.getSiblingDB("atlasmart");app.policy_orders.drop();app.policy_orders.insertOne({_id:"order-policy",version:0,status:"new"},{writeConcern:{w:"majority",wtimeout:5000}});printjson(rs.conf().members.map(m=>({host:m.host,tags:m.tags,votes:m.votes,delay:m.secondaryDelaySecs||0})));
3. Run the same order-version invariant under several policies
import math, timefrom pymongo import MongoClientfrom pymongo.monitoring import CommandListenerfrom pymongo.read_concern import ReadConcernfrom pymongo.write_concern import WriteConcernfrom pymongo.read_preferences import Primary, SecondaryURI="mongodb://127.0.0.1:27112,127.0.0.1:27113,127.0.0.1:27114/?replicaSet=atlasmart-rs15-l5"class Trace(CommandListener): def __init__(self): self.last_find_source=None def started(self,event): if event.command_name == "find": self.last_find_source=str(event.connection_id) def succeeded(self,event): pass def failed(self,event): passtrace=Trace()client=MongoClient(URI,serverSelectionTimeoutMS=30000,event_listeners=[trace])base=client["atlasmart"]["policy_orders"]def percentile(xs,p): ys=sorted(xs); return ys[min(len(ys)-1,max(0,math.ceil(p*len(ys))-1))]def run_policy(name,wc,rc,rp,use_session=False,loops=8): coll=base.with_options(write_concern=wc,read_concern=rc,read_preference=rp) write_ms=[]; read_ms=[]; observed=[]; sources=[] for i in range(loops): ctx=client.start_session(causal_consistency=True) if use_session else None try: t=time.perf_counter() coll.update_one({"_id":"order-policy"},{"$inc":{"version":1}},session=ctx) write_ms.append((time.perf_counter()-t)*1000) trace.last_find_source=None t=time.perf_counter() doc=coll.find_one({"_id":"order-policy"},session=ctx) read_ms.append((time.perf_counter()-t)*1000) observed.append(doc["version"] if doc else None) sources.append(trace.last_find_source) finally: if ctx: ctx.end_session() print(name,{"write_p50":percentile(write_ms,.50),"write_p95":percentile(write_ms,.95), "read_p50":percentile(read_ms,.50),"read_p95":percentile(read_ms,.95), "read_sources":sources,"observed_versions":observed})run_policy("w1-primary-local",WriteConcern(w=1),ReadConcern("local"),Primary())run_policy("majority-primary-majority",WriteConcern("majority",wtimeout=5000),ReadConcern("majority"),Primary())run_policy("majority-secondary-majority-no-session",WriteConcern("majority",wtimeout=5000),ReadConcern("majority"),Secondary([{"region":"west"}]))run_policy("majority-secondary-majority-causal",WriteConcern("majority",wtimeout=5000),ReadConcern("majority"),Secondary([{"region":"west"}]),use_session=True,loops=3)client.close()
The script intentionally reports measured values rather than expected fixed latency. Compare returned versions with the version just written in each iteration. The non-session west-secondary policy can lag; the causal policy may wait but should preserve its own dependency. Re-run several times and record p50/p95/p99 plus server-selection/replication evidence.
4. Add a linearizable guard only where the invariant demands it
const app=db.getSiblingDB("atlasmart");const t=Date.now();const r=app.runCommand({ find:"policy_orders", filter:{_id:"order-policy"}, readConcern:{level:"linearizable"}, maxTimeMS:5000});printjson({elapsedMs:Date.now()-t,result:r});
Do not place linearizable reads on every request because they sound “strong.” They can be significantly slower and less available when a majority cannot be confirmed. Use them for the narrow operation whose invariant needs real-time ordering of majority-acknowledged single-document state.
5. Decision record: correctness first, then operational fit
| Question | Evidence to collect | Rollback / alternative |
|---|---|---|
| What failure must the write survive? | Write-concern result, majority commit/lag, failover drill. | Lower concern only for explicitly weaker invariants. |
| How stale may the read be? | Returned version, member source, replication lag, max-staleness selection failures. | Primary read or causal dependency if freshness is required. |
| Does the workflow need read-your-writes? | Session operationTime/clusterTime and dependent read latency. | Keep read on primary if secondary wait is not worth complexity. |
| Is real-time order actually required? | Linearizable read latency/error rate under degraded majority. | Use majority/causal semantics if real-time order is unnecessary. |
| Can the latency budget tolerate the policy? | p50/p95/p99 under healthy and failed-member tests. | Change operation policy/model; do not weaken silently. |
| What happens during region/zone loss? | Production-like multi-zone drill; this lab cannot simulate it. | Document fallback read preference and invariant impact. |
Global defaults can be useful governance, but they should not replace an operation-level invariant review. Over-strengthening every read/write can create unnecessary tail latency and availability loss; under-strengthening can create rollback/staleness bugs. Record the intended default plus justified overrides and test both.
Production judgment. Make concerns/preferences part of API design and SLO review. Track write-concern timeouts, read source, replica lag, causal wait time, server-selection failures, and latency percentiles. Security still matters: member tags are routing metadata, not authorization; tenant filters are not access control; TLS/auth/network isolation remain required. Changing global defaults is a migration: inventory all callers, stage canaries, provide rollback, and verify old/new clients behave as intended.
Bridge to Chapter 16. Replica-set concerns
determine acknowledgement and visibility inside a replica set.
Sharding adds another dimension: mongos must route
operations across shard replica sets, making shard-key targeting
and cross-shard coordination part of the latency/correctness
story.
docker rm -f atlasmart-mongo-ch15-l5-a atlasmart-mongo-ch15-l5-b atlasmart-mongo-ch15-l5-c 2>/dev/null || truedocker volume rm atlasmart-mongo-ch15-l5-a-data atlasmart-mongo-ch15-l5-b-data atlasmart-mongo-ch15-l5-c-data 2>/dev/null || truedocker network rm atlasmart-ch15-l5-net 2>/dev/null || true
Check your understanding
- Why not use one strongest concern/preference globally?
- What policy gives durable read-your-writes across secondary reads?
- When is linearizable appropriate?
- Why is a single-host east/west tag lab not a regional test?
- What evidence belongs in a policy decision?
Review the answers
1. Different operations have different invariants and latency/availability budgets; over-strengthening can add cost without protecting additional correctness.
2. A causally consistent session with majority read concern and majority write concern, plus a read preference that may select the intended secondary.
3. When a primary-only read must reflect successful majority writes completed before the read begins in real-time order.
4. It cannot reproduce WAN RTT, bandwidth, correlated zone failures, routing infrastructure, or compliance boundaries.
5. Acknowledgement result, member source, returned version/staleness, replication state, error rate, p50/p95/p99 latency, and failure-drill observations.
Authoritative references
- MongoDB 8.3 release notes — Current server release line and patch-sensitive behavior.
-
Write Concern
—
w,j,wtimeout, majority acknowledgment, and journaling behavior. - Read Concern — Supported read-concern levels, operations, and transaction/session compatibility.
- Read Concern majority — Majority-commit visibility and rollback guarantees.
- Read Concern linearizable — Primary-only real-time ordering semantics and latency implications.
- Read Concern snapshot — Point-in-time majority-committed reads and snapshot-history limits.
- Read Preference — Primary/secondary routing modes and stale-read consequences.
- Server Selection Algorithm — Eligibility, latency windows, and per-operation member selection.
- Read Preference Tag Sets — Ordered tag-set matching for replica-set reads.
- maxStalenessSeconds — Coarse secondary-staleness filtering and the 90-second minimum.
- Causal Consistency and Concerns — Read-your-writes, monotonic reads/writes, and writes-follow-reads.
- Read Isolation, Consistency, and Recency — Session guarantees, visibility, and operation-time behavior.
- PyMongo CRUD configuration — Driver read preference, read concern, write concern, and tags.
- PyMongo sessions and causal consistency — ClientSession behavior and causal consistency.
- PyMongo release notes — Current 4.17 driver baseline.
- mongosh release notes — Current 2.10.0 shell baseline.