A production claim is credible only when the system can be measured, failed, recovered, secured, and operated from its runbook.
Capstone: Build, Load-Test, Fail, Recover, Secure, Tune, and Defend a Production MongoDB System
Integrate the entire course in a measurable AtlasMart capstone: build, load-test, fail, recover, harden, tune, observe, and defend the system with a runnable acceptance runbook.
Learning objectives
Assemble a runnable AtlasMart capstone with replica-set topology, indexes, driver configuration, observability, and bounded load.
Measure throughput and latency distributions without hiding queueing or errors.
Kill the current primary safely, observe election/recovery, and verify business invariants after failover.
Apply a least-privilege/authentication security gate and verify denied administrative access.
Produce a final architecture/runbook dossier with evidence, limits, recovery steps, and unresolved risks.
This final chapter pins
MongoDB Community Server 8.3.8 using
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0.
Labs use AtlasMart synthetic data, loopback-only Docker port
publishing, replica set name atlasmart-rs27 where
topology behavior matters, and explicit operation time budgets.
Authentication/TLS are disabled only for disposable mechanism
labs; the capstone security gate reuses Chapter 22
least-privilege/authentication requirements and treats TLS as
mandatory production acceptance. Default read preference is
primary and majority acknowledgement is used for business writes
unless a failure experiment explicitly states otherwise.
FCV is observed and never changed. Atlas,
Search, Vector Search, Enterprise Advanced, and KMS are optional
and are not required for mandatory work. Runtime load, failover,
pool, latency, retry, and recovery results were not executed in
the generation environment; learners must record their own
evidence instead of copying invented values.
1. Capstone acceptance contract
The final lab deliberately stays small enough for one workstation but uses the same evidence categories as production. It does not claim that three containers on one laptop reproduce independent failure domains, WAN latency, encrypted network transport, or cloud storage behavior. Those remain production acceptance items.
| Gate | Evidence |
|---|---|
| Build | healthy replica-set topology, indexes, users/config, seed data |
| Load | offered/completed/error counts and p50/p95/p99/max latency |
| Fail | primary stopped, election observed, client recovery measured |
| Recover | old member rejoins; majority writes and application invariants verify |
| Secure | loopback-only exposure plus least-privilege auth test; TLS required in production |
| Tune | one evidence-backed change followed by same workload/re-measurement |
| Defend | runbook, architecture decisions, known limits, rollback and restore paths |
2. Build and seed
docker rm -f atlasmart-rs27-lab 2>/dev/null || trueIMAGE=mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimdocker run -d --name atlasmart-rs27-lab \ -p 127.0.0.1:27209:27209 -p 127.0.0.1:27210:27210 -p 127.0.0.1:27211:27211 \ --entrypoint bash "$IMAGE" -lc 'set -emkdir -p /data/rs1 /data/rs2 /data/rs3mongod --dbpath /data/rs1 --port 27209 --replSet atlasmart-rs27 --bind_ip_all --fork --logpath /tmp/rs1.logmongod --dbpath /data/rs2 --port 27210 --replSet atlasmart-rs27 --bind_ip_all --fork --logpath /tmp/rs2.logmongod --dbpath /data/rs3 --port 27211 --replSet atlasmart-rs27 --bind_ip_all --fork --logpath /tmp/rs3.logtail -f /dev/null'sleep 4docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'rs.initiate({_id:"atlasmart-rs27",members:[ {_id:0,host:"localhost:27209"}, {_id:1,host:"localhost:27210"}, {_id:2,host:"localhost:27211"}]})'docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'while (rs.status().members.filter(m=>m.stateStr==="PRIMARY").length!==1 || rs.status().members.filter(m=>m.stateStr==="SECONDARY").length!==2) { sleep(1000) }printjson(rs.status().members.map(m=>({name:m.name,state:m.stateStr})))'
PRIMARY_PORT=$(docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'print(rs.status().members.find(m=>m.stateStr==="PRIMARY").name.split(":")[1])')docker exec atlasmart-rs27-lab mongosh --quiet --port "$PRIMARY_PORT" --eval 'use admindb.createUser({user:"atlasmart_lab_admin",pwd:"lab-admin-change-me",roles:[{role:"root",db:"admin"}]})use atlasmartdb.createUser({user:"atlasmart_app",pwd:"lab-only-change-me",roles:[{role:"readWrite",db:"atlasmart"}]})printjson(db.getUser("atlasmart_app"))'# Both passwords are synthetic and disposable. Production secrets belong in a secret manager, never source control.
python -c "import base64,secrets; print(base64.b64encode(secrets.token_bytes(512)).decode())" > rs27.keydocker cp rs27.key atlasmart-rs27-lab:/data/keyfilerm -f rs27.keydocker exec atlasmart-rs27-lab bash -lc 'chmod 400 /data/keyfilemongod --shutdown --dbpath /data/rs1mongod --shutdown --dbpath /data/rs2mongod --shutdown --dbpath /data/rs3mongod --dbpath /data/rs1 --port 27209 --replSet atlasmart-rs27 --bind_ip_all --keyFile /data/keyfile --fork --logpath /tmp/rs1-auth.logmongod --dbpath /data/rs2 --port 27210 --replSet atlasmart-rs27 --bind_ip_all --keyFile /data/keyfile --fork --logpath /tmp/rs2-auth.logmongod --dbpath /data/rs3 --port 27211 --replSet atlasmart-rs27 --bind_ip_all --keyFile /data/keyfile --fork --logpath /tmp/rs3-auth.log'sleep 5mongosh 'mongodb://atlasmart_lab_admin:lab-admin-change-me@localhost:27209,localhost:27210,localhost:27211/admin?replicaSet=atlasmart-rs27&authSource=admin' --quiet --eval 'printjson(rs.status().members.map(m=>({name:m.name,state:m.stateStr})))'
from pymongo import MongoClient, ASCENDING, DESCENDINGuri="mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart&w=majority&retryWrites=true"client=MongoClient(uri,serverSelectionTimeoutMS=5000,maxPoolSize=16,waitQueueTimeoutMS=1000,timeoutMS=3000)db=client.atlasmartorders=db.orders_capstoneorders.drop()orders.insert_many([{"_id":i,"tenantId":f"tenant-{i%20:02d}","status":"paid" if i%3 else "open","createdSeq":i,"totalCents":1200+i%9000} for i in range(30000)])orders.create_index([("tenantId",ASCENDING),("status",ASCENDING),("createdSeq",DESCENDING)],name="idx_tenant_status_created")print(db.command("ping"))print(orders.count_documents({}))client.close()
3. Load test with explicit offered load and percentiles
from concurrent.futures import ThreadPoolExecutorfrom statistics import medianfrom time import perf_counterfrom pymongo import MongoClient, DESCENDINGfrom pymongo.errors import PyMongoErroruri="mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart&w=majority&retryWrites=true"client=MongoClient(uri,maxPoolSize=16,waitQueueTimeoutMS=750,serverSelectionTimeoutMS=4000,timeoutMS=2500)c=client.atlasmart.orders_capstonedef percentile(xs,p): xs=sorted(xs); return xs[min(len(xs)-1,round((len(xs)-1)*p))]def op(i): t=perf_counter() try: if i%5: list(c.find({"tenantId":f"tenant-{i%20:02d}","status":"paid"},{"_id":1,"createdSeq":1}).sort("createdSeq",DESCENDING).limit(20)) else: c.update_one({"_id":i%30000},{"$inc":{"touches":1}}) return (perf_counter()-t)*1000,None except PyMongoError as e: return (perf_counter()-t)*1000,type(e).__name__N=4000started=perf_counter()with ThreadPoolExecutor(max_workers=32) as ex: rows=list(ex.map(op,range(N)))elapsed=perf_counter()-startedlat=[x for x,e in rows]; errors=[e for x,e in rows if e]print({"offered":N,"completed":N-len(errors),"errors":len(errors),"seconds":round(elapsed,3), "ops_per_s":round((N-len(errors))/elapsed,1),"p50_ms":round(percentile(lat,.50),2), "p95_ms":round(percentile(lat,.95),2),"p99_ms":round(percentile(lat,.99),2),"max_ms":round(max(lat),2)})client.close()
Run several warm-up and measured iterations, record environment and error distribution, and compare like-for-like before/after changes. This harness is intentionally finite; it does not claim to be a coordinated-omission-proof open-loop benchmark. For production performance certification, use the Chapter 26 scheduled-arrival methodology and include scheduler lag.
4. Fail the current primary and measure recovery
PRIMARY_PORT=$(docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'print(rs.status().members.find(m=>m.stateStr==="PRIMARY").name.split(":")[1])')echo "primary port before controlled failover: $PRIMARY_PORT"# Safe disposable failure injection: force the current primary to step down.docker exec atlasmart-rs27-lab mongosh --quiet --port "$PRIMARY_PORT" --eval 'try { rs.stepDown(20) } catch(e) { print(e.codeName || e.message) }' || true# Poll all three members. Record when exactly one reports isWritablePrimary=true.for i in $(seq 1 30); do date -u +%FT%TZ for p in 27209 27210 27211; do printf "%s " "$p" mongosh --quiet --host 127.0.0.1 --port "$p" --eval 'print(db.hello().isWritablePrimary)' 2>/dev/null || true done sleep 1done# Run the same application probe/load script during/after the election and record errors/latency.# Recovery is complete when all three members are healthy again and a majority write succeeds.
Do not promise a fixed election duration. The evidence is the measured interruption, driver error/selection behavior, first successful post-election majority write, and final replica-set health.
5. Security and least-privilege gate
The capstone user has readWrite only on
atlasmart. This verifies application authorization
but does not replace the Chapter 22 production requirements for
TLS, internal member authentication, secret rotation, network
isolation, administrative separation, and auditing where
applicable.
mongosh 'mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart' --quiet --eval 'print(db.orders_capstone.countDocuments({tenantId:"tenant-00"}))'# Administrative operation should fail for the application identity:mongosh 'mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/admin?replicaSet=atlasmart-rs27&authSource=atlasmart' --quiet --eval 'printjson(db.runCommand({replSetGetConfig:1}))' || echo 'expected authorization failure'# Admin identity can inspect replica configuration:mongosh 'mongodb://atlasmart_lab_admin:lab-admin-change-me@localhost:27209,localhost:27210,localhost:27211/admin?replicaSet=atlasmart-rs27&authSource=admin' --quiet --eval 'print(db.runCommand({replSetGetConfig:1}).ok)'# Verify Docker publishes only 127.0.0.1 ports:docker ps --filter name=atlasmart-rs27-lab --format '{{.Names}} {{.Ports}}'
Do not deploy the lab password. Require TLS with hostname validation, internal cluster authentication, secret-manager delivery/rotation, least privilege, private networking/firewall controls, and protected audit/log pipelines as specified in Chapter 22.
6. Tune one thing, then re-measure
Capture explain("executionStats"), pool/command
telemetry, cache/disk evidence, and load percentiles. Change
exactly one reviewed factor—for example, hide an unused index,
add a proven compound index, reduce an unbounded application
batch, or cap concurrency—then rerun the same workload. A tuning
change is accepted only if its target metric improves without
violating write latency, storage, correctness, or recovery
gates.
from pymongo import MongoClienturi="mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart"client=MongoClient(uri,serverSelectionTimeoutMS=5000,timeoutMS=3000)c=client.atlasmart.orders_capstoneassert c.count_documents({})==30000assert c.count_documents({"tenantId":{"$not":{"$regex":"^tenant-[0-9]{2}$"}}})==0print("documents",c.count_documents({}),"indexes",[i["name"] for i in c.list_indexes()])print(client.admin.command({"getParameter":1,"featureCompatibilityVersion":1}))client.close()
7. Final architecture/runbook dossier
| Artifact | Minimum contents |
|---|---|
| Architecture decision record | topology, concerns, shard/index strategy, security, backup, Search/encryption dependencies |
| Driver contract | URI/options, pool/deadline/retry rules, error taxonomy, idempotency strategy |
| Performance report | dataset, workload, concurrency/offered load, p50/p95/p99, errors, environment |
| Failure report | failure injection, election/recovery timeline, client-visible impact, invariant checks |
| Recovery runbook | backup/PIT sources, keys/secrets, restore order, RPO/RTO and cutover tests |
| Security evidence | positive/negative authorization tests, TLS/network/secret/audit controls |
| Operations runbook | SLO alerts, dashboards, first evidence, containment, stop/rollback conditions |
Check your understanding
- Why is three containers on one laptop not production HA?
- What should be measured during primary failure?
- Why is readWrite not enough for production security?
- How should a tuning change be validated?
- What is the final proof of this course?
Review the answers
1. They share host, power, disk/network and often failure domain; the lab proves mechanisms, not independent infrastructure.
2. Client-visible errors/latency, server selection/election timing, first successful post-election majority write, and final replica health.
3. Transport identity/confidentiality, internal member auth, secret management, private networking, admin separation, and logging/auditing are separate controls.
4. Change one factor, rerun the same workload, verify the target metric and guardrail metrics/correctness.
5. A reproducible evidence package showing the system can meet invariants, fail/recover, restore, secure, observe, and be operated from written runbooks.
8. Course completion and production judgment
You now have the complete MongoDB path from BSON and data modeling through queries, indexes, transactions, replication, consistency, sharding, CDC, time-series, geospatial, Search/Vector Search, security, encryption, WiredTiger, backup, operations, and resilient application drivers. The final standard is not memorizing commands. It is being able to state an invariant, choose a mechanism, observe what MongoDB and the driver actually did, inject a safe failure, recover, verify, and explain the production tradeoff. Re-run the capstone whenever server, driver, topology, schema, shard key, security, Search/encryption mode, or infrastructure changes materially.
Authoritative references
Driver defaults and deployment behavior evolve. Re-check the exact server patch, PyMongo release, topology, and managed-service tier before freezing production assumptions.
- PyMongo driver documentation
- Connect to MongoDB with PyMongo
- PyMongo connection pools
- PyMongo client-side operation timeout
- PyMongo monitoring
- PyMongo CRUD configuration / retries
- PyMongo transactions
- PyMongo bulk writes
- PyMongo release notes
- MongoDB connection strings
- Connection string options
- Retryable writes
- Retryable reads
- Transactions
- Change streams
- cursor.skip() and range pagination
- Explain results
- Read concern
- Write concern
- Read preference
- Replica sets
- Sharding
- Choose a shard key
- Security checklist
- Backup methods
- serverStatus
- MongoDB 8.3 release notes
- mongosh changelog