A production claim is credible only when the system can be measured, failed, recovered, secured, and operated from its runbook.

Capstone: Build, Load-Test, Fail, Recover, Secure, Tune, and Defend a Production MongoDB System

Integrate the entire course in a measurable AtlasMart capstone: build, load-test, fail, recover, harden, tune, observe, and defend the system with a runnable acceptance runbook.

Advanced120–240 minutesDriver/capstone labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Assemble a runnable AtlasMart capstone with replica-set topology, indexes, driver configuration, observability, and bounded load.

02

Measure throughput and latency distributions without hiding queueing or errors.

03

Kill the current primary safely, observe election/recovery, and verify business invariants after failover.

04

Apply a least-privilege/authentication security gate and verify denied administrative access.

05

Produce a final architecture/runbook dossier with evidence, limits, recovery steps, and unresolved risks.

Reproducible lab baseline

This final chapter pins MongoDB Community Server 8.3.8 using mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0. Labs use AtlasMart synthetic data, loopback-only Docker port publishing, replica set name atlasmart-rs27 where topology behavior matters, and explicit operation time budgets. Authentication/TLS are disabled only for disposable mechanism labs; the capstone security gate reuses Chapter 22 least-privilege/authentication requirements and treats TLS as mandatory production acceptance. Default read preference is primary and majority acknowledgement is used for business writes unless a failure experiment explicitly states otherwise. FCV is observed and never changed. Atlas, Search, Vector Search, Enterprise Advanced, and KMS are optional and are not required for mandatory work. Runtime load, failover, pool, latency, retry, and recovery results were not executed in the generation environment; learners must record their own evidence instead of copying invented values.

1. Capstone acceptance contract

The final lab deliberately stays small enough for one workstation but uses the same evidence categories as production. It does not claim that three containers on one laptop reproduce independent failure domains, WAN latency, encrypted network transport, or cloud storage behavior. Those remain production acceptance items.

Gate Evidence
Build healthy replica-set topology, indexes, users/config, seed data
Load offered/completed/error counts and p50/p95/p99/max latency
Fail primary stopped, election observed, client recovery measured
Recover old member rejoins; majority writes and application invariants verify
Secure loopback-only exposure plus least-privilege auth test; TLS required in production
Tune one evidence-backed change followed by same workload/re-measurement
Defend runbook, architecture decisions, known limits, rollback and restore paths

2. Build and seed

start/reuse three-member replica set
docker rm -f atlasmart-rs27-lab 2>/dev/null || trueIMAGE=mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimdocker run -d --name atlasmart-rs27-lab \  -p 127.0.0.1:27209:27209 -p 127.0.0.1:27210:27210 -p 127.0.0.1:27211:27211 \  --entrypoint bash "$IMAGE" -lc 'set -emkdir -p /data/rs1 /data/rs2 /data/rs3mongod --dbpath /data/rs1 --port 27209 --replSet atlasmart-rs27 --bind_ip_all --fork --logpath /tmp/rs1.logmongod --dbpath /data/rs2 --port 27210 --replSet atlasmart-rs27 --bind_ip_all --fork --logpath /tmp/rs2.logmongod --dbpath /data/rs3 --port 27211 --replSet atlasmart-rs27 --bind_ip_all --fork --logpath /tmp/rs3.logtail -f /dev/null'sleep 4docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'rs.initiate({_id:"atlasmart-rs27",members:[ {_id:0,host:"localhost:27209"}, {_id:1,host:"localhost:27210"}, {_id:2,host:"localhost:27211"}]})'docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'while (rs.status().members.filter(m=>m.stateStr==="PRIMARY").length!==1 || rs.status().members.filter(m=>m.stateStr==="SECONDARY").length!==2) { sleep(1000) }printjson(rs.status().members.map(m=>({name:m.name,state:m.stateStr})))' 
create an application user through the current primary
PRIMARY_PORT=$(docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'print(rs.status().members.find(m=>m.stateStr==="PRIMARY").name.split(":")[1])')docker exec atlasmart-rs27-lab mongosh --quiet --port "$PRIMARY_PORT" --eval 'use admindb.createUser({user:"atlasmart_lab_admin",pwd:"lab-admin-change-me",roles:[{role:"root",db:"admin"}]})use atlasmartdb.createUser({user:"atlasmart_app",pwd:"lab-only-change-me",roles:[{role:"readWrite",db:"atlasmart"}]})printjson(db.getUser("atlasmart_app"))'# Both passwords are synthetic and disposable. Production secrets belong in a secret manager, never source control.
enable internal + client authentication on the disposable replica set
python -c "import base64,secrets; print(base64.b64encode(secrets.token_bytes(512)).decode())" > rs27.keydocker cp rs27.key atlasmart-rs27-lab:/data/keyfilerm -f rs27.keydocker exec atlasmart-rs27-lab bash -lc 'chmod 400 /data/keyfilemongod --shutdown --dbpath /data/rs1mongod --shutdown --dbpath /data/rs2mongod --shutdown --dbpath /data/rs3mongod --dbpath /data/rs1 --port 27209 --replSet atlasmart-rs27 --bind_ip_all --keyFile /data/keyfile --fork --logpath /tmp/rs1-auth.logmongod --dbpath /data/rs2 --port 27210 --replSet atlasmart-rs27 --bind_ip_all --keyFile /data/keyfile --fork --logpath /tmp/rs2-auth.logmongod --dbpath /data/rs3 --port 27211 --replSet atlasmart-rs27 --bind_ip_all --keyFile /data/keyfile --fork --logpath /tmp/rs3-auth.log'sleep 5mongosh 'mongodb://atlasmart_lab_admin:lab-admin-change-me@localhost:27209,localhost:27210,localhost:27211/admin?replicaSet=atlasmart-rs27&authSource=admin' --quiet --eval 'printjson(rs.status().members.map(m=>({name:m.name,state:m.stateStr})))'
seed_capstone.py
from pymongo import MongoClient, ASCENDING, DESCENDINGuri="mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart&w=majority&retryWrites=true"client=MongoClient(uri,serverSelectionTimeoutMS=5000,maxPoolSize=16,waitQueueTimeoutMS=1000,timeoutMS=3000)db=client.atlasmartorders=db.orders_capstoneorders.drop()orders.insert_many([{"_id":i,"tenantId":f"tenant-{i%20:02d}","status":"paid" if i%3 else "open","createdSeq":i,"totalCents":1200+i%9000} for i in range(30000)])orders.create_index([("tenantId",ASCENDING),("status",ASCENDING),("createdSeq",DESCENDING)],name="idx_tenant_status_created")print(db.command("ping"))print(orders.count_documents({}))client.close()

3. Load test with explicit offered load and percentiles

load_capstone.py
from concurrent.futures import ThreadPoolExecutorfrom statistics import medianfrom time import perf_counterfrom pymongo import MongoClient, DESCENDINGfrom pymongo.errors import PyMongoErroruri="mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart&w=majority&retryWrites=true"client=MongoClient(uri,maxPoolSize=16,waitQueueTimeoutMS=750,serverSelectionTimeoutMS=4000,timeoutMS=2500)c=client.atlasmart.orders_capstonedef percentile(xs,p):    xs=sorted(xs); return xs[min(len(xs)-1,round((len(xs)-1)*p))]def op(i):    t=perf_counter()    try:        if i%5:            list(c.find({"tenantId":f"tenant-{i%20:02d}","status":"paid"},{"_id":1,"createdSeq":1}).sort("createdSeq",DESCENDING).limit(20))        else:            c.update_one({"_id":i%30000},{"$inc":{"touches":1}})        return (perf_counter()-t)*1000,None    except PyMongoError as e:        return (perf_counter()-t)*1000,type(e).__name__N=4000started=perf_counter()with ThreadPoolExecutor(max_workers=32) as ex: rows=list(ex.map(op,range(N)))elapsed=perf_counter()-startedlat=[x for x,e in rows]; errors=[e for x,e in rows if e]print({"offered":N,"completed":N-len(errors),"errors":len(errors),"seconds":round(elapsed,3),       "ops_per_s":round((N-len(errors))/elapsed,1),"p50_ms":round(percentile(lat,.50),2),       "p95_ms":round(percentile(lat,.95),2),"p99_ms":round(percentile(lat,.99),2),"max_ms":round(max(lat),2)})client.close()

Run several warm-up and measured iterations, record environment and error distribution, and compare like-for-like before/after changes. This harness is intentionally finite; it does not claim to be a coordinated-omission-proof open-loop benchmark. For production performance certification, use the Chapter 26 scheduled-arrival methodology and include scheduler lag.

4. Fail the current primary and measure recovery

bounded failure drill
PRIMARY_PORT=$(docker exec atlasmart-rs27-lab mongosh --quiet --port 27209 --eval 'print(rs.status().members.find(m=>m.stateStr==="PRIMARY").name.split(":")[1])')echo "primary port before controlled failover: $PRIMARY_PORT"# Safe disposable failure injection: force the current primary to step down.docker exec atlasmart-rs27-lab mongosh --quiet --port "$PRIMARY_PORT" --eval 'try { rs.stepDown(20) } catch(e) { print(e.codeName || e.message) }' || true# Poll all three members. Record when exactly one reports isWritablePrimary=true.for i in $(seq 1 30); do  date -u +%FT%TZ  for p in 27209 27210 27211; do    printf "%s " "$p"    mongosh --quiet --host 127.0.0.1 --port "$p" --eval 'print(db.hello().isWritablePrimary)' 2>/dev/null || true  done  sleep 1done# Run the same application probe/load script during/after the election and record errors/latency.# Recovery is complete when all three members are healthy again and a majority write succeeds.

Do not promise a fixed election duration. The evidence is the measured interruption, driver error/selection behavior, first successful post-election majority write, and final replica-set health.

5. Security and least-privilege gate

The capstone user has readWrite only on atlasmart. This verifies application authorization but does not replace the Chapter 22 production requirements for TLS, internal member authentication, secret rotation, network isolation, administrative separation, and auditing where applicable.

positive and negative authorization tests
mongosh 'mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart' --quiet --eval 'print(db.orders_capstone.countDocuments({tenantId:"tenant-00"}))'# Administrative operation should fail for the application identity:mongosh 'mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/admin?replicaSet=atlasmart-rs27&authSource=atlasmart' --quiet --eval 'printjson(db.runCommand({replSetGetConfig:1}))' || echo 'expected authorization failure'# Admin identity can inspect replica configuration:mongosh 'mongodb://atlasmart_lab_admin:lab-admin-change-me@localhost:27209,localhost:27210,localhost:27211/admin?replicaSet=atlasmart-rs27&authSource=admin' --quiet --eval 'print(db.runCommand({replSetGetConfig:1}).ok)'# Verify Docker publishes only 127.0.0.1 ports:docker ps --filter name=atlasmart-rs27-lab --format '{{.Names}} {{.Ports}}' 
Production security acceptance:

Do not deploy the lab password. Require TLS with hostname validation, internal cluster authentication, secret-manager delivery/rotation, least privilege, private networking/firewall controls, and protected audit/log pipelines as specified in Chapter 22.

6. Tune one thing, then re-measure

Capture explain("executionStats"), pool/command telemetry, cache/disk evidence, and load percentiles. Change exactly one reviewed factor—for example, hide an unused index, add a proven compound index, reduce an unbounded application batch, or cap concurrency—then rerun the same workload. A tuning change is accepted only if its target metric improves without violating write latency, storage, correctness, or recovery gates.

final invariant probe
from pymongo import MongoClienturi="mongodb://atlasmart_app:lab-only-change-me@localhost:27209,localhost:27210,localhost:27211/atlasmart?replicaSet=atlasmart-rs27&authSource=atlasmart"client=MongoClient(uri,serverSelectionTimeoutMS=5000,timeoutMS=3000)c=client.atlasmart.orders_capstoneassert c.count_documents({})==30000assert c.count_documents({"tenantId":{"$not":{"$regex":"^tenant-[0-9]{2}$"}}})==0print("documents",c.count_documents({}),"indexes",[i["name"] for i in c.list_indexes()])print(client.admin.command({"getParameter":1,"featureCompatibilityVersion":1}))client.close()

7. Final architecture/runbook dossier

Artifact Minimum contents
Architecture decision record topology, concerns, shard/index strategy, security, backup, Search/encryption dependencies
Driver contract URI/options, pool/deadline/retry rules, error taxonomy, idempotency strategy
Performance report dataset, workload, concurrency/offered load, p50/p95/p99, errors, environment
Failure report failure injection, election/recovery timeline, client-visible impact, invariant checks
Recovery runbook backup/PIT sources, keys/secrets, restore order, RPO/RTO and cutover tests
Security evidence positive/negative authorization tests, TLS/network/secret/audit controls
Operations runbook SLO alerts, dashboards, first evidence, containment, stop/rollback conditions

Check your understanding

  1. Why is three containers on one laptop not production HA?
  2. What should be measured during primary failure?
  3. Why is readWrite not enough for production security?
  4. How should a tuning change be validated?
  5. What is the final proof of this course?
Review the answers

1. They share host, power, disk/network and often failure domain; the lab proves mechanisms, not independent infrastructure.

2. Client-visible errors/latency, server selection/election timing, first successful post-election majority write, and final replica health.

3. Transport identity/confidentiality, internal member auth, secret management, private networking, admin separation, and logging/auditing are separate controls.

4. Change one factor, rerun the same workload, verify the target metric and guardrail metrics/correctness.

5. A reproducible evidence package showing the system can meet invariants, fail/recover, restore, secure, observe, and be operated from written runbooks.

8. Course completion and production judgment

You now have the complete MongoDB path from BSON and data modeling through queries, indexes, transactions, replication, consistency, sharding, CDC, time-series, geospatial, Search/Vector Search, security, encryption, WiredTiger, backup, operations, and resilient application drivers. The final standard is not memorizing commands. It is being able to state an invariant, choose a mechanism, observe what MongoDB and the driver actually did, inject a safe failure, recover, verify, and explain the production tradeoff. Re-run the capstone whenever server, driver, topology, schema, shard key, security, Search/encryption mode, or infrastructure changes materially.

Authoritative references

Driver defaults and deployment behavior evolve. Re-check the exact server patch, PyMongo release, topology, and managed-service tier before freezing production assumptions.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.