Production design is an explicit set of invariants, failure budgets, and operational responsibilities.
Production Design: Topology, Concerns, Shard Key, Index Budget, Backups, Security, Monitoring, and Cost
Defend a production MongoDB architecture by tying topology, consistency, sharding, index budget, recovery, security, observability, and cost to explicit AtlasMart invariants.
Learning objectives
Turn application invariants and SLOs into topology, concern, indexing, sharding, recovery, and security decisions.
Budget secondary indexes from read benefit plus write/storage/cache cost rather than collecting speculative indexes.
Choose a shard key from tenant isolation, routing, cardinality, frequency, monotonicity, and write distribution evidence.
Define recovery, security, observability, and operating-cost ownership before calling a system production-ready.
Produce a concise architecture decision record and acceptance checklist that an operator can defend.
This final chapter pins
MongoDB Community Server 8.3.8 using
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0.
Labs use AtlasMart synthetic data, loopback-only Docker port
publishing, replica set name atlasmart-rs27 where
topology behavior matters, and explicit operation time budgets.
Authentication/TLS are disabled only for disposable mechanism
labs; the capstone security gate reuses Chapter 22
least-privilege/authentication requirements and treats TLS as
mandatory production acceptance. Default read preference is
primary and majority acknowledgement is used for business writes
unless a failure experiment explicitly states otherwise.
FCV is observed and never changed. Atlas,
Search, Vector Search, Enterprise Advanced, and KMS are optional
and are not required for mandatory work. Runtime load, failover,
pool, latency, retry, and recovery results were not executed in
the generation environment; learners must record their own
evidence instead of copying invented values.
1. Start with invariants, not products
| AtlasMart invariant/SLO | Design consequence |
|---|---|
| No tenant may read another tenant's orders | Tenant predicate enforced in application authorization; tenant-aware indexes; security tests |
| Accepted order must survive a single-node failure | Replica set; majority business writes; driver failover/retry policy |
| Checkout p99 has a defined budget | Explicit client timeout; bounded pool; index-backed query shapes; load-test gate |
| Restore must meet measured RPO/RTO | Tested backup/PIT strategy; restore drills; key material included |
| Horizontal scale must preserve targeted tenant queries | Shard key begins with/routs by tenant where workload evidence supports it |
2. Reference production architecture
A plausible AtlasMart self-managed design is a three-member
replica set per shard, a three-member config-server replica set,
multiple mongos routers behind application
infrastructure, authenticated/TLS traffic, tenant-aware shard
keys, explicit read/write concerns per operation, independent
backups, and centralized metrics/logs. The exact node count,
regions, hardware, Atlas tier, and shard count are capacity
decisions—not curriculum constants.
| Concern | Decision record should state | Evidence before launch |
|---|---|---|
| Topology | replica/sharded layout and failure domains | election/failover drill |
| Concerns | which operations use majority/snapshot/local etc. | consistency/failure tests |
| Shard key | key and routing assumptions | analyzeShardKey/workload evidence |
| Indexes | approved index budget and owners | explain + write/storage impact |
| Backup | snapshot/PIT mechanism and key dependencies | clean-target restore drill |
| Security | authN/authZ/TLS/network/secret/audit boundaries | positive + negative access tests |
| Monitoring | SLOs, alerts, runbooks, on-call ownership | synthetic incident drill |
| Cost | compute/storage/egress/search/backup/operator effort | capacity model and growth scenarios |
3. Shard-key and index-budget example
For tenant-dominated access, a compound key such as
{tenantId:1, orderId:"hashed"} can preserve tenant
targeting while distributing writes inside a tenant, but it is
not universally correct: cross-tenant analytics may broadcast,
zones may require different key structure, and very large
tenants can still dominate workload. Evaluate candidate keys
with the Chapter 17 evidence loop rather than copying the
example.
from dataclasses import dataclass@dataclassclass Assumptions: docs:int=20_000_000 avg_doc_bytes:int=1800 secondary_index_bytes_per_doc:int=420 secondary_index_count:int=5 growth_per_month:float=.12a=Assumptions()data=a.docs*a.avg_doc_bytesindexes=a.docs*a.secondary_index_bytes_per_doc*a.secondary_index_countprint("logical data GiB",round(data/2**30,1))print("estimated secondary-index GiB",round(indexes/2**30,1))print("next-month logical multiplier",1+a.growth_per_month)print("MODEL ONLY: replace with measured compression/index/cache/replication/backup factors")
4. Deliberately wrong: architecture by feature checklist
“Replica set + sharding + Search + encryption + ten indexes + five retries” is not an architecture. More features can increase failure modes, write amplification, cost, and operational skill requirements. The repair is traceability: each component must serve an invariant or measured bottleneck, have an owner, have observability, and have a rollback/recovery story.
| Question | Reject the design if… |
|---|---|
| Can a new engineer operate it from the runbook? | critical commands, dependencies, or stop conditions are tribal knowledge |
| Can failures be reproduced safely? | no election/restore/security/load drill exists |
| Can costs be explained? | no growth, backup, index, Search, egress, or staffing assumptions exist |
| Can changes be rolled back? | binary/FCV/schema/index/shard-key/key changes have no compatibility plan |
5. Production acceptance checklist
- Driver compatibility and timeout/retry policy tested against the pinned server line.
- Major query shapes have executionStats evidence and index ownership.
- Read/write concerns map to named business invariants.
- Replica/shard failure drills and application recovery are measured.
- Backup is proven by restore, including encryption/KMS dependencies.
- Least privilege, TLS, network exposure, secret rotation, and audit/log handling are verified.
- Capacity headroom and growth assumptions are recorded.
- SLO alerts link to evidence-first incident runbooks.
Check your understanding
- Why should architecture start with invariants?
- Why is an index budget operationally important?
- Does sharding automatically improve every query?
- What proves a backup design?
- Why include operator skill in architecture cost?
Review the answers
1. They define what the system must preserve; features are only mechanisms for meeting those requirements.
2. Indexes consume storage/cache and amplify writes, so speculative indexes can degrade the workload they were meant to help.
3. No. Queries without usable shard-key information may scatter/gather and coordination adds operational cost.
4. A successful clean-target restore with integrity/application checks and measured RPO/RTO, not the existence of backup files.
5. Complex topologies and security/recovery systems require people who can diagnose and recover them safely.
6. Production judgment
The production design is a set of defended tradeoffs, not a reference diagram copied into every company. Keep the decision record tied to workload evidence and recovery/security ownership. Lesson 5 turns this design into the final acceptance exercise: build it, load it, fail it, recover it, secure it, tune it, and defend the result.
Authoritative references
Driver defaults and deployment behavior evolve. Re-check the exact server patch, PyMongo release, topology, and managed-service tier before freezing production assumptions.
- PyMongo driver documentation
- Connect to MongoDB with PyMongo
- PyMongo connection pools
- PyMongo client-side operation timeout
- PyMongo monitoring
- PyMongo CRUD configuration / retries
- PyMongo transactions
- PyMongo bulk writes
- PyMongo release notes
- MongoDB connection strings
- Connection string options
- Retryable writes
- Retryable reads
- Transactions
- Change streams
- cursor.skip() and range pagination
- Explain results
- Read concern
- Write concern
- Read preference
- Replica sets
- Sharding
- Choose a shard key
- Security checklist
- Backup methods
- serverStatus
- MongoDB 8.3 release notes
- mongosh changelog