Production design is an explicit set of invariants, failure budgets, and operational responsibilities.

Production Design: Topology, Concerns, Shard Key, Index Budget, Backups, Security, Monitoring, and Cost

Defend a production MongoDB architecture by tying topology, consistency, sharding, index budget, recovery, security, observability, and cost to explicit AtlasMart invariants.

Advanced120–240 minutesDriver/capstone labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Turn application invariants and SLOs into topology, concern, indexing, sharding, recovery, and security decisions.

02

Budget secondary indexes from read benefit plus write/storage/cache cost rather than collecting speculative indexes.

03

Choose a shard key from tenant isolation, routing, cardinality, frequency, monotonicity, and write distribution evidence.

04

Define recovery, security, observability, and operating-cost ownership before calling a system production-ready.

05

Produce a concise architecture decision record and acceptance checklist that an operator can defend.

Reproducible lab baseline

This final chapter pins MongoDB Community Server 8.3.8 using mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0. Labs use AtlasMart synthetic data, loopback-only Docker port publishing, replica set name atlasmart-rs27 where topology behavior matters, and explicit operation time budgets. Authentication/TLS are disabled only for disposable mechanism labs; the capstone security gate reuses Chapter 22 least-privilege/authentication requirements and treats TLS as mandatory production acceptance. Default read preference is primary and majority acknowledgement is used for business writes unless a failure experiment explicitly states otherwise. FCV is observed and never changed. Atlas, Search, Vector Search, Enterprise Advanced, and KMS are optional and are not required for mandatory work. Runtime load, failover, pool, latency, retry, and recovery results were not executed in the generation environment; learners must record their own evidence instead of copying invented values.

1. Start with invariants, not products

AtlasMart invariant/SLO Design consequence
No tenant may read another tenant's orders Tenant predicate enforced in application authorization; tenant-aware indexes; security tests
Accepted order must survive a single-node failure Replica set; majority business writes; driver failover/retry policy
Checkout p99 has a defined budget Explicit client timeout; bounded pool; index-backed query shapes; load-test gate
Restore must meet measured RPO/RTO Tested backup/PIT strategy; restore drills; key material included
Horizontal scale must preserve targeted tenant queries Shard key begins with/routs by tenant where workload evidence supports it

2. Reference production architecture

A plausible AtlasMart self-managed design is a three-member replica set per shard, a three-member config-server replica set, multiple mongos routers behind application infrastructure, authenticated/TLS traffic, tenant-aware shard keys, explicit read/write concerns per operation, independent backups, and centralized metrics/logs. The exact node count, regions, hardware, Atlas tier, and shard count are capacity decisions—not curriculum constants.

Concern Decision record should state Evidence before launch
Topology replica/sharded layout and failure domains election/failover drill
Concerns which operations use majority/snapshot/local etc. consistency/failure tests
Shard key key and routing assumptions analyzeShardKey/workload evidence
Indexes approved index budget and owners explain + write/storage impact
Backup snapshot/PIT mechanism and key dependencies clean-target restore drill
Security authN/authZ/TLS/network/secret/audit boundaries positive + negative access tests
Monitoring SLOs, alerts, runbooks, on-call ownership synthetic incident drill
Cost compute/storage/egress/search/backup/operator effort capacity model and growth scenarios

3. Shard-key and index-budget example

For tenant-dominated access, a compound key such as {tenantId:1, orderId:"hashed"} can preserve tenant targeting while distributing writes inside a tenant, but it is not universally correct: cross-tenant analytics may broadcast, zones may require different key structure, and very large tenants can still dominate workload. Evaluate candidate keys with the Chapter 17 evidence loop rather than copying the example.

small architecture calculator
from dataclasses import dataclass@dataclassclass Assumptions:    docs:int=20_000_000    avg_doc_bytes:int=1800    secondary_index_bytes_per_doc:int=420    secondary_index_count:int=5    growth_per_month:float=.12a=Assumptions()data=a.docs*a.avg_doc_bytesindexes=a.docs*a.secondary_index_bytes_per_doc*a.secondary_index_countprint("logical data GiB",round(data/2**30,1))print("estimated secondary-index GiB",round(indexes/2**30,1))print("next-month logical multiplier",1+a.growth_per_month)print("MODEL ONLY: replace with measured compression/index/cache/replication/backup factors")

4. Deliberately wrong: architecture by feature checklist

“Replica set + sharding + Search + encryption + ten indexes + five retries” is not an architecture. More features can increase failure modes, write amplification, cost, and operational skill requirements. The repair is traceability: each component must serve an invariant or measured bottleneck, have an owner, have observability, and have a rollback/recovery story.

Question Reject the design if…
Can a new engineer operate it from the runbook? critical commands, dependencies, or stop conditions are tribal knowledge
Can failures be reproduced safely? no election/restore/security/load drill exists
Can costs be explained? no growth, backup, index, Search, egress, or staffing assumptions exist
Can changes be rolled back? binary/FCV/schema/index/shard-key/key changes have no compatibility plan

5. Production acceptance checklist

  • Driver compatibility and timeout/retry policy tested against the pinned server line.
  • Major query shapes have executionStats evidence and index ownership.
  • Read/write concerns map to named business invariants.
  • Replica/shard failure drills and application recovery are measured.
  • Backup is proven by restore, including encryption/KMS dependencies.
  • Least privilege, TLS, network exposure, secret rotation, and audit/log handling are verified.
  • Capacity headroom and growth assumptions are recorded.
  • SLO alerts link to evidence-first incident runbooks.

Check your understanding

  1. Why should architecture start with invariants?
  2. Why is an index budget operationally important?
  3. Does sharding automatically improve every query?
  4. What proves a backup design?
  5. Why include operator skill in architecture cost?
Review the answers

1. They define what the system must preserve; features are only mechanisms for meeting those requirements.

2. Indexes consume storage/cache and amplify writes, so speculative indexes can degrade the workload they were meant to help.

3. No. Queries without usable shard-key information may scatter/gather and coordination adds operational cost.

4. A successful clean-target restore with integrity/application checks and measured RPO/RTO, not the existence of backup files.

5. Complex topologies and security/recovery systems require people who can diagnose and recover them safely.

6. Production judgment

The production design is a set of defended tradeoffs, not a reference diagram copied into every company. Keep the decision record tied to workload evidence and recovery/security ownership. Lesson 5 turns this design into the final acceptance exercise: build it, load it, fail it, recover it, secure it, tune it, and defend the result.

Authoritative references

Driver defaults and deployment behavior evolve. Re-check the exact server patch, PyMongo release, topology, and managed-service tier before freezing production assumptions.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.