Separate acknowledgement, journal persistence, checkpoint persistence, and crash recovery so durability claims are tied to the exact write concern and storage timeline.
Journaling and Checkpoints: Durability Timeline, Crash Recovery, and Write Concern Interaction
AtlasMart receives a successful checkout write immediately before a host crash. The team must know what “acknowledged,” “journaled,” and “checkpointed” mean, which writes crash recovery can replay, and why a checkpoint is not the same thing as a backup.
Learning objectives
Draw the durability timeline from in-memory modification to journal sync, checkpoint, and crash recovery.
Explain how j:true, w:1, and
w:"majority" interact without equating them.
Observe documented WiredTiger log and checkpoint metrics around writes.
Run a safe disposable hard-stop recovery drill that proves a journaled marker survives without claiming an unjournaled marker must be lost.
Avoid changing journal/checkpoint intervals as a first-line performance tweak.
This lesson pins
MongoDB Community Server 8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and
PyMongo 4.17.0 where a driver workload is useful.
Topology: disposable one-member replica set. The host publishes
only 127.0.0.1:27187. Authentication and TLS are
disabled only for this isolated disposable lab; production
security remains the Chapter 22 prerequisite. Default read/write
concern and primary read preference are used unless a comparison
says otherwise.
FCV is observed and never changed.
Atlas/Search/Vector Search/KMS/Enterprise capabilities are not
required. WiredTiger internals are treated as version-sensitive
implementation details; use supported MongoDB commands and
metrics instead of editing .wt files or
undocumented knobs. The one-member set is used only to expose
replica-set write-concern semantics; it has no production high
availability and no meaningful multi-member durability. Product
runtime labs were not executed in the generation environment, so
cache ratios, checkpoint durations, journal sync times, disk
bytes, and latency percentiles must be measured locally rather
than copied as invented values.
1. Four different moments can be hidden behind the word “written”
For AtlasMart, an application acknowledgement, a journal flush, a checkpoint, and replica-set majority acknowledgement answer different questions. WiredTiger uses a write-ahead log (the MongoDB journal) together with checkpoints. A checkpoint makes a consistent snapshot of data files. The journal records modifications that occurred after the last checkpoint so an unclean restart can replay durable records. A client’s write concern controls when the server replies; it does not redefine what a checkpoint is.
| Event | What it means | What it does not mean |
|---|---|---|
| Operation accepted in memory | The primary/storage engine has processed the update | Not necessarily journal-synced or majority replicated |
j:true acknowledged |
The write’s journal record has been flushed as required before reply | Not a backup; does not mean every data-file page is checkpointed |
w:"majority" acknowledged |
Replica-set majority acknowledgement condition is
satisfied; with
writeConcernMajorityJournalDefault:true,
majority implies journaling
|
Not linearizable read semantics and not a multi-region DR guarantee |
| Checkpoint completed | WiredTiger has a consistent on-disk checkpoint | Does not eliminate the need for journal entries after that checkpoint or independent backup |
| Crash recovery completed | WiredTiger restored a consistent checkpoint and replays durable journal work as needed |
Does not promise that an unjournaled
j:false write was lost; it may have reached
disk anyway
|
2. Start the replica-set lab and inspect the actual defaults instead of memorizing them
docker rm -f atlasmart-ch24-l2 2>/dev/null || truedocker volume rm atlasmart-ch24-l2-db 2>/dev/null || truedocker run -d --name atlasmart-ch24-l2 \ -p 127.0.0.1:27187:27017 \ -v atlasmart-ch24-l2-db:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim \ --bind_ip_all --port 27017 --replSet rs24until docker exec atlasmart-ch24-l2 mongosh --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donedocker exec atlasmart-ch24-l2 mongosh --quiet --eval ' try { rs.initiate({_id:"rs24",members:[{_id:0,host:"localhost:27017"}]}) } catch(e) { print(e.codeName) }'until docker exec atlasmart-ch24-l2 mongosh --quiet --eval 'db.hello().isWritablePrimary' 2>/dev/null | grep -q true; do sleep 1; donemongosh "mongodb://127.0.0.1:27187/admin?directConnection=true" --quiet --eval ' printjson(db.adminCommand({getParameter:1,featureCompatibilityVersion:1})); printjson(rs.status().members.map(m=>({name:m.name,stateStr:m.stateStr})));'
const admin=db.getSiblingDB("admin");printjson(admin.runCommand({getParameter:1,syncdelay:1,journalCommitInterval:1}));const cfg=rs.conf();printjson({writeConcernMajorityJournalDefault:cfg.writeConcernMajorityJournalDefault});const s=admin.serverStatus();printjson({ log:{ bytes:s.wiredTiger.log["log bytes written"], syncOps:s.wiredTiger.log["log sync operations"], syncMicros:s.wiredTiger.log["log sync time duration (usecs)"] }, checkpointMs:s.wiredTiger.transaction["transaction checkpoint most recent time (msecs)"]});
The documented defaults are a 60-second checkpoint/flush
interval (syncdelay) and a maximum 100 ms journal
commit interval, but the lesson reads the running server rather
than assuming a future release keeps those values. A write that
includes or implies j:true can trigger journal
synchronization sooner.
3. Compare acknowledgement policies with measured distributions
python -m venv /tmp/atlasmart-ch24-journal-venv. /tmp/atlasmart-ch24-journal-venv/bin/activatepython -m pip install --disable-pip-version-check "pymongo==4.17.0"python - <<'PY'from statistics import medianfrom time import perf_counterfrom pymongo import MongoClient, WriteConcernclient=MongoClient("mongodb://127.0.0.1:27187/?directConnection=true")base=client.atlasmart.get_collection("durability_ch24_l2")base.drop()policies={ "w1-jfalse": WriteConcern(w=1,j=False), "w1-jtrue": WriteConcern(w=1,j=True), "majority": WriteConcern(w="majority",wtimeout=5000),}for label,wc in policies.items(): c=base.with_options(write_concern=wc); samples=[] for i in range(40): t=perf_counter(); c.insert_one({"_id":f"{label}-{i}","policy":label,"seq":i}); samples.append((perf_counter()-t)*1000) q=sorted(samples) print(label,{"p50_ms":median(q),"p95_ms":q[int(0.95*(len(q)-1))],"max_ms":max(q)})PY
These tiny local timings are not a production benchmark. They reveal only the acknowledgement cost on this machine/topology. Record disk type, container limits, cache warmth, concurrency, and journal settings before comparing environments.
4. Safe crash-recovery drill: prove the guaranteed case, do not fabricate the uncertain one
from pymongo import MongoClient, WriteConcernclient=MongoClient("mongodb://127.0.0.1:27187/?directConnection=true")base=client.atlasmart.get_collection("crash_ch24_l2")base.drop()base.with_options(write_concern=WriteConcern(w=1,j=True)).insert_one({"_id":"journaled","claim":"must be durable when acknowledgement returns"})base.with_options(write_concern=WriteConcern(w=1,j=False)).insert_one({"_id":"not-proven","claim":"acknowledgement alone did not require a journal sync"})print(list(base.find({}, {"_id":1,"claim":1})))
docker kill --signal=KILL atlasmart-ch24-l2docker start atlasmart-ch24-l2until mongosh "mongodb://127.0.0.1:27187/admin?directConnection=true" --quiet --eval 'db.runCommand({ping:1}).ok' 2>/dev/null | grep -q 1; do sleep 1; donemongosh "mongodb://127.0.0.1:27187/atlasmart?directConnection=true" --quiet --eval ' printjson(db.crash_ch24_l2.find().sort({_id:1}).toArray()); printjson(db.getSiblingDB("admin").serverStatus().wiredTiger.log);'
The expected invariant is that the
journaled marker remains after an acknowledged
j:true write and recovery. The
not-proven marker may also survive because the
operating system/storage engine can persist it before the
crash. Its survival does not upgrade
j:false into a durability guarantee.
5. Checkpoint pressure and journal pressure are related but distinguishable
A rising
transaction checkpoint most recent time (msecs)
under steady load can indicate storage saturation, but a single
slow checkpoint is not enough evidence. Journal sync time
reflects synchronous durability work; dirty cache growth and
pages written from cache provide context for
checkpoint/reconciliation pressure. Host-level device latency
and queue depth are still required to prove a disk bottleneck.
6. Production judgment
Do not tune syncdelay or
journalCommitInterval to solve generic slowness.
MongoDB documentation explicitly advises leaving the checkpoint
interval alone in almost every production situation. Choose
write concern from the business invariant and failure model,
then measure the resulting latency. A checkpoint is a local
storage consistency mechanism, not a backup or replica. The next
lesson compares compression strategies and shows why smaller
storage can reduce I/O yet still increase CPU.
Check your understanding
-
What does
j:trueadd to aw:1write? - Does a checkpoint make the journal unnecessary?
-
Why can a
j:falsemarker survive the crash lab? - When does
w:"majority"imply journaling? - Why is changing the checkpoint interval a poor first response to slow writes?
Review the answers
1. The acknowledgement waits for the relevant journal persistence condition instead of only the primary acknowledgement.
2. No. The journal covers durable modifications after the last checkpoint and supports recovery between checkpoints.
3. The write may have reached durable storage anyway; lack of a journaling requirement means lack of a guarantee, not guaranteed loss.
4. When the replica-set setting
writeConcernMajorityJournalDefault is true,
which is the normal WiredTiger configuration and should be
verified.
5. It changes durability/I/O behavior without proving the root cause; first correlate journal, dirty-cache, checkpoint, host-I/O, and latency evidence.
Authoritative references
WiredTiger metrics and internal field names are implementation- and version-sensitive. The lesson uses documented MongoDB interfaces for evidence and requires re-checking the current server manual before relying on exact metric names or defaults in a later release.
- WiredTiger Storage Engine
- Storage FAQ
- serverStatus
- Self-Managed Diagnostics FAQ
- Journaling
- Configure Journaling
- Write Concern
- Write Operation Performance
- Configuration File Options
- Server Parameters
- mongod Options
- db.collection.stats()
- $collStats
- dbStats
- db.createCollection() Storage Engine Options
- Create Indexes and Storage Engine Options
- Performance Tuning
- Production Notes
- Log Messages
- MongoDB 8.3 Release Notes
- mongosh Release Notes
- PyMongo Release Notes