Treat transactions as bounded coordination: inspect lifetime and lock limits, safe failure modes, operation restrictions, and the extra coordination required when writes span shards.
Production Limits: Runtime, Locking, Sharded Costs, DDL Restrictions, and Abort Scenarios
Bound transaction runtime and lock behavior, test safe aborts/restrictions, and expose why multi-shard commit coordination is operationally more expensive.
Learning objectives
Read the transaction lifetime and lock-wait limits from the server instead of relying on folklore.
Produce a controlled lock/write-conflict failure and verify which transaction commits.
Distinguish supported CRUD work from restricted aggregation/DDL/parallel operations inside transactions.
Explain why long or large transactions consume more cache, history, lock, and retry budget.
Model the extra participant/coordinator path of multi-shard commits without pretending a single-node lab proves sharded behavior.
This lesson pins MongoDB Community Server
8.3.8 with
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo
4.17.0 where a driver example is used. Transactions
require a replica set or sharded cluster, so the mandatory local
lab runs a disposable replica set published only on loopback
beginning at 127.0.0.1:27082. Authentication and
TLS are disabled only for this isolated learning topology.
Feature Compatibility Version (FCV) is observed and never
changed. Unless a section explicitly overrides them, read/write
concern and read preference use the transaction/client defaults
described beside the example. Atlas, Search, Vector Search, KMS,
and Enterprise Advanced are not mandatory. The mandatory lab is
a one-member replica set for safe lock/restriction experiments.
Cross-shard coordination is taught with current server semantics
and a deterministic conceptual simulation; reproducing a full
sharded cluster is optional rather than mandatory. The product
commands were not executed in this generation environment
because Docker, mongod, mongosh, and PyMongo are unavailable
here; measured output must be recorded on the learner's machine.
1. Transactions are bounded coordination, not an unlimited batch primitive
MongoDB's default
transactionLifetimeLimitSeconds is 60 seconds. A
periodic cleanup process aborts expired transactions. By
default, a transaction waits up to
maxTransactionLockRequestTimeoutMillis=5
milliseconds for required locks before it aborts. Those are
defaults, not universal tuning targets; production changes
require workload evidence and, in Atlas, server-parameter
changes may require service-specific controls/support.
Large transactions also retain more WiredTiger history/cache state. Since MongoDB writes separate oplog entries as needed, there is no old 16 MiB total transaction oplog-entry limit, but each individual oplog entry remains subject to the 16 MiB BSON document limit.
docker rm -f atlasmart-mongo-ch13-l4 2>/dev/null || truedocker volume rm atlasmart-mongo-ch13-l4-data 2>/dev/null || truedocker run -d --name atlasmart-mongo-ch13-l4 \ -p 127.0.0.1:27082:27017 \ -v atlasmart-mongo-ch13-l4-data:/data/db \ mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs13-l4 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27082/admin?directConnection=true" --quiet --eval \'quit(db.runCommand({ping:1}).ok === 1 ? 0 : 1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27082/admin?directConnection=true" --quiet --eval \'rs.initiate({_id:"atlasmart-rs13-l4",members:[{_id:0,host:"localhost:27017"}]})'until mongosh "mongodb://127.0.0.1:27082/admin?directConnection=true&replicaSet=atlasmart-rs13-l4" --quiet --eval \'quit(db.hello().isWritablePrimary ? 0 : 1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27082/admin?directConnection=true&replicaSet=atlasmart-rs13-l4" --quiet --eval \'printjson({server:db.version(),setName:db.hello().setName,primary:db.hello().isWritablePrimary}); printjson(db.runCommand({getParameter:1,featureCompatibilityVersion:1,transactionLifetimeLimitSeconds:1,maxTransactionLockRequestTimeoutMillis:1}))'
2. Observe actual limits and create one hot document
const c=db.locks_ch13_l4;c.drop();c.insertOne({_id:"hot-doc",tenantId:"tenant-a",counter:0});printjson(db.getSiblingDB("admin").runCommand({ getParameter:1, transactionLifetimeLimitSeconds:1, maxTransactionLockRequestTimeoutMillis:1}));
3. Controlled lock/write-conflict experiment
Session 1 updates the document and remains open. Session 2 then tries to update the same document. Depending on the exact conflict path, the second operation should fail/abort rather than silently interleave both transaction writes. The first session commits, and the final counter should reflect only its increment.
const s1=db.getMongo().startSession();const s2=db.getMongo().startSession();const c1=s1.getDatabase("atlasmart").locks_ch13_l4;const c2=s2.getDatabase("atlasmart").locks_ch13_l4;s1.startTransaction({readConcern:{level:"snapshot"},writeConcern:{w:"majority"}});c1.updateOne({_id:"hot-doc"},{$inc:{counter:1}});printjson(db.getSiblingDB("admin").aggregate([ {$currentOp:{allUsers:true,idleSessions:true}}, {$match:{transaction:{$exists:true}}}, {$project:{_id:0,type:1,lsid:1,transaction:1}}]).toArray());s2.startTransaction({readConcern:{level:"snapshot"},writeConcern:{w:"majority"}});try { c2.updateOne({_id:"hot-doc"},{$inc:{counter:100}}); s2.commitTransaction(); print("unexpected: second transaction committed");} catch (e) { print("second transaction failed while first held the conflicting write:",e.code,e.message); try { s2.abortTransaction(); } catch (_) {}}s1.commitTransaction();printjson(db.locks_ch13_l4.findOne({_id:"hot-doc"}));s1.endSession(); s2.endSession();
$currentOp can expose the idle transaction's
lsid, txnNumber, read concern,
open/active/inactive times, and expiry time. On a real sharded
cluster it can also expose two-phase commit coordinator
information for transactions spanning shards.
4. Operation restrictions and DDL nuance
Transactions allow many CRUD operations across documents,
collections, databases, and shards, but not every operation is
legal. Aggregation stages such as $out,
$merge, $currentOp,
$planCacheStats, and several session-listing stages
are prohibited inside transactions. Parallel operations in one
transaction are unsupported.
DDL rules are nuanced rather than “DDL is always forbidden.”
Current MongoDB supports creating collections and indexes in
transactions in supported cases; explicit creation requires
transaction read concern local. A cross-shard write
transaction cannot create a new collection on another shard.
Metadata operations can also wait behind transactions and create
lock contention.
const s=db.getMongo().startSession();const sdb=s.getDatabase("atlasmart");s.startTransaction({readConcern:{level:"snapshot"},writeConcern:{w:"majority"}});try { sdb.locks_ch13_l4.aggregate([{$set:{copy:1}},{$out:"should_not_exist_ch13_l4"}]).toArray(); print("unexpected: $out succeeded inside transaction");} catch (e) { print("expected transaction restriction:",e.code,e.message); try { s.abortTransaction(); } catch (_) {}}print("target exists:",db.getCollectionNames().includes("should_not_exist_ch13_l4"));s.endSession();
5. Why sharded transactions cost more coordination
A replica-set transaction coordinates one replica-set participant. A transaction whose writes span shards has multiple participants and a transaction coordinator that drives a distributed commit decision. That introduces network hops, participant availability dependencies, coordinator state/recovery, and chunk-migration interactions. The following code is a conceptual state-path model only; it does not claim to execute MongoDB’s internal protocol.
function coordination(participants) { if (participants <= 1) { return {participants,conceptualPath:["execute writes","commit replica-set transaction"]}; } return { participants, conceptualPath:["execute writes on participant shards","prepare/coordinate decision","persist commit decision","notify participants","return according to commit semantics"], extraFailureSurface:["participant availability","coordinator recovery","chunk migration interaction","cross-shard latency"] };}printjson(coordination(1));printjson(coordination(3));
On real sharded clusters,
$currentOp.twoPhaseCommitCoordinator can expose
participant counts, coordinator state, and decision progress.
Chunk migration can block or abort transactions depending on
interleaving. Outside reads with snapshot/linearizable
or causal afterClusterTime behavior wait
appropriately during commit; weaker outside reads may continue
seeing before-transaction versions.
Production judgment. Keep transactions short
and narrow. Alert on rising abort/retry rates, long
timeOpenMicros, transaction cache pressure, lock
waits, and coordinator latency. Avoid changing global timeout
parameters just to mask a modeling or contention problem. Test
DDL deployment windows separately from transactional traffic. On
shards, measure transaction participant count and routing; a
transaction that unexpectedly fans out can turn an ordinary
write path into a distributed coordination hotspot.
Bridge. Lesson 5 takes a transaction-heavy order model and reduces the number of independent documents that need one commit.
docker rm -f atlasmart-mongo-ch13-l4docker volume rm atlasmart-mongo-ch13-l4-data
Check your understanding
- What is the default transaction lifetime limit?
- How long do transactions wait for required locks by default?
-
Can
$outrun inside a multi-document transaction? - Is all DDL categorically forbidden inside transactions?
- What additional component appears for a transaction whose writes span shards?
Review the answers
1. 60 seconds, controlled by
transactionLifetimeLimitSeconds.
2. Up to 5 ms by default, controlled by
maxTransactionLockRequestTimeoutMillis.
3. No.
4. No. Some collection/index creation is
supported in defined cases, with restrictions such as read
concern local; cross-shard write transactions
cannot create a new collection on another shard.
5. A distributed transaction coordinator and multiple shard participants, increasing latency and failure/operational surface.
Authoritative references
- MongoDB 8.3 release notes — Current 8.3 release line and patch-sensitive server behavior.
- Atomicity and transactions — Single-document atomicity and guidance to minimize unnecessary distributed transactions.
- Transactions — Sessions, transaction read/write concern, read preference, and transaction semantics.
- Drivers API for transactions — Callback versus core APIs and retry labels for transient transactions and ambiguous commits.
- Transaction production considerations — Runtime, locking, cache, DDL, conflicts, and operational constraints.
- Sharded transaction considerations — Cross-shard snapshot semantics, commit coordination, migrations, and outside reads.
- Transactions and operations — Operations permitted and prohibited inside transactions.
- $currentOp — Session/transaction observability including lsid, txnNumber, timing, and sharded coordinators.
- PyMongo transactions — PyMongo session, with_transaction, retry, and callback behavior.
- PyMongo release notes — Current PyMongo 4.17 behavior and session APIs.
- mongosh release notes — Current mongosh 2.10.0 release baseline.
- Server parameters — Current default lock-request timeout and transaction parameters.