Automatic database retries do not make arbitrary application side effects exactly once.

Retryable Reads/Writes, Transactions, Idempotency, Command Monitoring, and Error Classification

Make automatic retries, transaction retries, application idempotency, command monitoring, and error classification explicit so transient failures do not become duplicate business side effects.

Advanced120–240 minutesDriver/capstone labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Distinguish automatic retryable reads/writes from application retry loops and transaction callback retries.

02

Classify retryable, duplicate, timeout, validation, and unknown-result errors before choosing an action.

03

Use database idempotency keys and unique constraints to make repeated business requests safe.

04

Observe command attempts and error labels without logging sensitive payloads.

05

Keep non-database side effects outside transaction callbacks or make them independently idempotent.

Reproducible lab baseline

This final chapter pins MongoDB Community Server 8.3.8 using mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0. Labs use AtlasMart synthetic data, loopback-only Docker port publishing, replica set name atlasmart-rs27 where topology behavior matters, and explicit operation time budgets. Authentication/TLS are disabled only for disposable mechanism labs; the capstone security gate reuses Chapter 22 least-privilege/authentication requirements and treats TLS as mandatory production acceptance. Default read preference is primary and majority acknowledgement is used for business writes unless a failure experiment explicitly states otherwise. FCV is observed and never changed. Atlas, Search, Vector Search, Enterprise Advanced, and KMS are optional and are not required for mandatory work. Runtime load, failover, pool, latency, retry, and recovery results were not executed in the generation environment; learners must record their own evidence instead of copying invented values.

1. Retry layers solve different failures

Layer What may repeat? Required application rule
Retryable read A qualifying read after a transient network/server failure Read must tolerate repetition
Retryable write A qualifying single-document write is retried once by the driver Do not infer arbitrary external side effects are deduplicated
Transaction callback Commit and sometimes entire callback Callback may run multiple times; external effects need idempotency
Business retry Whole HTTP/job request Use an idempotency key and durable result record

Official drivers enable retryable reads/writes for supported operations by default. A generic while True: retry() loop is different: it can amplify overload, repeat non-retryable failures, and exceed user deadlines.

2. Idempotent AtlasMart order acceptance

reuse the Chapter 27 replica set
docker ps --filter name=atlasmart-rs27-lab --format "{{.Names}} {{.Status}}"# If absent, run Lesson 1 setup first.
idempotent_order.py
from pymongo import MongoClient, ASCENDINGfrom pymongo.errors import DuplicateKeyErrorfrom datetime import datetime, timezoneclient=MongoClient("mongodb://localhost:27209,localhost:27210,localhost:27211/?replicaSet=atlasmart-rs27&retryWrites=true&w=majority")db=client.atlasmartorders=db.orders_ch27_l2orders.drop()orders.create_index([("tenantId",ASCENDING),("idempotencyKey",ASCENDING)],unique=True,name="uq_tenant_idempotency")def place_order(tenant,key,total):    existing=orders.find_one({"tenantId":tenant,"idempotencyKey":key})    if existing: return existing["orderId"],True    doc={"tenantId":tenant,"idempotencyKey":key,"orderId":f"ord-{key}",         "totalCents":total,"createdAt":datetime.now(timezone.utc)}    try:        orders.insert_one(doc)        return doc["orderId"],False    except DuplicateKeyError:        # Concurrent replay won the race: return the durable first result.        return orders.find_one({"tenantId":tenant,"idempotencyKey":key})["orderId"],Trueprint(place_order("tenant-a","req-1001",4200))print(place_order("tenant-a","req-1001",4200))print(orders.count_documents({"tenantId":"tenant-a","idempotencyKey":"req-1001"}))client.close()

Verification is the invariant: both calls return the same order identifier and exactly one order exists for the tenant/idempotency key.

3. Transactions: database atomicity is not side-effect atomicity

ClientSession.with_transaction() may retry the commit or the entire callback. Therefore a callback that sends email, charges a card, or publishes a non-idempotent message can repeat those effects even when the database transaction ultimately commits once. Keep external effects after commit, or use an outbox/idempotent external API.

transaction callback with an outbox record
from pymongo import MongoClientfrom pymongo.read_concern import ReadConcernfrom pymongo.write_concern import WriteConcernfrom pymongo.read_preferences import ReadPreferenceclient=MongoClient("mongodb://localhost:27209,localhost:27210,localhost:27211/?replicaSet=atlasmart-rs27")db=client.atlasmartdef transfer(session):    db.wallets_ch27_l2.update_one({"_id":"buyer","cents":{"$gte":500}}, {"$inc":{"cents":-500}}, session=session)    db.wallets_ch27_l2.update_one({"_id":"seller"},{"$inc":{"cents":500}},session=session)    db.outbox_ch27_l2.update_one({"eventId":"pay-1001"},{"$setOnInsert":{"kind":"payment.captured","sent":False}},upsert=True,session=session)with client.start_session() as s:    s.with_transaction(transfer,read_concern=ReadConcern("snapshot"),                       write_concern=WriteConcern("majority"),read_preference=ReadPreference.PRIMARY)client.close()

4. Classify first; then decide retry, reconcile, or fail

bounded error classifier
from pymongo.errors import (DuplicateKeyError, NetworkTimeout, ServerSelectionTimeoutError,                            ExecutionTimeout, OperationFailure, PyMongoError)def classify(exc):    if isinstance(exc, DuplicateKeyError): return "conflict-not-retry"    if isinstance(exc, (NetworkTimeout, ServerSelectionTimeoutError, ExecutionTimeout)): return "deadline-or-transient"    if isinstance(exc, OperationFailure):        if exc.has_error_label("NoWritesPerformed"): return "safe-no-write-performed"        if exc.has_error_label("TransientTransactionError"): return "retry-whole-transaction-within-budget"        if exc.has_error_label("UnknownTransactionCommitResult"): return "retry-commit-or-reconcile"        return f"server-error-{exc.code}"    if isinstance(exc, PyMongoError): return "driver-error-requires-classification"    return "application-error"

Error labels are evidence, not permission for infinite retry. Every retry policy still needs an attempt/deadline budget, jitter/backoff where appropriate, and a reconciliation path for ambiguous outcomes.

5. Observe attempts without leaking commands

minimal command outcome listener
from pymongo import monitoringclass Outcomes(monitoring.CommandListener):    def started(self,e): print("start",e.command_name,e.request_id)    def succeeded(self,e): print("ok",e.command_name,e.request_id,round(e.duration_micros/1000,2))    def failed(self,e): print("fail",e.command_name,e.request_id,type(e.failure).__name__)

Production logs should correlate request/trace IDs with command name, duration, server address, retry/error classification, and result state. Do not dump authentication commands, full customer documents, encryption material, or payment data into logs.

Check your understanding

  1. Why is retryWrites=true not an exactly-once guarantee for an HTTP request?
  2. Why can a with_transaction callback be dangerous for sending email?
  3. What makes the order example idempotent?
  4. What does NoWritesPerformed help distinguish?
  5. Why bound application retries?
Review the answers

1. It only retries qualifying MongoDB writes with server/driver semantics; the larger business request can include other effects and repeated application calls.

2. The driver may invoke the callback multiple times while resolving transient transaction/commit failures.

3. A tenant-scoped unique idempotency key plus returning the durable first result on replay/race.

4. A failed retryable write/batch in which the server can state that no writes were performed.

5. Unbounded retries can violate deadlines and amplify an overloaded or permanently failing dependency.

6. Production judgment

Prefer driver retryability for supported transient database failures, database uniqueness/idempotency for business commands, and explicit reconciliation for ambiguous outcomes. Transactions protect MongoDB state, not arbitrary external effects. Lesson 3 turns the same discipline toward data-flow volume: pagination, batches, schema evolution, change streams, and backpressure.

Authoritative references

Driver defaults and deployment behavior evolve. Re-check the exact server patch, PyMongo release, topology, and managed-service tier before freezing production assumptions.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.