Prompt 18 · Lesson 01 · Supported CDC contract

Change Streams vs Tailing the Oplog: Stable API, Resume Tokens, and Deployment Requirements

Change streams expose durable MongoDB changes through a supported resumable API. Learn the event contract before attaching external side effects.

Intermediate–Advanced120–190 minutesCDC/resumability engineering labMongoDB 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning objectives

01

Explain why change streams are a supported CDC interface while raw oplog tailing couples applications to replication internals.

02

State the exact deployment/storage/protocol requirements and the Stable API boundary.

03

Interpret a raw change event, including resume token, operationType, namespace, documentKey, clusterTime, and updateDescription.

04

Explain majority-committed notification and why CDC events are not automatically business/domain events.

05

Checkpoint and resume a PyMongo stream without inventing exactly-once guarantees.

Reproducible lab baseline

This lesson pins MongoDB Community Server 8.3.8 with mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim, mongosh 2.10.0, and PyMongo 4.17.0. The mandatory lab uses a disposable single-member replica set on loopback port 27155; change streams require a replica set or sharded cluster, so a standalone is intentionally not used. A single-member replica set is enough to learn event/resume mechanics but does not demonstrate high availability, multi-node majority durability, or failover. Authentication and TLS are disabled only for this isolated lab. Feature Compatibility Version (FCV) is inspected. FCV is observed and never changed. Default read/write concern and primary read preference apply unless a command says otherwise. Change streams are in Stable API V1; showExpandedEvents is not. Time-series collections do not support change streams, a boundary that becomes important in Chapter 19. Atlas, Search, Vector Search, KMS, and Enterprise Advanced are not mandatory. Product commands were not executed in this generation environment because Docker, mongod, mongosh, and PyMongo are unavailable here; runtime timings and token values must be measured on the learner machine rather than copied as invented output.

1. AtlasMart needs durable change notification, not a private replication parser

AtlasMart wants catalog writes to trigger cache invalidation and search-index refresh. Polling repeatedly scans for work; tailing local.oplog.rs exposes an internal replication log whose representation is not the application-facing contract. A change stream is MongoDB's resumable Change Data Capture (CDC) interface: it converts majority-committed database changes into change-event documents and lets the client filter them with a restricted aggregation pipeline.

CDC answers what changed in the database. A domain event such as OrderPaid carries business intent and a deliberately designed schema. An update to orders.status may be evidence from which an application infers a business transition, but MongoDB does not know whether that write was a payment capture, data repair, replay, or operator correction.

Layer Stable contract Do not assume
Change stream Supported event/resume API over collection, database, or deployment changes. Exactly-once side effects or infinite history.
Oplog Replication history used by replica sets and by change streams internally. A public domain-event schema suitable for application parsing.
Domain event Application-owned business contract such as OrderPaid. That every low-level database mutation maps one-to-one to business intent.

2. Deployment and Stable API requirements

Change streams are available on replica sets and sharded clusters using WiredTiger and replica-set protocol version 1. They are not available on standalone servers. On a sharded cluster, clients open the stream through mongos; the router opens per-shard streams, sorts/filters results, and maintains a total ordering with the deployment logical clock. Change streams themselves are Stable API V1, but showExpandedEvents is outside Stable API V1 and therefore must not be smuggled into an API-strict contract.

start disposable replica set (l1)
docker rm -f atlasmart-ch18-l1 2>/dev/null || truedocker volume rm atlasmart-ch18-l1-data 2>/dev/null || truedocker run -d --name atlasmart-ch18-l1 \  -p 127.0.0.1:27155:27017 \  -v atlasmart-ch18-l1-data:/data/db \  mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim --replSet atlasmart-rs18-l1 --oplogSize 128 --bind_ip_alluntil mongosh "mongodb://127.0.0.1:27155/admin?directConnection=true" --quiet --eval 'quit(db.runCommand({ping:1}).ok===1?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27155/admin?directConnection=true" --quiet --eval 'rs.initiate({_id:"atlasmart-rs18-l1",members:[{_id:0,host:"atlasmart-ch18-l1:27017"}]})'until mongosh "mongodb://127.0.0.1:27155/admin?directConnection=true" --quiet --eval 'quit(db.hello().isWritablePrimary?0:1)'; do sleep 1; donemongosh "mongodb://127.0.0.1:27155/admin?replicaSet=atlasmart-rs18-l1" --quiet --eval 'printjson(db.version());printjson(db.runCommand({getParameter:1,featureCompatibilityVersion:1}).featureCompatibilityVersion);printjson(rs.status().members.map(m=>({name:m.name,stateStr:m.stateStr})));' 

3. Observe the event contract and resume token

capture insert and update events with PyMongo
from pprint import pprintfrom pymongo import MongoClienturi = "mongodb://127.0.0.1:27155/?replicaSet=atlasmart-rs18-l1"client = MongoClient(uri, serverSelectionTimeoutMS=5000)db = client.atlasmartproducts = db.products_ch18_l1products.drop()products.insert_one({"_id":"p-1801","name":"USB-C Hub","price":49,"stock":12})with products.watch(max_await_time_ms=1000) as stream:    products.update_one({"_id":"p-1801"},{"$set":{"price":45,"stock":10}})    event = stream.next()    pprint({        "resumeToken": event["_id"],        "operationType": event["operationType"],        "ns": event["ns"],        "documentKey": event["documentKey"],        "clusterTime": event["clusterTime"],        "updateDescription": event.get("updateDescription"),    })    checkpoint = stream.resume_tokenwith products.watch(resume_after=checkpoint, max_await_time_ms=1000) as resumed:    products.update_one({"_id":"p-1801"},{"$inc":{"stock":1}})    pprint(resumed.next())client.close()

The exact token payload is intentionally not printed as a fixed expected value: it depends on server state and FCV. The invariant is that the event's _id is the resume token and that the resumed stream starts after the checkpointed event while the required history still exists. Do not remove or rewrite event _id in the change-stream pipeline.

What “majority committed” proves

A collection watch notifies only after the underlying data change has persisted to a majority of data-bearing replica-set members. In this one-member teaching replica set that majority is one member, so the lab demonstrates the API but not real multi-node durability. It also does not make a downstream email, cache write, or HTTP request exactly once.

4. Deliberately wrong approach: treat oplog rows as domain events

A practitioner may tail local.oplog.rs, serialize whatever internal operation structure appears, and publish it as OrderPaid. That leaks internal storage/replication representation into external contracts and loses business intent. The safer design is: use change streams for durable database-change notification, then either transform changes into an explicitly versioned downstream schema or write an application-owned outbox/domain-event record in the same atomic aggregate update when the business invariant demands intent.

diagnostic only: compare oplog history to supported change API
const local = db.getSiblingDB("local");printjson(local.oplog.rs.find({ns:"atlasmart.products_ch18_l1"}).sort({$natural:-1}).limit(3).toArray());print("Do not build application contracts from this internal representation.");

5. Production judgment

Use change streams when downstream work must react to durable MongoDB changes with resumability. Size connection pools for the number of open streams because each stream can hold a connection while waiting on getMore. In authenticated deployments grant the narrow find + changeStream privileges for the watched scope. Monitor consumer lag, checkpoint age relative to the oplog window, reconnect/resume errors, event processing failures, and downstream reconciliation drift. On sharded deployments, cold or geographically distant shards can increase notification latency because mongos participates in global ordering.

Bridge. Lesson 2 widens the scope from one collection to database/deployment watches and uses server-side change-stream pipelines to discard irrelevant events before they cross the application boundary.

Cleanup/reset

Everything in this lesson is disposable. Remove only the chapter-specific container and volume:

cleanup
docker rm -f atlasmart-ch18-l1 2>/dev/null || truedocker volume rm atlasmart-ch18-l1-data 2>/dev/null || true

Check your understanding

  1. Why prefer change streams over direct oplog tailing?
  2. Is a change event automatically a domain event?
  3. What field is the event resume token?
  4. Can a standalone mongod provide change streams?
  5. Does a majority-committed change-stream event make a webhook exactly once?
Review the answers

1. Change streams are the supported resumable event API; direct oplog parsing couples applications to replication internals.

2. No. It describes a durable database change, not necessarily the business intent that caused it.

3. The complete change-event _id document.

4. No. Use a replica set or sharded cluster.

5. No. Downstream side effects need idempotency/deduplication and reconciliation.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.