Chapter 03 · CRUD Fundamentals: Insert, Find, Projection, Sort, Limit, and Delete
Cursor Behavior, Batch Size, Iteration, Exhaustion, and Client-Side Resource Management
Expose the cursor lifecycle underneath find: firstBatch, getMore, batch size, iteration, exhaustion, cancellation, and explicit client cleanup.
Learning outcomes
A query returning 50,000 products should not require the server
to serialize all 50,000 into one response or the client to keep
every document in memory. MongoDB therefore returns
multi-document reads through a cursor: a
server/client protocol state that delivers results in batches.
This lesson makes the cursor lifecycle visible rather than
treating for doc in collection.find() as magic.
Explain firstBatch, getMore, cursor id, batchSize, iteration, and exhaustion.
Observe a cursor directly with the find and getMore database commands.
Distinguish batch size from total query limit and from an application memory guarantee.
Use PyMongo iteration and context-manager/close patterns to release client/server resources.
Diagnose two bad patterns: materializing an unbounded cursor and abandoning long-lived cursors without cleanup.
Mandatory labs use a disposable loopback-only standalone
mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim
with a dedicated AtlasMart collection and explicit reset
commands. Driver examples pin pymongo==4.17.0.
The standalone is intentionally unauthenticated only for these
short-lived local exercises; do not publish it beyond
127.0.0.1. Docker commands use syntax accepted
directly by ordinary shells; if the first cleanup reports that
the container does not exist, that is harmless.
Retryable-write behavior that requires a replica set or
sharded cluster is explained but not claimed for this
standalone.
Docker, mongod, mongosh, and PyMongo are not available in this generation environment. Commands were checked against current official MongoDB Server and PyMongo documentation, but product commands were not executed here. Expected-output blocks describe stable fields and relationships to verify; they are not fabricated captured transcripts.
1. find returns a cursor, not necessarily all matching documents
At the wire-command level, a find response contains
a cursor document with a namespace, a cursor
identifier, and a firstBatch. If more results
remain, the cursor id is nonzero and the client requests more
with getMore. When the result is exhausted or
closed, the server-side cursor id becomes zero/invalid and
resources can be released.
The find command's default initial batch is bounded
by both document count and response size; current documentation
describes the initial batch as the lesser of 101 documents or 16
MiB, with later batches limited by 16 MiB. An explicit
batchSize can request fewer documents, but it
cannot make a response exceed protocol size limits.
use atlasmartconst r1 = db.runCommand({ find: "crud_products", filter: {active:true}, sort: {_id:1}, batchSize: 2});printjson({id:r1.cursor.id, firstBatch:r1.cursor.firstBatch});if (r1.cursor.id.toString() !== "0") { const r2 = db.runCommand({ getMore: r1.cursor.id, collection: "crud_products", batchSize: 2 }); printjson({id:r2.cursor.id, nextBatch:r2.cursor.nextBatch});}
This command exposes the mechanism drivers normally hide. The exact numeric cursor id is dynamic and must never be hard-coded into documentation or application logic. The stable evidence is whether it is zero/nonzero and which documents arrived in each batch.
2. batchSize changes transfer granularity; limit changes total result count
batchSize(2) does not mean “return exactly two
documents total.” It means request batches of up to that many
documents subject to byte limits. limit(2) means
return at most two documents total. Combining them is valid but
they solve different problems.
// Up to 6 documents total, transported in small batches.const c = db.crud_products.find({active:true}).sort({_id:1}).limit(6).batchSize(2);while (c.hasNext()) { printjson(c.next());}
Very small batches can increase round trips; very large batches can increase per-response memory and latency. There is no universal best batch size. Choose using document size, processing cost, network behavior, latency objectives, and measurement.
3. Iteration is streaming at the application boundary; list() is materialization
PyMongo's find() returns a Cursor. A
for loop consumes batches as needed. Converting the
cursor to list() asks Python to retain all returned
documents in memory. That may be convenient for a known-small
fixture but is dangerous for unbounded queries.
from pymongo import MongoClientclient = MongoClient("mongodb://127.0.0.1:27029/?directConnection=true", serverSelectionTimeoutMS=3000)coll = client.atlasmart.crud_products# Good: bounded, iterative consumption with deterministic ordering.with coll.find({"active": True}, {"_id":1,"sku":1}).sort("_id", 1).batch_size(2) as cursor: for doc in cursor: print(doc)# Deliberately risky for an unknown/unbounded result set:# everything = list(coll.find({}))client.close()
PyMongo closes an exhausted cursor automatically. Explicit
close() or a with statement is still
valuable when iteration can stop early due to a condition,
cancellation, or exception. The client itself owns sockets/pools
and should also be closed when the application lifecycle ends.
4. Idle cursors, noCursorTimeout, and session boundaries
Normal idle server cursors can time out. A
noCursorTimeout option is not an immortality
guarantee: MongoDB drivers and mongosh use sessions, and an idle
session can expire; current documentation notes that a 30-minute
idle session can cause associated cursors to be killed even when
noCursorTimeout was requested. Long-running batch
jobs should process steadily, checkpoint application progress,
and use explicit session-refresh patterns only when genuinely
needed.
A robust job should also be resumable at the application level. Cursor lifetime is not a durable progress record. If a process crashes after handling item 5000, a new cursor does not know what external side effects were completed. Store a continuation key/checkpoint or make processing idempotent.
5. AtlasMart cursor lab: prove batching, exhaustion, and cleanup
docker rm -f atlasmart-mongo-ch03-l3docker run --name atlasmart-mongo-ch03-l3 -p 127.0.0.1:27029:27017 -d mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimdocker logs atlasmart-mongo-ch03-l3 --tail 25
mongosh "mongodb://127.0.0.1:27029/atlasmart?directConnection=true" --quiet --eval 'db.crud_products.drop();db.crud_products.insertMany(Array.from({length:7}, (_,i)=>({_id:"p"+(i+1),sku:"sku-"+(i+1),active:true})));const first=db.runCommand({find:"crud_products",filter:{active:true},sort:{_id:1},batchSize:2});printjson({firstId:first.cursor.id, firstBatch:first.cursor.firstBatch.map(x=>x._id)});let id=first.cursor.id;while (id.toString() !== "0") { const more=db.runCommand({getMore:id,collection:"crud_products",batchSize:2}); printjson({nextId:more.cursor.id,batch:more.cursor.nextBatch.map(x=>x._id)}); id=more.cursor.id;}'
Verification checklist
- The first batch contains at most two documents and a nonzero cursor id while more data remains.
-
Subsequent
getMoreresponses advance through the deterministic_idorder. - The final response reaches an exhausted cursor state.
-
You can explain why
batchSize:2is not the same aslimit:2. - You can identify where explicit close/cancellation logic belongs in a real driver application.
Check your understanding
- What does a nonzero cursor id mean?
- Why can list(cursor) be dangerous?
- Does batchSize set the total number of documents returned?
- When should a driver cursor be explicitly closed?
- Is noCursorTimeout sufficient for a cursor idle for hours?
Review the answers
The server has cursor state that can provide additional results through getMore; the id is opaque and dynamic.
It materializes all results in client memory, which can exhaust memory for large or unbounded queries.
No. It controls transfer batch granularity; limit controls the total maximum result count.
When the application may stop before exhaustion or needs prompt resource release; context managers are a safe pattern.
No. Session idle timeout can still expire the session and kill associated cursors; long jobs need explicit lifecycle/checkpoint design.
docker rm -f atlasmart-mongo-ch03-l3
The next lesson applies the same observable-result discipline to destructive operations: precise deletes.
Authoritative references
- MongoDB release notes — Official current stable server series and patch notes.
- MongoDB 8.3 release notes — Official 8.3 patch history; 8.3.8 is the latest released patch at review time.
- MongoDB CRUD operations — Official CRUD overview and server semantics.
- PyMongo CRUD guides — Official Python driver CRUD behavior and result objects.
- find database command — First batch, batchSize, limit, skip, and cursor response semantics.
- getMore database command — Fetching subsequent cursor batches.
- PyMongo cursor access — Iteration, list materialization warning, explicit close, and context manager.
- noCursorTimeout() — Idle cursor/session timeout interaction.