Chapter 03 · CRUD Fundamentals: Insert, Find, Projection, Sort, Limit, and Delete

Cursor Behavior, Batch Size, Iteration, Exhaustion, and Client-Side Resource Management

Expose the cursor lifecycle underneath find: firstBatch, getMore, batch size, iteration, exhaustion, cancellation, and explicit client cleanup.

Beginner100–125 minutesCursor firstBatch/getMore + PyMongo cleanup labMongoDB Community Server 8.3.8 · mongosh 2.10.0 · PyMongo 4.17.0Last reviewed: September 2026

Learning outcomes

A query returning 50,000 products should not require the server to serialize all 50,000 into one response or the client to keep every document in memory. MongoDB therefore returns multi-document reads through a cursor: a server/client protocol state that delivers results in batches. This lesson makes the cursor lifecycle visible rather than treating for doc in collection.find() as magic.

01

Explain firstBatch, getMore, cursor id, batchSize, iteration, and exhaustion.

02

Observe a cursor directly with the find and getMore database commands.

03

Distinguish batch size from total query limit and from an application memory guarantee.

04

Use PyMongo iteration and context-manager/close patterns to release client/server resources.

05

Diagnose two bad patterns: materializing an unbounded cursor and abandoning long-lived cursors without cleanup.

Chapter 03 reproducible baseline

Mandatory labs use a disposable loopback-only standalone mongodb/mongodb-community-server:8.3.8-ubuntu2204-slim with a dedicated AtlasMart collection and explicit reset commands. Driver examples pin pymongo==4.17.0. The standalone is intentionally unauthenticated only for these short-lived local exercises; do not publish it beyond 127.0.0.1. Docker commands use syntax accepted directly by ordinary shells; if the first cleanup reports that the container does not exist, that is harmless. Retryable-write behavior that requires a replica set or sharded cluster is explained but not claimed for this standalone.

Generation-time execution note

Docker, mongod, mongosh, and PyMongo are not available in this generation environment. Commands were checked against current official MongoDB Server and PyMongo documentation, but product commands were not executed here. Expected-output blocks describe stable fields and relationships to verify; they are not fabricated captured transcripts.

1. find returns a cursor, not necessarily all matching documents

At the wire-command level, a find response contains a cursor document with a namespace, a cursor identifier, and a firstBatch. If more results remain, the cursor id is nonzero and the client requests more with getMore. When the result is exhausted or closed, the server-side cursor id becomes zero/invalid and resources can be released.

The find command's default initial batch is bounded by both document count and response size; current documentation describes the initial batch as the lesser of 101 documents or 16 MiB, with later batches limited by 16 MiB. An explicit batchSize can request fewer documents, but it cannot make a response exceed protocol size limits.

mongosh · observe find/firstBatch/getMore
use atlasmartconst r1 = db.runCommand({  find: "crud_products",  filter: {active:true},  sort: {_id:1},  batchSize: 2});printjson({id:r1.cursor.id, firstBatch:r1.cursor.firstBatch});if (r1.cursor.id.toString() !== "0") {  const r2 = db.runCommand({    getMore: r1.cursor.id,    collection: "crud_products",    batchSize: 2  });  printjson({id:r2.cursor.id, nextBatch:r2.cursor.nextBatch});}

This command exposes the mechanism drivers normally hide. The exact numeric cursor id is dynamic and must never be hard-coded into documentation or application logic. The stable evidence is whether it is zero/nonzero and which documents arrived in each batch.

2. batchSize changes transfer granularity; limit changes total result count

batchSize(2) does not mean “return exactly two documents total.” It means request batches of up to that many documents subject to byte limits. limit(2) means return at most two documents total. Combining them is valid but they solve different problems.

mongosh · total limit vs batch size
// Up to 6 documents total, transported in small batches.const c = db.crud_products.find({active:true}).sort({_id:1}).limit(6).batchSize(2);while (c.hasNext()) {  printjson(c.next());}

Very small batches can increase round trips; very large batches can increase per-response memory and latency. There is no universal best batch size. Choose using document size, processing cost, network behavior, latency objectives, and measurement.

3. Iteration is streaming at the application boundary; list() is materialization

PyMongo's find() returns a Cursor. A for loop consumes batches as needed. Converting the cursor to list() asks Python to retain all returned documents in memory. That may be convenient for a known-small fixture but is dangerous for unbounded queries.

Python · iterate and close a cursor explicitly
from pymongo import MongoClientclient = MongoClient("mongodb://127.0.0.1:27029/?directConnection=true", serverSelectionTimeoutMS=3000)coll = client.atlasmart.crud_products# Good: bounded, iterative consumption with deterministic ordering.with coll.find({"active": True}, {"_id":1,"sku":1}).sort("_id", 1).batch_size(2) as cursor:    for doc in cursor:        print(doc)# Deliberately risky for an unknown/unbounded result set:# everything = list(coll.find({}))client.close()

PyMongo closes an exhausted cursor automatically. Explicit close() or a with statement is still valuable when iteration can stop early due to a condition, cancellation, or exception. The client itself owns sockets/pools and should also be closed when the application lifecycle ends.

4. Idle cursors, noCursorTimeout, and session boundaries

Normal idle server cursors can time out. A noCursorTimeout option is not an immortality guarantee: MongoDB drivers and mongosh use sessions, and an idle session can expire; current documentation notes that a 30-minute idle session can cause associated cursors to be killed even when noCursorTimeout was requested. Long-running batch jobs should process steadily, checkpoint application progress, and use explicit session-refresh patterns only when genuinely needed.

A robust job should also be resumable at the application level. Cursor lifetime is not a durable progress record. If a process crashes after handling item 5000, a new cursor does not know what external side effects were completed. Store a continuation key/checkpoint or make processing idempotent.

5. AtlasMart cursor lab: prove batching, exhaustion, and cleanup

shell · start disposable MongoDB on 127.0.0.1:27029
docker rm -f atlasmart-mongo-ch03-l3docker run --name atlasmart-mongo-ch03-l3 -p 127.0.0.1:27029:27017 -d mongodb/mongodb-community-server:8.3.8-ubuntu2204-slimdocker logs atlasmart-mongo-ch03-l3 --tail 25
shell · observe all cursor batches
mongosh "mongodb://127.0.0.1:27029/atlasmart?directConnection=true" --quiet --eval 'db.crud_products.drop();db.crud_products.insertMany(Array.from({length:7}, (_,i)=>({_id:"p"+(i+1),sku:"sku-"+(i+1),active:true})));const first=db.runCommand({find:"crud_products",filter:{active:true},sort:{_id:1},batchSize:2});printjson({firstId:first.cursor.id, firstBatch:first.cursor.firstBatch.map(x=>x._id)});let id=first.cursor.id;while (id.toString() !== "0") {  const more=db.runCommand({getMore:id,collection:"crud_products",batchSize:2});  printjson({nextId:more.cursor.id,batch:more.cursor.nextBatch.map(x=>x._id)});  id=more.cursor.id;}' 

Verification checklist

  • The first batch contains at most two documents and a nonzero cursor id while more data remains.
  • Subsequent getMore responses advance through the deterministic _id order.
  • The final response reaches an exhausted cursor state.
  • You can explain why batchSize:2 is not the same as limit:2.
  • You can identify where explicit close/cancellation logic belongs in a real driver application.

Check your understanding

  1. What does a nonzero cursor id mean?
  2. Why can list(cursor) be dangerous?
  3. Does batchSize set the total number of documents returned?
  4. When should a driver cursor be explicitly closed?
  5. Is noCursorTimeout sufficient for a cursor idle for hours?
Review the answers

The server has cursor state that can provide additional results through getMore; the id is opaque and dynamic.

It materializes all results in client memory, which can exhaust memory for large or unbounded queries.

No. It controls transfer batch granularity; limit controls the total maximum result count.

When the application may stop before exhaustion or needs prompt resource release; context managers are a safe pattern.

No. Session idle timeout can still expire the session and kill associated cursors; long jobs need explicit lifecycle/checkpoint design.

bash · cleanup/reset
docker rm -f atlasmart-mongo-ch03-l3

The next lesson applies the same observable-result discipline to destructive operations: precise deletes.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.