Design explicit Search mappings and compound queries that keep eligibility filters separate from relevance scoring.
Atlas Search Index Mappings, Text Operators, Compound Queries, Highlighting, and Scoring
Model coordinates correctly before asking the database a spatial question.
Learning objectives
Design static Search mappings for analyzed prose, exact filter/facet values, numeric fields, and highlighted content.
Use text and compound queries with must/should/filter/mustNot and explain why filter clauses do not add relevance score.
Inspect searchScore and highlight metadata without treating score magnitudes as cross-query probabilities.
Run a facet query and distinguish top hits from metadata aggregation.
Diagnose a deliberately poor mapping/analyzer choice and rebuild the index safely.
This lesson pins MongoDB Server 8.3.8 through the
evaluation image
mongodb/mongodb-atlas-local:8.3.8-20260827T164726Z,
mongosh 2.10.0, and PyMongo
4.17.0 where client measurement is useful. The
current self-managed Search GA line is
mongot 1.70.1 or later for MongoDB 8.3; production
Community and Enterprise deployments operate
mongot as a separate process, while Atlas manages
it. The mandatory lab uses a disposable single-node replica-set
development image on loopback port 27171 named
atlasmart-ch21-l2. It is an evaluation topology,
not high availability. The local credential is synthetic and
must never be reused outside this lab.
FCV is observed and never changed. Atlas cloud,
KMS, external embedding APIs, and paid search nodes are not
required. Automated Embedding is not used; vectors are
deterministic teaching data so no model key or network call is
needed. Product labs were not executed in this generation
environment; index build time, search scores, recall, freshness
lag, and latency must be measured locally rather than copied as
invented output. Search-index updates/rebuilds are asynchronous;
the existing index definition can remain queryable while a
replacement generation builds, so status must be checked before
assuming a cutover.
1. Mapping is the retrieval schema
AtlasMart wants “usb c charger” to match useful products, exact tenant/category filters to remain exact, price ranges to remain numeric, and UI snippets to highlight matching text. That means the search index has its own schema. Static mapping states which fields are indexed and how. Dynamic mapping is convenient for exploration but can index more than the application needs and makes cost/security review harder.
An index analyzer processes stored text; a search analyzer processes the query. The default standard analyzer is often a sensible baseline, but language-specific stemming, stop words, autocomplete, synonyms, or custom tokenization can materially change retrieval.
2. Build the mapping and inspect lifecycle state
docker rm -f atlasmart-ch21-l2 2>/dev/null || truedocker volume rm atlasmart-ch21-l2-db atlasmart-ch21-l2-config atlasmart-ch21-l2-search 2>/dev/null || truedocker run -d --name atlasmart-ch21-l2 \ -p 127.0.0.1:27171:27017 \ -e MONGODB_INITDB_ROOT_USERNAME=atlaslab \ -e MONGODB_INITDB_ROOT_PASSWORD=local-only-change-me \ -e DO_NOT_TRACK=1 \ -v atlasmart-ch21-l2-db:/data/db \ -v atlasmart-ch21-l2-config:/data/configdb \ -v atlasmart-ch21-l2-search:/data/mongot \ mongodb/mongodb-atlas-local:8.3.8-20260827T164726Zuntil [ "$(docker inspect -f '{{.State.Health.Status}}' atlasmart-ch21-l2 2>/dev/null)" = "healthy" ]; do sleep 2; donemongosh "mongodb://atlaslab:local-only-change-me@127.0.0.1:27171/admin?directConnection=true" --eval 'db.runCommand({ping:1})'
db = db.getSiblingDB("atlasmart");const c = db.catalog_search_ch21;c.drop();c.insertMany([ {_id:"P01",tenantId:"tenant-a",active:true, category:"power", title:"Portable USB-C Power Bank", description:"Compact travel battery with fast USB-C charging", embedding:[0.96,0.10,0.05,0.02], priceCents:4900}, {_id:"P02",tenantId:"tenant-a",active:true, category:"power", title:"65W USB-C Travel Charger", description:"GaN wall charger for laptops phones and travel", embedding:[0.92,0.16,0.03,0.02], priceCents:5900}, {_id:"P03",tenantId:"tenant-a",active:true, category:"cables", title:"Braided USB-C Cable", description:"Two meter durable charging and data cable", embedding:[0.82,0.20,0.08,0.04], priceCents:1800}, {_id:"P04",tenantId:"tenant-a",active:true, category:"audio", title:"Noise Cancelling Travel Headphones", description:"Over ear headphones for flights and commuting", embedding:[0.08,0.92,0.12,0.03], priceCents:12900}, {_id:"P05",tenantId:"tenant-a",active:true, category:"travel", title:"Universal Travel Adapter", description:"International plug adapter with USB-C ports", embedding:[0.72,0.14,0.55,0.06], priceCents:3900}, {_id:"P06",tenantId:"tenant-a",active:false,category:"power", title:"Legacy Power Brick", description:"Discontinued high capacity portable battery", embedding:[0.90,0.06,0.03,0.02], priceCents:3500}, {_id:"P07",tenantId:"tenant-b",active:true, category:"power", title:"Tenant B Private Charger", description:"Private catalog USB-C charging device", embedding:[0.95,0.11,0.03,0.02], priceCents:5100}, {_id:"P08",tenantId:"tenant-b",active:true, category:"security", title:"Tenant B Security Token", description:"Private authentication hardware token", embedding:[0.05,0.04,0.09,0.97], priceCents:7600}, {_id:"P09",tenantId:"tenant-a",active:true, category:"bags", title:"Laptop Travel Backpack", description:"Carry-on backpack with laptop compartment", embedding:[0.10,0.22,0.91,0.05], priceCents:8900}, {_id:"P10",tenantId:"tenant-a",active:true, category:"power", title:"Wireless Charging Pad", description:"Desk charger for Qi compatible phones", embedding:[0.78,0.14,0.05,0.03], priceCents:3200}, {_id:"P11",tenantId:"tenant-a",active:true, category:"audio", title:"USB-C Earbuds", description:"Wired earbuds with USB-C connector", embedding:[0.35,0.82,0.04,0.03], priceCents:2900}, {_id:"P12",tenantId:"tenant-a",active:true, category:"travel", title:"Packing Cube Set", description:"Lightweight organizers for carry-on travel", embedding:[0.08,0.13,0.96,0.02], priceCents:2600}]);print("documents", c.countDocuments({}));c.createSearchIndex("catalog_text", { mappings:{dynamic:false,fields:{ title:{type:"string",analyzer:"lucene.english"}, description:{type:"string",analyzer:"lucene.english"}, tenantId:{type:"token",normalizer:"lowercase"}, active:{type:"boolean"}, category:{type:"token",normalizer:"lowercase"}, priceCents:{type:"number"} }}});printjson(c.getSearchIndexes());
const wanted = ["catalog_text"];for (let attempt=0; attempt<120; attempt++) { const m = new Map(c.getSearchIndexes().map(x => [x.name,x])); const ready = wanted.every(n => m.get(n) && m.get(n).status === "READY" && m.get(n).queryable === true); if (ready) { print("READY", wanted.join(",")); break; } sleep(1000);}printjson(c.getSearchIndexes());
3. Compound clauses separate eligibility from relevance
must means every clause must match and contributes
score. should can contribute score and, depending
on the structure, can express preferences.
mustNot excludes. filter is the
crucial production tool for tenant, visibility, availability, or
other eligibility predicates because it constrains results
without pretending the security/business condition is a
relevance signal.
printjson(c.aggregate([ {$search:{index:"catalog_text",compound:{ must:[{text:{query:"usb charger",path:["title","description"]}}], should:[{text:{query:"travel",path:"description",score:{boost:{value:1.5}}}}], filter:[ {equals:{path:"tenantId",value:"tenant-a"}}, {equals:{path:"active",value:true}}, {range:{path:"priceCents",lte:7000}} ] },highlight:{path:["title","description"]}}}, {$project:{_id:1,title:1,priceCents:1,score:{$meta:"searchScore"},highlights:{$meta:"searchHighlights"}}}, {$limit:5}]).toArray());
Inspect the returned IDs, scores, and highlight fragments. If a
document is absent because of filter, increasing a
text boost must not resurrect it. That property is important for
security and catalog eligibility.
4. Facets answer a different UI question than hits
A faceted UI needs metadata about the whole eligible result set,
not just the categories of the first page. Use
$searchMeta with a facet collector so the result is
metadata rather than ordinary documents.
printjson(c.aggregate([ {$searchMeta:{index:"catalog_text",facet:{ operator:{compound:{ must:[{text:{query:"travel",path:["title","description"]}}], filter:[{equals:{path:"tenantId",value:"tenant-a"}},{equals:{path:"active",value:true}}] }}, facets:{category:{type:"string",path:"category",numBuckets:10}} }}}]).toArray());
Facet counts describe the matching root documents under the current query definition. They are not a substitute for authorization and they can change as the search index catches up with source changes.
5. Wrong mapping: keyword-like analysis for human prose
A common failure is to choose tokenization that is too literal. If a long title/description is indexed as a single keyword-like token, a query for one meaningful word may not match. The inverse error—aggressive stemming or autocomplete on identifiers—can create false matches.
c.updateSearchIndex("catalog_text", { mappings:{dynamic:false,fields:{ title:{type:"string",analyzer:"lucene.english"}, description:{type:"string",analyzer:"lucene.english"}, tenantId:{type:"token",normalizer:"lowercase"}, active:{type:"boolean"}, category:{type:"token",normalizer:"lowercase"}, priceCents:{type:"number"} }}});printjson(c.getSearchIndexes("catalog_text"));
Updating the definition triggers a new build. Do not assume the
call means the new generation is active; observe
latestDefinitionVersion, status, and queryable
state. During rebuilds, current documentation allows an existing
queryable definition to continue serving until the replacement
is ready.
6. Scoring and highlighting boundaries
Relevance scores should be evaluated with judgement queries, click/satisfaction signals, or curated expected results. Score magnitudes are not calibrated probabilities and should not be compared across unrelated query/index definitions without validation. Highlighting returns snippets from indexed content; treat those snippets as presentation data that still requires ordinary output escaping.
7. Production judgment
Prefer explicit mappings for production search contracts: you can reason about cost, security, analyzer changes, and rebuild blast radius. Record index definition versions, test representative queries before and after a mapping change, and keep a rollback definition. Separate relevance clauses from mandatory filters.
Bridge. Lesson 3 adds vector embeddings, similarity functions, exact-nearest-neighbor ground truth, and approximate-nearest-neighbor recall.
docker rm -f atlasmart-ch21-l2 2>/dev/null || truedocker volume rm atlasmart-ch21-l2-db atlasmart-ch21-l2-config atlasmart-ch21-l2-search 2>/dev/null || true
Check your understanding
- Why use compound.filter for tenant eligibility?
- Why wait for READY after create/update?
- Can searchScore be interpreted as a probability?
- Why use $searchMeta for facets?
- What is a safe mapping rollout?
Review the answers
1. It constrains the candidate set without contributing relevance score, which keeps security/business eligibility separate from ranking.
2. Search index builds are asynchronous; command acknowledgement is not proof that the new index generation is serving queries.
3. No. Treat it as a query/index-specific ranking signal.
4. Facets describe metadata over the eligible search result set rather than only the returned hit page.
5. Create/update, monitor build state, run judgement queries and leakage tests, observe the new generation, and retain a rollback definition.
Authoritative references
- Self-Managed MongoDB Search and Vector Search — mongot architecture, Community/Enterprise deployment paths, and Atlas-managed responsibility boundary.
- mongot Compatibility and Requirements — MongoDB 8.3 / mongot 1.70.1 compatibility, platforms, and topology requirements.
- Local Development Quickstart — Atlas Local evaluation topology and search/vector index workflow.
- Self-Managed mongot Release Notes — mongot 1.70.1 GA baseline for Community and Enterprise self-managed search.
- Verify mongot Connection — health/readiness endpoints, index state, and end-to-end verification.
- Troubleshoot Self-Managed mongot — replication lag, stale indexes, rebuilds, disk pressure, and query failures.
- mongot Metrics Reference — per-index size, status, replication lag, and resource metrics.
- MongoDB Search Queries and Indexes — dynamic/static mappings, analyzers, query/index relationship.
- MongoDB Search text Operator — analyzed text matching, options, and scoring semantics.
- MongoDB Search compound Operator — must/should/filter/mustNot clauses and score-neutral filtering.
- Search Highlighting — highlight metadata and index requirements.
- Search Faceting — facet collector and current token-oriented string faceting guidance.
- Vector Search Index Fields — vector/filter fields, dimensions, similarity, and quantization options.
- MongoDB Search vectorSearch Operator — ANN/ENN, numCandidates, filters, dimensions, and vector scoring.
- Measure Vector Search Accuracy — ENN judgement lists, ANN recall/overlap, numCandidates tuning, and reranking.
- $rankFusion — reciprocal-rank fusion for multi-pipeline retrieval on MongoDB 8.0+.
- $scoreFusion — score normalization and weighted fusion, GA in MongoDB 8.3.
- getSearchIndexes() — search/vector index lifecycle states including READY, BUILDING, and STALE.
- MongoDB 8.3 Release Notes — current server line, 8.3.8 patch baseline, and scoreFusion GA.
- mongosh Release Notes — mongosh 2.10.0 baseline.
- PyMongo Release Notes — PyMongo 4.17 driver baseline.