Build an AtlasMart inverted index from analyzed product text, separate scoring from exact filtering, and expose refresh lag, segment-like visibility, facets, and analyzer mistakes.
Inverted Indexes, Tokens, Analyzers, Relevance, Filters, Facets, and Aggregations
Follow text from source documents through analysis into terms and postings, then reason about relevance, exact filters, aggregations, refresh visibility, and why search indexes are optimized derived views rather than transactional truth by default.
Learning outcomes
Follow text from source documents through analysis into terms and postings, then reason about relevance, exact filters, aggregations, refresh visibility, and why search indexes are optimized derived views rather than transactional truth by default.
Build the document→analysis→term→postings mental model of an inverted index.
Separate text relevance scoring from exact filters and facet/aggregation computation.
Explain analyzer compatibility, near-real-time refresh, segment/merge concepts, and stale-search windows.
Treat search indexes as retrieval-optimized views unless their transactional authority is explicitly designed.
Mandatory work uses Python 3.13+ standard library only in one local process. No Elasticsearch, OpenSearch, cloud service, Docker image, paid feature, network manipulation, or destructive failure injection is required. Current optional reference snapshots are Elasticsearch 9.5.2 (released August 20, 2026; default distribution under Elastic License 2.0) and OpenSearch 3.8.0 (released August 4, 2026; Apache License 2.0). Product-specific refresh, index, ANN, clustering, quota, security, and licensing semantics are examples—not universal database guarantees.
1. Invert the relationship between documents and terms
A source document naturally says “product p1 contains the terms waterproof, trail, shoe.” An inverted index reverses that relation so the index can answer “which documents contain trail?” with a postings list for the term. Each posting may carry document identifiers, term frequency, positions, offsets, payloads, or other statistics depending on the engine. This structure is optimized for text retrieval because a query can combine postings for analyzed terms instead of scanning every document's text. Lucene—the library underneath Elasticsearch and OpenSearch—organizes searchable index state into immutable segments that can later be merged.
2. Analysis defines the vocabulary
An analyzer transforms input text into indexed
terms through character handling, tokenization, normalization,
and optional token filters such as lowercasing, stemming,
stop-word removal, synonym expansion, or language-specific
rules. The lab shows a small but important failure: naive
whitespace leaves trail-shoe as one token, while
punctuation-aware analysis produces trail and
shoe. If index-time and query-time analyzers are
incompatible, valid documents can become unreachable or receive
distorted scores. Analyzer changes are therefore schema changes
for a search system and often require reindexing or carefully
managed search-analyzer updates.
3. AtlasMart lab: make the mechanism observable
Save the following as lesson2_inverted_index.py and
run it with python lesson2_inverted_index.py. It
uses only deterministic in-memory data and mutates no external
service.
import re, math
from collections import defaultdict, Counter
source = {
"p1":{"title":"Waterproof Trail Shoe","category":"footwear","price":120},
"p2":{"title":"Lightweight Trail Runner","category":"footwear","price":95},
"p3":{"title":"Wireless Studio Headphones","category":"audio","price":220},
"p4":{"title":"Pocket Wireless Earbuds","category":"audio","price":45},
"p5":{"title":"Trail Backpack 20L","category":"outdoor","price":80},
}
def analyze(text):
return re.findall(r"[a-z0-9]+", text.lower())
def naive(text):
return text.lower().split()
print("naive tokens:", naive("Waterproof trail-shoe"))
print("analyzed tokens:", analyze("Waterproof trail-shoe"))
opened = {}
pending = {}
def refresh():
opened.update({k:dict(v) for k,v in pending.items()})
pending.clear()
def index_doc(doc_id):
pending[doc_id] = dict(source[doc_id])
def postings(docs):
inv = defaultdict(dict)
for doc_id, doc in docs.items():
counts = Counter(analyze(doc["title"]))
for term, tf in counts.items():
inv[term][doc_id] = tf
return inv
def search(query, category=None):
inv = postings(opened)
qterms = analyze(query)
scores = defaultdict(float)
n = max(1, len(opened))
for term in qterms:
df = len(inv.get(term, {}))
if not df: continue
idf = math.log((n + 1) / (df + 1)) + 1
for doc_id, tf in inv[term].items():
if category and opened[doc_id]["category"] != category:
continue
scores[doc_id] += tf * idf
return sorted(scores.items(), key=lambda x:(-x[1], x[0]))
for doc_id in source: index_doc(doc_id)
print("search before first refresh:", search("trail"))
refresh()
print("trail results after refresh:", search("trail"))
print("trail + exact category=footwear:", search("trail", category="footwear"))
# Facet counts are aggregations over the matched result set.
matched = [doc_id for doc_id,_ in search("trail")]
facets = Counter(opened[d]["category"] for d in matched)
print("trail category facets:", dict(sorted(facets.items())))
# Authority changes, but the search view is stale until the next refresh.
source["p2"]["title"] = "Waterproof Lightweight Trail Runner"
index_doc("p2")
print("source p2 contains waterproof:", "waterproof" in analyze(source["p2"]["title"]))
print("search waterproof before refresh:", search("waterproof"))
refresh()
print("search waterproof after refresh:", search("waterproof"))
# Relevance ranking is not an integrity rule.
print("top trail result:", search("trail")[0][0])
print("authoritative price p1:", source["p1"]["price"])
Expected evidence: punctuation-aware analysis turns “trail-shoe” into separate searchable terms while naive whitespace tokenization does not; search returns nothing before the first refresh; trail results become visible after refresh; exact category filtering and facet counts operate on the result set; an authoritative p2 title update remains invisible to the opened search view until refresh; and relevance order is clearly separate from authoritative price/integrity data.
4. Relevance and filters answer different questions
A relevance query asks “which documents best match this text?”
and produces a ranking score. A filter asks an
exact Boolean question such as category equals
footwear, tenant equals t1, or price
below a threshold. Search systems often combine both: lexical
relevance narrows and ranks text matches while filters enforce
non-scoring constraints. Facets and
aggregations summarize the matched set—for
example counts by category or price band. These operations can
be expensive at high cardinality, and security filters must be
applied before exposing facets or counts that could leak hidden
tenant data.
5. Refresh visibility is not durability
Elasticsearch documents near real-time search: newly indexed changes become searchable after a refresh, which opens new segment state for search. A refresh is not the same concept as a durable commit, fsync, replicated acknowledgement, or backup. The lab models only visibility: p2 changes in the source authority, an indexing operation is pending, and the search result remains stale until refresh. Production diagnostics must distinguish ingest acknowledgement, translog/WAL durability, refresh lag, replica state, segment merge pressure, and end-to-end source-to-search freshness instead of calling all of them “indexing latency.”
6. Search indexes trade write work for retrieval work
Tokenizing and indexing one product can create many term postings, doc-values/columnar fields, stored fields, vectors, and metadata. Frequent updates may produce new segment data and deleted-document markers that are reclaimed during merges. Aggressive refresh can create many small segments and shift cost into indexing, searching, and merging; Elastic's current refresh documentation explicitly warns about that tradeoff. Search capacity planning therefore includes ingestion rate, analyzed-field count, term cardinality, refresh interval, merge I/O, cache/page-cache needs, query concurrency, aggregation memory, and replica count.
7. Production judgment
Use an inverted index when lexical retrieval, relevance, filters, facets, or aggregations justify the derived structure. Preserve a clear source of truth for business invariants unless the search engine is deliberately operated as the authority. Version analyzer configurations, validate index mappings, measure source-to-search lag, test delete propagation and tenant filters, and rehearse full rebuilds. Do not infer correctness from a high relevance score, and do not use search-document freshness as a substitute for transactional read-your-writes unless the product-specific semantics and request path explicitly provide it.
Wrong approach: change analysis rules and assume old postings reinterpret themselves
A team changes tokenization or synonyms in application code and expects all previously indexed documents to behave as if they had been analyzed with the new rules. Existing postings still reflect the old analysis. Queries become inconsistent across old/new documents or nodes. The repair is to version analyzer definitions, determine whether a search-only analyzer reload is sufficient for the intended change, and reindex whenever index-time terms must change; then verify sampled term/posting behavior and relevance before cutover.
Verification, cleanup, and production checklist
Verification is the deterministic program output plus the
conceptual checks below. Cleanup is deleting the local
lesson2_inverted_index.py file; the lab creates no
sockets, databases, containers, credentials, indexes, or cloud
resources. In production, additionally record
authoritative-versus-derived ownership, source/index versions,
refresh or projection lag, p95/p99 read/write latency, index
size and write amplification, shard/index-key skew, rebuild
throughput, ANN recall where applicable, tenant/authorization
tests, backup/rebuild evidence, current security advisories, and
edition/license constraints before relying on a product-specific
feature.
Check your understanding
- What does an inverted index map?
- Why are analyzers part of schema governance?
- What is the difference between a relevance query and an exact filter?
- What does a refresh prove?
- Why can frequent refresh be costly?
Review the answers
1. Terms to the documents/postings that contain those terms, rather than documents to their terms.
2. They determine which terms exist and therefore which documents are reachable and how queries match them.
3. Relevance produces a ranking score; a filter evaluates an exact predicate without needing relevance scoring.
4. That pending index changes are visible to search; it does not by itself prove durable commit, replication, or backup.
5. It can create many small segments and increase index, search, and merge work.
References
Foundational statements use primary research or standards where appropriate. Version-sensitive implementation examples use current official documentation and remain explicitly scoped to the cited product/version.
- Elasticsearch — Near real-time search — Current explanation of buffers, segments, refresh, and search visibility.
- Elasticsearch — Refresh parameter — Current refresh options and cost tradeoffs.
- Apache Lucene — Primary implementation family underlying Elasticsearch/OpenSearch inverted indexing.
- Elasticsearch 9.5.2 — Current version snapshot, released 2026-08-20.