Build an AtlasMart inverted index from analyzed product text, separate scoring from exact filtering, and expose refresh lag, segment-like visibility, facets, and analyzer mistakes.

Inverted Indexes, Tokens, Analyzers, Relevance, Filters, Facets, and Aggregations

Follow text from source documents through analysis into terms and postings, then reason about relevance, exact filters, aggregations, refresh visibility, and why search indexes are optimized derived views rather than transactional truth by default.

Intermediate100–125 minutesInverted-index + refresh labPython 3.13+ · standard libraryElasticsearch 9.5.2 / Lucene concepts optionalLast reviewed: August 2026

Learning outcomes

Follow text from source documents through analysis into terms and postings, then reason about relevance, exact filters, aggregations, refresh visibility, and why search indexes are optimized derived views rather than transactional truth by default.

01

Build the document→analysis→term→postings mental model of an inverted index.

02

Separate text relevance scoring from exact filters and facet/aggregation computation.

03

Explain analyzer compatibility, near-real-time refresh, segment/merge concepts, and stale-search windows.

04

Treat search indexes as retrieval-optimized views unless their transactional authority is explicitly designed.

Implementation snapshot

Mandatory work uses Python 3.13+ standard library only in one local process. No Elasticsearch, OpenSearch, cloud service, Docker image, paid feature, network manipulation, or destructive failure injection is required. Current optional reference snapshots are Elasticsearch 9.5.2 (released August 20, 2026; default distribution under Elastic License 2.0) and OpenSearch 3.8.0 (released August 4, 2026; Apache License 2.0). Product-specific refresh, index, ANN, clustering, quota, security, and licensing semantics are examples—not universal database guarantees.

1. Invert the relationship between documents and terms

A source document naturally says “product p1 contains the terms waterproof, trail, shoe.” An inverted index reverses that relation so the index can answer “which documents contain trail?” with a postings list for the term. Each posting may carry document identifiers, term frequency, positions, offsets, payloads, or other statistics depending on the engine. This structure is optimized for text retrieval because a query can combine postings for analyzed terms instead of scanning every document's text. Lucene—the library underneath Elasticsearch and OpenSearch—organizes searchable index state into immutable segments that can later be merged.

2. Analysis defines the vocabulary

An analyzer transforms input text into indexed terms through character handling, tokenization, normalization, and optional token filters such as lowercasing, stemming, stop-word removal, synonym expansion, or language-specific rules. The lab shows a small but important failure: naive whitespace leaves trail-shoe as one token, while punctuation-aware analysis produces trail and shoe. If index-time and query-time analyzers are incompatible, valid documents can become unreachable or receive distorted scores. Analyzer changes are therefore schema changes for a search system and often require reindexing or carefully managed search-analyzer updates.

3. AtlasMart lab: make the mechanism observable

Save the following as lesson2_inverted_index.py and run it with python lesson2_inverted_index.py. It uses only deterministic in-memory data and mutates no external service.

python · AtlasMart deterministic simulation
import re, math
from collections import defaultdict, Counter

source = {
 "p1":{"title":"Waterproof Trail Shoe","category":"footwear","price":120},
 "p2":{"title":"Lightweight Trail Runner","category":"footwear","price":95},
 "p3":{"title":"Wireless Studio Headphones","category":"audio","price":220},
 "p4":{"title":"Pocket Wireless Earbuds","category":"audio","price":45},
 "p5":{"title":"Trail Backpack 20L","category":"outdoor","price":80},
}

def analyze(text):
    return re.findall(r"[a-z0-9]+", text.lower())

def naive(text):
    return text.lower().split()

print("naive tokens:", naive("Waterproof trail-shoe"))
print("analyzed tokens:", analyze("Waterproof trail-shoe"))

opened = {}
pending = {}

def refresh():
    opened.update({k:dict(v) for k,v in pending.items()})
    pending.clear()

def index_doc(doc_id):
    pending[doc_id] = dict(source[doc_id])

def postings(docs):
    inv = defaultdict(dict)
    for doc_id, doc in docs.items():
        counts = Counter(analyze(doc["title"]))
        for term, tf in counts.items():
            inv[term][doc_id] = tf
    return inv

def search(query, category=None):
    inv = postings(opened)
    qterms = analyze(query)
    scores = defaultdict(float)
    n = max(1, len(opened))
    for term in qterms:
        df = len(inv.get(term, {}))
        if not df: continue
        idf = math.log((n + 1) / (df + 1)) + 1
        for doc_id, tf in inv[term].items():
            if category and opened[doc_id]["category"] != category:
                continue
            scores[doc_id] += tf * idf
    return sorted(scores.items(), key=lambda x:(-x[1], x[0]))

for doc_id in source: index_doc(doc_id)
print("search before first refresh:", search("trail"))
refresh()
print("trail results after refresh:", search("trail"))
print("trail + exact category=footwear:", search("trail", category="footwear"))

# Facet counts are aggregations over the matched result set.
matched = [doc_id for doc_id,_ in search("trail")]
facets = Counter(opened[d]["category"] for d in matched)
print("trail category facets:", dict(sorted(facets.items())))

# Authority changes, but the search view is stale until the next refresh.
source["p2"]["title"] = "Waterproof Lightweight Trail Runner"
index_doc("p2")
print("source p2 contains waterproof:", "waterproof" in analyze(source["p2"]["title"]))
print("search waterproof before refresh:", search("waterproof"))
refresh()
print("search waterproof after refresh:", search("waterproof"))

# Relevance ranking is not an integrity rule.
print("top trail result:", search("trail")[0][0])
print("authoritative price p1:", source["p1"]["price"])
Expected evidence

Expected evidence: punctuation-aware analysis turns “trail-shoe” into separate searchable terms while naive whitespace tokenization does not; search returns nothing before the first refresh; trail results become visible after refresh; exact category filtering and facet counts operate on the result set; an authoritative p2 title update remains invisible to the opened search view until refresh; and relevance order is clearly separate from authoritative price/integrity data.

4. Relevance and filters answer different questions

A relevance query asks “which documents best match this text?” and produces a ranking score. A filter asks an exact Boolean question such as category equals footwear, tenant equals t1, or price below a threshold. Search systems often combine both: lexical relevance narrows and ranks text matches while filters enforce non-scoring constraints. Facets and aggregations summarize the matched set—for example counts by category or price band. These operations can be expensive at high cardinality, and security filters must be applied before exposing facets or counts that could leak hidden tenant data.

5. Refresh visibility is not durability

Elasticsearch documents near real-time search: newly indexed changes become searchable after a refresh, which opens new segment state for search. A refresh is not the same concept as a durable commit, fsync, replicated acknowledgement, or backup. The lab models only visibility: p2 changes in the source authority, an indexing operation is pending, and the search result remains stale until refresh. Production diagnostics must distinguish ingest acknowledgement, translog/WAL durability, refresh lag, replica state, segment merge pressure, and end-to-end source-to-search freshness instead of calling all of them “indexing latency.”

6. Search indexes trade write work for retrieval work

Tokenizing and indexing one product can create many term postings, doc-values/columnar fields, stored fields, vectors, and metadata. Frequent updates may produce new segment data and deleted-document markers that are reclaimed during merges. Aggressive refresh can create many small segments and shift cost into indexing, searching, and merging; Elastic's current refresh documentation explicitly warns about that tradeoff. Search capacity planning therefore includes ingestion rate, analyzed-field count, term cardinality, refresh interval, merge I/O, cache/page-cache needs, query concurrency, aggregation memory, and replica count.

7. Production judgment

Use an inverted index when lexical retrieval, relevance, filters, facets, or aggregations justify the derived structure. Preserve a clear source of truth for business invariants unless the search engine is deliberately operated as the authority. Version analyzer configurations, validate index mappings, measure source-to-search lag, test delete propagation and tenant filters, and rehearse full rebuilds. Do not infer correctness from a high relevance score, and do not use search-document freshness as a substitute for transactional read-your-writes unless the product-specific semantics and request path explicitly provide it.

Wrong approach: change analysis rules and assume old postings reinterpret themselves

Failure injection / diagnosis

A team changes tokenization or synonyms in application code and expects all previously indexed documents to behave as if they had been analyzed with the new rules. Existing postings still reflect the old analysis. Queries become inconsistent across old/new documents or nodes. The repair is to version analyzer definitions, determine whether a search-only analyzer reload is sufficient for the intended change, and reindex whenever index-time terms must change; then verify sampled term/posting behavior and relevance before cutover.

Verification, cleanup, and production checklist

Verification is the deterministic program output plus the conceptual checks below. Cleanup is deleting the local lesson2_inverted_index.py file; the lab creates no sockets, databases, containers, credentials, indexes, or cloud resources. In production, additionally record authoritative-versus-derived ownership, source/index versions, refresh or projection lag, p95/p99 read/write latency, index size and write amplification, shard/index-key skew, rebuild throughput, ANN recall where applicable, tenant/authorization tests, backup/rebuild evidence, current security advisories, and edition/license constraints before relying on a product-specific feature.

Check your understanding

  1. What does an inverted index map?
  2. Why are analyzers part of schema governance?
  3. What is the difference between a relevance query and an exact filter?
  4. What does a refresh prove?
  5. Why can frequent refresh be costly?
Review the answers

1. Terms to the documents/postings that contain those terms, rather than documents to their terms.

2. They determine which terms exist and therefore which documents are reachable and how queries match them.

3. Relevance produces a ranking score; a filter evaluates an exact predicate without needing relevance scoring.

4. That pending index changes are visible to search; it does not by itself prove durable commit, replication, or backup.

5. It can create many small segments and increase index, search, and merge work.

References

Foundational statements use primary research or standards where appropriate. Version-sensitive implementation examples use current official documentation and remain explicitly scoped to the cited product/version.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.