Chapter 01 · Search Engine Foundations, Elasticsearch vs OpenSearch, Deployment Models, and Lab Setup
Inverted-Index Search vs Relational / Document / Vector Databases: Workload Fit and Architectural Tradeoffs
Decide when inverted-index search is the right serving architecture and prove the boundary with a tiny dual-platform AtlasMart fixture.
Learning outcomes
AtlasMart needs fast product discovery across names, descriptions, brands, categories, and attributes while its system of record still owns prices, inventory and orders. The architecture question is not “should everything move into Elasticsearch?” It is which read problem an inverted index solves better than a transactional database, where a document store or vector database is more appropriate, and what consistency and operational costs appear when search becomes a derived serving system.
Explain an inverted index, term dictionary, postings list, document, field, analyzer, query, score, shard, replica and Lucene segment without treating Elasticsearch/OpenSearch as generic JSON databases.
Separate search-system responsibilities from relational integrity, document-primary storage, and vector-nearest-neighbor retrieval.
Trace a tiny AtlasMart text query from indexed fields to matching documents and explain what the result does and does not guarantee.
Choose a search architecture from workload evidence: relevance, filtering, aggregation, freshness, durability, latency, write rate, schema and failure tolerance.
Run the same minimal mapping/index/query fixture on pinned Elasticsearch and OpenSearch while recording product/version/security differences.
This chapter pins Elasticsearch 9.5.3 (released 3
September 2026) and OpenSearch 3.8.0 (released 4
August 2026) for reproducible examples. OpenSearch 3.9.0 is
scheduled for 29 September 2026 and is therefore not treated
as current. Re-check both projects before reusing these
commands later. Elasticsearch and OpenSearch are independent
products: shared Lucene ancestry does not make their APIs,
plugins, security, lifecycle, vector features, clients, or
managed offerings interchangeable.
The environment used to generate this lesson does not provide
Docker, Elasticsearch, OpenSearch, Kibana, or OpenSearch
Dashboards. The commands and API shapes were reviewed against
the current official documentation but were not executed here.
Expected output is described by invariant and field shape
rather than presented as captured benchmark evidence. Every
destructive action is scoped to
atlasmart-* course containers, volumes, indices,
and local loopback ports.
1. Search engines optimize a different access path
A relational index is usually designed to find rows by ordered key values or to support a known predicate. A full-text search engine instead builds structures that answer “which documents contain terms related to this query?” efficiently. At the center is the inverted index: instead of reading every product description and asking whether it contains “wireless keyboard”, the engine records terms and the document identifiers in which those terms occur. A term dictionary organizes searchable terms; a postings list associates a term with matching documents and can carry frequency/position information used by scoring and phrase queries.
Elasticsearch and OpenSearch expose JSON over HTTP, but the JSON document is an input and retrieval abstraction—not proof that the engine behaves like a document database. During indexing, fields can be parsed by mappings; text fields can pass through an analyzer that produces a token stream; Lucene writes immutable segments; and a search can execute across one or many shards before a coordinating node reduces results. Later chapters unpack each mechanism. For now, the key rule is that search quality and cost are created at both write time and query time.
| System style | Good default question | Search-platform boundary |
|---|---|---|
| Relational database | “Which order has this primary key, and can this transaction preserve constraints?” | Keep authoritative transactions/invariants here unless search-specific evidence justifies duplication. |
| Document database | “Can I retrieve/update an aggregate-shaped document by known keys?” | A document model alone does not provide the same analyzed-text ranking and search execution model. |
| Lexical search engine | “Which documents best match these terms plus structured filters?” | Excellent for retrieval/analytics, but near-real-time visibility and distributed search semantics differ from OLTP transactions. |
| Vector database / vector index | “Which vectors are nearest under a similarity metric?” | Semantic similarity can complement lexical retrieval, but does not replace exact identifiers, filters, authorization or judged relevance. |
2. Workload fit is about questions, not product labels
AtlasMart’s catalog search is a strong search-engine candidate because users do not know exact product identifiers. They type noisy natural-language terms, filter by category and brand, sort or boost by business signals, and expect ranked results. A relational database can support some of this, but building language analysis, typo tolerance, distributed relevance, aggregations and later hybrid retrieval usually creates a separate search concern anyway.
By contrast, “decrement inventory exactly once when payment commits” is not a search-first problem. Search results may be near real time: an indexing request can be acknowledged before a normal search sees the new document, depending on refresh behavior. Search copies also have their own mappings, index versions, snapshots and recovery paths. Treating a search index as an effortless mirror hides synchronization and failure modes.
If the business requirement can be expressed only as “we need something fast,” the architecture is underspecified. Record query shapes, ranking needs, filters, aggregations, accepted freshness lag, update rate, document size, field cardinality, retention, failure objectives, tenant/security constraints, and the authoritative source of truth before selecting a platform.
3. Minimal common-denominator AtlasMart fixture
The following mapping deliberately uses conservative constructs
supported by both current platforms: one shard, zero replicas
for a single-node lab, keyword identifiers/facets,
text analyzed content, and a scaled numeric price.
It is a teaching fixture, not a production shard recommendation.
PUT /atlasmart-products-v1{ "settings": {"number_of_shards": 1, "number_of_replicas": 0}, "mappings": { "properties": { "product_id": {"type": "keyword"}, "name": {"type": "text", "fields": {"keyword": {"type": "keyword"}}}, "category": {"type": "keyword"}, "description": {"type": "text"}, "price": {"type": "scaled_float", "scaling_factor": 100} } }}
PUT /atlasmart-products-v1/_doc/P-1001?refresh=true{"product_id":"P-1001","name":"Quiet Wireless Keyboard","category":"keyboards","description":"compact wireless keyboard with quiet keys","price":49.90}PUT /atlasmart-products-v1/_doc/P-1002?refresh=true{"product_id":"P-1002","name":"Mechanical Gaming Keyboard","category":"keyboards","description":"wired mechanical keyboard with tactile switches","price":89.00}PUT /atlasmart-products-v1/_doc/P-1003?refresh=true{"product_id":"P-1003","name":"Wireless Travel Mouse","category":"mice","description":"compact wireless mouse for travel","price":29.50}
GET /atlasmart-products-v1/_search{ "query": { "bool": { "must": [{"match": {"description": "wireless keyboard"}}], "filter": [{"term": {"category": "keyboards"}}] } }, "_source": ["product_id", "name", "category", "price"]}
The expected invariant is that P-1001 matches
because its analyzed description contains terms corresponding to
the query and its exact category satisfies the filter. Do not
freeze an exact _score in course prose: score
depends on query structure, field statistics, analyzer output
and engine/version details. The
refresh=true parameter is used only to make this
tiny lesson deterministic; forcing refresh on every production
write can damage indexing throughput.
4. What the engine is doing
At index time the engine parses the JSON, validates it against
the mapping, analyzes the text fields, records
searchable terms and field-oriented structures, writes operation
state through its durability path, and eventually exposes the
change to search through refresh/segment mechanics. At query
time the match clause analyzes the query text; the
term-level category filter uses the exact indexed value;
shard-local searches produce candidates; and the coordinating
path combines them into the response.
That flow explains several common mistakes. A
keyword field is not analyzed like full text. A
successful index response is not the same event as snapshot
protection. A matching document does not prove the authoritative
product is still in stock. A high score is not a calibrated
probability that the result is “correct.” Each guarantee belongs
to a different mechanism.
Store orders, payments, inventory decrements, product-search documents and semantic embeddings only in the search cluster because it “can store JSON.” The concrete failure is ownership confusion: search refresh, mapping changes, reindexing, relevance experiments and shard recovery now collide with transactional invariants. Repair the design by naming the authoritative system for each fact, treating search as a purpose-built serving index where appropriate, and defining replay/reconciliation for synchronization.
5. Lab: prove fit and boundaries on both products
Run the same fixture against the two local endpoints created later in Lesson 4. Record the root version response, cluster name, the created mapping, document count, query result IDs, and any response-field differences. The acceptance criterion is not byte-for-byte equality; it is that you can separate the common retrieval idea from product-specific behavior.
curl "$ES_URL/"curl "$ES_URL/_cluster/health?pretty"curl "$ES_URL/atlasmart-products-v1/_mapping?pretty"curl "$OS_URL/"curl "$OS_URL/_cluster/health?pretty"curl "$OS_URL/atlasmart-products-v1/_mapping?pretty"
Do not publish a benchmark from three documents. For a future production comparison, hold the dataset, mappings/analyzers, shard topology, warmup, request mix and hardware constant; measure distributions such as p50/p95/p99 latency plus indexing freshness; and inspect failures as carefully as successful averages.
Production judgment
Choose Elasticsearch or OpenSearch for search/analytics because the workload benefits from their indexed retrieval and distributed execution—not because the APIs are convenient. Preserve an exit path through source data, reproducible mappings, index templates, fixtures, relevance judgments and migration tests. The next lesson examines why product choice itself is now an architectural variable.
Check your understanding
- Why is an inverted index called “inverted”?
- Why can an acknowledged indexing request still be a different event from ordinary search visibility?
- When is a relational or transactional database still the better authority even if Elasticsearch/OpenSearch serves reads?
- Why should exact BM25 scores not be frozen as universal expected values?
- What evidence would you hold constant for a fair Elasticsearch/OpenSearch workload comparison?
Review the answers
1. It reverses the document-to-terms view into a term-to-matching-documents access structure so retrieval can start from query terms instead of scanning every document.
2. Durability/acknowledgement and refresh/search visibility are separate mechanisms; normal search observes refreshed searchable structures, not merely receipt of a write request.
3. When the business requirement centers on transactional constraints, authoritative updates, joins/invariants or exact write semantics rather than ranked search and analytical retrieval.
4. Scores are query-relative and depend on analyzer output, document/field statistics, query structure and implementation/version details; they are ranking evidence, not probabilities.
5. Dataset, mapping/analyzers, shard/replica topology, cache/segment state, query mix, concurrency, resources, warmup, freshness targets and the same judged relevance set.
Summary and next step
Search engines are specialized serving systems whose value comes from index structures, analysis, ranking, filtering and distributed retrieval. Their storage capability does not erase transactional, freshness, security or recovery boundaries. Next, separate Elasticsearch from OpenSearch as products, ecosystems and operational contracts.
Authoritative references
- Elasticsearch 9.5.3 download page — Current release and default-distribution information.
- Elastic search documentation — Official entry point for Elasticsearch search concepts and current APIs.
- OpenSearch 3.8.0 version history — Current OpenSearch release history.
- OpenSearch documentation — Official current documentation entry point.
- Apache Lucene — Underlying Lucene project; use for implementation details, not as a substitute for product contracts.