Chapter 01 · Search Engine Foundations, Elasticsearch vs OpenSearch, Deployment Models, and Lab Setup

Inverted-Index Search vs Relational / Document / Vector Databases: Workload Fit and Architectural Tradeoffs

Decide when inverted-index search is the right serving architecture and prove the boundary with a tiny dual-platform AtlasMart fixture.

Intermediate100–120 minutesDual-platform mechanism labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart needs fast product discovery across names, descriptions, brands, categories, and attributes while its system of record still owns prices, inventory and orders. The architecture question is not “should everything move into Elasticsearch?” It is which read problem an inverted index solves better than a transactional database, where a document store or vector database is more appropriate, and what consistency and operational costs appear when search becomes a derived serving system.

01

Explain an inverted index, term dictionary, postings list, document, field, analyzer, query, score, shard, replica and Lucene segment without treating Elasticsearch/OpenSearch as generic JSON databases.

02

Separate search-system responsibilities from relational integrity, document-primary storage, and vector-nearest-neighbor retrieval.

03

Trace a tiny AtlasMart text query from indexed fields to matching documents and explain what the result does and does not guarantee.

04

Choose a search architecture from workload evidence: relevance, filtering, aggregation, freshness, durability, latency, write rate, schema and failure tolerance.

05

Run the same minimal mapping/index/query fixture on pinned Elasticsearch and OpenSearch while recording product/version/security differences.

Version baseline reviewed 10 September 2026

This chapter pins Elasticsearch 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 (released 4 August 2026) for reproducible examples. OpenSearch 3.9.0 is scheduled for 29 September 2026 and is therefore not treated as current. Re-check both projects before reusing these commands later. Elasticsearch and OpenSearch are independent products: shared Lucene ancestry does not make their APIs, plugins, security, lifecycle, vector features, clients, or managed offerings interchangeable.

Execution and safety note

The environment used to generate this lesson does not provide Docker, Elasticsearch, OpenSearch, Kibana, or OpenSearch Dashboards. The commands and API shapes were reviewed against the current official documentation but were not executed here. Expected output is described by invariant and field shape rather than presented as captured benchmark evidence. Every destructive action is scoped to atlasmart-* course containers, volumes, indices, and local loopback ports.

1. Search engines optimize a different access path

A relational index is usually designed to find rows by ordered key values or to support a known predicate. A full-text search engine instead builds structures that answer “which documents contain terms related to this query?” efficiently. At the center is the inverted index: instead of reading every product description and asking whether it contains “wireless keyboard”, the engine records terms and the document identifiers in which those terms occur. A term dictionary organizes searchable terms; a postings list associates a term with matching documents and can carry frequency/position information used by scoring and phrase queries.

Elasticsearch and OpenSearch expose JSON over HTTP, but the JSON document is an input and retrieval abstraction—not proof that the engine behaves like a document database. During indexing, fields can be parsed by mappings; text fields can pass through an analyzer that produces a token stream; Lucene writes immutable segments; and a search can execute across one or many shards before a coordinating node reduces results. Later chapters unpack each mechanism. For now, the key rule is that search quality and cost are created at both write time and query time.

System style Good default question Search-platform boundary
Relational database “Which order has this primary key, and can this transaction preserve constraints?” Keep authoritative transactions/invariants here unless search-specific evidence justifies duplication.
Document database “Can I retrieve/update an aggregate-shaped document by known keys?” A document model alone does not provide the same analyzed-text ranking and search execution model.
Lexical search engine “Which documents best match these terms plus structured filters?” Excellent for retrieval/analytics, but near-real-time visibility and distributed search semantics differ from OLTP transactions.
Vector database / vector index “Which vectors are nearest under a similarity metric?” Semantic similarity can complement lexical retrieval, but does not replace exact identifiers, filters, authorization or judged relevance.

2. Workload fit is about questions, not product labels

AtlasMart’s catalog search is a strong search-engine candidate because users do not know exact product identifiers. They type noisy natural-language terms, filter by category and brand, sort or boost by business signals, and expect ranked results. A relational database can support some of this, but building language analysis, typo tolerance, distributed relevance, aggregations and later hybrid retrieval usually creates a separate search concern anyway.

By contrast, “decrement inventory exactly once when payment commits” is not a search-first problem. Search results may be near real time: an indexing request can be acknowledged before a normal search sees the new document, depending on refresh behavior. Search copies also have their own mappings, index versions, snapshots and recovery paths. Treating a search index as an effortless mirror hides synchronization and failure modes.

Decision test

If the business requirement can be expressed only as “we need something fast,” the architecture is underspecified. Record query shapes, ranking needs, filters, aggregations, accepted freshness lag, update rate, document size, field cardinality, retention, failure objectives, tenant/security constraints, and the authoritative source of truth before selecting a platform.

3. Minimal common-denominator AtlasMart fixture

The following mapping deliberately uses conservative constructs supported by both current platforms: one shard, zero replicas for a single-node lab, keyword identifiers/facets, text analyzed content, and a scaled numeric price. It is a teaching fixture, not a production shard recommendation.

Dev Tools / REST · create a tiny product index
PUT /atlasmart-products-v1{  "settings": {"number_of_shards": 1, "number_of_replicas": 0},  "mappings": {    "properties": {      "product_id": {"type": "keyword"},      "name": {"type": "text", "fields": {"keyword": {"type": "keyword"}}},      "category": {"type": "keyword"},      "description": {"type": "text"},      "price": {"type": "scaled_float", "scaling_factor": 100}    }  }}
REST · index three deterministic documents
PUT /atlasmart-products-v1/_doc/P-1001?refresh=true{"product_id":"P-1001","name":"Quiet Wireless Keyboard","category":"keyboards","description":"compact wireless keyboard with quiet keys","price":49.90}PUT /atlasmart-products-v1/_doc/P-1002?refresh=true{"product_id":"P-1002","name":"Mechanical Gaming Keyboard","category":"keyboards","description":"wired mechanical keyboard with tactile switches","price":89.00}PUT /atlasmart-products-v1/_doc/P-1003?refresh=true{"product_id":"P-1003","name":"Wireless Travel Mouse","category":"mice","description":"compact wireless mouse for travel","price":29.50}
REST · lexical query plus structured filter
GET /atlasmart-products-v1/_search{  "query": {    "bool": {      "must": [{"match": {"description": "wireless keyboard"}}],      "filter": [{"term": {"category": "keyboards"}}]    }  },  "_source": ["product_id", "name", "category", "price"]}

The expected invariant is that P-1001 matches because its analyzed description contains terms corresponding to the query and its exact category satisfies the filter. Do not freeze an exact _score in course prose: score depends on query structure, field statistics, analyzer output and engine/version details. The refresh=true parameter is used only to make this tiny lesson deterministic; forcing refresh on every production write can damage indexing throughput.

4. What the engine is doing

At index time the engine parses the JSON, validates it against the mapping, analyzes the text fields, records searchable terms and field-oriented structures, writes operation state through its durability path, and eventually exposes the change to search through refresh/segment mechanics. At query time the match clause analyzes the query text; the term-level category filter uses the exact indexed value; shard-local searches produce candidates; and the coordinating path combines them into the response.

That flow explains several common mistakes. A keyword field is not analyzed like full text. A successful index response is not the same event as snapshot protection. A matching document does not prove the authoritative product is still in stock. A high score is not a calibrated probability that the result is “correct.” Each guarantee belongs to a different mechanism.

Deliberately wrong design

Store orders, payments, inventory decrements, product-search documents and semantic embeddings only in the search cluster because it “can store JSON.” The concrete failure is ownership confusion: search refresh, mapping changes, reindexing, relevance experiments and shard recovery now collide with transactional invariants. Repair the design by naming the authoritative system for each fact, treating search as a purpose-built serving index where appropriate, and defining replay/reconciliation for synchronization.

5. Lab: prove fit and boundaries on both products

Run the same fixture against the two local endpoints created later in Lesson 4. Record the root version response, cluster name, the created mapping, document count, query result IDs, and any response-field differences. The acceptance criterion is not byte-for-byte equality; it is that you can separate the common retrieval idea from product-specific behavior.

curl · evidence checklist (substitute secure auth options from Lesson 4)
curl "$ES_URL/"curl "$ES_URL/_cluster/health?pretty"curl "$ES_URL/atlasmart-products-v1/_mapping?pretty"curl "$OS_URL/"curl "$OS_URL/_cluster/health?pretty"curl "$OS_URL/atlasmart-products-v1/_mapping?pretty"

Do not publish a benchmark from three documents. For a future production comparison, hold the dataset, mappings/analyzers, shard topology, warmup, request mix and hardware constant; measure distributions such as p50/p95/p99 latency plus indexing freshness; and inspect failures as carefully as successful averages.

Production judgment

Choose Elasticsearch or OpenSearch for search/analytics because the workload benefits from their indexed retrieval and distributed execution—not because the APIs are convenient. Preserve an exit path through source data, reproducible mappings, index templates, fixtures, relevance judgments and migration tests. The next lesson examines why product choice itself is now an architectural variable.

Check your understanding

  1. Why is an inverted index called “inverted”?
  2. Why can an acknowledged indexing request still be a different event from ordinary search visibility?
  3. When is a relational or transactional database still the better authority even if Elasticsearch/OpenSearch serves reads?
  4. Why should exact BM25 scores not be frozen as universal expected values?
  5. What evidence would you hold constant for a fair Elasticsearch/OpenSearch workload comparison?
Review the answers

1. It reverses the document-to-terms view into a term-to-matching-documents access structure so retrieval can start from query terms instead of scanning every document.

2. Durability/acknowledgement and refresh/search visibility are separate mechanisms; normal search observes refreshed searchable structures, not merely receipt of a write request.

3. When the business requirement centers on transactional constraints, authoritative updates, joins/invariants or exact write semantics rather than ranked search and analytical retrieval.

4. Scores are query-relative and depend on analyzer output, document/field statistics, query structure and implementation/version details; they are ranking evidence, not probabilities.

5. Dataset, mapping/analyzers, shard/replica topology, cache/segment state, query mix, concurrency, resources, warmup, freshness targets and the same judged relevance set.

Summary and next step

Search engines are specialized serving systems whose value comes from index structures, analysis, ranking, filtering and distributed retrieval. Their storage capability does not erase transactional, freshness, security or recovery boundaries. Next, separate Elasticsearch from OpenSearch as products, ecosystems and operational contracts.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.