Chapter 02 · Documents, Indices, Data Streams, Shards, Replicas, Nodes, and Cluster Architecture
JSON Documents, _source, Metadata Fields, IDs, Versions/Sequence Numbers, and Optimistic Concurrency
Separate source JSON, engine metadata, document identity, optimistic concurrency and search visibility before relying on distributed behavior.
Learning outcomes
AtlasMart now has a search-serving copy of product data, but application developers still need to know what a “document” actually means operationally. If two catalog workers update the same product, a naive last-write-wins overwrite can lose fields. If an indexing call is acknowledged and a search immediately returns no hit, that is not necessarily data loss. This lesson separates stored source, indexed structures, metadata, identity, concurrency and search visibility before shards and routing are expanded in the next lesson.
Distinguish a JSON request body, the stored
_source, indexed field structures, metadata
fields and the search hit returned to a client.
Explain _index, _id,
_version, _seq_no,
_primary_term and _routing without
treating them as interchangeable identifiers.
Use sequence number plus primary term for optimistic concurrency control and diagnose a deliberate stale-write conflict.
Predict the difference between real-time GET behavior and near-real-time search visibility around refresh.
Keep AtlasMart business identity independent of engine-internal metadata so reindexing and migration remain possible.
This chapter continues the Chapter 01 lab baseline with
Elasticsearch 9.5.3 (released 3 September 2026)
and OpenSearch 3.8.0 (released 4 August 2026).
The existing local endpoints remain
https://localhost:9200 for Elasticsearch and
https://localhost:9201 for OpenSearch.
Elasticsearch requests use the Chapter 01 HTTP CA plus the
local elastic password; OpenSearch uses its local
demo TLS/admin path only for the disposable course lab.
Re-check versions, APIs and security defaults before reusing
these examples later.
The environment used to generate this chapter does not provide
Docker, Elasticsearch or OpenSearch, so commands were reviewed
against current official documentation but were not executed
here. Expected output is described by field shape and
invariant rather than fabricated as captured output. Failure
injection targets only disposable
atlasmart-* indices and the isolated Chapter 01
course containers; every destructive or allocation-changing
step includes an explicit reset path.
1. A document has several representations
Both Elasticsearch and OpenSearch accept JSON documents, but the
request body is not the only representation the engine
maintains. The _source field is
the source document retained for retrieval when source storage
is enabled. Searchable fields are separately encoded into Lucene
structures according to the mapping: analyzed text can become
tokens in postings; keyword/numeric/date fields can have
structures for filtering, sorting or aggregation; and metadata
such as document ID and routing participates in indexing and
lookup.
This distinction matters because _source itself is
not “the inverted index.” A value can be present in
_source yet not be searchable in the same way you
expect if the mapping disables indexing or chooses an
inappropriate field type. Conversely, derived indexed structures
can support search behavior that is not visible by merely
printing the original JSON.
| Surface | What it represents | Do not infer |
|---|---|---|
_source |
The source JSON retained for retrieval, subject to product/settings behavior. | That every field is indexed, sortable or aggregatable in the same form. |
_id |
Document identity within the target index/data stream write context. | A globally unique AtlasMart business key across every index or future migration. |
_version |
A document version counter exposed by document APIs. | That a numeric version alone captures primary failover identity or is the preferred modern read-modify-write guard. |
_seq_no |
An operation sequence number assigned by the primary shard. | A global cluster transaction number or stable value across reindexing. |
_primary_term |
Identity epoch for the primary shard generation used with sequence numbers. | A replica count, shard ID or document revision. |
_routing |
The routing value that selects a primary shard; default
derives from _id.
|
A security tenant boundary or proof that workload distribution is balanced. |
2. Business identity should survive reindexing
AtlasMart uses product_id such as
P-1001 as the durable domain identifier. The engine
also needs a document _id. For a simple catalog,
using the same stable product key as _id is
convenient because retries can target the same document instead
of creating duplicates. That is an application design choice,
not a requirement that all search documents mirror business keys
one-for-one.
Keep the domain key in the document body as well. Reindexing, aliases, data-stream generations and future Elasticsearch ↔ OpenSearch migrations operate on index resources whose metadata can change. A query, export or downstream event should not have to recover business identity from a physical index name or shard location.
DELETE /atlasmart-products-v2PUT /atlasmart-products-v2{ "settings": { "number_of_shards": 2, "number_of_replicas": 0, "refresh_interval": "30s" }, "mappings": { "properties": { "product_id": {"type": "keyword"}, "name": {"type": "text"}, "category": {"type": "keyword"}, "price": {"type": "double"}, "updated_at": {"type": "date"} } }}
The two-primary-shard setting is deliberate for routing experiments in this chapter. The replica count stays zero initially because the Chapter 01 baseline has one node per product. Replica behavior is introduced explicitly rather than leaving every one-node lab yellow and asking learners to ignore it.
3. Sequence number + primary term protect read-modify-write
Distributed write systems can process concurrent requests in
different orders. Elasticsearch and OpenSearch assign each
changing operation a sequence number at the
primary shard. A primary term identifies the
current primary generation and changes when a new primary
generation is established. Together, the current
_seq_no and _primary_term provide a
compare-and-set style guard for document APIs.
Suppose two catalog editors read product P-2001.
Editor A changes the price; Editor B changes the name. If B
blindly replaces the entire document after A, B can erase A’s
newer price. Instead, B can say “apply this replacement only if
the document is still exactly the revision I read.” A mismatch
returns a version conflict rather than silently losing the other
update.
PUT /atlasmart-products-v2/_doc/P-2001{ "product_id": "P-2001", "name": "Atlas Wireless Keyboard", "category": "keyboards", "price": 79.00, "updated_at": "2026-09-10T16:00:00Z"}GET /atlasmart-products-v2/_doc/P-2001# Copy the actual _seq_no and _primary_term from the GET response.PUT /atlasmart-products-v2/_doc/P-2001?if_seq_no=<SEQ>&if_primary_term=<TERM>{ "product_id": "P-2001", "name": "Atlas Wireless Keyboard", "category": "keyboards", "price": 74.00, "updated_at": "2026-09-10T16:05:00Z"}
Do not hard-code the example sequence number. The correct values are whatever your own GET returns. A successful guarded write proves only that the document had not changed relative to those observed metadata values at the conditional check.
After the successful guarded update, repeat the request using
the old if_seq_no/if_primary_term.
Expect an HTTP 409 version-conflict style failure rather than
a silent overwrite. Diagnose it by GETting the current
document and metadata, deciding whether to merge/retry or
reject at the application level, then issue a new guarded
operation only if that business action is still valid.
4. GET and search have different visibility contracts
Search is near real time because newly indexed operations become visible to search after a refresh opens new searchable segment state. A document GET is designed as a document lookup and is real-time by default in Elasticsearch; OpenSearch document APIs likewise expose the current stored document independently of ordinary search refresh timing. Therefore, “GET found it but search did not” can be a refresh-boundary observation rather than corruption.
The lab sets refresh_interval to 30 seconds
precisely so the distinction is easier to observe. Index a new
product without forcing refresh, immediately compare GET and
search, then call the Refresh API and repeat the search. Do not
promise that a pre-refresh search will always miss—an automatic
refresh may occur between requests. The lesson is to inspect the
mechanism, not manufacture a race outcome.
PUT /atlasmart-products-v2/_doc/P-2002{ "product_id": "P-2002", "name": "Atlas Mechanical Keyboard", "category": "keyboards", "price": 99.00, "updated_at": "2026-09-10T16:10:00Z"}GET /atlasmart-products-v2/_doc/P-2002GET /atlasmart-products-v2/_search{ "query": {"term": {"product_id": "P-2002"}}}POST /atlasmart-products-v2/_refreshGET /atlasmart-products-v2/_search{ "query": {"term": {"product_id": "P-2002"}}}
If a workflow must index and then search the new state, prefer
the product’s supported
refresh=wait_for semantics where appropriate
rather than forcing refresh=true on every write.
Forced refreshes can create inefficient small-segment churn.
Later chapters quantify refresh/merge costs.
5. Observe metadata without confusing it with guarantees
GET /atlasmart-products-v2/_search?seq_no_primary_term=true{ "query": {"match_all": {}}, "sort": [{"product_id": "asc"}]}
Inspect _index, _id,
_seq_no, _primary_term and
_source on each hit. These fields help correlate a
request with a document revision. They do not prove that
replicas exist, snapshots are valid, the search was authorized
correctly for every tenant, or that the document is the
authoritative business record.
Elasticsearch 9.x also has version-sensitive source-storage
options such as synthetic _source in supported
subscription contexts. Do not teach that storage optimization as
if it were a universal OpenSearch feature or as if reconstructed
source were always byte-for-byte identical to the original
request. This chapter uses ordinary source storage so document
semantics stay clear.
Verification checklist
-
atlasmart-products-v2has exactly two primary shards and zero replicas for the initial one-node-per-product lab. -
Both engines return document identity and current
sequence/primary-term metadata for
P-2001. - A stale conditional update is rejected rather than overwriting the newer revision.
-
After an explicit refresh, search returns
P-2002. -
The business key remains in
_source.product_idand is not derived from physical shard placement.
Production judgment
Use engine metadata to enforce technical concurrency, but keep business conflict policy in the application. A 409 tells AtlasMart that the precondition became stale; it does not decide whether to merge, retry, abandon or compensate. Stable IDs improve idempotency, but custom routing and multi-tenant identity can change the correct key strategy. Source-storage choices trade retrieval capability and disk cost. Refresh requirements trade freshness against indexing efficiency. Those constraints must be explicit before Chapter 03 turns field mappings into a schema contract.
Check your understanding
-
Why is
_sourcenot the same thing as the inverted index? -
What do
_seq_noand_primary_termprotect together? - Why can GET succeed before an ordinary search returns the document?
-
Why keep
product_idin the JSON body even if it is also used as_id? - What should an application do after a stale OCC write receives a 409?
Review the answers
1. It is the retained source representation used for retrieval; searchable structures are separately built from mappings and analysis.
2. They identify the document revision/primary generation used by optimistic concurrency control so a stale read-modify-write can be rejected.
3. Document lookup and search visibility have different contracts; search depends on refresh opening searchable segment state.
4. The domain key survives reindexing, alias changes, exports and product migration without depending on physical engine metadata.
5. Read the current state, apply business conflict logic, and retry with current metadata only if the business action is still valid; do not blindly loop every 409.
Summary and next step
A search document is a source payload plus indexed structures and engine metadata, not merely “JSON in a database.” Sequence numbers and primary terms make stale writes observable; refresh separates acknowledged indexing from search visibility. Next, place these documents inside indices, aliases, data streams and shard-routing structures so you can predict where each request goes.
Authoritative references
- Elasticsearch <code>_source</code> field — Official distinction between retained source and indexed/searchable structures, including version-sensitive source options.
- Elasticsearch optimistic concurrency control — Current sequence-number and primary-term conditional-write semantics.
- Elasticsearch near real-time search — Refresh and search-visibility model.
- OpenSearch document APIs — Current document metadata, sequence-number/primary-term and replication overview.
-
OpenSearch Get Document API
— Current response metadata including
_source, routing and concurrency fields. - OpenSearch Index Document API — Conditional indexing, refresh and active-shard controls.