Chapter 02 · Documents, Indices, Data Streams, Shards, Replicas, Nodes, and Cluster Architecture

JSON Documents, _source, Metadata Fields, IDs, Versions/Sequence Numbers, and Optimistic Concurrency

Separate source JSON, engine metadata, document identity, optimistic concurrency and search visibility before relying on distributed behavior.

Intermediate100–125 minutesDistributed document/shard labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart now has a search-serving copy of product data, but application developers still need to know what a “document” actually means operationally. If two catalog workers update the same product, a naive last-write-wins overwrite can lose fields. If an indexing call is acknowledged and a search immediately returns no hit, that is not necessarily data loss. This lesson separates stored source, indexed structures, metadata, identity, concurrency and search visibility before shards and routing are expanded in the next lesson.

01

Distinguish a JSON request body, the stored _source, indexed field structures, metadata fields and the search hit returned to a client.

02

Explain _index, _id, _version, _seq_no, _primary_term and _routing without treating them as interchangeable identifiers.

03

Use sequence number plus primary term for optimistic concurrency control and diagnose a deliberate stale-write conflict.

04

Predict the difference between real-time GET behavior and near-real-time search visibility around refresh.

05

Keep AtlasMart business identity independent of engine-internal metadata so reindexing and migration remain possible.

Version baseline reviewed 10 September 2026

This chapter continues the Chapter 01 lab baseline with Elasticsearch 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 (released 4 August 2026). The existing local endpoints remain https://localhost:9200 for Elasticsearch and https://localhost:9201 for OpenSearch. Elasticsearch requests use the Chapter 01 HTTP CA plus the local elastic password; OpenSearch uses its local demo TLS/admin path only for the disposable course lab. Re-check versions, APIs and security defaults before reusing these examples later.

Execution and safety note

The environment used to generate this chapter does not provide Docker, Elasticsearch or OpenSearch, so commands were reviewed against current official documentation but were not executed here. Expected output is described by field shape and invariant rather than fabricated as captured output. Failure injection targets only disposable atlasmart-* indices and the isolated Chapter 01 course containers; every destructive or allocation-changing step includes an explicit reset path.

1. A document has several representations

Both Elasticsearch and OpenSearch accept JSON documents, but the request body is not the only representation the engine maintains. The _source field is the source document retained for retrieval when source storage is enabled. Searchable fields are separately encoded into Lucene structures according to the mapping: analyzed text can become tokens in postings; keyword/numeric/date fields can have structures for filtering, sorting or aggregation; and metadata such as document ID and routing participates in indexing and lookup.

This distinction matters because _source itself is not “the inverted index.” A value can be present in _source yet not be searchable in the same way you expect if the mapping disables indexing or chooses an inappropriate field type. Conversely, derived indexed structures can support search behavior that is not visible by merely printing the original JSON.

Surface What it represents Do not infer
_source The source JSON retained for retrieval, subject to product/settings behavior. That every field is indexed, sortable or aggregatable in the same form.
_id Document identity within the target index/data stream write context. A globally unique AtlasMart business key across every index or future migration.
_version A document version counter exposed by document APIs. That a numeric version alone captures primary failover identity or is the preferred modern read-modify-write guard.
_seq_no An operation sequence number assigned by the primary shard. A global cluster transaction number or stable value across reindexing.
_primary_term Identity epoch for the primary shard generation used with sequence numbers. A replica count, shard ID or document revision.
_routing The routing value that selects a primary shard; default derives from _id. A security tenant boundary or proof that workload distribution is balanced.

2. Business identity should survive reindexing

AtlasMart uses product_id such as P-1001 as the durable domain identifier. The engine also needs a document _id. For a simple catalog, using the same stable product key as _id is convenient because retries can target the same document instead of creating duplicates. That is an application design choice, not a requirement that all search documents mirror business keys one-for-one.

Keep the domain key in the document body as well. Reindexing, aliases, data-stream generations and future Elasticsearch ↔ OpenSearch migrations operate on index resources whose metadata can change. A query, export or downstream event should not have to recover business identity from a physical index name or shard location.

REST · create the Chapter 02 product fixture
DELETE /atlasmart-products-v2PUT /atlasmart-products-v2{  "settings": {    "number_of_shards": 2,    "number_of_replicas": 0,    "refresh_interval": "30s"  },  "mappings": {    "properties": {      "product_id": {"type": "keyword"},      "name": {"type": "text"},      "category": {"type": "keyword"},      "price": {"type": "double"},      "updated_at": {"type": "date"}    }  }}

The two-primary-shard setting is deliberate for routing experiments in this chapter. The replica count stays zero initially because the Chapter 01 baseline has one node per product. Replica behavior is introduced explicitly rather than leaving every one-node lab yellow and asking learners to ignore it.

3. Sequence number + primary term protect read-modify-write

Distributed write systems can process concurrent requests in different orders. Elasticsearch and OpenSearch assign each changing operation a sequence number at the primary shard. A primary term identifies the current primary generation and changes when a new primary generation is established. Together, the current _seq_no and _primary_term provide a compare-and-set style guard for document APIs.

Suppose two catalog editors read product P-2001. Editor A changes the price; Editor B changes the name. If B blindly replaces the entire document after A, B can erase A’s newer price. Instead, B can say “apply this replacement only if the document is still exactly the revision I read.” A mismatch returns a version conflict rather than silently losing the other update.

REST · create, read and perform guarded update
PUT /atlasmart-products-v2/_doc/P-2001{  "product_id": "P-2001",  "name": "Atlas Wireless Keyboard",  "category": "keyboards",  "price": 79.00,  "updated_at": "2026-09-10T16:00:00Z"}GET /atlasmart-products-v2/_doc/P-2001# Copy the actual _seq_no and _primary_term from the GET response.PUT /atlasmart-products-v2/_doc/P-2001?if_seq_no=<SEQ>&if_primary_term=<TERM>{  "product_id": "P-2001",  "name": "Atlas Wireless Keyboard",  "category": "keyboards",  "price": 74.00,  "updated_at": "2026-09-10T16:05:00Z"}

Do not hard-code the example sequence number. The correct values are whatever your own GET returns. A successful guarded write proves only that the document had not changed relative to those observed metadata values at the conditional check.

Deliberate stale-write failure

After the successful guarded update, repeat the request using the old if_seq_no/if_primary_term. Expect an HTTP 409 version-conflict style failure rather than a silent overwrite. Diagnose it by GETting the current document and metadata, deciding whether to merge/retry or reject at the application level, then issue a new guarded operation only if that business action is still valid.

4. GET and search have different visibility contracts

Search is near real time because newly indexed operations become visible to search after a refresh opens new searchable segment state. A document GET is designed as a document lookup and is real-time by default in Elasticsearch; OpenSearch document APIs likewise expose the current stored document independently of ordinary search refresh timing. Therefore, “GET found it but search did not” can be a refresh-boundary observation rather than corruption.

The lab sets refresh_interval to 30 seconds precisely so the distinction is easier to observe. Index a new product without forcing refresh, immediately compare GET and search, then call the Refresh API and repeat the search. Do not promise that a pre-refresh search will always miss—an automatic refresh may occur between requests. The lesson is to inspect the mechanism, not manufacture a race outcome.

REST · compare document lookup and search visibility
PUT /atlasmart-products-v2/_doc/P-2002{  "product_id": "P-2002",  "name": "Atlas Mechanical Keyboard",  "category": "keyboards",  "price": 99.00,  "updated_at": "2026-09-10T16:10:00Z"}GET /atlasmart-products-v2/_doc/P-2002GET /atlasmart-products-v2/_search{  "query": {"term": {"product_id": "P-2002"}}}POST /atlasmart-products-v2/_refreshGET /atlasmart-products-v2/_search{  "query": {"term": {"product_id": "P-2002"}}}
Production pattern

If a workflow must index and then search the new state, prefer the product’s supported refresh=wait_for semantics where appropriate rather than forcing refresh=true on every write. Forced refreshes can create inefficient small-segment churn. Later chapters quantify refresh/merge costs.

5. Observe metadata without confusing it with guarantees

REST · ask search to return concurrency metadata
GET /atlasmart-products-v2/_search?seq_no_primary_term=true{  "query": {"match_all": {}},  "sort": [{"product_id": "asc"}]}

Inspect _index, _id, _seq_no, _primary_term and _source on each hit. These fields help correlate a request with a document revision. They do not prove that replicas exist, snapshots are valid, the search was authorized correctly for every tenant, or that the document is the authoritative business record.

Elasticsearch 9.x also has version-sensitive source-storage options such as synthetic _source in supported subscription contexts. Do not teach that storage optimization as if it were a universal OpenSearch feature or as if reconstructed source were always byte-for-byte identical to the original request. This chapter uses ordinary source storage so document semantics stay clear.

Verification checklist

  • atlasmart-products-v2 has exactly two primary shards and zero replicas for the initial one-node-per-product lab.
  • Both engines return document identity and current sequence/primary-term metadata for P-2001.
  • A stale conditional update is rejected rather than overwriting the newer revision.
  • After an explicit refresh, search returns P-2002.
  • The business key remains in _source.product_id and is not derived from physical shard placement.

Production judgment

Use engine metadata to enforce technical concurrency, but keep business conflict policy in the application. A 409 tells AtlasMart that the precondition became stale; it does not decide whether to merge, retry, abandon or compensate. Stable IDs improve idempotency, but custom routing and multi-tenant identity can change the correct key strategy. Source-storage choices trade retrieval capability and disk cost. Refresh requirements trade freshness against indexing efficiency. Those constraints must be explicit before Chapter 03 turns field mappings into a schema contract.

Check your understanding

  1. Why is _source not the same thing as the inverted index?
  2. What do _seq_no and _primary_term protect together?
  3. Why can GET succeed before an ordinary search returns the document?
  4. Why keep product_id in the JSON body even if it is also used as _id?
  5. What should an application do after a stale OCC write receives a 409?
Review the answers

1. It is the retained source representation used for retrieval; searchable structures are separately built from mappings and analysis.

2. They identify the document revision/primary generation used by optimistic concurrency control so a stale read-modify-write can be rejected.

3. Document lookup and search visibility have different contracts; search depends on refresh opening searchable segment state.

4. The domain key survives reindexing, alias changes, exports and product migration without depending on physical engine metadata.

5. Read the current state, apply business conflict logic, and retry with current metadata only if the business action is still valid; do not blindly loop every 409.

Summary and next step

A search document is a source payload plus indexed structures and engine metadata, not merely “JSON in a database.” Sequence numbers and primary terms make stale writes observable; refresh separates acknowledged indexing from search visibility. Next, place these documents inside indices, aliases, data streams and shard-routing structures so you can predict where each request goes.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.