Chapter 02 · Documents, Indices, Data Streams, Shards, Replicas, Nodes, and Cluster Architecture
Trace One Document from Index Request Through Refresh to Search Result and Explain Every Distributed Step
Trace one AtlasMart document from index request through routing, revision metadata and refresh to a verified search result.
Learning outcomes
The chapter closes by tracing one AtlasMart document end to end. The goal is not another API recipe: it is to predict each state transition and prove it using metadata. You will choose a business ID and routing key, observe the selected primary-shard group, capture concurrency metadata, separate write acknowledgement from search refresh, then search and correlate the returned hit with the original write.
Trace one product from client request through routing, primary execution, sequence metadata, refresh and search result.
Correlate document metadata with shard placement using GET,
search, _search_shards and CAT/health evidence.
Demonstrate that write acknowledgement, durability/replication preconditions and search visibility are distinct concepts.
Use a data-stream event beside the mutable product index to compare two legitimate resource models.
Produce a Chapter 02 evidence bundle and reset procedure that leaves the Chapter 01 lab ready for Chapter 03 mappings.
This chapter continues the Chapter 01 lab baseline with
Elasticsearch 9.5.3 (released 3 September 2026)
and OpenSearch 3.8.0 (released 4 August 2026).
The existing local endpoints remain
https://localhost:9200 for Elasticsearch and
https://localhost:9201 for OpenSearch.
Elasticsearch requests use the Chapter 01 HTTP CA plus the
local elastic password; OpenSearch uses its local
demo TLS/admin path only for the disposable course lab.
Re-check versions, APIs and security defaults before reusing
these examples later.
The environment used to generate this chapter does not provide
Docker, Elasticsearch or OpenSearch, so commands were reviewed
against current official documentation but were not executed
here. Expected output is described by field shape and
invariant rather than fabricated as captured output. Failure
injection targets only disposable
atlasmart-* indices and the isolated Chapter 01
course containers; every destructive or allocation-changing
step includes an explicit reset path.
1. Define the trace before sending the request
The trace document is P-TRACE-01 for tenant
TENANT-TRACE. The product index is
atlasmart-products-v2 behind alias
atlasmart-products. It has two primary shards and
zero replicas in the mandatory one-node baseline. The
search-event stream is atlasmart-search-events.
These names are not arbitrary: they make every cleanup command
narrow and every observation attributable to the course.
| Trace element | Chosen value | Why it is explicit |
|---|---|---|
| Business/document ID | P-TRACE-01 |
Makes idempotent replacement and evidence correlation straightforward. |
| Routing key | TENANT-TRACE |
Allows deterministic shard selection to be inspected. |
| Mutable resource | atlasmart-products-v2 |
Product state can be replaced/updated by stable ID. |
| Stable app read name | atlasmart-products alias |
Separates application resource name from index generation. |
| Append-only event resource | atlasmart-search-events data stream |
Captures the search event independently of mutable product state. |
| Visibility control | Explicit refresh only at the trace checkpoint | Keeps acknowledgement and searchability separate. |
2. Step A — resolve routing, then index the document
Ask the cluster which shard group the routing value maps to before writing. Then send the document with the same routing value. The receiving node coordinates the request; routing chooses one primary shard; the primary executes the operation and assigns sequence metadata; configured active replicas receive the operation according to the replication model. In the mandatory baseline there are no replicas, so the response should not be misrepresented as replicated high availability.
GET /_search_shards/atlasmart-products-v2?routing=TENANT-TRACEGET /_cat/shards/atlasmart-products-v2?v=true&h=index,shard,prirep,state,node
PUT /atlasmart-products-v2/_doc/P-TRACE-01?routing=TENANT-TRACE{ "product_id": "P-TRACE-01", "name": "Atlas Trace Keyboard", "category": "keyboards", "price": 88.00, "updated_at": "2026-09-10T17:30:00Z"}
Capture the actual response fields _index,
_id, _version, _seq_no,
_primary_term, result and
_shards. Do not substitute the example values from
documentation; your trace is evidence only if it records the
engine’s response.
3. Step B — prove current document state before proving search visibility
Immediately perform a routed GET. The correct routing key is part of the lookup because the document was custom-routed. The GET should return the current source and concurrency metadata. Then perform a routed search for the exact product key before the explicit refresh. Depending on whether an automatic refresh occurred, it may already be visible. Record the observation either way; the contract is that refresh governs ordinary search visibility, not that your manual timing always wins a race.
GET /atlasmart-products-v2/_doc/P-TRACE-01?routing=TENANT-TRACEGET /atlasmart-products-v2/_search?routing=TENANT-TRACE&seq_no_primary_term=true{ "query": {"term": {"product_id": "P-TRACE-01"}}}POST /atlasmart-products-v2/_refreshGET /atlasmart-products-v2/_search?routing=TENANT-TRACE&seq_no_primary_term=true{ "query": {"term": {"product_id": "P-TRACE-01"}}}
After the explicit refresh completes, the routed search should
see the indexed document unless another failure changed the
state. Correlate the hit’s _id,
_seq_no, _primary_term and
_source.product_id with the GET/write evidence. If
they differ because the document was updated between steps, that
is a useful concurrency observation, not a reason to edit the
evidence.
4. Step C — make a guarded change and record the resulting search event
Take the current sequence number and primary term from the routed GET and perform one conditional price change. This demonstrates that the trace can protect against stale replacement while still preserving the same business ID/routing key. Then append a search event to the data stream. The two resources now show why one course needs both mutable-index and append-only data-stream semantics.
PUT /atlasmart-products-v2/_doc/P-TRACE-01?routing=TENANT-TRACE&if_seq_no=<CURRENT_SEQ>&if_primary_term=<CURRENT_TERM>{ "product_id": "P-TRACE-01", "name": "Atlas Trace Keyboard", "category": "keyboards", "price": 84.00, "updated_at": "2026-09-10T17:35:00Z"}POST /atlasmart-products-v2/_refresh
POST /atlasmart-search-events/_doc{ "@timestamp": "2026-09-10T17:36:00Z", "event_id": "SE-TRACE-01", "customer_id": "C-TRACE", "query": "atlas trace keyboard", "result_count": 1}POST /atlasmart-search-events/_refreshGET /atlasmart-search-events/_search{ "query": {"term": {"event_id": "SE-TRACE-01"}}}
The product document is updated in place by stable identity in the current index; the event stream records a new timestamped observation. Later chapters will add templates, lifecycle policies, ingest pipelines and zero-downtime index changes without collapsing these workloads into one generic “JSON collection.”
5. Build the evidence bundle: what each observation proves
| Evidence | What it proves | What it does not prove |
|---|---|---|
| Root + node/version response | Which product/version/node answered the endpoint. | Relevance correctness, data completeness, backup health or failover readiness. |
_search_shards?routing=... |
Which shard groups/copies are candidates for that routed search on the current topology. | That tenant routing is secure or balanced under future traffic. |
| Index response metadata | The result and revision metadata reported for that write. | Immediate ordinary-search visibility or independent backup. |
| Routed GET | Current document state addressed by ID+routing under document-lookup semantics. | That every search shard has refreshed the revision. |
| Refresh response | Target shards completed a refresh request according to API semantics. | That data is fsynced as a snapshot, replicated off-host or immune to deletion. |
Search hit + _shards |
The visible matching document and shard-level search completion state. | That another query/routing key would return the same set. |
| Data stream inventory | Backing/write-index relationship for the stream. | That ILM/ISM retention, rollover and disaster recovery are configured. |
A useful course lab preserves these observations as text/JSON alongside the time and commands used. Do not publish secret-bearing headers, passwords, CA private keys or cloud credentials. Redact only secrets; keep version, index, shard and error metadata that makes diagnosis possible.
Index with custom routing, GET without routing, force
refresh=true on every write, ignore
_shards.failed, and claim
“Elasticsearch/OpenSearch lost the document” when one request
returns no hit. Repair by reproducing the original
target/routing, separating GET from search visibility,
inspecting shard selection and failures, and using refresh
controls intentionally.
6. Reset Chapter 02 without damaging the Chapter 01 environment
The chapter reset removes only the new product index/alias, data stream and its template. It leaves the pinned containers, persistent volumes, network and Chapter 01 certificate/credentials in place for Chapter 03. Because deleting a data stream also deletes its backing indices, that operation is intentionally destructive to the stream fixture.
DELETE /atlasmart-products-v2DELETE /_data_stream/atlasmart-search-eventsDELETE /_index_template/atlasmart-search-events-template# Verify only the named Chapter 02 resources are gone.GET /_cat/indices/atlasmart-products-v2?v=trueGET /_data_stream/atlasmart-search-eventsGET /_index_template/atlasmart-search-events-template
If you want to continue directly into Chapter 03 instead of
resetting, keep atlasmart-products-v2 and the
stream, but record their exact mappings/settings first. Chapter
03 will deliberately revisit field types, and the course must
know whether it is evolving an existing index or creating a
clean mapping fixture.
Chapter 02 acceptance checklist
- All document metadata terms are defined and demonstrated with actual response shapes.
- OCC stale writes fail visibly with current sequence/primary-term semantics.
- Alias, index, data stream and backing index are kept distinct.
- Custom routing is correlated with shard selection and its missing-routing failure mode.
- Replica count is varied safely and allocation failures are explained, not ignored.
-
Elasticsearch
masterand OpenSearchcluster_managerterminology remains product-specific. - Write and search paths are traced through coordination, primary/replica and scatter/gather mechanisms.
- Partial-search handling is an explicit application correctness choice.
- No benchmark, latency, recovery or failover number is invented.
- All cleanup operations name only Chapter 02 course resources.
Production judgment
A durable search architecture must make five contracts explicit: business document identity, shard/routing topology, replication/failure tolerance, search freshness, and completeness under partial failure. Changing any one can alter latency, recovery time, write cost and correctness. Data streams solve a particular append-heavy lifecycle problem; aliases solve a different naming/cutover problem. Replica copies improve availability but are not backups. Routing can reduce fan-out but can create hot shards. The next chapter turns mappings and field types into the schema layer that determines how these documents are actually indexed and queried.
Check your understanding
-
What is the complete path of
P-TRACE-01from HTTP request to visible search hit? - Which step makes the document searchable, and which step establishes its write revision metadata?
- Why is the search event stored in a data stream while the product uses a mutable index?
- What evidence would you check before claiming a search response is complete?
- Why does Chapter 02 cleanup leave the Docker volumes and credentials in place?
Review the answers
1. Receiving node coordinates → routing selects a primary-shard group → primary executes/assigns revision metadata → active replicas are handled by the replication model → refresh opens searchable segment state → a routed search executes on the relevant shard copy and the coordinator returns the hit.
2. Refresh governs ordinary search visibility; the primary-shard write operation assigns the sequence/primary-term metadata for the revision.
3. Products are mutable current-state entities addressed by stable ID, while search events are append-heavy timestamped observations that fit backing-index/rollover semantics.
4. Inspect the response shard summary, failures, timeout/partial-result indicators, target index/routing and application-specific completeness policy.
5. Those are the reusable Chapter 01 platform baseline. Chapter 02 cleanup should remove only the chapter’s data resources so later chapters can continue without recreating unrelated infrastructure.
Summary and next step
You can now follow one document from JSON source and stable business identity through routing, primary-shard sequencing, replica policy, refresh and distributed search evidence. You can also distinguish mutable product indices from append-only data streams and explain node/control-plane roles. Chapter 03 uses this architecture to answer the next question: how should each field be mapped so indexing and queries have the intended semantics?
Authoritative references
- Elasticsearch optimistic concurrency control — Sequence-number/primary-term revision control.
- Elasticsearch near real-time search — Refresh-to-searchability mechanism.
- Elasticsearch routing field — Shard selection and custom-routing requirements.
- Elasticsearch data streams — Backing/write-index semantics for append-heavy data.
- OpenSearch document APIs — Document metadata and replication model.
- OpenSearch Refresh Index API — Current refresh/search visibility and production guidance.
- OpenSearch data streams — Backing indices, current write index and time-series constraints.
- OpenSearch search shard routing — Shard selection and routing for searches.