Chapter 02 · Documents, Indices, Data Streams, Shards, Replicas, Nodes, and Cluster Architecture
Indices, Aliases, Data Streams, Backing Indices, Shards, Replicas, and Routing
Connect indices, aliases, data streams, backing indices, shard copies and routing to observable document placement and failure signals.
Learning outcomes
AtlasMart’s catalog and search-event workloads now need two different resource shapes. Products are mutable entities that benefit from a versioned index plus a stable alias. Search events are append-heavy time-series documents that fit a data stream. Both are physically partitioned into shards, and document routing determines which primary shard owns each write. This lesson connects those abstractions without pretending an index is a single file or a data stream is just an alias with a fancy name.
Distinguish index, alias, data stream, backing index, primary shard and replica shard as separate layers.
Predict default and custom routing behavior and show why a routing key can reduce fan-out while creating skew risk.
Explain why a replica is an availability/search-capacity copy rather than an independent backup.
Build an append-only AtlasMart search-event data stream and identify its write backing index.
Use CAT/structured shard inventory and explicit refresh to prove placement and visibility rather than guessing from names.
This chapter continues the Chapter 01 lab baseline with
Elasticsearch 9.5.3 (released 3 September 2026)
and OpenSearch 3.8.0 (released 4 August 2026).
The existing local endpoints remain
https://localhost:9200 for Elasticsearch and
https://localhost:9201 for OpenSearch.
Elasticsearch requests use the Chapter 01 HTTP CA plus the
local elastic password; OpenSearch uses its local
demo TLS/admin path only for the disposable course lab.
Re-check versions, APIs and security defaults before reusing
these examples later.
The environment used to generate this chapter does not provide
Docker, Elasticsearch or OpenSearch, so commands were reviewed
against current official documentation but were not executed
here. Expected output is described by field shape and
invariant rather than fabricated as captured output. Failure
injection targets only disposable
atlasmart-* indices and the isolated Chapter 01
course containers; every destructive or allocation-changing
step includes an explicit reset path.
1. Index name is a logical resource; shards are the distributed storage units
An index groups documents under one mapping/settings contract, but it is not one physical Lucene directory shared by every node. At creation time the index has a fixed number of primary shards. Each primary shard is an independent Lucene index and owns a subset of documents. A configured replica shard is another copy of a primary shard’s replication group placed on another eligible node when possible.
The primary shard count influences parallelism, per-shard overhead, routing distribution, recovery and search fan-out. It is not a knob to maximize. The replica count can be changed dynamically and affects redundancy/read capacity, but replicas consume storage and indexing work. A replica cannot protect against accidental index deletion, destructive credentials or repository-wide corruption in the way an independent tested snapshot can.
GET /_cat/indices/atlasmart-products-v2?v=true&h=health,status,index,pri,rep,docs.count,store.sizeGET /_cat/shards/atlasmart-products-v2?v=true&h=index,shard,prirep,state,docs,store,nodeGET /atlasmart-products-v2/_settings?flat_settings=true&filter_path=*.settings.index.number_of_shards,*.settings.index.number_of_replicas
CAT endpoints are excellent operator evidence, but applications
should prefer stable structured APIs rather than parsing table
formatting. The row marked p is a primary;
r is a replica. On the Chapter 01 single-node
topology, the initial replica count is zero, so both primaries
can allocate on the one data-capable node.
2. Aliases decouple application names from versioned indices
AtlasMart should not force every caller to know whether the
current catalog lives in atlasmart-products-v1 or
atlasmart-products-v2. An alias is
a named indirection over one or more indices. A write alias can
nominate a write index, which later enables reindex/cutover
workflows without changing application configuration. Aliases
can also carry filters or routing, but those features must be
treated as part of the data-access contract and tested for
authorization assumptions.
POST /_aliases{ "actions": [ {"add": {"index": "atlasmart-products-v2", "alias": "atlasmart-products", "is_write_index": true}} ]}GET /_alias/atlasmart-productsGET /atlasmart-products/_search{ "query": {"match_all": {}}}
Chapter 01 did not establish a stable alias, so this chapter
adds atlasmart-products only to v2. If
your local workspace has additional experimental alias
memberships, inspect and remove them explicitly rather than
relying on product-specific ignore/must-exist options. Later
zero-downtime chapters will use an atomic alias action for
controlled cutover and rollback.
3. Data streams add time-series write/backing-index semantics
A data stream is a logical resource over one or more hidden backing indices optimized for continuously generated time-series data. Reads target the stream and fan across its backing indices. New writes go to the current write backing index. Both Elasticsearch and OpenSearch require a matching data-stream-enabled index template and timestamped documents, but lifecycle and advanced time-series features diverge across products.
AtlasMart search events fit this shape because they are append-heavy observations: who searched, what they typed, when the query happened and how many results were returned. They are not mutable product records expected to be overwritten repeatedly by ID.
PUT /_index_template/atlasmart-search-events-template{ "index_patterns": ["atlasmart-search-events*"], "priority": 500, "data_stream": {}, "template": { "settings": {"number_of_shards": 1, "number_of_replicas": 0}, "mappings": { "properties": { "@timestamp": {"type": "date"}, "event_id": {"type": "keyword"}, "customer_id": {"type": "keyword"}, "query": {"type": "text"}, "result_count": {"type": "integer"} } } }}PUT /_data_stream/atlasmart-search-eventsPOST /atlasmart-search-events/_doc{ "@timestamp": "2026-09-10T16:30:00Z", "event_id": "SE-0001", "customer_id": "C-007", "query": "wireless keyboard", "result_count": 18}POST /atlasmart-search-events/_refreshGET /_data_stream/atlasmart-search-eventsGET /atlasmart-search-events/_search{ "query": {"term": {"customer_id": "C-007"}}}
Try to index an event without @timestamp. The
data stream’s timestamp contract should reject the document
rather than quietly inventing an event time. Repair the
producer so event time has clear semantics; do not “solve” the
failure by converting every event stream into an unconstrained
ordinary index.
4. Routing chooses a primary shard before the write executes
By default, document routing is derived from _id. A
hash of the routing value selects one primary shard. Custom
routing replaces the default value with an application-supplied
key, such as tenant_id. It can reduce search
fan-out when queries provide the same routing values, but it
also concentrates every document for a hot tenant onto one shard
and can make documents appear “missing” if GET/update/delete
omits the original routing value.
For newly created Elasticsearch 9.4+ indices, current
documentation uses
shard_num = hash(_routing) % num_primary_shards.
Elasticsearch indices created before 9.4 retain the legacy
routing function based on num_routing_shards.
OpenSearch documents the direct modulo formula. This means
reindexing an old Elasticsearch index into a new 9.5 index can
redistribute routing values even if the application routing keys
are unchanged. Teach the contract—same routing key
deterministically selects a shard for that index—not an eternal
implementation formula.
PUT /atlasmart-products-v2/_doc/P-T1-001?routing=TENANT-A{ "product_id": "P-T1-001", "name": "Tenant A Keyboard", "category": "keyboards", "price": 61.00, "updated_at": "2026-09-10T16:35:00Z"}GET /_search_shards/atlasmart-products-v2?routing=TENANT-AGET /atlasmart-products-v2/_doc/P-T1-001?routing=TENANT-A# Deliberately omit routing. Depending on which default shard _id maps to,# this lookup does not address the same routed document location.GET /atlasmart-products-v2/_doc/P-T1-001
A routing key is a placement optimization, not an authorization boundary. A caller that can search the whole index may still see other tenants unless product-specific security privileges and application authorization prevent it. Later security chapters keep routing and authorization separate.
5. Replica count makes health signals meaningful
On a one-node cluster, setting one replica asks for a second copy that cannot be colocated with its primary on the same node. The primary remains available, so the index can become yellow: primary shards are assigned, but at least one replica is unassigned. This is a useful deliberate failure because it proves that a green/yellow/red status is derived from shard allocation, not a cosmetic service light.
PUT /atlasmart-products-v2/_settings{"index": {"number_of_replicas": 1}}GET /_cluster/health/atlasmart-products-v2?prettyGET /_cat/shards/atlasmart-products-v2?v=true# Why can the replica not allocate?GET /_cluster/allocation/explain{ "index": "atlasmart-products-v2", "shard": 0, "primary": false}# Reset the single-node lab to its intentional baseline.PUT /atlasmart-products-v2/_settings{"index": {"number_of_replicas": 0}}GET /_cluster/health/atlasmart-products-v2?wait_for_status=green&timeout=30s
When you later add a second eligible node, the same replica setting can allocate a copy on the other node and health can become green. That demonstrates redundancy only against certain node/shard failures; it is still not a backup or a substitute for snapshot restore tests.
Verification checklist
-
The stable alias resolves to
atlasmart-products-v2. - The product index has two primaries and initially zero replicas.
- The search-events stream has at least one hidden backing index and a current write index.
- A timestamped event is searchable after refresh; a missing-timestamp event is rejected.
-
Custom-routed
P-T1-001is retrievable with the same routing key. - Setting one replica on the single-node product cluster produces an unassigned-replica explanation, then resetting replicas to zero restores the intended green baseline.
Production judgment
Shard count fixes part of the index topology at creation time, so it belongs in capacity design rather than ad-hoc tuning. Replica count trades write/storage work for resilience/read capacity. Custom routing trades search fan-out for skew and correctness obligations. Aliases help versioned mutable data; data streams help append-heavy time series. None of these are interchangeable, and none is a backup. The next lesson maps those shards to node roles and cluster-state responsibilities.
Check your understanding
- Why is an index not equivalent to one physical Lucene directory?
- What makes a data stream different from an ordinary alias?
- Why can one replica make a healthy single-node index yellow?
- What correctness obligation does custom routing add to GET/update/delete?
- Why might an old Elasticsearch index and a new 9.5 index distribute the same routing keys differently?
Review the answers
1. The logical index is partitioned into primary shards, each of which is its own Lucene index and can have replica copies on different nodes.
2. A data stream has time-series timestamp semantics and managed backing/write-index behavior; it is not merely a name resolving to arbitrary indices.
3. The replica cannot be placed on the same node as its primary, so primaries stay available while a replica remains unassigned.
4. The caller must provide the same routing value used when the document was indexed or it may address a different shard and fail to find the document.
5. Elasticsearch 9.4+ changed the routing function for newly created indices while indices created before 9.4 retain the legacy formula.
Summary and next step
Indices, aliases and data streams are logical access layers; primary and replica shards are distributed storage/execution units; routing selects the replication group for a document. Next, inspect the nodes and cluster state that decide where those shards can live and which node coordinates each client request.
Authoritative references
- Elasticsearch data streams — Current backing-index, write-index, timestamp and read/write semantics.
- Elasticsearch <code>_routing</code> field — Current routing formula, custom routing, and the 9.4 creation-version distinction.
- Elasticsearch aliases — Alias and write-index behavior.
- OpenSearch data streams — Current backing-index/write-index behavior and time-series use.
- OpenSearch routing metadata — Default/custom routing and lookup requirements.
- OpenSearch shard allocation filtering — Allocation controls used later for controlled shard-placement experiments.