Chapter 02 · Documents, Indices, Data Streams, Shards, Replicas, Nodes, and Cluster Architecture

Indices, Aliases, Data Streams, Backing Indices, Shards, Replicas, and Routing

Connect indices, aliases, data streams, backing indices, shard copies and routing to observable document placement and failure signals.

Intermediate110–140 minutesDistributed document/shard labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart’s catalog and search-event workloads now need two different resource shapes. Products are mutable entities that benefit from a versioned index plus a stable alias. Search events are append-heavy time-series documents that fit a data stream. Both are physically partitioned into shards, and document routing determines which primary shard owns each write. This lesson connects those abstractions without pretending an index is a single file or a data stream is just an alias with a fancy name.

01

Distinguish index, alias, data stream, backing index, primary shard and replica shard as separate layers.

02

Predict default and custom routing behavior and show why a routing key can reduce fan-out while creating skew risk.

03

Explain why a replica is an availability/search-capacity copy rather than an independent backup.

04

Build an append-only AtlasMart search-event data stream and identify its write backing index.

05

Use CAT/structured shard inventory and explicit refresh to prove placement and visibility rather than guessing from names.

Version baseline reviewed 10 September 2026

This chapter continues the Chapter 01 lab baseline with Elasticsearch 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 (released 4 August 2026). The existing local endpoints remain https://localhost:9200 for Elasticsearch and https://localhost:9201 for OpenSearch. Elasticsearch requests use the Chapter 01 HTTP CA plus the local elastic password; OpenSearch uses its local demo TLS/admin path only for the disposable course lab. Re-check versions, APIs and security defaults before reusing these examples later.

Execution and safety note

The environment used to generate this chapter does not provide Docker, Elasticsearch or OpenSearch, so commands were reviewed against current official documentation but were not executed here. Expected output is described by field shape and invariant rather than fabricated as captured output. Failure injection targets only disposable atlasmart-* indices and the isolated Chapter 01 course containers; every destructive or allocation-changing step includes an explicit reset path.

1. Index name is a logical resource; shards are the distributed storage units

An index groups documents under one mapping/settings contract, but it is not one physical Lucene directory shared by every node. At creation time the index has a fixed number of primary shards. Each primary shard is an independent Lucene index and owns a subset of documents. A configured replica shard is another copy of a primary shard’s replication group placed on another eligible node when possible.

The primary shard count influences parallelism, per-shard overhead, routing distribution, recovery and search fan-out. It is not a knob to maximize. The replica count can be changed dynamically and affects redundancy/read capacity, but replicas consume storage and indexing work. A replica cannot protect against accidental index deletion, destructive credentials or repository-wide corruption in the way an independent tested snapshot can.

REST · inspect index and shard inventory
GET /_cat/indices/atlasmart-products-v2?v=true&h=health,status,index,pri,rep,docs.count,store.sizeGET /_cat/shards/atlasmart-products-v2?v=true&h=index,shard,prirep,state,docs,store,nodeGET /atlasmart-products-v2/_settings?flat_settings=true&filter_path=*.settings.index.number_of_shards,*.settings.index.number_of_replicas

CAT endpoints are excellent operator evidence, but applications should prefer stable structured APIs rather than parsing table formatting. The row marked p is a primary; r is a replica. On the Chapter 01 single-node topology, the initial replica count is zero, so both primaries can allocate on the one data-capable node.

2. Aliases decouple application names from versioned indices

AtlasMart should not force every caller to know whether the current catalog lives in atlasmart-products-v1 or atlasmart-products-v2. An alias is a named indirection over one or more indices. A write alias can nominate a write index, which later enables reindex/cutover workflows without changing application configuration. Aliases can also carry filters or routing, but those features must be treated as part of the data-access contract and tested for authorization assumptions.

REST · create the stable product alias
POST /_aliases{  "actions": [    {"add": {"index": "atlasmart-products-v2", "alias": "atlasmart-products", "is_write_index": true}}  ]}GET /_alias/atlasmart-productsGET /atlasmart-products/_search{  "query": {"match_all": {}}}

Chapter 01 did not establish a stable alias, so this chapter adds atlasmart-products only to v2. If your local workspace has additional experimental alias memberships, inspect and remove them explicitly rather than relying on product-specific ignore/must-exist options. Later zero-downtime chapters will use an atomic alias action for controlled cutover and rollback.

3. Data streams add time-series write/backing-index semantics

A data stream is a logical resource over one or more hidden backing indices optimized for continuously generated time-series data. Reads target the stream and fan across its backing indices. New writes go to the current write backing index. Both Elasticsearch and OpenSearch require a matching data-stream-enabled index template and timestamped documents, but lifecycle and advanced time-series features diverge across products.

AtlasMart search events fit this shape because they are append-heavy observations: who searched, what they typed, when the query happened and how many results were returned. They are not mutable product records expected to be overwritten repeatedly by ID.

REST · create a portable Chapter 02 data-stream template
PUT /_index_template/atlasmart-search-events-template{  "index_patterns": ["atlasmart-search-events*"],  "priority": 500,  "data_stream": {},  "template": {    "settings": {"number_of_shards": 1, "number_of_replicas": 0},    "mappings": {      "properties": {        "@timestamp": {"type": "date"},        "event_id": {"type": "keyword"},        "customer_id": {"type": "keyword"},        "query": {"type": "text"},        "result_count": {"type": "integer"}      }    }  }}PUT /_data_stream/atlasmart-search-eventsPOST /atlasmart-search-events/_doc{  "@timestamp": "2026-09-10T16:30:00Z",  "event_id": "SE-0001",  "customer_id": "C-007",  "query": "wireless keyboard",  "result_count": 18}POST /atlasmart-search-events/_refreshGET /_data_stream/atlasmart-search-eventsGET /atlasmart-search-events/_search{  "query": {"term": {"customer_id": "C-007"}}}
Deliberate data-stream failure

Try to index an event without @timestamp. The data stream’s timestamp contract should reject the document rather than quietly inventing an event time. Repair the producer so event time has clear semantics; do not “solve” the failure by converting every event stream into an unconstrained ordinary index.

4. Routing chooses a primary shard before the write executes

By default, document routing is derived from _id. A hash of the routing value selects one primary shard. Custom routing replaces the default value with an application-supplied key, such as tenant_id. It can reduce search fan-out when queries provide the same routing values, but it also concentrates every document for a hot tenant onto one shard and can make documents appear “missing” if GET/update/delete omits the original routing value.

For newly created Elasticsearch 9.4+ indices, current documentation uses shard_num = hash(_routing) % num_primary_shards. Elasticsearch indices created before 9.4 retain the legacy routing function based on num_routing_shards. OpenSearch documents the direct modulo formula. This means reindexing an old Elasticsearch index into a new 9.5 index can redistribute routing values even if the application routing keys are unchanged. Teach the contract—same routing key deterministically selects a shard for that index—not an eternal implementation formula.

REST · route tenant-scoped documents and inspect the selected shard
PUT /atlasmart-products-v2/_doc/P-T1-001?routing=TENANT-A{  "product_id": "P-T1-001",  "name": "Tenant A Keyboard",  "category": "keyboards",  "price": 61.00,  "updated_at": "2026-09-10T16:35:00Z"}GET /_search_shards/atlasmart-products-v2?routing=TENANT-AGET /atlasmart-products-v2/_doc/P-T1-001?routing=TENANT-A# Deliberately omit routing. Depending on which default shard _id maps to,# this lookup does not address the same routed document location.GET /atlasmart-products-v2/_doc/P-T1-001
Routing is not tenancy security

A routing key is a placement optimization, not an authorization boundary. A caller that can search the whole index may still see other tenants unless product-specific security privileges and application authorization prevent it. Later security chapters keep routing and authorization separate.

5. Replica count makes health signals meaningful

On a one-node cluster, setting one replica asks for a second copy that cannot be colocated with its primary on the same node. The primary remains available, so the index can become yellow: primary shards are assigned, but at least one replica is unassigned. This is a useful deliberate failure because it proves that a green/yellow/red status is derived from shard allocation, not a cosmetic service light.

REST · controlled replica-allocation failure and repair
PUT /atlasmart-products-v2/_settings{"index": {"number_of_replicas": 1}}GET /_cluster/health/atlasmart-products-v2?prettyGET /_cat/shards/atlasmart-products-v2?v=true# Why can the replica not allocate?GET /_cluster/allocation/explain{  "index": "atlasmart-products-v2",  "shard": 0,  "primary": false}# Reset the single-node lab to its intentional baseline.PUT /atlasmart-products-v2/_settings{"index": {"number_of_replicas": 0}}GET /_cluster/health/atlasmart-products-v2?wait_for_status=green&timeout=30s

When you later add a second eligible node, the same replica setting can allocate a copy on the other node and health can become green. That demonstrates redundancy only against certain node/shard failures; it is still not a backup or a substitute for snapshot restore tests.

Verification checklist

  • The stable alias resolves to atlasmart-products-v2.
  • The product index has two primaries and initially zero replicas.
  • The search-events stream has at least one hidden backing index and a current write index.
  • A timestamped event is searchable after refresh; a missing-timestamp event is rejected.
  • Custom-routed P-T1-001 is retrievable with the same routing key.
  • Setting one replica on the single-node product cluster produces an unassigned-replica explanation, then resetting replicas to zero restores the intended green baseline.

Production judgment

Shard count fixes part of the index topology at creation time, so it belongs in capacity design rather than ad-hoc tuning. Replica count trades write/storage work for resilience/read capacity. Custom routing trades search fan-out for skew and correctness obligations. Aliases help versioned mutable data; data streams help append-heavy time series. None of these are interchangeable, and none is a backup. The next lesson maps those shards to node roles and cluster-state responsibilities.

Check your understanding

  1. Why is an index not equivalent to one physical Lucene directory?
  2. What makes a data stream different from an ordinary alias?
  3. Why can one replica make a healthy single-node index yellow?
  4. What correctness obligation does custom routing add to GET/update/delete?
  5. Why might an old Elasticsearch index and a new 9.5 index distribute the same routing keys differently?
Review the answers

1. The logical index is partitioned into primary shards, each of which is its own Lucene index and can have replica copies on different nodes.

2. A data stream has time-series timestamp semantics and managed backing/write-index behavior; it is not merely a name resolving to arbitrary indices.

3. The replica cannot be placed on the same node as its primary, so primaries stay available while a replica remains unassigned.

4. The caller must provide the same routing value used when the document was indexed or it may address a different shard and fail to find the document.

5. Elasticsearch 9.4+ changed the routing function for newly created indices while indices created before 9.4 retain the legacy formula.

Summary and next step

Indices, aliases and data streams are logical access layers; primary and replica shards are distributed storage/execution units; routing selects the replication group for a document. Next, inspect the nodes and cluster state that decide where those shards can live and which node coordinates each client request.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.