Chapter 13 · Shard Sizing, Routing, Allocation, Awareness, and Cluster Topology
Allocation Deciders, Disk Watermarks, Awareness/Forced Awareness, Data Tiers, and Zone Failure
Read allocation decisions as evidence: connect disk watermarks, awareness, forced awareness, filters and tier placement to the shard state seen during a zone failure.
Learning outcomes
AtlasMart expands to three availability zones. During a zone incident, operators see yellow health, a node above the high disk watermark, and replicas that refuse to allocate. “Just reroute the shard” is not a diagnosis. Allocation is the result of multiple deciders—independent rules that can say yes, no or throttle for a candidate node.
Read allocation explain output and identify the decider blocking or throttling placement.
Explain low/high/flood-stage disk watermarks as safety controls, not target utilization levels.
Differentiate awareness from forced awareness during a zone failure.
Use allocation filters/data tiers without accidentally making a legal placement impossible.
Design a zone-loss test that proves capacity headroom without filling disks or modifying host networking.
Examples target self-managed
Elasticsearch 9.5.3 / Kibana 9.5.3 and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain
https://localhost:9200 for Elasticsearch using
its copied CA and https://localhost:9201 for the
disposable OpenSearch demo certificate. Use each
distribution's bundled JVM for this lab unless its current
support matrix says otherwise. The initial Chapter 01
single-node containers are useful for API inspection but
cannot demonstrate replica placement or zone resilience, so
this chapter uses deterministic traces and an optional
disposable multi-node topology. No moving
latest tags and no production allocation changes
are used.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Allocation is a constraint-satisfaction problem
The elected master node in Elasticsearch—or cluster-manager node in OpenSearch terminology—maintains cluster metadata and decides shard placement. A candidate node must satisfy all applicable allocation rules. Typical deciders cover same-shard separation, disk thresholds, awareness, filters/tier requirements, recovery throttling and other index/cluster constraints.
A shard that is unassigned is therefore a symptom. Ask the allocator why before changing settings.
GET _cluster/health?level=indices
GET _cat/shards?v&h=index,shard,prirep,state,unassigned.reason,node
POST _cluster/allocation/explain
{
"index":"atlasmart-products-v2",
"shard":0,
"primary":false
}
Look for the specific decider, its
YES/NO/THROTTLE decision and explanation. Do not
disable the disk allocator or awareness globally just because
one shard is unassigned.
2. Disk watermarks are guardrails with cascading behavior
In current Elasticsearch and OpenSearch self-managed defaults, the familiar percentage thresholds are 85% low, 90% high and 95% flood stage. Treat these as current defaults to inspect—not as universal tuning recommendations. Percentage thresholds refer to used disk; byte thresholds refer to free disk, which reverses intuition and must not be mixed inconsistently.
At the low watermark, additional allocations are restricted. Above the high watermark, the allocator tries to move shards away. At flood stage, indices with shards on the affected node can receive a read-only-allow-delete block to prevent complete disk exhaustion. If every node is above the low watermark, there may be nowhere legal to move shards.
GET _cluster/settings?include_defaults=true&filter_path=*.cluster.routing.allocation.disk.*,*.cluster.info.update.interval
GET _cat/allocation?v&h=node,shards,disk.indices,disk.used,disk.avail,disk.total,disk.percent
| Signal | Meaning | Safe response pattern |
|---|---|---|
| Low watermark crossed | Node should stop receiving more ordinary shard allocations | Add/restore headroom; inspect growth and other nodes. |
| High watermark crossed | Allocator attempts relocations away | Confirm legal targets exist and recovery is progressing. |
| Flood stage crossed | Write-protection last resort may apply | Free/add disk and resolve root cause; do not merely clear the block. |
| All nodes near threshold | Allocator has no escape route | Capacity incident: scale/delete/restore strategy, not a reroute trick. |
3. Awareness and forced awareness make different failure choices
Awareness tells the allocator about a
failure-domain attribute such as zone. With two or
three zones present, shard copies are spread when possible. If a
whole zone disappears, ordinary awareness may place missing
replicas into the remaining zones to restore copy count.
Forced awareness lists all expected attribute values. If one expected zone is missing, the cluster may intentionally leave replicas unassigned rather than concentrating all copies and load in the surviving zone. Yellow health can therefore be the designed safety outcome, not a broken allocator.
PUT _cluster/settings
{
"persistent":{
"cluster.routing.allocation.awareness.attributes":"zone",
"cluster.routing.allocation.awareness.force.zone.values":"zone-a,zone-b,zone-c"
}
}
# Each self-managed node must expose a corresponding node.attr.zone value.
# Managed services may control zone awareness through provider topology instead.
Forced awareness protects against overloading the survivors only if the application can tolerate temporarily unassigned replicas. Capacity planning must specify what one-zone-down traffic and recovery should look like.
4. Data tiers and filters add more placement constraints
Elasticsearch has formal data-tier roles such as
content/hot/warm/cold/frozen and automatically applies tier
allocation rules for managed lifecycle placement. OpenSearch
supports allocation attributes and a built-in
_tier filter surface in current cluster settings,
but tiering/lifecycle mechanisms and managed-service behavior
differ. Do not copy an ILM tier rule into ISM or vice versa.
Cluster-level include, require and
exclude filters combine with index-level filters
and awareness. A shard must satisfy all active constraints.
During decommissioning, an exclude rule can be
appropriate; forgetting to remove it later can silently reduce
capacity.
GET _cluster/settings?include_defaults=true&flat_settings=true
GET atlasmart-products*/_settings?flat_settings=true&include_defaults=true
GET _nodes?filter_path=nodes.*.name,nodes.*.roles,nodes.*.attributes
5. Deliberately wrong approach: defeat the decider
When a shard is unassigned because the only surviving node is in the wrong zone and above the high watermark, disabling awareness or disk thresholds may force allocation into the exact failure domain that the safety rules were designed to avoid. It also makes the next disk or zone failure harder to survive.
Repair the incident by identifying the blocking constraint, then restoring legal capacity: bring the zone back, add a node in a valid zone/tier, free disk, reduce a conflicting filter, or temporarily accept an unassigned replica if forced awareness deliberately chooses that state. Any temporary setting change requires an explicit rollback step.
PUT _cluster/settings
{
"persistent":{
"cluster.routing.allocation.awareness.attributes":null,
"cluster.routing.allocation.awareness.force.zone.values":null
}
}
6. AtlasMart deterministic zone-failure trace
If you do not have three disposable data nodes, use a trace rather than pretending a single-node cluster demonstrates zone behavior. Model six data nodes, two per zone, one replica per primary. Remove both nodes in zone-b from the trace and compare ordinary awareness with forced awareness.
Expected reasoning: ordinary awareness is permitted to rebuild missing replicas in zones a/c if other constraints and capacity allow; forced awareness can leave those replicas unassigned because one declared zone is absent. The exact health transitions and timings depend on delayed allocation, node roles, shard sizes and current settings, so record them if you run a real multi-node test.
Before failure:
zones: a,b,c
data nodes: 2 per zone
replicas: 1
requirement: primary and replica not on same node; awareness=zone
Inject:
remove both zone-b nodes in disposable lab OR mark them absent in the trace
Observe/model:
- cluster health and unassigned replicas
- allocation explain for one replica
- surviving-zone CPU/disk headroom
- recovery queue and bytes remaining
Compare:
A) awareness only
B) forced awareness with a,b,c declared
Acceptance:
explain why each unassigned/relocating shard state is legal or illegal.
Production judgment
Operate to headroom below safety thresholds, not to the thresholds themselves. A cluster that needs every node at steady state cannot absorb recovery work. Zone topology and allocation constraints belong in infrastructure-as-code and deployment review, with a game-day that proves the surviving zones can serve the agreed workload.
Check your understanding
- What does allocation explain answer?
- Why can yellow health be intentional with forced awareness?
- What happens at flood-stage disk pressure?
- Why is clearing a flood-stage block alone incomplete?
- Why inspect both cluster- and index-level allocation rules?
Review the answers
1. Why a specific shard can, cannot, or is throttled from allocating on candidate nodes, including the responsible deciders.
2. Replicas may remain unassigned while an expected zone is absent so survivors are not overloaded or all copies concentrated in one domain.
3. The system can apply read-only-allow-delete protection to affected indices as a last-resort disk safety measure.
4. The block is a symptom; without restoring disk headroom the allocator will re-enter the unsafe state.
5. They combine; a shard must satisfy all active constraints, and one hidden filter can make placement impossible.
Summary and next step
You can now explain why a shard is placed, blocked or intentionally left unassigned. Lesson 4 follows the shard after placement changes: recovery, relocation and rebalancing consume real disk/network/CPU resources and need their own headroom and throttling policy.
Authoritative references
- Elastic shard allocation settings — Current allocation, rebalance, disk-watermark and recovery controls.
- Elastic shard allocation awareness — Awareness and forced-awareness behavior across failure domains.
- Elastic allocation explain API — Explain why a shard is or is not allocatable.
- Elastic index recovery API — Recovery stage and byte/file progress.
- OpenSearch cluster settings — Current routing/allocation, awareness and recovery settings.
- OpenSearch cluster tuning — Shard allocation awareness and forced awareness concepts.
- OpenSearch CAT shards — Shard state and placement inspection.
- OpenSearch cluster allocation explain — Allocation diagnostics and decider evidence.
- Elastic disk watermark troubleshooting — Current low/high/flood-stage behavior and recovery guidance.
- OpenSearch cluster allocation settings — Current awareness/filter/recovery settings and built-in attributes.