Chapter 13 · Shard Sizing, Routing, Allocation, Awareness, and Cluster Topology

Allocation Deciders, Disk Watermarks, Awareness/Forced Awareness, Data Tiers, and Zone Failure

Read allocation decisions as evidence: connect disk watermarks, awareness, forced awareness, filters and tier placement to the shard state seen during a zone failure.

Intermediate115–145 minutesShard topology & failure-domain labElasticsearch 9.5.3 · OpenSearch 3.8.0Last reviewed: September 2026

Learning outcomes

AtlasMart expands to three availability zones. During a zone incident, operators see yellow health, a node above the high disk watermark, and replicas that refuse to allocate. “Just reroute the shard” is not a diagnosis. Allocation is the result of multiple deciders—independent rules that can say yes, no or throttle for a candidate node.

01

Read allocation explain output and identify the decider blocking or throttling placement.

02

Explain low/high/flood-stage disk watermarks as safety controls, not target utilization levels.

03

Differentiate awareness from forced awareness during a zone failure.

04

Use allocation filters/data tiers without accidentally making a legal placement impossible.

05

Design a zone-loss test that proves capacity headroom without filling disks or modifying host networking.

Chapter baseline reviewed 11 September 2026

Examples target self-managed Elasticsearch 9.5.3 / Kibana 9.5.3 and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0. The established AtlasMart endpoints remain https://localhost:9200 for Elasticsearch using its copied CA and https://localhost:9201 for the disposable OpenSearch demo certificate. Use each distribution's bundled JVM for this lab unless its current support matrix says otherwise. The initial Chapter 01 single-node containers are useful for API inspection but cannot demonstrate replica placement or zone resilience, so this chapter uses deterministic traces and an optional disposable multi-node topology. No moving latest tags and no production allocation changes are used.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Allocation is a constraint-satisfaction problem

The elected master node in Elasticsearch—or cluster-manager node in OpenSearch terminology—maintains cluster metadata and decides shard placement. A candidate node must satisfy all applicable allocation rules. Typical deciders cover same-shard separation, disk thresholds, awareness, filters/tier requirements, recovery throttling and other index/cluster constraints.

A shard that is unassigned is therefore a symptom. Ask the allocator why before changing settings.

Dev Tools · evidence first
GET _cluster/health?level=indices
GET _cat/shards?v&h=index,shard,prirep,state,unassigned.reason,node
POST _cluster/allocation/explain
{
  "index":"atlasmart-products-v2",
  "shard":0,
  "primary":false
}
Interpret the deciders

Look for the specific decider, its YES/NO/THROTTLE decision and explanation. Do not disable the disk allocator or awareness globally just because one shard is unassigned.

2. Disk watermarks are guardrails with cascading behavior

In current Elasticsearch and OpenSearch self-managed defaults, the familiar percentage thresholds are 85% low, 90% high and 95% flood stage. Treat these as current defaults to inspect—not as universal tuning recommendations. Percentage thresholds refer to used disk; byte thresholds refer to free disk, which reverses intuition and must not be mixed inconsistently.

At the low watermark, additional allocations are restricted. Above the high watermark, the allocator tries to move shards away. At flood stage, indices with shards on the affected node can receive a read-only-allow-delete block to prevent complete disk exhaustion. If every node is above the low watermark, there may be nowhere legal to move shards.

Inspect configured and default thresholds
GET _cluster/settings?include_defaults=true&filter_path=*.cluster.routing.allocation.disk.*,*.cluster.info.update.interval
GET _cat/allocation?v&h=node,shards,disk.indices,disk.used,disk.avail,disk.total,disk.percent
Signal Meaning Safe response pattern
Low watermark crossed Node should stop receiving more ordinary shard allocations Add/restore headroom; inspect growth and other nodes.
High watermark crossed Allocator attempts relocations away Confirm legal targets exist and recovery is progressing.
Flood stage crossed Write-protection last resort may apply Free/add disk and resolve root cause; do not merely clear the block.
All nodes near threshold Allocator has no escape route Capacity incident: scale/delete/restore strategy, not a reroute trick.

3. Awareness and forced awareness make different failure choices

Awareness tells the allocator about a failure-domain attribute such as zone. With two or three zones present, shard copies are spread when possible. If a whole zone disappears, ordinary awareness may place missing replicas into the remaining zones to restore copy count.

Forced awareness lists all expected attribute values. If one expected zone is missing, the cluster may intentionally leave replicas unassigned rather than concentrating all copies and load in the surviving zone. Yellow health can therefore be the designed safety outcome, not a broken allocator.

Self-managed example · awareness then forced awareness
PUT _cluster/settings
{
  "persistent":{
    "cluster.routing.allocation.awareness.attributes":"zone",
    "cluster.routing.allocation.awareness.force.zone.values":"zone-a,zone-b,zone-c"
  }
}

# Each self-managed node must expose a corresponding node.attr.zone value.
# Managed services may control zone awareness through provider topology instead.
Failure-domain capacity

Forced awareness protects against overloading the survivors only if the application can tolerate temporarily unassigned replicas. Capacity planning must specify what one-zone-down traffic and recovery should look like.

4. Data tiers and filters add more placement constraints

Elasticsearch has formal data-tier roles such as content/hot/warm/cold/frozen and automatically applies tier allocation rules for managed lifecycle placement. OpenSearch supports allocation attributes and a built-in _tier filter surface in current cluster settings, but tiering/lifecycle mechanisms and managed-service behavior differ. Do not copy an ILM tier rule into ISM or vice versa.

Cluster-level include, require and exclude filters combine with index-level filters and awareness. A shard must satisfy all active constraints. During decommissioning, an exclude rule can be appropriate; forgetting to remove it later can silently reduce capacity.

Inspect allocation constraints before changing them
GET _cluster/settings?include_defaults=true&flat_settings=true
GET atlasmart-products*/_settings?flat_settings=true&include_defaults=true
GET _nodes?filter_path=nodes.*.name,nodes.*.roles,nodes.*.attributes

5. Deliberately wrong approach: defeat the decider

When a shard is unassigned because the only surviving node is in the wrong zone and above the high watermark, disabling awareness or disk thresholds may force allocation into the exact failure domain that the safety rules were designed to avoid. It also makes the next disk or zone failure harder to survive.

Repair the incident by identifying the blocking constraint, then restoring legal capacity: bring the zone back, add a node in a valid zone/tier, free disk, reduce a conflicting filter, or temporarily accept an unassigned replica if forced awareness deliberately chooses that state. Any temporary setting change requires an explicit rollback step.

Reset only lab-specific awareness settings
PUT _cluster/settings
{
  "persistent":{
    "cluster.routing.allocation.awareness.attributes":null,
    "cluster.routing.allocation.awareness.force.zone.values":null
  }
}

6. AtlasMart deterministic zone-failure trace

If you do not have three disposable data nodes, use a trace rather than pretending a single-node cluster demonstrates zone behavior. Model six data nodes, two per zone, one replica per primary. Remove both nodes in zone-b from the trace and compare ordinary awareness with forced awareness.

Expected reasoning: ordinary awareness is permitted to rebuild missing replicas in zones a/c if other constraints and capacity allow; forced awareness can leave those replicas unassigned because one declared zone is absent. The exact health transitions and timings depend on delayed allocation, node roles, shard sizes and current settings, so record them if you run a real multi-node test.

Trace worksheet
Before failure:
  zones: a,b,c
  data nodes: 2 per zone
  replicas: 1
  requirement: primary and replica not on same node; awareness=zone

Inject:
  remove both zone-b nodes in disposable lab OR mark them absent in the trace

Observe/model:
  - cluster health and unassigned replicas
  - allocation explain for one replica
  - surviving-zone CPU/disk headroom
  - recovery queue and bytes remaining

Compare:
  A) awareness only
  B) forced awareness with a,b,c declared

Acceptance:
  explain why each unassigned/relocating shard state is legal or illegal.

Production judgment

Operate to headroom below safety thresholds, not to the thresholds themselves. A cluster that needs every node at steady state cannot absorb recovery work. Zone topology and allocation constraints belong in infrastructure-as-code and deployment review, with a game-day that proves the surviving zones can serve the agreed workload.

Check your understanding

  1. What does allocation explain answer?
  2. Why can yellow health be intentional with forced awareness?
  3. What happens at flood-stage disk pressure?
  4. Why is clearing a flood-stage block alone incomplete?
  5. Why inspect both cluster- and index-level allocation rules?
Review the answers

1. Why a specific shard can, cannot, or is throttled from allocating on candidate nodes, including the responsible deciders.

2. Replicas may remain unassigned while an expected zone is absent so survivors are not overloaded or all copies concentrated in one domain.

3. The system can apply read-only-allow-delete protection to affected indices as a last-resort disk safety measure.

4. The block is a symptom; without restoring disk headroom the allocator will re-enter the unsafe state.

5. They combine; a shard must satisfy all active constraints, and one hidden filter can make placement impossible.

Summary and next step

You can now explain why a shard is placed, blocked or intentionally left unassigned. Lesson 4 follows the shard after placement changes: recovery, relocation and rebalancing consume real disk/network/CPU resources and need their own headroom and throttling policy.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.