Turn AtlasMart cluster health into an operational decision tree: distinguish yellow from red, explain unassigned shards before acting, inspect voting membership, and preserve quorum while one node is maintained.

Cluster Health, Unassigned Shards, Allocation Explain, Voting/Master Eligibility, and Quorum Safety

Operate Elasticsearch and OpenSearch with explicit cluster-health, quorum, compatibility, recovery and application gates instead of maintenance folklore.

Intermediate → Advanced120–170 minutesCluster Operations · Chapter 21 · Lesson 01Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

AtlasMart has a maintenance window, one node reports disk pressure, and an operator sees a yellow cluster. The dangerous response is to “make it green” with a reroute command before understanding whether the missing shard is a harmless replica, a primary with no valid copy, or a deliberate allocation restriction. Cluster health is evidence, not a repair instruction.

01

Distinguish green/yellow/red health from the underlying primary/replica availability state.

02

Use allocation explain to diagnose an unassigned shard before changing allocation or reroute settings.

03

Explain Elasticsearch master-eligible and OpenSearch cluster-manager-eligible voting majorities without treating node count as quorum by itself.

04

Inspect the current voting configuration and identify when planned removal requires voting exclusions or a different maintenance order.

05

Define explicit maintenance stop gates that protect data availability and application correctness.

Pinned maintenance baseline

Examples are frozen to Elasticsearch 9.5.3 / Kibana 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0 (released 4 August 2026). Use each distribution's bundled JVM in the lab and record the actual runtime with GET /_nodes/jvm; do not carry forward arbitrary JVM flags from an older installation. The earlier atlasmart-es and atlasmart-os single-node containers remain useful for API syntax, but this chapter's quorum/recovery exercises require either a disposable three-node cluster or the deterministic trace path supplied below. TLS/authentication from Chapters 19–20 stays enabled. Managed services may hide node ordering, voting, plugin, filesystem, or allocation controls; their provider runbook overrides self-managed commands.

Stop gate

A maintenance gate is intentionally conservative: stop if a primary shard is unassigned, the cluster cannot elect/publish through a voting majority, recovery or relocation is still in progress beyond the runbook budget, a required plugin/client is incompatible, a snapshot/recovery point is not verified, or application correctness/error-rate criteria fail. Yellow can be acceptable only when the exact cause is understood and explicitly permitted by the runbook; never normalize unexpected yellow/red health.

Execution note

The generation environment cannot run the multi-node Docker game day. Commands and response fragments are deterministic instructions and expected invariants, not claimed captures. Collect your own timestamps, node/version/plugin inventories, voting configuration, shard movement, recovery duration, smoke-test latency/error rate, snapshot status, and p95/p99 latency before/after each maintenance gate.

1. Health color is a summary, not a diagnosis

Green means every configured primary and replica shard is assigned. Yellow means all primaries are assigned but at least one replica is not, so data can still be served but redundancy is below the intended state. Red means at least one primary is unassigned; some indexed data is unavailable. A single-node lab with replicas can be predictably yellow, while the same yellow state in a three-node production cluster can indicate a failed node, allocation filter, disk watermark, awareness constraint, or version rule.

Baseline evidence — run before touching a node
# Elasticsearch 9.5.3 or OpenSearch 3.8.0; authenticate normally.
GET /
GET /_cluster/health?pretty
GET /_cat/nodes?v&h=name,ip,node.role,master,version,uptime
GET /_cat/shards?v&h=index,shard,prirep,state,node,unassigned.reason
GET /_cat/recovery?v&active_only=true
GET /_cluster/pending_tasks
GET /_nodes/jvm,plugins,settings?filter_path=nodes.*.name,nodes.*.version,nodes.*.jvm.version,nodes.*.plugins.name,nodes.*.settings.node.roles

# OpenSearch: inspect the committed voting configuration.
GET /_cluster/state?filter_path=metadata.cluster_coordination.last_committed_config,cluster_manager_node

# Record application smoke baselines too: search correctness, write/read-back,
# p95/p99 latency, error rate, indexing lag and authentication success.
Signal Question it answers What it does not prove
`/_cluster/health` Are primaries/replicas assigned and is cluster work pending? Why a shard is unassigned or whether the data copy is recoverable.
`/_cat/shards` Which shard copies are STARTED/RELOCATING/INITIALIZING/UNASSIGNED? Which allocation decider is blocking the shard.
`/_cluster/allocation/explain` Why this shard can/cannot allocate or rebalance to candidate nodes. That manually forcing a shard is safe.
`/_cat/recovery` Is peer/snapshot recovery moving bytes and stages? That application correctness is restored.
Application smoke test Can the application perform its required read/write/security contract? That the cluster has redundancy or upgrade compatibility.

2. Allocation explain before reroute

Allocation is decided by a set of deciders: disk thresholds, allocation enablement, awareness, include/require/exclude filters, same-shard placement, tier/role eligibility, version compatibility and other constraints. The explain API exposes those decisions. Repair the cause that says “NO”; do not bypass a real no-valid-copy condition.

Explain an unassigned shard
GET /_cluster/allocation/explain?include_disk_info=true&include_yes_decisions=false
{
  "index": "atlasmart-products-v2",
  "shard": 0,
  "primary": false
}

# Expected categories include:
# current_state: "unassigned"
# can_allocate: "no" | "throttled" | "allocation_delayed" | ...
# node_allocation_decisions[].deciders[] with a concrete reason.
Unsafe shortcut

Never use forced stale-primary allocation merely to change red to green. If there is no valid in-sync copy, the correct recovery path can be to return the missing node or restore a verified snapshot. A forced stale primary can intentionally discard acknowledged writes and is therefore outside the mandatory lab.

3. Voting configuration is the availability boundary

Both products use majority-based cluster coordination, but terminology differs. Elasticsearch calls eligible coordinators master-eligible nodes; OpenSearch 2.x+ prefers cluster-manager-eligible nodes while retaining some legacy master field names for compatibility. A decision requires a majority of the current voting configuration, not merely “one of three machines.”

Scenario Voting set Safe simultaneous loss Reason
3 eligible voters 3 1 2 of 3 remains a majority.
5 eligible voters 5 2 3 of 5 remains a majority.
2 eligible voters 2 0 Both are required for majority; this is poor maintenance tolerance.
Planned multi-node removal Inspect current config first Depends Use product voting-exclusion workflow when automatic reconfiguration cannot safely complete before shutdown.
Inspect coordination evidence
# Elasticsearch/OpenSearch cluster state fields vary by product/version.
GET /_cluster/state?filter_path=master_node,cluster_manager_node,metadata.cluster_coordination.last_committed_config,metadata.cluster_coordination.last_accepted_config

# Node-role view. OpenSearch still exposes `master` in some CAT output for compatibility.
GET /_cat/nodes?v&h=name,node.role,master,version

4. Wrong approach: restart every eligible coordinator together

If AtlasMart stops all voting nodes at once, no majority remains to publish cluster state. Even if data nodes are still alive, the cluster cannot safely coordinate metadata changes. The repair is procedural: maintain a voting majority, change one node at a time, wait for it to rejoin and stabilize, then proceed. For a permanent removal, follow the product-specific voting exclusion/removal procedure instead of relying on timeout luck.

Bootstrap trap

Do not re-add `cluster.initial_master_nodes` / `cluster.initial_cluster_manager_nodes` while restarting an established cluster. Bootstrap settings are for initial cluster formation, not rolling maintenance.

5. AtlasMart deterministic quorum game

Use a disposable three-voter topology: atlasmart-search-1, -2, -3. Two nodes also hold data; every test index has at least one replica so a single-node failure can be tolerated. If local resources are insufficient, use this deterministic trace and mark it as simulation evidence rather than executing destructive steps on the shared single-node lab.

Game-day state table
Gate 0: voters = {1,2,3}; active leader = 1; green; app smoke PASS
Action: stop node 3 only
Gate 1: voters reachable = {1,2}; majority preserved; leader remains/elects; app smoke PASS
Expected transient: replica may be UNASSIGNED/INITIALIZING depending on allocation delay
Action: restart node 3
Gate 2: node 3 rejoins; recovery completes; green; app smoke PASS
STOP if: primary unassigned, no leader, voting majority lost, incompatible version/plugin,
         recovery exceeds budget, p99/error rate violates gate, or smoke test fails.

The deterministic trace proves the reasoning but not your infrastructure timings. A real rehearsal must collect recovery duration, shard bytes, CPU/disk pressure and client latency.

6. Production judgment

Quorum safety is necessary but not sufficient. A cluster can retain a leader while a huge shard recovery saturates disk and makes p99 search unacceptable. The maintenance runbook therefore gates on coordination, shard state, recovery headroom, application correctness, latency/error rate, security/authentication and recovery-point readiness. Replicas are availability copies, not backups; Chapter 18’s snapshot evidence remains the rollback artifact.

Check your understanding

  1. Does yellow always mean maintenance may continue?
  2. Why use allocation explain before reroute?
  3. What determines quorum?
  4. Why are bootstrap settings dangerous during restart?
  5. What is the final gate after cluster health?
Review the answers

1. No. It means primaries are assigned but redundancy is incomplete; continue only if the exact cause is expected and permitted by the runbook.

2. It exposes the allocation deciders blocking the shard so you can repair the cause instead of bypassing a safety rule.

3. A majority of the current voting configuration, not raw server count.

4. They are for initial cluster formation and can create coordination mistakes if reused on an established cluster.

5. Application smoke/correctness and measured operational SLO evidence, including latency and error rate.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.