Run a controlled AtlasMart maintenance game day with explicit stop gates, inject one disposable node failure, verify quorum/shards/application behavior, and produce evidence for recovery or rollback decisions instead of improvising under pressure.
Execute a Maintenance Game Day with a Node Failure During Upgrade and Prove Recovery Procedures
Operate Elasticsearch and OpenSearch with explicit cluster-health, quorum, compatibility, recovery and application gates instead of maintenance folklore.
Learning outcomes
This capstone turns the chapter into an AtlasMart maintenance game day. The team will use a disposable three-node cluster or the deterministic trace below, execute one rolling maintenance step, inject one controlled node failure, and produce a decision log that proves whether to continue, repair forward, or invoke recovery. Success is not “all commands returned 200”; success is a defensible sequence of evidence.
Prepare a complete pre-maintenance evidence bundle and declare stop/rollback thresholds.
Execute one node-at-a-time maintenance while preserving the voting majority and application service.
Inject a safe failure only in disposable infrastructure and diagnose it with health/allocation/recovery evidence.
Use node/version/plugin/snapshot/application gates to choose continue, repair forward or rebuild/restore.
Produce a reusable runbook with measured recovery and p95/p99 behavior rather than undocumented operator intuition.
Examples are frozen to
Elasticsearch 9.5.3 / Kibana 9.5.3 (released
3 September 2026) and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0
(released 4 August 2026). Use each distribution's bundled JVM
in the lab and record the actual runtime with
GET /_nodes/jvm; do not carry forward arbitrary
JVM flags from an older installation. The earlier
atlasmart-es and
atlasmart-os single-node containers remain useful
for API syntax, but this chapter's quorum/recovery exercises
require either a disposable three-node cluster or the
deterministic trace path supplied below. TLS/authentication
from Chapters 19–20 stays enabled. Managed services may hide
node ordering, voting, plugin, filesystem, or allocation
controls; their provider runbook overrides self-managed
commands.
Run destructive node-stop/version/failure injection only on disposable multi-node containers or an isolated staging cluster. The mandatory alternative is the deterministic trace. Never inject primary-data loss or force stale-primary allocation on the shared course cluster.
A maintenance gate is intentionally conservative: stop if a primary shard is unassigned, the cluster cannot elect/publish through a voting majority, recovery or relocation is still in progress beyond the runbook budget, a required plugin/client is incompatible, a snapshot/recovery point is not verified, or application correctness/error-rate criteria fail. Yellow can be acceptable only when the exact cause is understood and explicitly permitted by the runbook; never normalize unexpected yellow/red health.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Scenario and success criteria
AtlasMart has three voting-capable nodes, enough replicas to tolerate one node loss, and a product-search workload. A maintenance change is planned for data node 3. During its return, the exercise simulates a join failure caused by an intentionally wrong plugin/config/version artifact. The team must keep nodes 1–2 coordinated, avoid removing healthy copies, repair the failed node, and restore full redundancy before touching another node.
| Metric/evidence | Baseline | Stop threshold |
|---|---|---|
| Primary availability | 0 unassigned primaries | any unexpected unassigned primary |
| Voting/leader | majority reachable, leader elected | no leader or majority at risk |
| Recovery | 0 active at gate | recovery not stabilizing within declared budget |
| App correctness | search + authenticated write/read-back pass | any correctness/security failure |
| Tail latency/error | record measured p95/p99/error rate | declared SLO breach |
| Snapshot/recovery point | verified | missing/unverified before version change |
2. Preflight evidence package
# Elasticsearch 9.5.3 or OpenSearch 3.8.0; authenticate normally.
GET /
GET /_cluster/health?pretty
GET /_cat/nodes?v&h=name,ip,node.role,master,version,uptime
GET /_cat/shards?v&h=index,shard,prirep,state,node,unassigned.reason
GET /_cat/recovery?v&active_only=true
GET /_cluster/pending_tasks
GET /_nodes/jvm,plugins,settings?filter_path=nodes.*.name,nodes.*.version,nodes.*.jvm.version,nodes.*.plugins.name,nodes.*.settings.node.roles
# OpenSearch: inspect the committed voting configuration.
GET /_cluster/state?filter_path=metadata.cluster_coordination.last_committed_config,cluster_manager_node
# Record application smoke baselines too: search correctness, write/read-back,
# p95/p99 latency, error rate, indexing lag and authentication success.
# Use the repository established in Chapter 18; names are examples.
GET /_snapshot/atlasmart-dr/_verify
GET /_snapshot/atlasmart-dr/atlasmart-pre-maintenance-20260911
# Export or version-control product configuration outside the cluster:
# - elasticsearch.yml / opensearch.yml overrides
# - JVM overrides
# - plugin inventory
# - TLS/security configuration references (never raw secrets in Git)
# - application/client versions and endpoint configuration
3. Choose the correct role order
For a same-version restart, node order is mostly about quorum and workload. For an upgrade, use the exact product role order: Elasticsearch data tiers first and master-eligible last; OpenSearch data first and cluster-manager last. The exercise changes only one data node, so the voting majority remains available and the current elected coordinator is untouched.
4. Execute gate A: prepare and stop one node
PUT /_cluster/settings
{
"persistent": {"cluster.routing.allocation.enable": "primaries"}
}
POST /_flush
# stop ONLY atlasmart-node-3 using your container/service command
# then verify from another node:
GET /_cluster/health?pretty
GET /_cat/nodes?v&h=name,node.role,master,version
GET /_cat/shards?v&h=index,shard,prirep,state,node,unassigned.reason
# Expected invariant: voting majority still exists and all primaries remain available.
If the cluster loses primary availability or coordination, stop the exercise and restore the node immediately. Do not continue to the injected join failure.
5. Inject one controlled join failure
On the disposable node only, introduce exactly one reversible mismatch: for example, omit a required version-compatible plugin or use a copied configuration with a wrong cluster/discovery/TLS setting. Start the node and observe that it does not join. The goal is to diagnose without touching the healthy nodes.
# From a healthy node:
GET /_cat/nodes?v
GET /_cluster/health?pretty
GET /_cluster/pending_tasks
GET /_cat/recovery?v&active_only=true
# On failed node: inspect local service/container logs.
# Expected decision: STOP rollout. Keep healthy majority. Repair node 3 only.
Do not restart nodes 1 and 2 “to make discovery work,” do not clear cluster metadata, and do not bootstrap a new cluster. Those actions can turn one-node join failure into a quorum/data incident.
6. Repair forward and rejoin
Restore the known-good configuration or install the exact compatible plugin, then start node 3. Verify it joins the same cluster UUID and expected version. Restore normal allocation and wait for recovery before resuming indexing pressure or moving to another node.
GET /
GET /_cat/nodes?v&h=name,node.role,master,version,uptime
PUT /_cluster/settings
{
"persistent": {"cluster.routing.allocation.enable": null}
}
GET /_cluster/health?wait_for_status=green&wait_for_no_relocating_shards=true&timeout=10m
GET /_cat/recovery?v&active_only=true
GET /_cat/shards?v&h=index,shard,prirep,state,node
7. Application and security acceptance
Use the least-privilege reader/ingester credentials from Chapters 19–20 rather than administrator credentials. Run a representative search, an idempotent write/read-back, one expected 403, and record p95/p99/error rate compared with the baseline. This catches a class of “green cluster, broken application” failures.
READ: atlasmart-products-read search returns expected IDs/facets
WRITE: controlled idempotent event/product write succeeds through intended principal
READ-BACK: written document is visible after normal refresh semantics
DENY: reader cannot delete index or call security administration
LATENCY: record p50/p95/p99 and errors; compare to preflight
FRESHNESS: record indexing/search lag
8. Decision tree: continue, repair forward, or recover
| Evidence | Decision |
|---|---|
| Node 3 fixed, cluster green, recovery done, app gates pass | Continue to next planned node only if runbook says so. |
| Node 3 not joining but healthy majority/data intact | Stop rollout; repair node 3 only. Do not increase blast radius. |
| Target version incompatible but old node has not been modified | Abort change and preserve old node/version. |
| Upgrade crossed no-downgrade boundary; cluster repairable | Complete supported upgrade/repair forward according to vendor guidance. |
| Cluster/data cannot be recovered in-place | Provision compatible recovery cluster and restore verified snapshot; measure RTO/RPO. |
9. Post-game-day evidence record
change_id: ATLASMART-SEARCH-MAINT-001
product: Elasticsearch | OpenSearch
source_version: ...
target_version: ...
node_changed: ...
pre_snapshot: ... VERIFIED
pre_health: ...
voting_config_before: ...
failure_injected: one reversible join incompatibility
failure_detected_by: logs + node inventory + health
repair: ...
rejoin_time_seconds: measured
recovery_time_seconds: measured
bytes_recovered: measured
p95_before/after: measured
p99_before/after: measured
errors_before/after: measured
security_smoke: PASS/FAIL
final_health: ...
rollback_boundary_crossed: yes/no
continue_authorized_by_gate: yes/no
follow_up: ...
10. Bridge to Chapter 22
Single-cluster maintenance is now evidence-driven. Chapter 22 extends the same discipline across remote clusters and regions, where WAN latency, remote-cluster trust, cross-cluster search, replication lag, write ownership and RPO/RTO introduce new failure modes. A local green cluster does not imply the remote system is safe to fail over.
Check your understanding
- What is the purpose of the injected join failure?
- Why run application smoke tests after green health?
- When may the next node be touched?
- What is the rollback plan after the no-downgrade boundary?
- What evidence makes the game day reusable?
Review the answers
1. To prove the team stops the rollout, contains blast radius, diagnoses one node and repairs it without destabilizing the healthy voting majority.
2. Cluster allocation can be healthy while API compatibility, security, mappings, clients or application behavior is broken.
3. Only after the prior node rejoins, allocation/recovery stabilizes and every declared cluster/application/security gate passes.
4. Rebuild a compatible cluster and restore a verified compatible snapshot if repair-forward cannot satisfy recovery objectives.
5. Exact versions/config/plugins, voting/shard state, measured recovery and tail latency, failure/repair actions, stop criteria and final authorization decision.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic 9.5.3 release notes — Current Elasticsearch release baseline.
- Elastic self-managed Elasticsearch upgrade — Rolling order, allocation restriction, flush, plugin upgrade, recovery and no-downgrade boundary.
- Elastic plan/prepare upgrade — Upgrade path, compatibility and preparation guidance.
- Elastic add/remove nodes — Node removal, shard draining and master-eligible considerations.
- Elastic diagnose unassigned shards — Health meaning and allocation-explain workflow.
- Elastic shard allocation/routing settings — Allocation controls, filters, awareness and decommissioning primitives.
- OpenSearch 3.8 version history — Current OpenSearch release baseline.
- OpenSearch rolling upgrade — Health gate, allocation restriction, flush and role ordering.
- OpenSearch migrate or upgrade — Upgrade methods, plugin compatibility, configuration backup and snapshots.
- OpenSearch voting and quorum — Voting majority and maintenance safety.
- OpenSearch voting configuration — Voting membership and automatic shrink behavior.
- OpenSearch allocation explain API — Allocation diagnosis.
- OpenSearch cluster settings — Allocation include/require/exclude and recovery settings.
- OpenSearch nodes info API — Version/JVM/plugin/node-role inventory.