Run a controlled AtlasMart maintenance game day with explicit stop gates, inject one disposable node failure, verify quorum/shards/application behavior, and produce evidence for recovery or rollback decisions instead of improvising under pressure.

Execute a Maintenance Game Day with a Node Failure During Upgrade and Prove Recovery Procedures

Operate Elasticsearch and OpenSearch with explicit cluster-health, quorum, compatibility, recovery and application gates instead of maintenance folklore.

Intermediate → Advanced120–170 minutesCluster Operations · Chapter 21 · Lesson 05Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

This capstone turns the chapter into an AtlasMart maintenance game day. The team will use a disposable three-node cluster or the deterministic trace below, execute one rolling maintenance step, inject one controlled node failure, and produce a decision log that proves whether to continue, repair forward, or invoke recovery. Success is not “all commands returned 200”; success is a defensible sequence of evidence.

01

Prepare a complete pre-maintenance evidence bundle and declare stop/rollback thresholds.

02

Execute one node-at-a-time maintenance while preserving the voting majority and application service.

03

Inject a safe failure only in disposable infrastructure and diagnose it with health/allocation/recovery evidence.

04

Use node/version/plugin/snapshot/application gates to choose continue, repair forward or rebuild/restore.

05

Produce a reusable runbook with measured recovery and p95/p99 behavior rather than undocumented operator intuition.

Pinned maintenance baseline

Examples are frozen to Elasticsearch 9.5.3 / Kibana 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0 (released 4 August 2026). Use each distribution's bundled JVM in the lab and record the actual runtime with GET /_nodes/jvm; do not carry forward arbitrary JVM flags from an older installation. The earlier atlasmart-es and atlasmart-os single-node containers remain useful for API syntax, but this chapter's quorum/recovery exercises require either a disposable three-node cluster or the deterministic trace path supplied below. TLS/authentication from Chapters 19–20 stays enabled. Managed services may hide node ordering, voting, plugin, filesystem, or allocation controls; their provider runbook overrides self-managed commands.

Game-day safety

Run destructive node-stop/version/failure injection only on disposable multi-node containers or an isolated staging cluster. The mandatory alternative is the deterministic trace. Never inject primary-data loss or force stale-primary allocation on the shared course cluster.

Stop gate

A maintenance gate is intentionally conservative: stop if a primary shard is unassigned, the cluster cannot elect/publish through a voting majority, recovery or relocation is still in progress beyond the runbook budget, a required plugin/client is incompatible, a snapshot/recovery point is not verified, or application correctness/error-rate criteria fail. Yellow can be acceptable only when the exact cause is understood and explicitly permitted by the runbook; never normalize unexpected yellow/red health.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Scenario and success criteria

AtlasMart has three voting-capable nodes, enough replicas to tolerate one node loss, and a product-search workload. A maintenance change is planned for data node 3. During its return, the exercise simulates a join failure caused by an intentionally wrong plugin/config/version artifact. The team must keep nodes 1–2 coordinated, avoid removing healthy copies, repair the failed node, and restore full redundancy before touching another node.

Metric/evidence Baseline Stop threshold
Primary availability 0 unassigned primaries any unexpected unassigned primary
Voting/leader majority reachable, leader elected no leader or majority at risk
Recovery 0 active at gate recovery not stabilizing within declared budget
App correctness search + authenticated write/read-back pass any correctness/security failure
Tail latency/error record measured p95/p99/error rate declared SLO breach
Snapshot/recovery point verified missing/unverified before version change

2. Preflight evidence package

Baseline evidence — run before touching a node
# Elasticsearch 9.5.3 or OpenSearch 3.8.0; authenticate normally.
GET /
GET /_cluster/health?pretty
GET /_cat/nodes?v&h=name,ip,node.role,master,version,uptime
GET /_cat/shards?v&h=index,shard,prirep,state,node,unassigned.reason
GET /_cat/recovery?v&active_only=true
GET /_cluster/pending_tasks
GET /_nodes/jvm,plugins,settings?filter_path=nodes.*.name,nodes.*.version,nodes.*.jvm.version,nodes.*.plugins.name,nodes.*.settings.node.roles

# OpenSearch: inspect the committed voting configuration.
GET /_cluster/state?filter_path=metadata.cluster_coordination.last_committed_config,cluster_manager_node

# Record application smoke baselines too: search correctness, write/read-back,
# p95/p99 latency, error rate, indexing lag and authentication success.
Snapshot and configuration evidence
# Use the repository established in Chapter 18; names are examples.
GET /_snapshot/atlasmart-dr/_verify
GET /_snapshot/atlasmart-dr/atlasmart-pre-maintenance-20260911

# Export or version-control product configuration outside the cluster:
# - elasticsearch.yml / opensearch.yml overrides
# - JVM overrides
# - plugin inventory
# - TLS/security configuration references (never raw secrets in Git)
# - application/client versions and endpoint configuration

3. Choose the correct role order

For a same-version restart, node order is mostly about quorum and workload. For an upgrade, use the exact product role order: Elasticsearch data tiers first and master-eligible last; OpenSearch data first and cluster-manager last. The exercise changes only one data node, so the voting majority remains available and the current elected coordinator is untouched.

4. Execute gate A: prepare and stop one node

Gate A
PUT /_cluster/settings
{
  "persistent": {"cluster.routing.allocation.enable": "primaries"}
}
POST /_flush

# stop ONLY atlasmart-node-3 using your container/service command
# then verify from another node:
GET /_cluster/health?pretty
GET /_cat/nodes?v&h=name,node.role,master,version
GET /_cat/shards?v&h=index,shard,prirep,state,node,unassigned.reason

# Expected invariant: voting majority still exists and all primaries remain available.

If the cluster loses primary availability or coordination, stop the exercise and restore the node immediately. Do not continue to the injected join failure.

5. Inject one controlled join failure

On the disposable node only, introduce exactly one reversible mismatch: for example, omit a required version-compatible plugin or use a copied configuration with a wrong cluster/discovery/TLS setting. Start the node and observe that it does not join. The goal is to diagnose without touching the healthy nodes.

Evidence when node 3 fails to join
# From a healthy node:
GET /_cat/nodes?v
GET /_cluster/health?pretty
GET /_cluster/pending_tasks
GET /_cat/recovery?v&active_only=true

# On failed node: inspect local service/container logs.
# Expected decision: STOP rollout. Keep healthy majority. Repair node 3 only.
Wrong response

Do not restart nodes 1 and 2 “to make discovery work,” do not clear cluster metadata, and do not bootstrap a new cluster. Those actions can turn one-node join failure into a quorum/data incident.

6. Repair forward and rejoin

Restore the known-good configuration or install the exact compatible plugin, then start node 3. Verify it joins the same cluster UUID and expected version. Restore normal allocation and wait for recovery before resuming indexing pressure or moving to another node.

Gate B — rejoin and recover
GET /
GET /_cat/nodes?v&h=name,node.role,master,version,uptime

PUT /_cluster/settings
{
  "persistent": {"cluster.routing.allocation.enable": null}
}

GET /_cluster/health?wait_for_status=green&wait_for_no_relocating_shards=true&timeout=10m
GET /_cat/recovery?v&active_only=true
GET /_cat/shards?v&h=index,shard,prirep,state,node

7. Application and security acceptance

Use the least-privilege reader/ingester credentials from Chapters 19–20 rather than administrator credentials. Run a representative search, an idempotent write/read-back, one expected 403, and record p95/p99/error rate compared with the baseline. This catches a class of “green cluster, broken application” failures.

Smoke-test contract
READ: atlasmart-products-read search returns expected IDs/facets
WRITE: controlled idempotent event/product write succeeds through intended principal
READ-BACK: written document is visible after normal refresh semantics
DENY: reader cannot delete index or call security administration
LATENCY: record p50/p95/p99 and errors; compare to preflight
FRESHNESS: record indexing/search lag

8. Decision tree: continue, repair forward, or recover

Evidence Decision
Node 3 fixed, cluster green, recovery done, app gates pass Continue to next planned node only if runbook says so.
Node 3 not joining but healthy majority/data intact Stop rollout; repair node 3 only. Do not increase blast radius.
Target version incompatible but old node has not been modified Abort change and preserve old node/version.
Upgrade crossed no-downgrade boundary; cluster repairable Complete supported upgrade/repair forward according to vendor guidance.
Cluster/data cannot be recovered in-place Provision compatible recovery cluster and restore verified snapshot; measure RTO/RPO.

9. Post-game-day evidence record

Maintenance evidence template
change_id: ATLASMART-SEARCH-MAINT-001
product: Elasticsearch | OpenSearch
source_version: ...
target_version: ...
node_changed: ...
pre_snapshot: ... VERIFIED
pre_health: ...
voting_config_before: ...
failure_injected: one reversible join incompatibility
failure_detected_by: logs + node inventory + health
repair: ...
rejoin_time_seconds: measured
recovery_time_seconds: measured
bytes_recovered: measured
p95_before/after: measured
p99_before/after: measured
errors_before/after: measured
security_smoke: PASS/FAIL
final_health: ...
rollback_boundary_crossed: yes/no
continue_authorized_by_gate: yes/no
follow_up: ...

10. Bridge to Chapter 22

Single-cluster maintenance is now evidence-driven. Chapter 22 extends the same discipline across remote clusters and regions, where WAN latency, remote-cluster trust, cross-cluster search, replication lag, write ownership and RPO/RTO introduce new failure modes. A local green cluster does not imply the remote system is safe to fail over.

Check your understanding

  1. What is the purpose of the injected join failure?
  2. Why run application smoke tests after green health?
  3. When may the next node be touched?
  4. What is the rollback plan after the no-downgrade boundary?
  5. What evidence makes the game day reusable?
Review the answers

1. To prove the team stops the rollout, contains blast radius, diagnoses one node and repairs it without destabilizing the healthy voting majority.

2. Cluster allocation can be healthy while API compatibility, security, mappings, clients or application behavior is broken.

3. Only after the prior node rejoins, allocation/recovery stabilizes and every declared cluster/application/security gate passes.

4. Rebuild a compatible cluster and restore a verified compatible snapshot if repair-forward cannot satisfy recovery objectives.

5. Exact versions/config/plugins, voting/shard state, measured recovery and tail latency, failure/repair actions, stop criteria and final authorization decision.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.