Execute rolling maintenance as a gated state machine: baseline, optional replica-allocation restriction, flush, one-node change, rejoin, recovery, smoke tests, and only then advance to the next node.

Rolling Restart/Upgrade Sequencing, Shard Allocation Control, Synced Recovery Concepts, and Validation

Operate Elasticsearch and OpenSearch with explicit cluster-health, quorum, compatibility, recovery and application gates instead of maintenance folklore.

Intermediate → Advanced120–170 minutesCluster Operations · Chapter 21 · Lesson 02Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

AtlasMart needs a certificate refresh and then a version upgrade. The operational temptation is to treat both as “restart one node at a time.” The reliable procedure is more explicit: establish a baseline, suppress unnecessary replica churn where appropriate, flush, change exactly one node, preserve version-compatible leadership, restore allocation, wait for recovery, validate the application, and only then move to the next node.

01

Build a product-specific rolling order for Elasticsearch 9.5.3 and OpenSearch 3.8.0.

02

Use `cluster.routing.allocation.enable` deliberately and restore it after every node gate.

03

Explain what a current flush does and why historical synced-flush folklore must not be copied blindly.

04

Validate node rejoin, shard recovery, plugin/JVM/version state and application behavior before advancing.

05

Distinguish a same-version rolling restart from a rolling upgrade with irreversible version boundaries.

Pinned maintenance baseline

Examples are frozen to Elasticsearch 9.5.3 / Kibana 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0 (released 4 August 2026). Use each distribution's bundled JVM in the lab and record the actual runtime with GET /_nodes/jvm; do not carry forward arbitrary JVM flags from an older installation. The earlier atlasmart-es and atlasmart-os single-node containers remain useful for API syntax, but this chapter's quorum/recovery exercises require either a disposable three-node cluster or the deterministic trace path supplied below. TLS/authentication from Chapters 19–20 stays enabled. Managed services may hide node ordering, voting, plugin, filesystem, or allocation controls; their provider runbook overrides self-managed commands.

Stop gate

A maintenance gate is intentionally conservative: stop if a primary shard is unassigned, the cluster cannot elect/publish through a voting majority, recovery or relocation is still in progress beyond the runbook budget, a required plugin/client is incompatible, a snapshot/recovery point is not verified, or application correctness/error-rate criteria fail. Yellow can be acceptable only when the exact cause is understood and explicitly permitted by the runbook; never normalize unexpected yellow/red health.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Rolling maintenance is a state machine

Gate Evidence required Advance only when
Baseline health, nodes, shards, plugins/JVM, snapshot, app p95/p99/errors known-good baseline stored
Pre-stop next node role/order, allocation setting, voting majority quorum and capacity remain after node stops
Node down node absent; leader still present; app still correct expected degraded state only
Node rejoin correct version/config/plugins/certs; node joined intended cluster no duplicate/new cluster identity
Recovery allocation restored; recovery/relocation complete green or explicitly accepted state
Application read/write/search/auth smoke; tail latency/error rate all acceptance criteria pass

2. Product-specific role order

In Elasticsearch self-managed rolling upgrades, current guidance upgrades data nodes tier-by-tier from frozen → cold → warm → hot → other data/content, then non-data/non-master roles such as ML, ingest, coordinating, transform and remote-cluster-client nodes, and finally master/voting-only master-eligible nodes. This ordering keeps the older-version master compatible with nodes joining during the upgrade.

OpenSearch rolling upgrades use data nodes first, then ingest/ML/coordinating nodes, and cluster-manager nodes last. Identify the currently elected cluster manager and restart it last within that group.

Same intent, different exact runbook

Never copy Elasticsearch role names/order verbatim into OpenSearch or vice versa. Record the node roles actually present in your cluster and map them to the current product documentation.

3. Restrict replica allocation to avoid wasteful churn

When a data node is expected to return quickly, temporarily setting allocation to primaries prevents immediate replica recreation elsewhere. This is not “turning off recovery”; primary allocation can still occur. The setting must be restored after the node rejoins. Forgetting it is a common self-inflicted yellow-cluster incident.

Allocation guard around one data node
# Before stopping one data node:
PUT /_cluster/settings
{
  "persistent": {
    "cluster.routing.allocation.enable": "primaries"
  }
}

# Optional: reduce nonessential indexing and flush.
POST /_flush

# ... stop/change/start exactly one node ...

# After it rejoins, restore normal allocation.
PUT /_cluster/settings
{
  "persistent": {
    "cluster.routing.allocation.enable": null
  }
}

GET /_cluster/health?wait_for_status=green&wait_for_no_relocating_shards=true&timeout=10m
GET /_cat/recovery?v&active_only=true

4. “Synced recovery” means understand recovery—not call an old endpoint

Older Elasticsearch operations literature often mentions a synced flush. In modern Elasticsearch the explicit synced-flush API was removed; a regular flush has had the same recovery-relevant effect since 7.6. Current Elasticsearch 9.5 and OpenSearch 3.8 rolling procedures use POST /_flush. A flush performs a Lucene commit and trims translog requirements; it does not guarantee instant recovery, and it does not make a missing shard copy magically valid.

Peer recovery can reuse existing local segment files and replay retained operations depending on the source/target histories, sequence-number state and retained recovery data. Measure /_cat/recovery rather than assuming “flush means zero-byte recovery.”

Incorrect command pattern

Do not carry `POST /_flush/synced` into a 9.5 Elasticsearch runbook. Likewise, do not invent a cross-product “synced recovery” API. Use current flush and recovery APIs, and validate the actual bytes/stages.

5. Version upgrade is not a reversible restart

Elasticsearch explicitly does not support downgrading an in-place upgraded cluster; after upgraded nodes participate, cluster metadata may advance. OpenSearch similarly states that nodes cannot be downgraded. The recovery boundary is therefore snapshot/rebuild, not “install the old binary back over the data directory.” Mixed versions are a temporary rolling-upgrade state only.

Node-by-node validation bundle
GET /
GET /_cat/nodes?v&h=name,node.role,master,version,uptime
GET /_nodes/plugins,jvm?filter_path=nodes.*.name,nodes.*.version,nodes.*.plugins.name,nodes.*.jvm.version
GET /_cluster/health?pretty
GET /_cat/shards?v&s=state
GET /_cat/recovery?v&active_only=true
GET /_cluster/pending_tasks

# App checks (examples, use least-privilege credentials):
GET /atlasmart-products-read/_search?size=1
# write/read-back only through the intended write principal/index

6. Wrong approach: suppress every symptom and keep going

An operator sees relocation and sets allocation to none, sees a plugin failure and removes the plugin, sees a yellow cluster and continues because search still works. Each action erases evidence and expands risk. The repair is one variable at a time: restore documented allocation state, install the target-compatible plugin, wait for recovery, and compare to the baseline. If the gate fails, stop the rollout.

7. AtlasMart rolling sequence

Runbook pseudocode
for node in PRODUCT_SPECIFIC_ORDER:
    assert snapshot_verified
    assert voting_majority_after_stop(node)
    assert cluster_gate_passes()
    if node.is_data:
        allocation = "primaries"
    flush_if_useful()
    stop(node)
    assert cluster_still_coordinated()
    upgrade_or_maintain(node)
    install_exact_compatible_plugins(node)
    start(node)
    assert node_joined_expected_cluster_and_version()
    allocation = "all/default"
    wait_for_recovery()
    assert app_smoke_and_p99_gate()
# never run next node merely because previous process started successfully

Check your understanding

  1. Why restrict allocation to primaries during a short data-node restart?
  2. What replaced historical synced flush in current Elasticsearch runbooks?
  3. Why are master/cluster-manager nodes upgraded last?
  4. Can an in-place upgrade be rolled back by reinstalling an older binary?
  5. What does a node process starting prove?
Review the answers

1. To avoid unnecessary replica relocation while still permitting primaries; restore normal allocation after the node rejoins.

2. A regular flush; the explicit synced-flush API is removed.

3. Older-version coordinators can generally lead while newer nodes join; the reverse is not always compatible during the rolling window.

4. No. Plan rollback as rebuild plus restore from a compatible snapshot/recovery point.

5. Only that it started; it still must join the intended cluster with correct version/plugins/config and pass recovery and application gates.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.