Execute rolling maintenance as a gated state machine: baseline, optional replica-allocation restriction, flush, one-node change, rejoin, recovery, smoke tests, and only then advance to the next node.
Rolling Restart/Upgrade Sequencing, Shard Allocation Control, Synced Recovery Concepts, and Validation
Operate Elasticsearch and OpenSearch with explicit cluster-health, quorum, compatibility, recovery and application gates instead of maintenance folklore.
Learning outcomes
AtlasMart needs a certificate refresh and then a version upgrade. The operational temptation is to treat both as “restart one node at a time.” The reliable procedure is more explicit: establish a baseline, suppress unnecessary replica churn where appropriate, flush, change exactly one node, preserve version-compatible leadership, restore allocation, wait for recovery, validate the application, and only then move to the next node.
Build a product-specific rolling order for Elasticsearch 9.5.3 and OpenSearch 3.8.0.
Use `cluster.routing.allocation.enable` deliberately and restore it after every node gate.
Explain what a current flush does and why historical synced-flush folklore must not be copied blindly.
Validate node rejoin, shard recovery, plugin/JVM/version state and application behavior before advancing.
Distinguish a same-version rolling restart from a rolling upgrade with irreversible version boundaries.
Examples are frozen to
Elasticsearch 9.5.3 / Kibana 9.5.3 (released
3 September 2026) and
OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0
(released 4 August 2026). Use each distribution's bundled JVM
in the lab and record the actual runtime with
GET /_nodes/jvm; do not carry forward arbitrary
JVM flags from an older installation. The earlier
atlasmart-es and
atlasmart-os single-node containers remain useful
for API syntax, but this chapter's quorum/recovery exercises
require either a disposable three-node cluster or the
deterministic trace path supplied below. TLS/authentication
from Chapters 19–20 stays enabled. Managed services may hide
node ordering, voting, plugin, filesystem, or allocation
controls; their provider runbook overrides self-managed
commands.
A maintenance gate is intentionally conservative: stop if a primary shard is unassigned, the cluster cannot elect/publish through a voting majority, recovery or relocation is still in progress beyond the runbook budget, a required plugin/client is incompatible, a snapshot/recovery point is not verified, or application correctness/error-rate criteria fail. Yellow can be acceptable only when the exact cause is understood and explicitly permitted by the runbook; never normalize unexpected yellow/red health.
Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.
1. Rolling maintenance is a state machine
| Gate | Evidence required | Advance only when |
|---|---|---|
| Baseline | health, nodes, shards, plugins/JVM, snapshot, app p95/p99/errors | known-good baseline stored |
| Pre-stop | next node role/order, allocation setting, voting majority | quorum and capacity remain after node stops |
| Node down | node absent; leader still present; app still correct | expected degraded state only |
| Node rejoin | correct version/config/plugins/certs; node joined intended cluster | no duplicate/new cluster identity |
| Recovery | allocation restored; recovery/relocation complete | green or explicitly accepted state |
| Application | read/write/search/auth smoke; tail latency/error rate | all acceptance criteria pass |
2. Product-specific role order
In Elasticsearch self-managed rolling upgrades, current guidance upgrades data nodes tier-by-tier from frozen → cold → warm → hot → other data/content, then non-data/non-master roles such as ML, ingest, coordinating, transform and remote-cluster-client nodes, and finally master/voting-only master-eligible nodes. This ordering keeps the older-version master compatible with nodes joining during the upgrade.
OpenSearch rolling upgrades use data nodes first, then ingest/ML/coordinating nodes, and cluster-manager nodes last. Identify the currently elected cluster manager and restart it last within that group.
Never copy Elasticsearch role names/order verbatim into OpenSearch or vice versa. Record the node roles actually present in your cluster and map them to the current product documentation.
3. Restrict replica allocation to avoid wasteful churn
When a data node is expected to return quickly, temporarily
setting allocation to primaries prevents immediate
replica recreation elsewhere. This is not “turning off
recovery”; primary allocation can still occur. The setting must
be restored after the node rejoins. Forgetting it is a common
self-inflicted yellow-cluster incident.
# Before stopping one data node:
PUT /_cluster/settings
{
"persistent": {
"cluster.routing.allocation.enable": "primaries"
}
}
# Optional: reduce nonessential indexing and flush.
POST /_flush
# ... stop/change/start exactly one node ...
# After it rejoins, restore normal allocation.
PUT /_cluster/settings
{
"persistent": {
"cluster.routing.allocation.enable": null
}
}
GET /_cluster/health?wait_for_status=green&wait_for_no_relocating_shards=true&timeout=10m
GET /_cat/recovery?v&active_only=true
4. “Synced recovery” means understand recovery—not call an old endpoint
Older Elasticsearch operations literature often mentions a
synced flush. In modern Elasticsearch the
explicit synced-flush API was removed; a regular flush has had
the same recovery-relevant effect since 7.6. Current
Elasticsearch 9.5 and OpenSearch 3.8 rolling procedures use
POST /_flush. A flush performs a Lucene commit and
trims translog requirements; it does not guarantee instant
recovery, and it does not make a missing shard copy magically
valid.
Peer recovery can reuse existing local segment files and replay
retained operations depending on the source/target histories,
sequence-number state and retained recovery data. Measure
/_cat/recovery rather than assuming “flush means
zero-byte recovery.”
Do not carry `POST /_flush/synced` into a 9.5 Elasticsearch runbook. Likewise, do not invent a cross-product “synced recovery” API. Use current flush and recovery APIs, and validate the actual bytes/stages.
5. Version upgrade is not a reversible restart
Elasticsearch explicitly does not support downgrading an in-place upgraded cluster; after upgraded nodes participate, cluster metadata may advance. OpenSearch similarly states that nodes cannot be downgraded. The recovery boundary is therefore snapshot/rebuild, not “install the old binary back over the data directory.” Mixed versions are a temporary rolling-upgrade state only.
GET /
GET /_cat/nodes?v&h=name,node.role,master,version,uptime
GET /_nodes/plugins,jvm?filter_path=nodes.*.name,nodes.*.version,nodes.*.plugins.name,nodes.*.jvm.version
GET /_cluster/health?pretty
GET /_cat/shards?v&s=state
GET /_cat/recovery?v&active_only=true
GET /_cluster/pending_tasks
# App checks (examples, use least-privilege credentials):
GET /atlasmart-products-read/_search?size=1
# write/read-back only through the intended write principal/index
6. Wrong approach: suppress every symptom and keep going
An operator sees relocation and sets allocation to
none, sees a plugin failure and removes the plugin,
sees a yellow cluster and continues because search still works.
Each action erases evidence and expands risk. The repair is one
variable at a time: restore documented allocation state, install
the target-compatible plugin, wait for recovery, and compare to
the baseline. If the gate fails, stop the rollout.
7. AtlasMart rolling sequence
for node in PRODUCT_SPECIFIC_ORDER:
assert snapshot_verified
assert voting_majority_after_stop(node)
assert cluster_gate_passes()
if node.is_data:
allocation = "primaries"
flush_if_useful()
stop(node)
assert cluster_still_coordinated()
upgrade_or_maintain(node)
install_exact_compatible_plugins(node)
start(node)
assert node_joined_expected_cluster_and_version()
allocation = "all/default"
wait_for_recovery()
assert app_smoke_and_p99_gate()
# never run next node merely because previous process started successfully
Check your understanding
- Why restrict allocation to primaries during a short data-node restart?
- What replaced historical synced flush in current Elasticsearch runbooks?
- Why are master/cluster-manager nodes upgraded last?
- Can an in-place upgrade be rolled back by reinstalling an older binary?
- What does a node process starting prove?
Review the answers
1. To avoid unnecessary replica relocation while still permitting primaries; restore normal allocation after the node rejoins.
2. A regular flush; the explicit synced-flush API is removed.
3. Older-version coordinators can generally lead while newer nodes join; the reverse is not always compatible during the rolling window.
4. No. Plan rollback as rebuild plus restore from a compatible snapshot/recovery point.
5. Only that it started; it still must join the intended cluster with correct version/plugins/config and pass recovery and application gates.
Summary and next step
Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.
References
- Elastic 9.5.3 release notes — Current Elasticsearch release baseline.
- Elastic self-managed Elasticsearch upgrade — Rolling order, allocation restriction, flush, plugin upgrade, recovery and no-downgrade boundary.
- Elastic plan/prepare upgrade — Upgrade path, compatibility and preparation guidance.
- Elastic add/remove nodes — Node removal, shard draining and master-eligible considerations.
- Elastic diagnose unassigned shards — Health meaning and allocation-explain workflow.
- Elastic shard allocation/routing settings — Allocation controls, filters, awareness and decommissioning primitives.
- OpenSearch 3.8 version history — Current OpenSearch release baseline.
- OpenSearch rolling upgrade — Health gate, allocation restriction, flush and role ordering.
- OpenSearch migrate or upgrade — Upgrade methods, plugin compatibility, configuration backup and snapshots.
- OpenSearch voting and quorum — Voting majority and maintenance safety.
- OpenSearch voting configuration — Voting membership and automatic shrink behavior.
- OpenSearch allocation explain API — Allocation diagnosis.
- OpenSearch cluster settings — Allocation include/require/exclude and recovery settings.
- OpenSearch nodes info API — Version/JVM/plugin/node-role inventory.