Replace or retire AtlasMart nodes without gambling on rebalance: drain shard ownership, preserve voting safety, validate disk/tier constraints, add capacity before removing it, and prove the old node is no longer required.

Node Replacement, Disk Expansion, Tier Migration, Rebalancing, and Decommissioning

Operate Elasticsearch and OpenSearch with explicit cluster-health, quorum, compatibility, recovery and application gates instead of maintenance folklore.

Intermediate → Advanced120–170 minutesCluster Operations · Chapter 21 · Lesson 03Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

AtlasMart must replace a data node whose disk is near its operational headroom limit. A safe replacement is not “add disk, then delete the old VM.” It is a capacity transfer: add/validate destination capacity, move shard responsibility intentionally, preserve voting membership, observe recovery, and only then remove the source from discovery and infrastructure.

01

Distinguish online disk expansion from node replacement and from tier migration.

02

Drain shards with allocation filters instead of shutting down a still-required data holder.

03

Coordinate master/cluster-manager voting changes separately from data-shard movement.

04

Validate rebalancing/recovery headroom and avoid decommissioning while the cluster is still dependent on the node.

05

Choose add-before-remove, replace-in-place, or migrate-to-new-tier based on failure domain, capacity and rollback needs.

Pinned maintenance baseline

Examples are frozen to Elasticsearch 9.5.3 / Kibana 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0 (released 4 August 2026). Use each distribution's bundled JVM in the lab and record the actual runtime with GET /_nodes/jvm; do not carry forward arbitrary JVM flags from an older installation. The earlier atlasmart-es and atlasmart-os single-node containers remain useful for API syntax, but this chapter's quorum/recovery exercises require either a disposable three-node cluster or the deterministic trace path supplied below. TLS/authentication from Chapters 19–20 stays enabled. Managed services may hide node ordering, voting, plugin, filesystem, or allocation controls; their provider runbook overrides self-managed commands.

Stop gate

A maintenance gate is intentionally conservative: stop if a primary shard is unassigned, the cluster cannot elect/publish through a voting majority, recovery or relocation is still in progress beyond the runbook budget, a required plugin/client is incompatible, a snapshot/recovery point is not verified, or application correctness/error-rate criteria fail. Yellow can be acceptable only when the exact cause is understood and explicitly permitted by the runbook; never normalize unexpected yellow/red health.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Four operations that look similar but have different risk

Operation Data movement Coordination concern Rollback
Disk expansion on same node May be none node/process restart may still be needed depending platform usually platform/storage rollback, not shard rollback
Replace node in place local data may be preserved or rebuilt same cluster identity/roles/config/certs required old host may be fallback only before metadata/data-format changes
Add new node then drain old peer recovery/relocation voting if old/new are eligible coordinators clear exclusion and move shards back if old still healthy
Tier migration potentially many shard relocations allocation/tier capacity and lifecycle rules reverse only if source tier/capacity remains available

2. Add capacity before subtracting it

For a permanent data-node replacement, first start the replacement with the intended exact server version, node roles, TLS trust, discovery settings, JVM policy and compatible plugins. Confirm it joins the existing cluster—never bootstrap a new cluster. Then use allocation filters to drain the old node.

Drain one old node — portable allocation intent
# Inspect first.
GET /_cat/allocation?v
GET /_cat/shards?v&s=node,index,shard
GET /_cluster/health?pretty

# Exclude a node from shard allocation. Both products support built-in _name filtering.
PUT /_cluster/settings
{
  "persistent": {
    "cluster.routing.allocation.exclude._name": "atlasmart-data-old"
  }
}

# Observe; do not stop the node yet.
GET /_cat/recovery?v&active_only=true
GET /_cat/shards?v&h=index,shard,prirep,state,node
GET /_cluster/allocation/explain?include_disk_info=true

# After the node owns no required shards and the destination state is healthy,
# clear the temporary filter.
PUT /_cluster/settings
{
  "persistent": {
    "cluster.routing.allocation.exclude._name": null
  }
}

3. Data drain and voting removal are separate workflows

A node can be empty of shards but still be part of the coordination voting configuration. Conversely, a data-only node can be drained without any voting change. For a permanent removal of multiple master/cluster-manager-eligible nodes, inspect and use the product-specific voting exclusion/reconfiguration mechanism so a majority remains throughout. Remove one eligible node at a time when possible.

Do not equate “0 shards” with “safe to power off”

A coordinating/master-eligible node may hold no data shards and still be essential to quorum. A data node may also host persistent tasks or application traffic. Validate every responsibility before removal.

4. Disk expansion does not erase watermarks or recovery physics

Expanding a volume may relieve a disk watermark, but the cluster still needs enough free space for merges, shard relocation, snapshots/restores and failure recovery. After expansion, verify the operating system sees the new filesystem size, then verify node filesystem stats and allocation decisions. Do not lower watermarks merely to postpone capacity work.

Capacity evidence after storage change
GET /_cat/allocation?v&bytes=gb
GET /_nodes/stats/fs?filter_path=nodes.*.name,nodes.*.fs.total
GET /_cluster/settings?include_defaults=true&flat_settings=true
GET /_cluster/allocation/explain?include_disk_info=true

5. Tier migration is policy plus capacity

Elasticsearch data tiers and OpenSearch allocation/tier patterns are not API-identical. The invariant is: the destination tier must have role/attribute eligibility, enough disk/IOPS/CPU, awareness placement and recovery headroom before policy or filters move shards. A lifecycle action that “should move warm” cannot manufacture warm capacity.

6. Wrong approach: decommission while recovery is still active

If AtlasMart shuts down atlasmart-data-old as soon as relocation begins, source reads may disappear mid-copy and the cluster can restart recovery from another copy, increasing network/disk load. If it was the only valid primary copy, the outcome can be worse. The correct gate is no required shards/tasks/coordination dependency plus healthy destination copies and application smoke success.

Decommission acceptance record
node: atlasmart-data-old
replacement: atlasmart-data-new
server/version: EXACT
plugins: EXACT COMPATIBLE SET
voting role: yes/no
shards_before: N
shards_after_drain: 0 required copies
active_recoveries: 0
cluster_health: green (or approved explicit state)
app_smoke: PASS
p99_delta: measured
snapshot/recovery_point: VERIFIED
seed_hosts/config updated: YES
old_node_poweroff_authorized: YES

7. Failure injection: replacement node cannot join

In a disposable lab, intentionally give the replacement one wrong discovery/TLS/plugin setting. The correct response is not to remove the old node anyway. Stop, inspect logs and node inventory, repair the single incompatibility, and require the new node to join before draining. This makes “add-before-remove” a real rollback boundary.

8. Bridge to compatibility planning

Node replacement can be a same-version infrastructure event or part of a version upgrade. The latter adds index-format, plugin, client and no-downgrade constraints. Lesson 4 freezes those dependencies before any binary changes.

Check your understanding

  1. Why add the replacement before draining the old node?
  2. Does zero shards mean a master/cluster-manager-eligible node is safe to remove?
  3. What proves a disk expansion succeeded operationally?
  4. Why wait for recovery before decommissioning?
  5. What is the safe response if the replacement cannot join?
Review the answers

1. It proves capacity, discovery, TLS, plugin and role compatibility while the old node is still available as a safe fallback.

2. No. Voting/coordination responsibility is independent of data-shard ownership.

3. Host filesystem plus node fs stats, allocation decisions, recovered headroom and workload behavior—not just cloud-control-plane status.

4. Stopping a source mid-recovery can restart or lengthen recovery and can threaten availability if valid copies are limited.

5. Keep the old node, stop the change, diagnose the exact compatibility/configuration issue, and retry only after the join gate passes.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.