Convert “upgrade the cluster” into a compatibility project that inventories server/JVM/plugins/clients, freezes an exact upgrade path, creates a tested recovery point, and defines the point beyond which rollback means rebuild-and-restore rather than downgrade.

Version Compatibility, Breaking Changes, Plugin Compatibility, Snapshot Before Upgrade, and Rollback Boundaries

Operate Elasticsearch and OpenSearch with explicit cluster-health, quorum, compatibility, recovery and application gates instead of maintenance folklore.

Intermediate → Advanced120–170 minutesCluster Operations · Chapter 21 · Lesson 04Elasticsearch/Kibana 9.5.3 · OpenSearch/Dashboards 3.8.0 · bundled JVMsLast reviewed: September 2026

Learning outcomes

AtlasMart’s upgrade board has an appealing but invalid plan: “take snapshot, install newest version everywhere, downgrade if needed.” A real plan must prove the supported source→target path, breaking/deprecation state, index compatibility, plugin and client compatibility, JVM/runtime assumptions, managed-service restrictions, and the exact rollback boundary before the first node changes.

01

Freeze exact source/target versions and a supported upgrade path rather than using `latest`.

02

Inventory and validate every plugin, client, ingest component, JVM/config override and security dependency.

03

Create and verify a pre-upgrade snapshot/recovery point without assuming arbitrary cross-version/product restore compatibility.

04

Explain why both Elasticsearch and OpenSearch treat downgrade as rebuild/restore rather than an in-place binary rollback.

05

Define explicit stop/continue checkpoints for breaking changes, mixed-version state and application compatibility.

Pinned maintenance baseline

Examples are frozen to Elasticsearch 9.5.3 / Kibana 9.5.3 (released 3 September 2026) and OpenSearch 3.8.0 / OpenSearch Dashboards 3.8.0 (released 4 August 2026). Use each distribution's bundled JVM in the lab and record the actual runtime with GET /_nodes/jvm; do not carry forward arbitrary JVM flags from an older installation. The earlier atlasmart-es and atlasmart-os single-node containers remain useful for API syntax, but this chapter's quorum/recovery exercises require either a disposable three-node cluster or the deterministic trace path supplied below. TLS/authentication from Chapters 19–20 stays enabled. Managed services may hide node ordering, voting, plugin, filesystem, or allocation controls; their provider runbook overrides self-managed commands.

No-downgrade boundary

Elasticsearch 9.5 documentation explicitly states that in-place downgrade is unsupported and a mixed-version cluster is valid only during rolling upgrade. OpenSearch likewise documents that nodes cannot be downgraded; rollback requires a new installation and snapshot restore. Treat this as an architectural constraint, not a last-minute surprise.

Execution and safety note

Run mutating, destructive, security, lifecycle, snapshot, failure-injection, and load-test commands only in the disposable AtlasMart lab or an equivalently isolated environment. Verify the target cluster, index, tenant, credentials, and rollback path before execution; treat shown output as an expected invariant unless the lesson explicitly labels it as captured evidence.

1. Build the compatibility ledger

Pre-upgrade manifest
cluster = atlasmart-search
source_server = <exact version>
target_server = <exact supported version>
release_dates_checked = yes
upgrade_path_checked = yes
nodes_and_roles = <inventory>
bundled_jvm = <GET _nodes/jvm>
config_overrides = <diff from shipped defaults>
plugins = <name + exact version per node>
clients = <language/client version + compatibility mode if applicable>
ingest_components = <exact versions>
security = <TLS CA/cert expiry + auth plugins/realms/backends>
index_compatibility = <checked>
breaking_changes = <resolved or accepted>
snapshot = <repository/snapshot + verification + restore drill reference>
rollback = rebuild exact source version + restore compatible snapshot

2. Exact versions matter in both ecosystems

For Elasticsearch, current 9.5.3 guidance includes source-version prerequisites for major upgrades; for example, upgrades to 9.1+ from 8.x require 8.19.x first. Elasticsearch also has release-date ordering rules for out-of-order minor releases. For OpenSearch, rolling upgrades support adjacent major versions and the current documentation states that upgrades to 3.x require at least 2.19.0. These are product/version rules, not semantic-version intuition.

Check Elasticsearch 9.5.3 OpenSearch 3.8.0
Upgrade path Use current Elastic upgrade planner/assistant and exact release rules Rolling supports adjacent majors; 3.x minimum documented source is 2.19.0
Node role order data tiers → other non-master → master/voting-only data → ingest/ML/coordinating → cluster-manager
Plugins upgrade every installed plugin on each upgraded node plugin major/minor/patch must match OpenSearch compatibility guidance
Downgrade unsupported in-place unsupported in-place
Recovery rebuild compatible old cluster + restore snapshot new installation + restore compatible snapshot

3. Snapshot before upgrade is necessary, not sufficient

A snapshot is useful only if the repository is readable, the snapshot completed successfully, the target recovery cluster is compatible, required feature/security state is understood, and the restore procedure has been rehearsed. Do not call a snapshot “rollback” until Chapter 18’s data/schema/application validation has been exercised.

Recovery-point gate
GET /_snapshot/<repo>/_verify
GET /_snapshot/<repo>/<pre-upgrade-snapshot>

# Verify status is SUCCESS and record included indices/data streams/state.
# Then reference a prior isolated restore drill and measured RTO.
# Never assume an Elasticsearch snapshot restores into OpenSearch, or vice versa.

4. Plugin compatibility is a hard join constraint

A server binary can be valid while a node still cannot join because an installed plugin is missing or incompatible. Inventory plugins from every node; do not assume uniformity because one node looks correct. Elasticsearch’s upgrade procedure requires upgrading installed plugins on the stopped node. OpenSearch advises matching plugin compatibility to the exact target line and separately checking third-party plugin support.

Inventory plugins on every node
GET /_nodes/plugins?filter_path=nodes.*.name,nodes.*.version,nodes.*.plugins.name,nodes.*.plugins.version,nodes.*.plugins.description

# Compare the set node-by-node and against target-version compatibility docs.
# Treat missing/extra plugins as a change item, not a surprise during restart.

5. Client compatibility and server compatibility are separate

Even if every node joins, AtlasMart can fail because a removed API, stricter validation or client serialization changed. Record client and ingest-component versions and exercise real requests against a staging/mixed-version rehearsal. REST compatibility can reduce migration pressure in some Elastic major-upgrade paths, but it is not permission to ignore deprecations indefinitely. OpenSearch clients/plugins have their own matrices.

6. Wrong approach: “downgrade if smoke tests fail”

Once a new-version node has participated, the old binary is not a supported rollback mechanism. The safe decision must happen earlier: if a gate fails while the old node is still untouched, stop there. If a failure occurs after the upgrade crossed the no-downgrade boundary, either complete the supported rolling path and repair forward or invoke the documented rebuild-and-restore recovery plan.

Mixed-version clusters are temporary

Do not operate a deliberately mixed-version cluster as steady state. Version skew exists only to complete the rolling upgrade.

7. Upgrade decision matrix

Condition Decision
Unsupported source→target path STOP; select supported intermediates or migration method.
Unresolved breaking/deprecation blocker STOP; remediate in staging/source version first.
Snapshot unverified / restore untested STOP; create and test recovery point.
Plugin has no target-compatible build STOP; remove/migrate feature safely or postpone.
Node rejoined but app smoke failed STOP next node; diagnose before increasing version skew.
All gates pass Advance exactly one node in documented role order.

8. Bridge to the game day

Lesson 5 combines coordination, recovery, node order, compatibility and application gates under one deliberately injected failure. The goal is not to make the cluster fail spectacularly; it is to prove that the team knows exactly when to stop and what evidence authorizes recovery.

Check your understanding

  1. Why pin exact source and target versions?
  2. Does a successful snapshot automatically prove rollback?
  3. Why inventory plugins on every node?
  4. What is the normal downgrade strategy?
  5. When should the rollout stop after one node?
Review the answers

1. Upgrade eligibility, release-date rules, plugin compatibility and breaking changes are version-specific.

2. No. Repository access, version/product compatibility, included state and an isolated restore/application validation must also be proven.

3. A single incompatible or missing plugin can prevent a node from starting/joining and plugin sets can drift between nodes.

4. There is no supported in-place downgrade; rebuild a compatible older cluster and restore a compatible snapshot if recovery requires going back.

5. Whenever cluster, recovery, compatibility, security or application acceptance gates fail.

Summary and next step

Preserve the evidence, assumptions, version boundaries, and safety checks established in this lesson. Carry them into the next lesson—or, at the end of the capstone, into the production runbook—rather than treating this lesson as an isolated recipe.

References

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.