Chapter 18 · RAC and Clustered Oracle Concepts

Rolling Maintenance, Instance Failure, Service Relocation, and Operational Resilience

Turn RAC availability into an operating procedure: Clusterware resource state, service draining/relocation, failed-instance recovery, reconfiguration evidence, rolling patch compatibility and rollback/fallback boundaries.

Advanced125–145 minutesDrain/relocate/reconfiguration maintenance runbookRU 23.26.3 RAC reconfiguration views includedRolling eligibility is patch-specificLast reviewed: August 2026

Learning outcomes

ServiceHub needs a quarterly database/Grid Infrastructure update without taking the whole cluster down. The old runbook runs shutdown abort on node 1, patches it, then repeats node 2. Users see dropped transactions and a thundering herd on the surviving node. RAC's availability advantage requires planned service draining, node-by-node resource changes, patch-specific rolling compatibility and explicit verification of cluster reconfiguration/recovery.

01

Explain instance failure/eviction, surviving-instance recovery and cluster reconfiguration at a mechanism level.

02

Drain/relocate services before planned instance maintenance and verify session/service state.

03

Use current 23.26.3 RAC reconfiguration history/histogram views as post-event evidence.

04

Apply Grid Infrastructure/database version compatibility and patch rolling-eligibility rules before maintenance.

05

Build a Free maintenance timeline/runbook while keeping actual node rolling operations on entitled RAC.

Generation-time baseline, licensing, topology, and tooling boundary

This chapter was reviewed against Oracle AI Database 26ai RU 23.26.3, SQL Developer 26.2, and SQLcl 26.2.1. Oracle AI Database Free is limited to 2 foreground CPU cores, 2 GB combined SGA/PGA RAM, 12 GB user data, one installation per logical environment, and receives no Release Update patches or Oracle Support service requests. Current 26ai licensing marks Oracle Real Application Clusters (RAC) unavailable in Free, SE2-ODA, BaseDB SE, BaseDB EE, and BaseDB EE-HP; RAC is an extra-cost option on EE/EE-ES and included with BaseDB EE-EP and ExaDB. Actual RAC requires Oracle Grid Infrastructure/Clusterware, multiple cluster nodes, shared database storage, a low-latency private interconnect, public/VIP/SCAN networking, and a certified platform. Mandatory Free labs therefore validate single-instance workload/service/storage facts and RAC design decisions; they never attempt to form a RAC cluster. Dynamic V$/GV$ diagnostics shown for entitled RAC systems avoid AWR/ASH so the lesson does not require Diagnostics Pack.

1. Node/instance failure triggers membership and recovery work

If an RAC instance fails or Clusterware evicts a node to preserve cluster integrity, surviving members reconfigure global resources. Locks/enqueues held by the failed instance are cleaned up and its redo thread is used for instance recovery so committed changes survive and uncommitted transactions are rolled back, just as Chapter 14's single-instance recovery mechanics require.

During reconfiguration, applications may see brief pauses while global resource ownership is remastered/recovered. A second node being “up” does not mean zero interruption.

sql · entitled RAC: instance/redo-thread state
SELECT  inst_id,  instance_name,  host_name,  status,  thread#,  startup_timeFROM gv$instanceORDER BY inst_id;SELECT  thread#,  status,  enabled,  instanceFROM gv$threadORDER BY thread#;

2. 23.26.3 adds direct RAC reconfiguration history

Oracle AI Database 26ai RU 23.26.3 introduces V$RAC_RECONFIGURATION_HISTORY and V$RAC_RECONFIGURATION_HISTOGRAM. The history view records completed reconfiguration events with elapsed time and JSON timing/statistics detail; the histogram summarizes elapsed-time buckets by reconfiguration type.

sql · entitled RAC at RU 23.26.3+
SELECT  inc#,  timestamp,  elapsed_time,  timing_stats,  statsFROM v$rac_reconfiguration_historyORDER BY timestamp DESCFETCH FIRST 20 ROWS ONLY;SELECT *FROM v$rac_reconfiguration_histogramORDER BY type;

These views are event evidence, not an application SLA. Correlate their timestamps with service/FAN/client latency and alert/Clusterware logs.

3. Planned maintenance starts with service movement, not instance kill

text · entitled cluster: verify placement and session load
srvctl status service   -db servicehub_rac   -service servicehub_oltp   -verbose
text · drain and relocate before stopping the instance
srvctl relocate service   -db servicehub_rac   -service servicehub_oltp   -oldinst shrac1   -newinst shrac2   -drain_timeout 120   -stopoption TRANSACTIONALsrvctl status service   -db servicehub_rac   -service servicehub_oltp   -verbose

120 seconds is an example only. The drain window should be measured from actual request/transaction durations. With FAN-aware pools, planned events can stop new work reaching the drained instance while existing requests finish gradually.

4. Then stop the instance through Clusterware-aware tooling

text · entitled cluster
srvctl stop instance   -db servicehub_rac   -instance shrac1   -stopoption TRANSACTIONAL   -drain_timeout 120srvctl status database -db servicehub_rac -verbose

Using SRVCTL keeps Clusterware resource state coherent. Direct SQL shutdown is sometimes valid for database administration, but clustered maintenance runbooks should avoid fighting Clusterware placement/restart policy.

5. Rolling does not mean every patch can be applied independently to every node

Before a rolling maintenance window, read the exact Release Update/one-off patch README and Oracle Support certification for whether the patch is rolling-compatible and whether Grid Infrastructure, database home and drivers must be upgraded in a particular order. Oracle Grid Infrastructure in the cluster must be the same version as or newer than the highest RAC database release it runs.

For 26ai quarterly RUs, Oracle recommends out-of-place patching with Gold Images. OPatch/OPatchAuto remain available for in-place patching; using OPatch/OPatchAuto for out-of-place patching is deprecated. Oracle AI Database Free does not receive RU patching and cannot reproduce this lab.

6. Generic in-place node-by-node shape

text · root/Grid owner shell — only if the exact patch is documented as rolling-compatible
# Drain database services/instance first according to the runbook.$GRID_HOME/OPatch/opatchauto apply /stage/patch_ID# Verify Clusterware/node health before moving to the next node.crsctl check crscrsctl stat res -t# After database service returns:srvctl status database -db servicehub_rac -verbosesrvctl status service -db servicehub_rac -verbose

Do not copy a patch ID, home order, or rollback sequence from another RU. Patch tooling and rolling support are patch-specific.

7. Failure handling differs from planned draining

Event Desired mechanism Expected client effect
Planned node maintenance FAN notification + request/session drain + service relocation Minimal/new work directed away; in-flight work given time to finish
Instance crash Cluster reconfiguration + instance recovery + service restart elsewhere Connections to failed process break; pools reconnect/replay where supported
Node/interconnect fault requiring eviction Clusterware fencing/eviction to preserve membership integrity Failed node removed; surviving capacity carries workload
Shared-storage failure Not solved merely by another RAC instance Can affect every instance; storage HA/backup/DR required

8. Deliberately wrong: kill instance 1 first and hope service relocation is graceful

An abrupt stop proves failure recovery, not planned maintenance quality. In-flight sessions die and all reconnect pressure can hit the surviving node at once. The repair is to relocate/drain services, verify capacity and session movement, then stop/patch the instance. Separately test true crash/eviction behavior in an approved failure drill.

9. Free design-validation lab: maintenance state machine

sql · create a local evidence runbook
CREATE TABLE servicehub_rac_maintenance_plan (  step_no NUMBER PRIMARY KEY,  step_name VARCHAR2(80) NOT NULL,  required_evidence VARCHAR2(500) NOT NULL,  rollback_gate VARCHAR2(500) NOT NULL);INSERT INTO servicehub_rac_maintenance_plan VALUES (10,'Capacity precheck',  'Surviving node capacity and service targets verified',  'Abort before drain if capacity insufficient');INSERT INTO servicehub_rac_maintenance_plan VALUES (20,'Drain and relocate',  'No new sessions on maintenance instance; requests drain',  'Relocate service back if drain/client errors exceed policy');INSERT INTO servicehub_rac_maintenance_plan VALUES (30,'Patch one node',  'Patch inventory success and Clusterware resources online',  'Follow patch-specific rollback/old-home procedure');INSERT INTO servicehub_rac_maintenance_plan VALUES (40,'Return service',  'Service/client/business checks pass',  'Keep service on healthy node if patched node fails validation');SELECT * FROM servicehub_rac_maintenance_plan ORDER BY step_no;DROP TABLE servicehub_rac_maintenance_plan PURGE;

Free cannot execute the cluster operations, but the lab forces every maintenance step to have evidence and a rollback gate.

10. Production judgment

RAC rolling maintenance is an availability workflow, not merely a patch command. Verify exact GI/database compatibility, patch rolling capability, node capacity, service drain behavior, client FAN/replay readiness, post-patch SQL/component state and reconfiguration timings. Maintain an explicit fallback path to the previous home/patch level where supported.

RU 23.26.3 is the chapter baseline and introduces the RAC reconfiguration history views used here. Free is not patchable with RUs and cannot host RAC. Lesson 5 closes the chapter by deciding when the operational/licensing complexity is justified—or when Data Guard, Globally Distributed Database or application-level scaling solves the actual problem more directly.

Check your understanding

  1. What happens to failed-instance transaction/resource state after an RAC instance failure?
  2. What is the purpose of planned service draining?
  3. What new RU 23.26.3 views expose RAC reconfiguration timing/history?
  4. Can every database/Grid patch be assumed rolling-compatible?
  5. Why doesn't a second RAC node solve shared-storage failure by itself?
Review the answers

Surviving members reconfigure global resources and perform instance recovery/cleanup using the failed instance's redo/transaction state.

It stops/directs new work away while giving existing requests/transactions a measured window to finish before the instance is stopped.

V$RAC_RECONFIGURATION_HISTORY and V$RAC_RECONFIGURATION_HISTOGRAM.

No. Rolling support, order and rollback procedures are specific to the exact RU/one-off and topology.

All RAC instances open the same database storage; a storage failure domain can affect every instance unless storage HA/DR protects it separately.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.