Chapter 18 · RAC and Clustered Oracle Concepts
Rolling Maintenance, Instance Failure, Service Relocation, and Operational Resilience
Turn RAC availability into an operating procedure: Clusterware resource state, service draining/relocation, failed-instance recovery, reconfiguration evidence, rolling patch compatibility and rollback/fallback boundaries.
Learning outcomes
ServiceHub needs a quarterly database/Grid Infrastructure update
without taking the whole cluster down. The old runbook runs
shutdown abort on node 1, patches it, then repeats
node 2. Users see dropped transactions and a thundering herd on
the surviving node. RAC's availability advantage requires
planned service draining, node-by-node resource
changes, patch-specific rolling compatibility and explicit
verification of cluster reconfiguration/recovery.
Explain instance failure/eviction, surviving-instance recovery and cluster reconfiguration at a mechanism level.
Drain/relocate services before planned instance maintenance and verify session/service state.
Use current 23.26.3 RAC reconfiguration history/histogram views as post-event evidence.
Apply Grid Infrastructure/database version compatibility and patch rolling-eligibility rules before maintenance.
Build a Free maintenance timeline/runbook while keeping actual node rolling operations on entitled RAC.
This chapter was reviewed against Oracle AI Database 26ai RU 23.26.3, SQL Developer 26.2, and SQLcl 26.2.1. Oracle AI Database Free is limited to 2 foreground CPU cores, 2 GB combined SGA/PGA RAM, 12 GB user data, one installation per logical environment, and receives no Release Update patches or Oracle Support service requests. Current 26ai licensing marks Oracle Real Application Clusters (RAC) unavailable in Free, SE2-ODA, BaseDB SE, BaseDB EE, and BaseDB EE-HP; RAC is an extra-cost option on EE/EE-ES and included with BaseDB EE-EP and ExaDB. Actual RAC requires Oracle Grid Infrastructure/Clusterware, multiple cluster nodes, shared database storage, a low-latency private interconnect, public/VIP/SCAN networking, and a certified platform. Mandatory Free labs therefore validate single-instance workload/service/storage facts and RAC design decisions; they never attempt to form a RAC cluster. Dynamic V$/GV$ diagnostics shown for entitled RAC systems avoid AWR/ASH so the lesson does not require Diagnostics Pack.
1. Node/instance failure triggers membership and recovery work
If an RAC instance fails or Clusterware evicts a node to preserve cluster integrity, surviving members reconfigure global resources. Locks/enqueues held by the failed instance are cleaned up and its redo thread is used for instance recovery so committed changes survive and uncommitted transactions are rolled back, just as Chapter 14's single-instance recovery mechanics require.
During reconfiguration, applications may see brief pauses while global resource ownership is remastered/recovered. A second node being “up” does not mean zero interruption.
SELECT inst_id, instance_name, host_name, status, thread#, startup_timeFROM gv$instanceORDER BY inst_id;SELECT thread#, status, enabled, instanceFROM gv$threadORDER BY thread#;
2. 23.26.3 adds direct RAC reconfiguration history
Oracle AI Database 26ai RU 23.26.3 introduces
V$RAC_RECONFIGURATION_HISTORY and
V$RAC_RECONFIGURATION_HISTOGRAM. The history view
records completed reconfiguration events with elapsed time and
JSON timing/statistics detail; the histogram summarizes
elapsed-time buckets by reconfiguration type.
SELECT inc#, timestamp, elapsed_time, timing_stats, statsFROM v$rac_reconfiguration_historyORDER BY timestamp DESCFETCH FIRST 20 ROWS ONLY;SELECT *FROM v$rac_reconfiguration_histogramORDER BY type;
These views are event evidence, not an application SLA. Correlate their timestamps with service/FAN/client latency and alert/Clusterware logs.
3. Planned maintenance starts with service movement, not instance kill
srvctl status service -db servicehub_rac -service servicehub_oltp -verbose
srvctl relocate service -db servicehub_rac -service servicehub_oltp -oldinst shrac1 -newinst shrac2 -drain_timeout 120 -stopoption TRANSACTIONALsrvctl status service -db servicehub_rac -service servicehub_oltp -verbose
120 seconds is an example only. The drain window
should be measured from actual request/transaction durations.
With FAN-aware pools, planned events can stop new work reaching
the drained instance while existing requests finish gradually.
4. Then stop the instance through Clusterware-aware tooling
srvctl stop instance -db servicehub_rac -instance shrac1 -stopoption TRANSACTIONAL -drain_timeout 120srvctl status database -db servicehub_rac -verbose
Using SRVCTL keeps Clusterware resource state
coherent. Direct SQL shutdown is sometimes valid for database
administration, but clustered maintenance runbooks should avoid
fighting Clusterware placement/restart policy.
5. Rolling does not mean every patch can be applied independently to every node
Before a rolling maintenance window, read the exact Release Update/one-off patch README and Oracle Support certification for whether the patch is rolling-compatible and whether Grid Infrastructure, database home and drivers must be upgraded in a particular order. Oracle Grid Infrastructure in the cluster must be the same version as or newer than the highest RAC database release it runs.
For 26ai quarterly RUs, Oracle recommends out-of-place patching with Gold Images. OPatch/OPatchAuto remain available for in-place patching; using OPatch/OPatchAuto for out-of-place patching is deprecated. Oracle AI Database Free does not receive RU patching and cannot reproduce this lab.
6. Generic in-place node-by-node shape
# Drain database services/instance first according to the runbook.$GRID_HOME/OPatch/opatchauto apply /stage/patch_ID# Verify Clusterware/node health before moving to the next node.crsctl check crscrsctl stat res -t# After database service returns:srvctl status database -db servicehub_rac -verbosesrvctl status service -db servicehub_rac -verbose
Do not copy a patch ID, home order, or rollback sequence from another RU. Patch tooling and rolling support are patch-specific.
7. Failure handling differs from planned draining
| Event | Desired mechanism | Expected client effect |
|---|---|---|
| Planned node maintenance | FAN notification + request/session drain + service relocation | Minimal/new work directed away; in-flight work given time to finish |
| Instance crash | Cluster reconfiguration + instance recovery + service restart elsewhere | Connections to failed process break; pools reconnect/replay where supported |
| Node/interconnect fault requiring eviction | Clusterware fencing/eviction to preserve membership integrity | Failed node removed; surviving capacity carries workload |
| Shared-storage failure | Not solved merely by another RAC instance | Can affect every instance; storage HA/backup/DR required |
8. Deliberately wrong: kill instance 1 first and hope service relocation is graceful
An abrupt stop proves failure recovery, not planned maintenance quality. In-flight sessions die and all reconnect pressure can hit the surviving node at once. The repair is to relocate/drain services, verify capacity and session movement, then stop/patch the instance. Separately test true crash/eviction behavior in an approved failure drill.
9. Free design-validation lab: maintenance state machine
CREATE TABLE servicehub_rac_maintenance_plan ( step_no NUMBER PRIMARY KEY, step_name VARCHAR2(80) NOT NULL, required_evidence VARCHAR2(500) NOT NULL, rollback_gate VARCHAR2(500) NOT NULL);INSERT INTO servicehub_rac_maintenance_plan VALUES (10,'Capacity precheck', 'Surviving node capacity and service targets verified', 'Abort before drain if capacity insufficient');INSERT INTO servicehub_rac_maintenance_plan VALUES (20,'Drain and relocate', 'No new sessions on maintenance instance; requests drain', 'Relocate service back if drain/client errors exceed policy');INSERT INTO servicehub_rac_maintenance_plan VALUES (30,'Patch one node', 'Patch inventory success and Clusterware resources online', 'Follow patch-specific rollback/old-home procedure');INSERT INTO servicehub_rac_maintenance_plan VALUES (40,'Return service', 'Service/client/business checks pass', 'Keep service on healthy node if patched node fails validation');SELECT * FROM servicehub_rac_maintenance_plan ORDER BY step_no;DROP TABLE servicehub_rac_maintenance_plan PURGE;
Free cannot execute the cluster operations, but the lab forces every maintenance step to have evidence and a rollback gate.
10. Production judgment
RAC rolling maintenance is an availability workflow, not merely a patch command. Verify exact GI/database compatibility, patch rolling capability, node capacity, service drain behavior, client FAN/replay readiness, post-patch SQL/component state and reconfiguration timings. Maintain an explicit fallback path to the previous home/patch level where supported.
RU 23.26.3 is the chapter baseline and introduces the RAC reconfiguration history views used here. Free is not patchable with RUs and cannot host RAC. Lesson 5 closes the chapter by deciding when the operational/licensing complexity is justified—or when Data Guard, Globally Distributed Database or application-level scaling solves the actual problem more directly.
Check your understanding
- What happens to failed-instance transaction/resource state after an RAC instance failure?
- What is the purpose of planned service draining?
- What new RU 23.26.3 views expose RAC reconfiguration timing/history?
- Can every database/Grid patch be assumed rolling-compatible?
- Why doesn't a second RAC node solve shared-storage failure by itself?
Review the answers
Surviving members reconfigure global resources and perform instance recovery/cleanup using the failed instance's redo/transaction state.
It stops/directs new work away while giving existing requests/transactions a measured window to finish before the instance is stopped.
V$RAC_RECONFIGURATION_HISTORY and V$RAC_RECONFIGURATION_HISTOGRAM.
No. Rolling support, order and rollback procedures are specific to the exact RU/one-off and topology.
All RAC instances open the same database storage; a storage failure domain can affect every instance unless storage HA/DR protects it separately.
Authoritative references
- Administering RAC Instances and Cluster Databases — SRVCTL and clustered instance administration
- Ensuring Application Continuity — planned service draining
- V$RAC_RECONFIGURATION_HISTORY — 23.26.3 reconfiguration history
- Grid Infrastructure Installation/Upgrade Guide — rolling patching and current 26ai patch-tool guidance
- Oracle Clusterware Administration and Deployment Guide — Grid Infrastructure version compatibility and cluster operations