Chapter 17 · Data Guard, Standby Databases, Broker, and High Availability
Failure Injection, Observer/Fast-Start Failover Concepts, Split-Brain Avoidance, and DR Testing
Design controlled DR failures around Broker Fast-Start Failover, observer quorum, fencing, client reconnect, role verification, reinstatement and measured RPO/RTO so 'automatic failover' becomes a tested operating procedure.
Learning outcomes
A real outage is not “the primary process stopped.” It can be a host crash, storage failure, network partition, listener outage, stalled archiver, corrupted control file, or a database that is alive but unreachable from clients. Automatic promotion without quorum can create two writable primaries. Oracle Data Guard Broker Fast-Start Failover (FSFO) uses an independent observer, a target standby, health/lag rules and primary-shutdown/fencing behavior to automate promotion while protecting database consistency.
Define observer, Fast-Start Failover target, threshold, lag limit, conditions, auto-reinstate, fencing and split brain.
Use observe-only FSFO as the preferred preproduction automation rehearsal on an entitled topology.
Design failure injections that distinguish database, network, transport/apply and client-routing failures.
Measure RPO and RTO from committed markers, target apply/transport state, service reconnect and business validation.
Require old-primary fencing/reinstate/rebuild and fallback evidence before declaring a DR drill complete.
The chapter was reviewed against Oracle AI Database 26ai RU 23.26.3, SQL Developer 26.2, and SQLcl 26.2.1. Oracle AI Database Free remains limited to 2 foreground CPU cores, 2 GB combined SGA/PGA memory, 12 GB user data, and one installation per logical environment, and receives no Oracle patches or Support service requests. Crucially, the current 26ai licensing matrix marks Oracle Data Guard Redo Apply, SQL Apply, Snapshot Standby, and Oracle Active Data Guard unavailable in Free. Therefore the mandatory Free exercises are architecture/design-validation labs using the existing FREE/FREEPDB1 database; they do not create a standby or enable Data Guard. Actual Data Guard commands are labeled for Oracle Enterprise Edition / EE-ES or entitled cloud offerings with at least two independent database systems plus Oracle Net connectivity. Active Data Guard features and Far Sync have their own option/offerings boundaries. No AWR/ASH/Diagnostics Pack is required for mandatory monitoring examples.
1. The observer provides independent reachability/health judgment
The Broker observer is a lightweight DGMGRL process placed on a third system with independent connectivity to the primary and target standby. FSFO uses the observer plus the target standby to decide whether the primary should be failed over. Oracle describes this as quorum behavior designed to preserve one-primary consistency.
Do not run the observer on the primary host you expect it to adjudicate. A site-wide failure would remove both the database and the observer's independent perspective.
2. Configure FSFO from explicit data-loss and detection policy
EDIT CONFIGURATION SET PROPERTY FastStartFailoverThreshold = 30;EDIT CONFIGURATION SET PROPERTY FastStartFailoverLagLimit = 0;EDIT CONFIGURATION SET PROPERTY FastStartFailoverAutoReinstate = TRUE;SHOW FAST_START FAILOVER;
FastStartFailoverThreshold is detection time after
loss/condition, not total application RTO. A nonzero
FastStartFailoverLagLimit can permit bounded data
loss in supported protection/transport configurations. Choose
both from business policy and tested network behavior, not
example numbers.
3. Observe-only mode rehearses decisions without changing roles
Current Broker supports observe-only FSFO. When trigger conditions occur, Broker records that a failover would have happened but does not promote the standby. This is a high-value staging step for validating observer placement, thresholds, health conditions and false-positive behavior.
ENABLE FAST_START FAILOVER OBSERVE ONLY;START OBSERVER IN BACKGROUND;SHOW FAST_START FAILOVER;SHOW CONFIGURATION;
Use observer/broker logs to confirm which failures would trigger promotion. After observe-only testing, enable actual FSFO only through a controlled change with application/fencing tests complete.
4. Normal shutdown is not an automatic-failover test
Oracle explicitly states that normal primary shutdown
(NORMAL, IMMEDIATE, or
TRANSACTIONAL) does not trigger FSFO. That is
desirable: planned maintenance should use switchover or
controlled shutdown semantics, not be mistaken for catastrophe.
To test database crash detection use an isolated crash process/host scenario; to test network partition, isolate paths; to test client routing, break service access without killing the database. One failure type cannot prove every DR behavior.
5. Split-brain avoidance requires fencing the old primary
Split brain means two sites believe they are primary and accept divergent writes. FSFO uses Broker/observer/target coordination and can shut down a primary under configured unsafe/stalled conditions. Operational DR plans still need external controls: storage/network fencing, service/DNS/load-balancer exclusivity, and confirmation that an isolated former primary cannot accept application writes.
SELECT db_unique_name, database_role, open_mode, protection_mode, protection_level, fs_failover_status, fs_failover_current_target, fs_failover_observer_presentFROM v$database;
Never rely only on a load balancer saying “primary endpoint healthy.” Query the database role itself and route writable traffic only to the verified primary.
6. Measure RPO with committed business markers
Before a drill, create a monotonically identified transaction marker on the primary and record its commit timestamp/SCN. After failover, query the highest marker visible on the new primary. The missing committed-marker interval is observed data loss for the drill.
CREATE TABLE servicehub_dr_marker ( marker_id NUMBER PRIMARY KEY, committed_at TIMESTAMP WITH TIME ZONE DEFAULT SYSTIMESTAMP NOT NULL, marker_scn NUMBER NOT NULL, marker_text VARCHAR2(100) NOT NULL);INSERT INTO servicehub_dr_marker( marker_id,marker_scn,marker_text)VALUES( 170001, DBMS_FLASHBACK.GET_SYSTEM_CHANGE_NUMBER, 'pre-failure marker');COMMIT;
For repeated drills, use a sequence/application generator rather than hard-coded IDs and preserve marker evidence outside the database. A committed marker tests application data currency more meaningfully than only comparing redo sequence numbers.
7. Measure RTO from failure declaration to validated service
| Milestone | Evidence |
|---|---|
| T0 failure injected/detected | Observer/broker/app timestamp |
| Promotion decision | Observer/broker log |
| New primary opened | V$DATABASE.DATABASE_ROLE='PRIMARY' |
| Writable service available | Listener/service + role-based service state |
| Client reconnect | Connection pool obtains new-primary session |
| Business validation | Marker/latest critical rows and transaction test pass |
| T_end | RTO = T_end − T0 |
Broker failover time is only one component of end-to-end RTO.
8. Free DR-testing lab: rehearse the measurement harness without Data Guard
Use Free to build the marker/runbook/stopwatch discipline, but do not claim a standby failover occurred.
ALTER SESSION SET CONTAINER=FREEPDB1;CREATE TABLE servicehub_dr_marker_free ( marker_id NUMBER PRIMARY KEY, committed_at TIMESTAMP WITH TIME ZONE DEFAULT SYSTIMESTAMP NOT NULL, marker_scn NUMBER NOT NULL);INSERT INTO servicehub_dr_marker_free(marker_id,marker_scn)VALUES(1,DBMS_FLASHBACK.GET_SYSTEM_CHANGE_NUMBER);COMMIT;SELECT *FROM servicehub_dr_marker_free;SELECT db_unique_name, database_role, open_modeFROM v$database;
$start = Get-Date# inject the approved failure / wait for client reconnection / run validation$end = Get-Date$end - $start
DROP TABLE servicehub_dr_marker_free PURGE;
On a real entitled topology, run the same timing/marker process on the promoted target and compare the last committed source marker to the last visible new-primary marker.
9. Reintegration is part of failover, not a later housekeeping detail
After failover, the former primary is divergent or failed. If Flashback Database history supports it, Broker can reinstate it as a standby; otherwise rebuild it. Verify that redo flows in the new direction and that the new standby is a viable future target before declaring HA restored.
SHOW CONFIGURATION;SHOW DATABASE VERBOSE 'servicehub_stby';SHOW DATABASE VERBOSE 'servicehub_pri';SHOW FAST_START FAILOVER;-- If Broker says the former primary is reinstateable:REINSTATE DATABASE 'servicehub_pri';
10. Deliberately wrong: test “automatic failover” by unplugging random cables in production
Uncontrolled failure injection can isolate only some network paths, leave the old primary writable, break unrelated systems, or create an event no one can reproduce. DR tests need an approved fault matrix, start/abort criteria, fencing plan, observer logs, client metrics, rollback/fallback, and designated authority.
11. Failure matrix
| Failure | Expected database behavior | Separate client test |
|---|---|---|
| Primary instance crash | Observer/target evaluate FSFO condition after threshold | Reconnect to new primary service |
| Primary host/site loss | Failover if target/observer/quorum/lag policy allow | DNS/LB/service path survives site loss |
| Primary↔standby transport break | Lag grows; protection status changes according to mode | Primary may remain serviceable |
| Observer-only loss | No reckless automatic promotion; observer state visible | Application should remain on valid primary |
| Standby apply stall | Apply lag grows while transport may remain healthy | Failover readiness/RTO degrades |
| Client-network failure only | Database roles unchanged | Application connectivity layer handles path failure |
12. Production judgment and chapter close
Enable automatic failover only after observe-only and manual drills show that topology, lag policy, fencing, services and clients behave correctly. Keep the observer independent, monitor it, and treat a stale/lagging target as a protection degradation. Record the exact RPO/RTO observed for every scenario, not the marketing target.
Data Guard Redo Apply is unavailable in Free; actual Broker/FSFO drills need an entitled Data Guard topology. Far Sync and real-time-query/automatic-block-repair features require Active Data Guard under the current offering rules. The observer host itself does not require a separate database license, but the protected databases do. This completes Chapter 17's progression from redo physics to tested HA/DR operations.
Check your understanding
- What independent component does FSFO use to help decide whether failover is safe?
- Will SHUTDOWN IMMEDIATE normally trigger FSFO?
- What is split brain?
- How should RPO be measured in a DR drill?
- Why is failover not complete immediately after the target becomes PRIMARY?
Review the answers
The Data Guard Broker observer, together with the target standby and Broker state.
No. Oracle explicitly suppresses FSFO for normal planned shutdowns.
Two databases/sites accept divergent writes because both behave as primary; fencing/routing/quorum must prevent it.
Compare the last source commit/business marker that should exist with the latest marker actually present on the new primary.
Clients/services must reconnect and validate, the former primary must be fenced/reinstated/rebuilt, and protection must be restored for the next failure.
Authoritative references
- Managing Fast-Start Failover — thresholds, lag limits, fencing and auto-reinstate
- Fast-Start Failover — observe-only mode and reinstatement
- SHOW FAST_START FAILOVER — observer/target/condition evidence
- Configure and Deploy Data Guard — observer placement and split-brain avoidance
- Licensing Information — Data Guard/ADG/observer-host licensing boundaries