Chapter 17 · Data Guard, Standby Databases, Broker, and High Availability

Failure Injection, Observer/Fast-Start Failover Concepts, Split-Brain Avoidance, and DR Testing

Design controlled DR failures around Broker Fast-Start Failover, observer quorum, fencing, client reconnect, role verification, reinstatement and measured RPO/RTO so 'automatic failover' becomes a tested operating procedure.

Advanced130–150 minutesFSFO/observer/fencing + measured DR drillObserve-only mode preferred before automatic failoverRPO/RTO measured, split-brain explicitly fencedLast reviewed: August 2026

Learning outcomes

A real outage is not “the primary process stopped.” It can be a host crash, storage failure, network partition, listener outage, stalled archiver, corrupted control file, or a database that is alive but unreachable from clients. Automatic promotion without quorum can create two writable primaries. Oracle Data Guard Broker Fast-Start Failover (FSFO) uses an independent observer, a target standby, health/lag rules and primary-shutdown/fencing behavior to automate promotion while protecting database consistency.

01

Define observer, Fast-Start Failover target, threshold, lag limit, conditions, auto-reinstate, fencing and split brain.

02

Use observe-only FSFO as the preferred preproduction automation rehearsal on an entitled topology.

03

Design failure injections that distinguish database, network, transport/apply and client-routing failures.

04

Measure RPO and RTO from committed markers, target apply/transport state, service reconnect and business validation.

05

Require old-primary fencing/reinstate/rebuild and fallback evidence before declaring a DR drill complete.

Generation-time baseline, licensing, topology, and safety boundary

The chapter was reviewed against Oracle AI Database 26ai RU 23.26.3, SQL Developer 26.2, and SQLcl 26.2.1. Oracle AI Database Free remains limited to 2 foreground CPU cores, 2 GB combined SGA/PGA memory, 12 GB user data, and one installation per logical environment, and receives no Oracle patches or Support service requests. Crucially, the current 26ai licensing matrix marks Oracle Data Guard Redo Apply, SQL Apply, Snapshot Standby, and Oracle Active Data Guard unavailable in Free. Therefore the mandatory Free exercises are architecture/design-validation labs using the existing FREE/FREEPDB1 database; they do not create a standby or enable Data Guard. Actual Data Guard commands are labeled for Oracle Enterprise Edition / EE-ES or entitled cloud offerings with at least two independent database systems plus Oracle Net connectivity. Active Data Guard features and Far Sync have their own option/offerings boundaries. No AWR/ASH/Diagnostics Pack is required for mandatory monitoring examples.

1. The observer provides independent reachability/health judgment

The Broker observer is a lightweight DGMGRL process placed on a third system with independent connectivity to the primary and target standby. FSFO uses the observer plus the target standby to decide whether the primary should be failed over. Oracle describes this as quorum behavior designed to preserve one-primary consistency.

Do not run the observer on the primary host you expect it to adjudicate. A site-wide failure would remove both the database and the observer's independent perspective.

2. Configure FSFO from explicit data-loss and detection policy

text · entitled broker configuration — example policy, not universal values
EDIT CONFIGURATION SET PROPERTY FastStartFailoverThreshold = 30;EDIT CONFIGURATION SET PROPERTY FastStartFailoverLagLimit = 0;EDIT CONFIGURATION SET PROPERTY FastStartFailoverAutoReinstate = TRUE;SHOW FAST_START FAILOVER;

FastStartFailoverThreshold is detection time after loss/condition, not total application RTO. A nonzero FastStartFailoverLagLimit can permit bounded data loss in supported protection/transport configurations. Choose both from business policy and tested network behavior, not example numbers.

3. Observe-only mode rehearses decisions without changing roles

Current Broker supports observe-only FSFO. When trigger conditions occur, Broker records that a failover would have happened but does not promote the standby. This is a high-value staging step for validating observer placement, thresholds, health conditions and false-positive behavior.

text · entitled topology
ENABLE FAST_START FAILOVER OBSERVE ONLY;START OBSERVER IN BACKGROUND;SHOW FAST_START FAILOVER;SHOW CONFIGURATION;

Use observer/broker logs to confirm which failures would trigger promotion. After observe-only testing, enable actual FSFO only through a controlled change with application/fencing tests complete.

4. Normal shutdown is not an automatic-failover test

Oracle explicitly states that normal primary shutdown (NORMAL, IMMEDIATE, or TRANSACTIONAL) does not trigger FSFO. That is desirable: planned maintenance should use switchover or controlled shutdown semantics, not be mistaken for catastrophe.

Failure injection must match the failure

To test database crash detection use an isolated crash process/host scenario; to test network partition, isolate paths; to test client routing, break service access without killing the database. One failure type cannot prove every DR behavior.

5. Split-brain avoidance requires fencing the old primary

Split brain means two sites believe they are primary and accept divergent writes. FSFO uses Broker/observer/target coordination and can shut down a primary under configured unsafe/stalled conditions. Operational DR plans still need external controls: storage/network fencing, service/DNS/load-balancer exclusivity, and confirmation that an isolated former primary cannot accept application writes.

sql · role verification after any transition
SELECT  db_unique_name,  database_role,  open_mode,  protection_mode,  protection_level,  fs_failover_status,  fs_failover_current_target,  fs_failover_observer_presentFROM v$database;

Never rely only on a load balancer saying “primary endpoint healthy.” Query the database role itself and route writable traffic only to the verified primary.

6. Measure RPO with committed business markers

Before a drill, create a monotonically identified transaction marker on the primary and record its commit timestamp/SCN. After failover, query the highest marker visible on the new primary. The missing committed-marker interval is observed data loss for the drill.

sql · application marker table on the primary — design for entitled DR drill
CREATE TABLE servicehub_dr_marker (  marker_id       NUMBER PRIMARY KEY,  committed_at    TIMESTAMP WITH TIME ZONE DEFAULT SYSTIMESTAMP NOT NULL,  marker_scn      NUMBER NOT NULL,  marker_text     VARCHAR2(100) NOT NULL);INSERT INTO servicehub_dr_marker(  marker_id,marker_scn,marker_text)VALUES(  170001,  DBMS_FLASHBACK.GET_SYSTEM_CHANGE_NUMBER,  'pre-failure marker');COMMIT;

For repeated drills, use a sequence/application generator rather than hard-coded IDs and preserve marker evidence outside the database. A committed marker tests application data currency more meaningfully than only comparing redo sequence numbers.

7. Measure RTO from failure declaration to validated service

Milestone Evidence
T0 failure injected/detected Observer/broker/app timestamp
Promotion decision Observer/broker log
New primary opened V$DATABASE.DATABASE_ROLE='PRIMARY'
Writable service available Listener/service + role-based service state
Client reconnect Connection pool obtains new-primary session
Business validation Marker/latest critical rows and transaction test pass
T_end RTO = T_end − T0

Broker failover time is only one component of end-to-end RTO.

8. Free DR-testing lab: rehearse the measurement harness without Data Guard

Use Free to build the marker/runbook/stopwatch discipline, but do not claim a standby failover occurred.

sql · create a local drill marker and capture source evidence
ALTER SESSION SET CONTAINER=FREEPDB1;CREATE TABLE servicehub_dr_marker_free (  marker_id NUMBER PRIMARY KEY,  committed_at TIMESTAMP WITH TIME ZONE DEFAULT SYSTIMESTAMP NOT NULL,  marker_scn NUMBER NOT NULL);INSERT INTO servicehub_dr_marker_free(marker_id,marker_scn)VALUES(1,DBMS_FLASHBACK.GET_SYSTEM_CHANGE_NUMBER);COMMIT;SELECT *FROM servicehub_dr_marker_free;SELECT  db_unique_name,  database_role,  open_modeFROM v$database;
powershell · PowerShell stopwatch skeleton for a later real DR drill
$start = Get-Date# inject the approved failure / wait for client reconnection / run validation$end = Get-Date$end - $start
sql · cleanup
DROP TABLE servicehub_dr_marker_free PURGE;

On a real entitled topology, run the same timing/marker process on the promoted target and compare the last committed source marker to the last visible new-primary marker.

9. Reintegration is part of failover, not a later housekeeping detail

After failover, the former primary is divergent or failed. If Flashback Database history supports it, Broker can reinstate it as a standby; otherwise rebuild it. Verify that redo flows in the new direction and that the new standby is a viable future target before declaring HA restored.

text · entitled post-failover Broker evidence
SHOW CONFIGURATION;SHOW DATABASE VERBOSE 'servicehub_stby';SHOW DATABASE VERBOSE 'servicehub_pri';SHOW FAST_START FAILOVER;-- If Broker says the former primary is reinstateable:REINSTATE DATABASE 'servicehub_pri';

10. Deliberately wrong: test “automatic failover” by unplugging random cables in production

Uncontrolled failure injection can isolate only some network paths, leave the old primary writable, break unrelated systems, or create an event no one can reproduce. DR tests need an approved fault matrix, start/abort criteria, fencing plan, observer logs, client metrics, rollback/fallback, and designated authority.

11. Failure matrix

Failure Expected database behavior Separate client test
Primary instance crash Observer/target evaluate FSFO condition after threshold Reconnect to new primary service
Primary host/site loss Failover if target/observer/quorum/lag policy allow DNS/LB/service path survives site loss
Primary↔standby transport break Lag grows; protection status changes according to mode Primary may remain serviceable
Observer-only loss No reckless automatic promotion; observer state visible Application should remain on valid primary
Standby apply stall Apply lag grows while transport may remain healthy Failover readiness/RTO degrades
Client-network failure only Database roles unchanged Application connectivity layer handles path failure

12. Production judgment and chapter close

Enable automatic failover only after observe-only and manual drills show that topology, lag policy, fencing, services and clients behave correctly. Keep the observer independent, monitor it, and treat a stale/lagging target as a protection degradation. Record the exact RPO/RTO observed for every scenario, not the marketing target.

Data Guard Redo Apply is unavailable in Free; actual Broker/FSFO drills need an entitled Data Guard topology. Far Sync and real-time-query/automatic-block-repair features require Active Data Guard under the current offering rules. The observer host itself does not require a separate database license, but the protected databases do. This completes Chapter 17's progression from redo physics to tested HA/DR operations.

Check your understanding

  1. What independent component does FSFO use to help decide whether failover is safe?
  2. Will SHUTDOWN IMMEDIATE normally trigger FSFO?
  3. What is split brain?
  4. How should RPO be measured in a DR drill?
  5. Why is failover not complete immediately after the target becomes PRIMARY?
Review the answers

The Data Guard Broker observer, together with the target standby and Broker state.

No. Oracle explicitly suppresses FSFO for normal planned shutdowns.

Two databases/sites accept divergent writes because both behave as primary; fencing/routing/quorum must prevent it.

Compare the last source commit/business marker that should exist with the latest marker actually present on the new primary.

Clients/services must reconnect and validate, the former primary must be fenced/reinstated/rebuilt, and protection must be restored for the next failure.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.