Chapter 14 · Redo, Undo, Control Files, Checkpoints, and Instance Recovery
Redo Log Buffer, Online Redo Logs, Log Groups/Members, Log Switches, and Archiving
Trace Oracle change vectors from the redo log buffer through LGWR into online redo groups and archived redo, while separating group size, member multiplexing, archive destinations, and switch cadence.
Learning outcomes
ServiceHub commits thousands of small changes. One operator assumes each online redo file is a separate backup; another adds a second member on the same fragile volume and calls the configuration fully redundant. The missing model is the redo pipeline: database work creates change vectors, server processes place redo entries in the System Global Area (SGA) redo log buffer, and Log Writer (LGWR) writes them sequentially to every member of the current online redo log group. In ARCHIVELOG mode, completed log sequences are copied to archive destinations before reuse.
Trace redo entries from the redo log buffer through LGWR into online redo groups and members.
Distinguish group size from member multiplexing and identify real versus cosmetic failure-domain redundancy.
Interpret V$LOG/V$LOGFILE status, sequence number, archive state, and current/active/inactive groups.
Explain natural and manual log switches without treating forced switching as normal tuning.
Inspect ARCHIVELOG mode and archive destinations without changing archiving mode in the mandatory lab.
Mandatory examples target Oracle AI Database Free 26ai and were reviewed against RU 23.26.3, SQL Developer 26.2, and SQLcl 26.2.1. Free is limited to 2 foreground CPUs, 2 GB combined SGA/PGA memory, 12 GB user data, and one installation per logical environment. Oracle provides no patches or Support service requests for Free, including no security patches. The course baseline uses CDB FREE and application PDB FREEPDB1. Redo logs, control files, checkpoints, and instance recovery are CDB/instance-level concerns; application DML remains in FREEPDB1. Mandatory diagnostics use ordinary dynamic performance views and do not require AWR, ASH, Diagnostics Pack, Tuning Pack, RAC, Data Guard, or Exadata.
1. Redo protects database changes for recovery
Redo records change vectors needed to reapply database changes after failure. Server processes copy redo entries into the redo log buffer; LGWR writes them to the current online redo log. Online redo contains records for committed and uncommitted transactions. Changes to undo blocks for permanent objects also generate redo, so redo and undo are complementary rather than substitutes.
SELECT name,value,isdefault,issys_modifiable,ispdb_modifiableFROM v$parameter WHERE name='log_buffer';SELECT name,log_mode,force_logging,checkpoint_change#FROM v$database;
LOG_BUFFER is static and not PDB-modifiable. A slow
commit is not evidence that this buffer should be enlarged.
2. Groups are logical; members are physical copies
LGWR cycles through online redo groups. A group may contain multiple identical members; LGWR writes the same redo to each. Multiplexing helps only to the extent that members occupy independent storage failure domains.
SELECT group#,thread#,sequence#,bytes,members,archived,statusFROM v$log ORDER BY group#;SELECT group#,type,status,member,is_recovery_dest_fileFROM v$logfile ORDER BY group#,member;
| Status | Meaning |
|---|---|
CURRENT |
LGWR is writing this group. |
ACTIVE |
Still needed for instance/crash recovery. |
INACTIVE |
No longer needed for instance recovery and reusable when archive rules permit. |
UNUSED |
Not yet written since creation or RESETLOGS. |
3. Log switches create recovery sequence boundaries
A log switch moves LGWR to the next reusable group. Normally the current group fills naturally. Oracle assigns a new log sequence number each time a group becomes current; recovery uses these sequence numbers in order.
SELECT * FROM ( SELECT thread#,sequence#,first_change#,first_time,next_change# FROM v$log_history ORDER BY first_time DESC) FETCH FIRST 20 ROWS ONLY;
Switch history proves switch cadence, not the root cause of latency. Correlate it with waits, DBWR/checkpoint progress, redo rate, and archive destination health.
4. ARCHIVELOG changes reuse rules
In ARCHIVELOG mode, ARCn processes archive completed online redo sequences. LGWR cannot reuse a required group until it is no longer needed for instance recovery and archiving requirements are satisfied. Archived redo supports media recovery, standby apply, and LogMiner.
SELECT log_mode FROM v$database;SELECT dest_id,status,target,destination,errorFROM v$archive_destWHERE status <> 'INACTIVE'ORDER BY dest_id;SELECT process,status,log_sequenceFROM v$archive_processesORDER BY process;
The lab does not switch the CDB between NOARCHIVELOG and ARCHIVELOG because that changes backup/recovery semantics and requires planned instance-level administration.
5. Deliberately wrong: force a switch every minute
ALTER SYSTEM SWITCH LOGFILE is a legitimate
administrative command, but frequent scheduled switching can
create unnecessary checkpoint/archive/control-file work and
expose log reuse pressure. One observation switch is enough to
understand the mechanism.
SELECT group#,sequence#,bytes,status,archived FROM v$log ORDER BY group#;ALTER SYSTEM SWITCH LOGFILE;SELECT group#,sequence#,bytes,status,archived FROM v$log ORDER BY group#;
Do not turn the observation into a scheduler job. If the next group cannot be reused, sessions can wait on checkpoint- or archiving-related log-switch events.
6. Measure redo generated by a real transaction
SHOW CON_NAME-- Ensure this is FREEPDB1 before creating application objects.BEGIN EXECUTE IMMEDIATE 'DROP TABLE servicehub_redo_case PURGE';EXCEPTION WHEN OTHERS THEN IF SQLCODE != -942 THEN RAISE; END IF;END;/CREATE TABLE servicehub_redo_case( case_id NUMBER PRIMARY KEY, status_code VARCHAR2(12) NOT NULL, payload VARCHAR2(500));
SELECT n.name,m.valueFROM v$mystat m JOIN v$statname n ON n.statistic#=m.statistic#WHERE n.name IN ('redo entries','redo size','redo synch writes')ORDER BY n.name;
INSERT INTO servicehub_redo_caseSELECT LEVEL, CASE MOD(LEVEL,3) WHEN 0 THEN 'OPEN' WHEN 1 THEN 'CLOSED' ELSE 'HOLD' END, RPAD('r',200,'r')FROM dual CONNECT BY LEVEL <= 5000;COMMIT;DROP TABLE servicehub_redo_case PURGE;
Snapshot the session counters before and after the insert.
redo size includes database-change overhead; it is
not merely payload/table bytes.
7. Production judgment
Size redo from measured generation rate, switch/checkpoint
behavior, archive throughput, storage latency, and recovery
objectives. Place multiplexed members on genuinely independent
failure domains where possible. Monitor archive destinations for
space/errors. Do not change LOG_BUFFER, force
switch cadence, or add members without a failure/performance
mechanism.
No option, pack, restart, or COMPATIBLE change is
required for the observations. Lesson 2 turns to undo—the
historical block-version mechanism that enables rollback, read
consistency, transaction recovery, and Flashback Query.
Check your understanding
- What is the difference between a redo group and a member?
- Can online redo include uncommitted work?
- What normally triggers a log switch?
- When can an ARCHIVELOG group be reused?
- Why is a forced switch every minute a poor generic fix?
Review the answers
A group is the logical circular redo unit; members are identical physical copies in that group.
Yes. Redo records changes whether the transaction later commits or rolls back.
Normally the current group fills and LGWR advances to a reusable group.
After it is no longer required for instance recovery and required archiving has completed.
It creates extra switch/checkpoint/archive work and can expose reuse pressure without fixing the real I/O, commit, or sizing cause.
Authoritative references
- Physical Storage Structures — Online Redo Log — redo purpose, reuse and switch mechanics
- Log Writer Process (LGWR) — redo-buffer writes and multiplexed members
- Managing the Redo Log — groups, members and switches
- Managing Archived Redo Logs — ARCHIVELOG and destinations
- V$LOG — online redo metadata/status