Chapter 01 · Data Warehouse Foundations: OLTP vs OLAP, Analytical Workloads, Architecture, and Lab Dataset
Batch, Micro-Batch, Streaming/CDC, and Near-Real-Time Warehousing: Freshness vs Complexity Tradeoffs
Define batch, micro-batch, streaming, CDC, and near-real-time by observable freshness and recovery semantics, including duplicate and late-event failure injection.
Learning outcomes
AtlasMart’s executives ask for “real-time sales,” while the ERP team can expose either 15-minute extracts or a change log. Those are not the same requirement. This lesson defines batch, micro-batch, streaming, CDC, and near-real-time by mechanics and measurable freshness, then injects duplicate and late events to show why faster transport does not automatically produce correct analytical state.
Distinguish batch, micro-batch, streaming processing, and change data capture (CDC).
Define near-real-time as a measurable service objective rather than a specific tool.
Separate event time, source commit time, capture time, processing time, and consumer availability.
Explain duplicate delivery, out-of-order changes, restart, and watermark/checkpoint concerns at a conceptual level.
Run a deterministic local CDC-envelope simulation that proves naive arrival-order application can regress state.
The lab simulates CDC envelopes with local Python data. It does not emulate a database transaction log, Kafka, Debezium, or a managed streaming service. Those technologies appear in later courses. Here the evidence is about ordering, deduplication, and freshness semantics.
1. Batch and micro-batch are schedule shapes
Batch processing collects a bounded set of data and processes it as a unit: for example, yesterday’s completed orders at 02:00 UTC. Micro-batch shortens the interval—perhaps every 5 or 15 minutes—but still processes bounded chunks. The distinction is operational, not a universal numeric threshold.
A five-minute schedule does not guarantee five-minute freshness. Source extraction may lag, a job may queue, downstream transformations may take longer, and a BI cache may refresh later.
2. Streaming is continuous processing; CDC is a change-capture mechanism
Streaming describes continuous or near-continuous processing of an unbounded event sequence. Change data capture (CDC) describes how inserts, updates, and deletes are observed from a source—often from a database log, sometimes from triggers or query-based mechanisms. CDC records can be processed continuously, in micro-batches, or even archived for later batch work.
Therefore “we use CDC” does not specify latency, and “we use streaming” does not specify source correctness. The pipeline must still define transaction boundaries, ordering, duplicates, delete semantics, schema evolution, checkpoints, and recovery.
3. Near-real-time is a contract
| AtlasMart path | Fictional target | Why |
|---|---|---|
| ERP orders → certified sales | 95% within 30 minutes of source commit | Operations wants same-shift visibility without a streaming platform requirement. |
| Inventory snapshot → certified inventory | Within 45 minutes of source snapshot | Stock planning tolerates bounded lag. |
| CRM customer profile → certified customer | By 03:00 UTC after 02:00 source extract | CRM only publishes a daily file in the case study. |
| Executive dashboard | Within 60 minutes of certified sales publication | BI refresh adds its own serving delay. |
These are course-case-study requirements, not generic recommendations. If the business later needs fraud decisions within two seconds, the architecture, measurement points, and failure response must be redesigned.
4. The timestamp chain makes hidden latency visible
For a source change, keep at least the timestamps needed to explain latency: event/business time, source commit position/time, capture time, processing/load time, and consumer availability. An event can be old but processed quickly, or recent but delayed in capture. Those are different failures.
When measuring freshness, also state the population. “95% within 30 minutes” requires a denominator, a time window, and treatment for failed or deliberately excluded records.
5. Controlled failure: apply changes in arrival order
Networks retry. Connectors restart. A snapshot and a log stream can overlap. If AtlasMart simply applies every change in the order it arrives, a late older update can overwrite a newer state. If it increments metrics for every delivery, a duplicate can double-count.
The repair depends on the source contract: stable event IDs or source positions for deduplication, deterministic per-key ordering, idempotent writes, checkpoints advanced only after durable processing, and replay tests. Claiming “exactly once” end to end requires evidence across every boundary; this course instead proves the concrete deduplication/restart behavior of each design.
6. Hands-on lab — inject duplicate and late CDC envelopes
Save the script below as cdc_simulation.py and run
it with Python. It uses only in-memory synthetic events.
from datetime import datetime, timezonedef ts(s): return datetime.fromisoformat(s.replace("Z", "+00:00"))events = [ {"event_id":"e100","seq":100,"order_id":"O2001","status":"created", "source_commit":"2026-09-20T10:00:00Z","available":"2026-09-20T10:00:18Z"}, {"event_id":"e101","seq":101,"order_id":"O2001","status":"paid", "source_commit":"2026-09-20T10:02:00Z","available":"2026-09-20T10:02:14Z"}, {"event_id":"e101","seq":101,"order_id":"O2001","status":"paid", "source_commit":"2026-09-20T10:02:00Z","available":"2026-09-20T10:02:16Z"}, # retry {"event_id":"e102","seq":102,"order_id":"O2001","status":"shipped", "source_commit":"2026-09-20T10:04:30Z","available":"2026-09-20T10:06:00Z"}, {"event_id":"e099","seq":99,"order_id":"O2001","status":"created", "source_commit":"2026-09-20T09:59:30Z","available":"2026-09-20T10:06:20Z"} # late old change]naive = Nonefor e in events: naive = e["status"]print("naive final status:", naive)seen = set()state = {}for e in events: if e["event_id"] in seen: continue seen.add(e["event_id"]) old = state.get(e["order_id"]) if old is None or e["seq"] > old["seq"]: state[e["order_id"]] = {"seq": e["seq"], "status": e["status"]}print("guarded final state:", state["O2001"])for e in events: lag = int((ts(e["available"]) - ts(e["source_commit"])).total_seconds()) print(e["event_id"], "freshness_seconds=", lag)
Expected evidence begins with
naive final status: created, proving that arrival
order regressed the state. The guarded state must be
{'seq': 102, 'status': 'shipped'}. The duplicate
e101 is ignored, and the late sequence 99 event
cannot replace sequence 102.
The freshness lags are 18, 14, 16, 90, and 410 seconds for the five delivered envelopes. The last event is dramatically stale even though it was processed immediately on arrival. This illustrates why event age and processing speed are separate metrics.
Verification: rerun the script and confirm
identical output; move the late event earlier in the Python list
and confirm sequence guarding still yields 102/shipped; remove
the guard and observe the regression.
Cleanup: delete only
cdc_simulation.py.
7. Batch, micro-batch, and streaming are trade spaces
| Question | Batch | Micro-batch | Continuous/streaming |
|---|---|---|---|
| Operational complexity | Usually lowest | Moderate | Usually highest because state/recovery are continuous |
| Typical latency envelope | Longer, scheduled | Shorter bounded intervals | Potentially seconds/sub-seconds, but not guaranteed |
| Replay reasoning | Bounded partitions/batches | Bounded mini-partitions | Offsets/checkpoints/state plus replay |
| Cost behavior | Burst compute | More frequent startup/compute | Long-lived or continuously billed resources depending on platform |
| Best fit | Freshness allows scheduled delivery | Minutes-level need without full streaming complexity | Requirements truly need continuous low-latency updates |
No row in this table is a universal winner. AtlasMart should start with the simplest mechanism that meets measured freshness and recovery requirements, then change when evidence shows it does not.
8. Production judgment
Do not buy a streaming stack to satisfy an undefined “real-time” request. Define start/end timestamps, percentile or maximum lag, business consequence of staleness, acceptable data loss/replay, and operator response. Then choose batch, micro-batch, CDC, or continuous processing that satisfies the contract with tolerable complexity.
The final lesson in Chapter 01 turns these decisions into the AtlasMart case-study contract: sources, processes, KPIs, SLAs/SLOs, security, and quality expectations.
Knowledge check
Check your understanding
- Why does CDC not imply streaming?
- Why can a five-minute micro-batch schedule miss a five-minute freshness target?
- What failure did the late sequence-99 event cause in the naive consumer?
- What evidence is required before claiming exactly-once behavior?
- Why is a 410-second-old event different from a slow processor?
Review the answers
1. CDC defines how source changes are captured; those changes can be consumed continuously, in micro-batches, or later in batch.
2. Extraction, queueing, transformation, commit, and BI refresh all add lag beyond the schedule interval.
3. Arrival-order application overwrote the newer shipped state with an older created state.
4. Every relevant boundary—capture, transport, processing, storage, retries, restart, and downstream effects—must prove its deduplication/atomicity contract; a tool label alone is insufficient.
5. The processor may handle the event quickly after arrival, but the source change was already stale before it reached the consumer; capture/transport lag caused the age.
Summary and next step
Batch, micro-batch, streaming, and CDC are mechanisms. Near-real-time is a measurable requirement. Preserve source order/identity, test duplicate and late delivery, and measure freshness across named timestamps.
Next: Define a Course Case Study with Source Systems, Business Processes, KPIs, SLAs, Security, and Data Quality Expectations.
Authoritative references
- Debezium documentation — Official CDC project documentation used as a later-course implementation reference; not required for this simulation.
- Big Data Academy — course curriculum — The current warehouse course scope and CDC/incremental-loading progression.
- Kimball Group — Dimensional Modeling Techniques — Business-process/history context that downstream CDC loads must preserve.
- Python documentation — datetime — Runtime used to calculate deterministic freshness intervals in the lab.