Chapter 14 · ETL vs ELT Architecture: Staging, Raw, Integration, Presentation, and Transform Ownership

Transformation Modularity, Reusable Intermediate Models, Dependency Graphs, and Environment Promotion

Decompose transformations into reusable contracts with an acyclic dependency graph and environment-neutral configuration so the same logic can be promoted without hard-coded dev/test/prod identifiers.

Intermediate → Advanced120–140 minutesModularity + promotion labPython 3 stdlib + sqlite3 · local/syntheticLast reviewed: September 2026

Learning outcomes

AtlasMart now has correct layer boundaries, but a single 900-line procedure can still make them operationally inseparable. If customer resolution, revision selection, quality handling, fact construction, and daily aggregation are all hidden in one procedure, a presentation change can force a risky full rebuild and failures are difficult to localize.

01

Split transformations at stable semantic boundaries rather than building one monolithic procedure.

02

Model transform dependencies as a directed acyclic graph and distinguish logical dependency from scheduler implementation.

03

Choose reusable intermediate models from semantic reuse and failure isolation, not from convenience alone.

04

Promote the same transform logic across environments through configuration rather than hard-coded database/path identifiers.

05

Explain how modularity improves testing and rollback without implying that more models are always better.

Chapter 14 continuity contract

Chapter 14 preserves the accepted AtlasMart production state established through Chapters 11–13: eight current paid order-line facts, five paid orders, ten units, 690 USD paid GMV, 425 USD cost-at-sale, 265 USD gross profit, and the governed inventory snapshot of 137 units. Chapter 13's Q-prefixed malformed training batch remains isolated quality-test evidence. This chapter changes where transformations execute and which intermediate states are persisted; it does not silently redefine sales grain, correction history, metric formulas, or customer identity.

Execution and scope boundary

The mandatory lab is synthetic, local, and free. It uses Python 3 standard library plus its bundled sqlite3 module. Generation-time validation ran with Python 3.13.5 and SQLite 3.46.1; learners should record their own python --version and sqlite3.sqlite_version because behavior and optimizer details can vary. The lab demonstrates layering, replay, lineage, and deterministic controls—not cloud pricing, distributed exactly-once guarantees, production durability, or vendor-specific warehouse performance.

1. Modularity starts at semantic contracts

A module is useful when it has a stable input grain, output grain, owner, tests, and reuse/failure boundary. Splitting every SELECT into a separate persisted table is not modularity; it is fragmentation. In the Chapter 14 lab, the reusable boundary is int_sales_current: one current governed revision per order line. Daily presentation models can depend on it without reimplementing correction selection.

Node Input contract Output grain Persist?
raw load one source revision record per CSV row same source revision grain yes as loaded audit table inside run DB
stg_sales raw loaded rows same grain; normalized status/types TEMP
int_sales_current staged revision rows one current paid row per (order_id,line_no) yes
mart_sales_daily current sales rows one row per UTC sales date yes

2. Dependency graph before scheduler

logical_dag.json
{  "raw_sales_loaded": [],  "stg_sales": ["raw_sales_loaded"],  "int_sales_current": ["stg_sales"],  "mart_sales_daily": ["int_sales_current"]}

This is a logical DAG: dependencies are acyclic and express data requirements. Airflow, dbt, a shell script, or a managed scheduler may later execute it, but those tools are not the dependency semantics themselves. Chapter 16 will handle orchestration, retries, backfills, and operational scheduling in depth.

3. Environment promotion: move code, parameterize location

Hard-coding prod.analytics.fact_sales into development SQL makes promotion a text-rewrite exercise. Instead, keep business logic stable and inject environment-specific database/path/credential settings outside the transform. The exact mechanism varies by platform; the principle is that environment identity is deployment configuration, not business meaning.

environment_config.py
from dataclasses import dataclassfrom pathlib import Path@dataclass(frozen=True)class EnvConfig:    name: str    root: Path    database: Pathcfg = EnvConfig(    name=os.environ.get("ATLASMART_ENV", "dev"),    root=Path(os.environ.get("ATLASMART_LAB_HOME", "atlasmart_ch14_lab")),    database=Path(os.environ.get("ATLASMART_DB", "atlasmart_ch14_lab/warehouse.db")),)# SQL still says int_sales_current and mart_sales_daily;# deployment decides where the database lives.

4. Controlled failure: monolith + environment literals

A monolithic script that reads a production path, embeds credentials, performs quality coercions, selects current revisions, aggregates revenue, and publishes a dashboard table has multiple independent reasons to change. A small presentation change then redeploys identity and correction logic. Hard-coded environment names also make test and production logic diverge through manual edits.

The repair is not “create fifty models.” It is to isolate durable semantic boundaries, parameterize deployment context, test nodes independently, and preserve end-to-end reconciliation. If an intermediate has one consumer, is cheap to recompute, and provides no recovery value, keep it ephemeral.

5. Failure injection and rollback reasoning

Inject a syntax error into the presentation transform: raw and integration evidence should remain intact, and rerunning presentation should restore the same checksum. Inject a different raw file under the existing batch ID: promotion should fail before semantic models change. These two failures require different rollback surfaces because one is transform failure and the other is evidence-identity violation.

Promotion should compare schema/contracts, tests, control totals, and transform versions before switching consumers. This chapter does not claim one universal blue/green or transactional deployment technique; those are engine/platform dependent.

Knowledge check

Check your understanding

  1. What makes an intermediate model worth naming as a module?
  2. Is a DAG the same as Airflow?
  3. Why parameterize environment identifiers?
  4. Does more modularity always mean more persisted tables?
  5. What should survive a presentation-only failure?
Review the answers

1. A stable input/output grain, semantic responsibility, owner/tests, and meaningful reuse or failure boundary.

2. No. A DAG is the dependency graph; Airflow is one possible orchestration implementation.

3. So the same business logic can be promoted without manual text edits that cause drift.

4. No. Ephemeral modules can be appropriate when recomputation is cheap and no recovery/reuse contract needs persistence.

5. Immutable raw evidence, integration state, lineage/batch metadata, and prior certified presentation until replacement passes.

Summary and next step

This lesson established the mechanism and production boundaries for Transformation Modularity, Reusable Intermediate Models, Dependency Graphs, and Environment Promotion while preserving AtlasMart’s declared grain, governed metrics, history, and reconciliation evidence. Continue to Design a Layered Pipeline for the Case Study and Justify Every Persisted Intermediate Dataset with those contracts unchanged unless an explicit, tested migration says otherwise.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.