Build an AtlasMart metadata system that separates machine-readable technical facts from business meaning, connects dictionaries, schema contracts, metric definitions, and glossary terms, and rejects documentation that is merely a list of column names.

Technical vs Business Metadata, Data Dictionary, Schema Registry, Metric Catalog, and Glossary

Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.

Intermediate → Advanced150–190 minutesMetadata catalog / glossary labDictionary + schema + metric + glossary checksLast reviewed: September 2026

Learning outcomes

01

Separate technical metadata from business metadata and explain why neither is sufficient alone.

02

Distinguish a data dictionary, schema registry, metric catalog, and business glossary by mechanism and consumer.

03

Bind AtlasMart fields and governed metrics to owners, definitions, versions, and lineage.

04

Detect the failure mode where auto-generated names/types are mistaken for complete documentation.

05

Preserve warehouse controls while adding metadata evidence.

Continuity: metadata does not redefine the warehouse

Chapter 23 begins from the governed and secured AtlasMart state produced by earlier chapters: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current paid order line; the Chapter 20 metric contracts remain authoritative; Chapter 21 certified marts remain dependent on those contracts; and Chapter 22 security/privacy controls remain in force. Metadata and lineage describe, govern, and make change safer. They do not silently create a second business definition.

Executed local lab contract

Runtime: Python 3.13.5 + SQLite 3.46.1. Mode: local, in-memory SQLite plus deterministic Python, no cloud account and no paid feature. Time: UTC. Currency: USD. Security: synthetic identities and non-secret metadata only; metadata can itself reveal sensitive topology, so production catalog access still requires policy. History: catalog entries and deprecation notices are versioned/effective-dated rather than overwritten without evidence. Non-guarantee: the lab proves graph/catalog logic for this fixture; it does not prove any vendor's automatic lineage completeness.

1. The realistic problem: three teams see the same column and infer three meanings

AtlasMart exposes fact_sales.line_amount_usd. A finance analyst treats it as recognized revenue; marketing calls it campaign-attributed revenue; an engineer assumes it includes pending orders because the name says only “amount.” The table is technically valid, yet the business interpretation is already drifting. Metadata is data about data. Technical metadata describes structures and execution facts such as type, nullability, location, schema version, job, and lineage. Business metadata describes meaning, ownership, approved use, exclusions, quality expectations, and lifecycle.

The decision is whether a consumer can use a field/metric without reverse-engineering semantics from SQL. The source event is one ERP paid order line. The declared warehouse grain is one current paid order line. The failure mode is not a broken SQL parser; it is a valid query that answers the wrong business question.

2. Dictionary, schema registry, metric catalog, and glossary are different contracts

Construct Primary question AtlasMart example Failure if missing
Data dictionary What does this field/table mean here? fact_sales.line_amount_usd: additive paid line amount, USD, one paid line grain Column names become the documentation
Schema registry What structural contract/version may producers and consumers exchange? erp.order_lines.v3: field names, types, nullability, status Breaking producer change arrives without compatibility decision
Metric catalog How is a governed analytical result computed and versioned? gross_revenue_usd.v1: paid current lines, UTC order-date window Dashboard authors reimplement filters independently
Business glossary What business concept do people mean independent of one table? Gross revenue: governed business term linked to metric version Teams use the same phrase for incompatible calculations

A single platform may store all four, but the logical responsibilities remain distinct. A schema registry can tell you a field is DECIMAL(12,2); it cannot safely infer whether that number is gross, net, booked, refunded, tax-inclusive, or campaign-attributed.

3. Observable catalog state

Executed fixture: catalog inventory
consumer       3contract       1glossary       1job            1mart           3measure        1metric         3model_column   2raw_column     1report         3source_column  1

The counts prove which kinds of assets were registered in this fixture. They do not prove the metadata is semantically complete, current, or correct. Completeness requires mandatory fields, lineage coverage, owner/steward gates, consumer inventory, and review.

4. Controlled failure: generated schema facts are mistaken for documentation

Machine-extracted facts versus human-reviewed context
AUTO_DOC  {"field":"line_amount_usd","nullable":false,"type":"DECIMAL(12,2)"}HUMAN_CONTEXT {"definition":"Gross paid line amount in USD; not net of returns",               "grain":"one paid order line",               "exclusions":["pending orders"],               "owner":"commerce-platform",               "steward":"orders-steward"}

The automatic record is useful and reproducible. It is also insufficient. It cannot infer whether returns reduce the measure, which order states qualify, whether restatements are allowed, or who can approve a semantic change. The repair is not to reject automation; it is to merge machine-derived evidence with reviewable human context and make missing human context observable.

5. A minimal metadata schema that can be tested

Executed SQLite schema excerpt
CREATE TABLE catalog_asset (  asset_id TEXT PRIMARY KEY,  asset_type TEXT NOT NULL,  technical_owner TEXT,  business_steward TEXT,  certification TEXT NOT NULL,  sla_minutes INTEGER,  definition TEXT,  version TEXT,  updated_at TEXT NOT NULL);CREATE TABLE lineage_edge (  upstream_id TEXT NOT NULL,  downstream_id TEXT NOT NULL,  edge_kind TEXT NOT NULL,  transform_note TEXT,  PRIMARY KEY (upstream_id, downstream_id, edge_kind));

This is deliberately small. Production catalogs can carry tags, classifications, owners, contacts, terms, contracts, incidents, quality scores, query popularity, and more. Persist a field because it supports a decision or control, not because a catalog diagram has a box for it.

6. Business term → metric → warehouse lineage

Layer Asset Meaning
Glossary glossary.gross_revenue Human business term
Metric catalog metric.gross_revenue_usd.v1 Governed formula/filter/time/currency contract
Semantic measure measure.paid_line_amount_usd Reusable additive input
Warehouse column fact_sales.line_amount_usd Atomic accepted value at paid-line grain
Raw/source raw.erp_order_lines.line_amount_usd → ERP source field Evidence and producer contract

The arrows are not decorative. They let a consumer move from a business term to a concrete metric implementation and then back to source evidence.

7. Correctness, freshness, replay, and governance

Correctness: metadata is evidence about the data system, not proof that the business definition is true. Freshness: stale lineage can be worse than missing lineage because it creates false confidence. Replay/idempotency: catalog extraction should upsert stable asset identities/versioned observations rather than duplicate rows on every scan. Quality/reconciliation: metadata changes must not alter the 820/495/325 warehouse controls. Security: catalogs can reveal PII locations, service identities, and infrastructure topology; least privilege still applies. Performance/cost: no universal catalog size/scan interval is asserted. Migration: version/deprecate metadata contracts before destructive source changes.

8. Verification checklist

  1. Confirm the warehouse controls remain 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit.
  2. Confirm every certified metric/mart/report has a business definition and owner/steward.
  3. Confirm schema entries record type/nullability without pretending those fields encode business meaning.
  4. Confirm gross_revenue_usd.v1 links to the glossary term and lineage to source evidence.
  5. Confirm catalog extraction can be rerun without creating duplicate asset identities.

9. Bridge to Lesson 2

A catalog becomes operationally useful when dependency edges can answer change questions. Lesson 2 makes lineage traversable from source field through jobs/models/metrics to reports and named consumers.

Knowledge check

Acceptance questions

  1. Why can a schema registry not replace a metric catalog?
  2. What business fact is missing from DECIMAL(12,2)?
  3. Why should metadata have freshness/ownership controls?
  4. Does a catalog entry prove metric accuracy?
Review the answers

1. Structural compatibility and governed analytical meaning are different contracts.

2. Population, grain, currency/unit meaning, exclusions, timing, and ownership.

3. Stale or ownerless metadata cannot be trusted during change/incident decisions.

4. No; it documents the contract/evidence path, while tests and reconciliation establish observed correctness.

Authoritative references

10. Lab cleanup/reset

The chapter lab uses an in-memory SQLite database. Close the process and rerun the embedded script to rebuild the catalog from deterministic fixtures. No repository, cloud catalog, or production warehouse is modified.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.