Generate technical documentation from code and schemas while preserving the human-reviewed business context that automation cannot infer safely, then test for stale or missing annotations.
Automated Documentation from Code/Models vs Human Business Context and Review
Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.
Learning outcomes
Identify which metadata can be generated reliably from schemas/code and which requires human business context.
Build generated documentation that is reviewable rather than authoritative by assumption.
Detect stale annotations when code/schema versions change.
Preserve lineage and ownership when documentation is regenerated.
Avoid “documentation drift” caused by copy-pasted wiki pages.
Chapter 23 begins from the governed and secured AtlasMart state produced by earlier chapters: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current paid order line; the Chapter 20 metric contracts remain authoritative; Chapter 21 certified marts remain dependent on those contracts; and Chapter 22 security/privacy controls remain in force. Metadata and lineage describe, govern, and make change safer. They do not silently create a second business definition.
Runtime: Python 3.13.5 + SQLite 3.46.1. Mode: local, in-memory SQLite plus deterministic Python, no cloud account and no paid feature. Time: UTC. Currency: USD. Security: synthetic identities and non-secret metadata only; metadata can itself reveal sensitive topology, so production catalog access still requires policy. History: catalog entries and deprecation notices are versioned/effective-dated rather than overwritten without evidence. Non-guarantee: the lab proves graph/catalog logic for this fixture; it does not prove any vendor's automatic lineage completeness.
1. The realistic problem: the docs build is green, but the meaning is wrong
AtlasMart automatically extracts table names, column types,
nullability, and SQL dependencies. The generated page for
line_amount_usd is syntactically correct, but it
does not say that pending orders are excluded, returns are not
netted here, the grain is one paid order line, or finance owns
the metric definition. Automation has documented the shape, not
the decision semantics.
Generated documentation is metadata derived deterministically from code, schemas, manifests, or runtime events. Human context is reviewed meaning that cannot be inferred safely: business definition, rationale, exclusions, ownership, intended use, caveats, policy interpretation, migration guidance.
2. What can be generated safely?
| Metadata | Automate? | Reason |
|---|---|---|
| Column names/types/nullability | Yes | Declared in schema/contract |
| SQL/model dependencies | Often | Parse/compiler/runtime lineage can emit evidence |
| Job/run IDs/timestamps | Yes | Runtime facts |
| Metric SQL text/hash | Yes | Code artifact |
| Business purpose | No, not safely | Requires stakeholder intent |
| Correct grain meaning | Partially | Structure may hint; contract must state it |
| Why pending orders are excluded | No | Business rule/rationale |
| Owner/steward approval | No | Accountability decision |
| Legal/privacy retention interpretation | No | Organization/jurisdiction policy |
3. Controlled failure: column names become the dictionary
AUTO_DOC{"field":"line_amount_usd","nullable":false,"type":"DECIMAL(12,2)"}MISSING IF WE STOP THERE- grain: one paid order line- definition: gross paid line amount in USD; not net of returns- exclusion: pending orders- technical owner: commerce-platform- business steward: orders-steward
The repair is a layered documentation model: regenerate technical facts on every relevant code/schema change, preserve reviewed annotations under stable asset IDs, and require re-review when the technical fingerprint or semantic contract version changes.
4. Merge generated and curated metadata without overwriting either
def render_asset(generated, curated): return { "asset_id": generated["asset_id"], "schema": generated["schema"], "lineage": generated["lineage"], "code_version": generated["code_version"], "definition": curated["definition"], "grain": curated["grain"], "owner": curated["owner"], "steward": curated["steward"], "approved_against_version": curated["approved_against_version"], "review_required": curated["approved_against_version"] != generated["code_version"], }
Do not regenerate a wiki page by deleting the curated definition. Do not keep a human description forever while the underlying code version changes. Store both provenance streams and make mismatch visible.
5. Generated docs still need acceptance tests
| Test | Failure it detects |
|---|---|
| Every certified asset has non-empty definition | Names/types presented as documentation |
| Curated approval version = current contract/model version | Stale human context after code change |
| Every metric links to source/model lineage | Metric catalog disconnected from implementation |
| Every report has owner + consumer inventory | Orphan dashboards |
| Deprecated field has replacement/timeline | Destructive change without migration path |
| Security classification present for sensitive asset | Catalog exposes location without access policy |
6. Documentation freshness is not data freshness
Warehouse data may be fresh while the documentation is stale, or vice versa. Track separate timestamps/versions: source contract effective time, job/code version, metadata extraction time, human review/approval time, metric version, and report certification time. Do not collapse them into one ambiguous “last updated.”
7. Metadata fingerprint for reproducibility
METADATA_SHA2562819b6f63a8f9801d9092556d5e90b10039c0faa049ec23df48c0f734884db83
The SHA-256 proves that this exact ordered fixture metadata state can be identified reproducibly. It does not prove the metadata is correct or complete; it lets tests detect unexpected state changes and compare reruns.
8. Production judgment
Automation: maximize deterministic extraction of technical facts. Human review: focus scarce attention on meaning, policy, ownership, and exceptions. Idempotency: stable IDs prevent duplicate pages on regeneration. Observability: alert on stale approvals, missing owners, and lineage gaps. Security: documentation can expose sensitive resource names and topology. Performance/cost: do not scrape every system continuously without a change/freshness objective. Migration: generated docs should show deprecated/replacement versions side by side during transition.
9. Bridge to Lesson 5
With catalog entries, lineage, owners, and documentation in place, AtlasMart can evaluate a proposed breaking source change before code is modified. Lesson 5 runs the complete impact workflow and turns the graph into a migration decision.
Knowledge check
Acceptance questions
- Why should generated docs not overwrite curated definitions?
- What does a metadata fingerprint prove?
- Why track review version separately from extraction time?
- Which metadata is least safe to infer automatically?
Review the answers
1. Technical extraction and business meaning have different provenance/authority.
2. Exact fixture state identity, not truth/completeness.
3. Fresh extraction can still carry human context approved against an older semantic version.
4. Business purpose, exclusions/rationale, policy/legal interpretation, and accountability.
Authoritative references
- OpenLineage — Lineage Dataset FacetCurrent specification for expressing dataset/job/field dependencies; useful for portable lineage concepts without making OpenLineage a prerequisite.
- OpenLineage — Column Level Lineage Dataset FacetShows fine-grained field dependencies and distinguishes identity, transformation, aggregation, join, filter, sort, window, and conditional influence.
- W3C — Data Catalog Vocabulary (DCAT) Version 3A standard vocabulary for describing cataloged datasets/data services and their metadata; referenced as an interoperability model, not as a required implementation.
- W3C — PROV-OStable provenance vocabulary for entities, activities, and agents; useful for reasoning about lineage/provenance boundaries.
- SQLite — WITH / recursive common-table expressionsThe local lab uses a recursive CTE to traverse downstream lineage deterministically.
- Python — hashlibUsed to fingerprint the deterministic metadata state after catalog repair and change registration.
10. Lab cleanup/reset
Because the fixture is in-memory, reset by rerunning it. A production documentation generator should support dry-run/diff so regenerated technical facts can be reviewed before publication.