Chapter 13 · Data Quality Engineering: Validation, Standardization, Matching, Reconciliation, and Quarantine
Standardize Names, Codes, Units, Time Zones, Encodings, and Reference Data
Standardize AtlasMart text, codes, units, timestamps, encodings, and reference values without erasing raw evidence or confusing canonical representation with business truth.
Learning outcomes
Standardize text and Unicode representation while retaining raw values and avoiding destructive transliteration.
Map codes and units through versioned reference data and reject unproved conversions.
Require explicit timestamp offsets and normalize valid instants to UTC without guessing local time.
Explain why canonical form improves comparison but does not establish business accuracy.
Produce a repeatable canonical record with raw-evidence hashes for AtlasMart.
1. Standardization narrows representation, not meaning
AtlasMart sees " Ada Retail ",
"ADA RETAIL", region values such as
" north ", status values such as
"PAID", and item-unit aliases
EA/each. These differences can prevent
joins or inflate distinct counts.
Standardization maps multiple approved
representations to a governed canonical representation while
retaining raw evidence. It must never invent missing business
meaning.
Non-destructive rule
Keep the raw field and a canonical derivative. Never overwrite raw evidence merely because the canonical value is easier to query.
2. Names and encodings: normalize carefully
Unicode permits visually equivalent text to have different
code-point sequences. The lab normalizes to NFC and collapses
whitespace. The decomposed Béna Shop becomes
Béna Shop. For matching, a separate case-folded
token can be created. The display name should not be forced to
uppercase, stripped of accents, or transliterated as a general
rule: those transformations can destroy meaningful distinctions
across languages.
from __future__ import annotationsimport copy, hashlib, json, re, unicodedatafrom datetime import datetime, timezonedef canonical_json(obj): return json.dumps(obj, sort_keys=True, separators=(",", ":"), ensure_ascii=False)def sha(obj): return hashlib.sha256(canonical_json(obj).encode("utf-8")).hexdigest()def clean_space(value): return re.sub(r"\s+", " ", value.strip())def text_nfc(value): return unicodedata.normalize("NFC", clean_space(value))def match_token(value): return text_nfc(value).casefold()def parse_aware(value): if value.endswith("Z"): value = value[:-1] + "+00:00" dt = datetime.fromisoformat(value) if dt.tzinfo is None: raise ValueError("timestamp has no UTC offset") return dt.astimezone(timezone.utc).isoformat().replace("+00:00", "Z")raw_name = "Be\u0301na Shop"canonical = text_nfc(raw_name)match_key = match_token(raw_name)print(canonical) # Béna Shopprint(match_key) # béna shopassert canonical == "Béna Shop"
3. Codes and reference data need versions and owners
| Raw value | Reference domain | Canonical result | Decision |
|---|---|---|---|
| PAID | order_status v1 | paid | accepted alias |
| WEB | channel v1 | web | accepted alias |
| EA | unit v1 | item | accepted alias |
| CASE | unit v1 | — | quarantine: conversion factor absent |
| north | region v1 | North | accepted code normalization |
A reference-data table is part of lineage. If a steward later
adds CASE → 12 item, that is not a cosmetic update:
it can change quantities and historical measures. Version it,
test it, and decide whether prior quarantined rows may be
reprocessed.
4. Time zones: reject ambiguity, then normalize the instant
2026-09-24T10:00:00+00:00 and
2026-09-24T10:00:00Z identify the same UTC instant.
2026-09-24 12:00 does not state an offset or named
zone. The lab rejects it. Assuming the server’s local time would
make results machine-dependent and could shift dates around
daylight-saving transitions.
from __future__ import annotationsimport copy, hashlib, json, re, unicodedatafrom datetime import datetime, timezonedef canonical_json(obj): return json.dumps(obj, sort_keys=True, separators=(",", ":"), ensure_ascii=False)def sha(obj): return hashlib.sha256(canonical_json(obj).encode("utf-8")).hexdigest()def clean_space(value): return re.sub(r"\s+", " ", value.strip())def text_nfc(value): return unicodedata.normalize("NFC", clean_space(value))def match_token(value): return text_nfc(value).casefold()def parse_aware(value): if value.endswith("Z"): value = value[:-1] + "+00:00" dt = datetime.fromisoformat(value) if dt.tzinfo is None: raise ValueError("timestamp has no UTC offset") return dt.astimezone(timezone.utc).isoformat().replace("+00:00", "Z")print(parse_aware("2026-09-24T10:00:00+00:00"))try: parse_aware("2026-09-24 12:00")except ValueError as exc: print("QUARANTINE:", exc)
5. End-to-end canonicalization evidence
| Raw field | Canonical field | Why separate? |
|---|---|---|
| name | name_canonical | Preserves original spelling/spacing while giving users a stable display representation. |
| name | name_match | Case-folded search/matching token is not suitable as the authoritative display value. |
| email_canonical | Synthetic lab uses lowercased trimmed email as deterministic identity evidence. | |
| order_ts | order_ts_utc | Canonical instant supports comparisons; raw timestamp remains available for audit. |
| unit | unit | Only approved reference mappings are accepted; unknown CASE is not coerced. |
| raw payload | raw_hash | Fingerprint supports evidence identity and rerun tracing, not semantic truth. |
6. Controlled failure: defaults and “helpful” coercion
A common shortcut replaces an unparseable quantity with zero, a missing region with Unknown, or CASE with item. These choices make the pipeline green but change business facts. Unknown members are appropriate only when the model’s policy says the business value is genuinely unknown; they are not a license to convert invalid source evidence into valid facts. In this lab, unresolved required semantics are quarantined with reason codes.
Boundary
Standardization may change representation only where a contract authorizes equivalence. Conversion that changes measurement meaning requires an explicit, versioned business rule.
7. Lab verification and reset
- Q-C007 normalizes from decomposed Unicode to NFC Béna Shop.
- QO1 status/channel/unit normalize to paid/web/item.
- QO2 remains quarantined because its time zone is not stated.
- QO5 remains quarantined because CASE has no conversion factor.
- Delete generated local artifacts to reset; raw fixtures remain reproducible in the lesson script.
8. Production judgment and bridge
Canonical fields simplify joins and rules, but they increase governance surface: reference maps, Unicode policy, time-zone interpretation, and unit conversion all need ownership. Observability should count each normalization path so a sudden rise in aliases or rejects is visible. Lesson 3 now separates true duplicate delivery from uncertain entity resolution.
Knowledge check
Check your understanding
- Why keep raw and canonical values together?
- Is lowercasing every human name a good canonical display strategy?
- Why is CASE quarantined in the lab?
- Why reject a parseable timestamp with no offset?
- Does Unicode normalization prove the name is accurate?
Review the answers
1. The raw value preserves evidence while the canonical value makes approved equivalences repeatable and queryable.
2. No. Case folding may support matching, but display values can carry language and identity semantics that destructive normalization would lose.
3. The contract lacks a governed conversion factor from CASE to item, so any quantity conversion would be invented.
4. Its instant is ambiguous; using machine-local time would make the result environment-dependent.
5. No. It only standardizes representation; accuracy needs trusted external evidence.
Authoritative references
- ISO/IEC 25012 — Data quality modelOfficial ISO catalogue entry for a general data-quality model; use it as background, not as a substitute for AtlasMart-specific acceptance rules.
- Unicode Standard Annex #15 — Unicode Normalization FormsAuthoritative normalization guidance used to explain NFC normalization without transliterating or overwriting the raw value.
- RFC 3339 — Date and Time on the InternetAuthoritative timestamp syntax reference supporting the rule that source timestamps carry an explicit offset before normalization to UTC.
- IANA — Time Zone DatabaseAuthoritative source for named civil time-zone identifiers when business semantics require a region-based zone rather than a fixed offset.
- Python documentation — unicodedataStandard-library Unicode database access used by the free/local lab.
- Python documentation — hashlibStandard-library hashing used for evidence fingerprints; hashes demonstrate identity of serialized evidence, not semantic correctness.