Chapter 13 · Data Quality Engineering: Validation, Standardization, Matching, Reconciliation, and Quarantine
Deduplication and Entity Matching: Deterministic/Probabilistic Concepts and Stewardship Boundaries
Separate deterministic deduplication from uncertain entity matching, retain match evidence, and define when automation must stop and a data steward must decide.
Learning outcomes
Distinguish duplicate records from distinct records that may refer to the same real entity.
Apply deterministic matching only when governed strong evidence is present.
Use probabilistic/similarity scores as evidence with calibrated, dataset-specific decision boundaries—not universal truth.
Preserve match evidence, survivor rules, lineage, and steward decisions.
Explain false-positive and false-negative costs before automating an entity merge.
Use only the lesson’s synthetic/local fixture for destructive setup, reset, correction, backfill, or cleanup steps. Record the stated runtime/version and assumptions, preserve the AtlasMart control totals, and never point cleanup commands at production data or credentials.
1. Duplicate delivery and entity matching solve different problems
QO1/1 is delivered twice under the same business event key. That is a record duplicate: the pipeline can deterministically keep one accepted fact and retain duplicate evidence. Q-C001 and Q-C001-DUP have different source customer IDs but the same normalized synthetic email. That is an entity-resolution problem: two source records may represent one business entity. Treating these as the same mechanism can accidentally merge real people or double-count events.
2. Deterministic rules need explicit evidence
For this synthetic lab only, exact normalized email is declared a strong matching key. Both Q-C001 rows map to one durable DQ customer ID. This is not a universal identity rule: shared email addresses, recycled addresses, privacy rules, and source-system semantics can invalidate it in production.
| Source ID | Normalized email | Durable ID | Evidence / action |
|---|---|---|---|
| Q-C001 | ada.retail@example.invalid | DQ-fcbf05dc86 | exact normalized email |
| Q-C001-DUP | ada.retail@example.invalid | DQ-fcbf05dc86 | exact normalized email |
3. Probabilistic concepts: score evidence, not reality
When strong identifiers are absent, matching systems may combine similarities across names, addresses, dates, phones, or other features. A probabilistic model estimates how evidence behaves for matching versus non-matching pairs; the score must be calibrated and validated on representative labeled data. This lesson uses a simple string-similarity score only to show workflow mechanics, not to claim probabilistic rigor.
| Candidate left | Candidate right | Lab similarity | Decision |
|---|---|---|---|
| Ben Home | Ben Homes | 0.941 | steward_review |
The pair is queued for stewardship. A threshold that works for one population, language, or error distribution can be unsafe elsewhere, so the chapter deliberately avoids a universal auto-merge cutoff.
4. Controlled failure: fuzzy match everything above 0.9
A developer sees “Ben Home” and “Ben Homes” score highly and auto-merges them. The result may be a false positive that combines separate households, corrupts history, and creates privacy exposure. The safe repair is to define features, labeled evaluation data, costs of false merge versus missed merge, confidence bands, and a steward queue for uncertain cases. Automatic merges must be reproducible and reversible through a match ledger.
from difflib import SequenceMatcherleft, right = "ben home", "ben homes"score = SequenceMatcher(None, left, right).ratio()record = { "left": left, "right": right, "score": round(score, 3), "decision": "steward_review", "reason": "name similarity alone is insufficient identity evidence",}print(record)assert record["decision"] == "steward_review"
5. Survivor selection is not identity truth
After two source rows are mapped to one durable entity, a survivorship rule chooses which attributes feed a current mastered representation: perhaps most recent verified phone, highest-authority source for legal name, or non-null value from a steward-approved system. That rule must be field-specific and versioned. “Take the newest row” can overwrite a more authoritative value with a lower-quality feed.
| Ledger field | Purpose |
|---|---|
| left_source_id / right_source_id | Reconstruct which records were compared. |
| features / normalized values | Explain what evidence was considered. |
| model/rule version | Make the decision reproducible after logic changes. |
| score / rule result | Quantify evidence under that version. |
| decision + actor | Distinguish automatic rule from human stewardship. |
| effective/rollback metadata | Support correction if the merge was wrong. |
6. Security and fairness boundary
Identity data can be sensitive. Production match features should be purpose-limited, access-controlled, retained only as authorized, and tested for population-specific error. Do not infer protected or sensitive attributes to “improve” matching. This synthetic lab contains no real personal information.
7. Lab verification and bridge
- QO1/1 duplicate delivery collapses to one accepted business event.
- Q-C001 and Q-C001-DUP map deterministically only because this lab explicitly governs normalized email as a strong key.
- The Ben Home/Homes pair remains unmerged and enters steward review.
- Every accepted/duplicate/match decision retains evidence and rule version.
Lesson 4 applies the same evidence discipline to invalid records: quarantine first, then authorized repair and reprocess.
Knowledge check
Check your understanding
- What is the difference between deduplication and entity resolution?
- Why is exact email matching not universally safe?
- What should happen to uncertain match candidates?
- Why is a similarity threshold dataset-specific?
- What makes a merge reversible?
Review the answers
1. Deduplication removes repeated representations of the same record/event key; entity resolution decides whether distinct source identities refer to the same real entity.
2. Emails can be shared, recycled, missing, mistyped, or governed differently across systems and jurisdictions.
3. Retain evidence and route them to a defined review/decision process rather than forcing an automated merge.
4. Feature distributions and false-positive/false-negative costs differ by population, language, source, and business use.
5. A match ledger with source identities, evidence, rule/model version, decision actor, and rollback/effective metadata.
Summary and next step
This lesson established the mechanism and production boundaries for Deduplication and Entity Matching: Deterministic/Probabilistic Concepts and Stewardship Boundaries while preserving AtlasMart’s declared grain, governed metrics, history, and reconciliation evidence. Continue to Quarantine Bad Records, Preserve Raw Evidence, Repair/Reprocess, and Avoid Silent Coercion with those contracts unchanged unless an explicit, tested migration says otherwise.
Authoritative references
- ISO/IEC 25012 — Data quality modelOfficial ISO catalogue entry for a general data-quality model; use it as background, not as a substitute for AtlasMart-specific acceptance rules.
- Unicode Standard Annex #15 — Unicode Normalization FormsAuthoritative normalization guidance used to explain NFC normalization without transliterating or overwriting the raw value.
- RFC 3339 — Date and Time on the InternetAuthoritative timestamp syntax reference supporting the rule that source timestamps carry an explicit offset before normalization to UTC.
- IANA — Time Zone DatabaseAuthoritative source for named civil time-zone identifiers when business semantics require a region-based zone rather than a fixed offset.
- Python documentation — unicodedataStandard-library Unicode database access used by the free/local lab.
- Python documentation — hashlibStandard-library hashing used for evidence fingerprints; hashes demonstrate identity of serialized evidence, not semantic correctness.