Chapter 13 · Data Quality Engineering: Validation, Standardization, Matching, Reconciliation, and Quarantine

Deduplication and Entity Matching: Deterministic/Probabilistic Concepts and Stewardship Boundaries

Separate deterministic deduplication from uncertain entity matching, retain match evidence, and define when automation must stop and a data steward must decide.

Intermediate → Advanced120–140 minutesMatching/stewardship labPython 3 stdlib · local/syntheticLast reviewed: September 2026

Learning outcomes

01

Distinguish duplicate records from distinct records that may refer to the same real entity.

02

Apply deterministic matching only when governed strong evidence is present.

03

Use probabilistic/similarity scores as evidence with calibrated, dataset-specific decision boundaries—not universal truth.

04

Preserve match evidence, survivor rules, lineage, and steward decisions.

05

Explain false-positive and false-negative costs before automating an entity merge.

Execution and safety note

Use only the lesson’s synthetic/local fixture for destructive setup, reset, correction, backfill, or cleanup steps. Record the stated runtime/version and assumptions, preserve the AtlasMart control totals, and never point cleanup commands at production data or credentials.

1. Duplicate delivery and entity matching solve different problems

QO1/1 is delivered twice under the same business event key. That is a record duplicate: the pipeline can deterministically keep one accepted fact and retain duplicate evidence. Q-C001 and Q-C001-DUP have different source customer IDs but the same normalized synthetic email. That is an entity-resolution problem: two source records may represent one business entity. Treating these as the same mechanism can accidentally merge real people or double-count events.

2. Deterministic rules need explicit evidence

For this synthetic lab only, exact normalized email is declared a strong matching key. Both Q-C001 rows map to one durable DQ customer ID. This is not a universal identity rule: shared email addresses, recycled addresses, privacy rules, and source-system semantics can invalidate it in production.

Source ID Normalized email Durable ID Evidence / action
Q-C001 ada.retail@example.invalid DQ-fcbf05dc86 exact normalized email
Q-C001-DUP ada.retail@example.invalid DQ-fcbf05dc86 exact normalized email

3. Probabilistic concepts: score evidence, not reality

When strong identifiers are absent, matching systems may combine similarities across names, addresses, dates, phones, or other features. A probabilistic model estimates how evidence behaves for matching versus non-matching pairs; the score must be calibrated and validated on representative labeled data. This lesson uses a simple string-similarity score only to show workflow mechanics, not to claim probabilistic rigor.

Candidate left Candidate right Lab similarity Decision
Ben Home Ben Homes 0.941 steward_review

The pair is queued for stewardship. A threshold that works for one population, language, or error distribution can be unsafe elsewhere, so the chapter deliberately avoids a universal auto-merge cutoff.

4. Controlled failure: fuzzy match everything above 0.9

A developer sees “Ben Home” and “Ben Homes” score highly and auto-merges them. The result may be a false positive that combines separate households, corrupts history, and creates privacy exposure. The safe repair is to define features, labeled evaluation data, costs of false merge versus missed merge, confidence bands, and a steward queue for uncertain cases. Automatic merges must be reproducible and reversible through a match ledger.

matching_evidence.py
from difflib import SequenceMatcherleft, right = "ben home", "ben homes"score = SequenceMatcher(None, left, right).ratio()record = {    "left": left,    "right": right,    "score": round(score, 3),    "decision": "steward_review",    "reason": "name similarity alone is insufficient identity evidence",}print(record)assert record["decision"] == "steward_review"

5. Survivor selection is not identity truth

After two source rows are mapped to one durable entity, a survivorship rule chooses which attributes feed a current mastered representation: perhaps most recent verified phone, highest-authority source for legal name, or non-null value from a steward-approved system. That rule must be field-specific and versioned. “Take the newest row” can overwrite a more authoritative value with a lower-quality feed.

Ledger field Purpose
left_source_id / right_source_id Reconstruct which records were compared.
features / normalized values Explain what evidence was considered.
model/rule version Make the decision reproducible after logic changes.
score / rule result Quantify evidence under that version.
decision + actor Distinguish automatic rule from human stewardship.
effective/rollback metadata Support correction if the merge was wrong.

6. Security and fairness boundary

Identity data can be sensitive. Production match features should be purpose-limited, access-controlled, retained only as authorized, and tested for population-specific error. Do not infer protected or sensitive attributes to “improve” matching. This synthetic lab contains no real personal information.

7. Lab verification and bridge

  • QO1/1 duplicate delivery collapses to one accepted business event.
  • Q-C001 and Q-C001-DUP map deterministically only because this lab explicitly governs normalized email as a strong key.
  • The Ben Home/Homes pair remains unmerged and enters steward review.
  • Every accepted/duplicate/match decision retains evidence and rule version.

Lesson 4 applies the same evidence discipline to invalid records: quarantine first, then authorized repair and reprocess.

Knowledge check

Check your understanding

  1. What is the difference between deduplication and entity resolution?
  2. Why is exact email matching not universally safe?
  3. What should happen to uncertain match candidates?
  4. Why is a similarity threshold dataset-specific?
  5. What makes a merge reversible?
Review the answers

1. Deduplication removes repeated representations of the same record/event key; entity resolution decides whether distinct source identities refer to the same real entity.

2. Emails can be shared, recycled, missing, mistyped, or governed differently across systems and jurisdictions.

3. Retain evidence and route them to a defined review/decision process rather than forcing an automated merge.

4. Feature distributions and false-positive/false-negative costs differ by population, language, source, and business use.

5. A match ledger with source identities, evidence, rule/model version, decision actor, and rollback/effective metadata.

Summary and next step

This lesson established the mechanism and production boundaries for Deduplication and Entity Matching: Deterministic/Probabilistic Concepts and Stewardship Boundaries while preserving AtlasMart’s declared grain, governed metrics, history, and reconciliation evidence. Continue to Quarantine Bad Records, Preserve Raw Evidence, Repair/Reprocess, and Avoid Silent Coercion with those contracts unchanged unless an explicit, tested migration says otherwise.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.