Define and measure end-to-end AtlasMart data SLOs from source availability through certified data products, and prove why task success is not the same as meeting a consumer-facing reliability objective.
End-to-End Data SLOs from Source Availability to Certified Dashboard/Data Product
Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.
Learning outcomes
Distinguish service-level indicators, objectives, and agreements in a data context.
Define SLOs from consumer-visible data availability rather than scheduler completion.
Measure freshness, completeness, and correctness together.
Interpret SLO misses without inventing arbitrary universal targets.
Use the SLO to drive incident response and communication.
Chapter 25 begins from the accepted AtlasMart state established through Chapters 01–24: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current accepted paid order line, source progress remains committed through sequence 208, Chapter 20 metric contracts remain authoritative, Chapter 21 certified marts remain dependent on conformed assets, Chapter 22 security/privacy controls remain in force, Chapter 23 lineage/ownership metadata supplies blast-radius context, and Chapter 24 tests remain the correctness gates. Observability adds continuous evidence and incident handling; it must not silently reinterpret business rules merely to make a dashboard look healthy.
Runtime: Python 3.13.5 + SQLite 3.46.1. Storage: local in-memory structures plus SQLite semantics where transactions matter. Clock: UTC with explicit timestamps. Security: synthetic identifiers only. Cost: free/local. Healthy run: 10 records in/out, 4,820 bytes, 7.0 minutes, zero rejects/retries, 860 ms fixture CPU, 34 MB fixture peak memory, and certification 12 minutes after source readiness. Important limitation: these CPU/memory numbers are fixture measurements used to teach signal relationships; they are not performance recommendations for any warehouse engine or cloud service.
1. The realistic problem: a successful run that fails the data product
AtlasMart’s late-source scenario completes successfully at 09:14 and certifies at 09:18. If the finance product promises certified data by 08:30, the pipeline is operationally “green” but the product is unreliable for its consumer. An end-to-end data SLO therefore spans producer readiness, data processing, quality/reconciliation, certification, and the consumer-visible product.
2. SLI, SLO, and SLA are not interchangeable
| Term | Meaning in this chapter | Example |
|---|---|---|
| SLI | Measured indicator of service behavior | certified_at minus partition deadline; completeness percentage |
| SLO | Internal target/range for an SLI | certified by 08:30 UTC for daily finance partition |
| SLA | External/formal commitment with business consequences, if one exists | Not defined by this lab |
The mandatory lab defines SLOs only. It does not create a legal/customer SLA or claim that its example times are broadly appropriate.
3. AtlasMart fixture SLO contract
| Objective | Fixture target | Measurement |
|---|---|---|
| Source availability | manifest ready by 08:10 UTC | producer readiness timestamp |
| Certified freshness | certified product by 08:30 UTC | certified_at for partition |
| Completeness | 100% of expected accepted rows for this fixed fixture | manifest/control comparison |
| Correctness | Chapter 24 blockers pass and gross revenue reconciles to 820 USD | test + reconciliation evidence |
| Availability to consumer | certified artifact/semantic result exists after security policy | certification registry / consumer query |
These values are deliberately fixture-specific. Production teams should choose SLOs from consumer needs, source capabilities, cost, and operational tradeoffs.
4. Compute the end-to-end result, not only task duration
from datetime import datetimedef ts(s): return datetime.fromisoformat(s.replace("Z", "+00:00"))certified_deadline = ts("2026-09-21T08:30:00Z")actual_certified = ts("2026-09-21T09:18:00Z")miss_minutes = (actual_certified - certified_deadline).total_seconds() / 60assert miss_minutes == 48.0
The late run itself lasts only eight minutes, but waiting for the source dominates the consumer-visible delay. Optimizing SQL by one minute would not fix the incident mechanism.
5. Executed SLO evaluation
HEALTHYsource ready 08:04 <= 08:10 PASScertified 08:16 <= 08:30 PASScompleteness 100% PASSrevenue control 820 USD PASSSOURCE-LATE INCIDENTsource ready 09:05 FAIL by 55 mincertified 09:18 FAIL by 48 mincompleteness 100% after run PASSrevenue control 820 USD PASSpipeline status SUCCESS (not sufficient)
This is the key distinction: reliability can fail through lateness even when the eventual data is complete and correct.
6. Completeness requires a denominator
“100% complete” must answer 100% of what: source manifest rows, required business entities, expected partitions, or all events through a watermark? AtlasMart’s fixed lab uses the manifest/control set for the partition. In real systems, the source may not expose a trustworthy denominator; then completeness becomes a weaker estimate and that uncertainty must be documented.
7. Correctness is not reducible to one checksum
For AtlasMart, correctness includes schema contracts, declared grain, customer relationships, SCD/event-time semantics, idempotency/restart behavior, and governed metric reconciliation. A matching 820 USD total could still hide wrong historical customer attribution; a perfect schema could still hide the 665 USD dashboard semantic bug. The SLO can reference a bundle of critical test indicators rather than pretending one number proves all correctness.
8. Error budgets and measurement windows are policy choices
An SLO may be evaluated per partition or over a window such as a month. Some teams use an error budget to decide how much unreliability is tolerable before prioritizing reliability work. The lab does not impose a target percentage or budget. The important mechanism is that the objective, window, exclusions, and action are explicit before an incident occurs.
9. Different consumers can need different objectives
| Consumer | Possible emphasis | Why one SLO may not fit |
|---|---|---|
| Finance close | correctness + reconciliation + fixed deadline | late/correct can still miss close process |
| Marketing exploration | freshness may be looser; broad access restricted | exploration can tolerate some delay but not policy bypass |
| Operations | nearer-real-time freshness + completeness | late order/fulfillment state affects actionability |
Do not weaken a shared metric’s semantics for a faster consumer. Instead expose appropriate products/latencies with explicit contracts.
10. Controlled failure: define SLO from task completion
Wrong: “95% of transform tasks finish within 10 minutes.” This can be useful as an internal performance objective but it says nothing about whether data arrived, tests passed, certification happened, or consumers can query it. Repair: include source readiness and certified-product availability; retain task duration as a diagnostic SLI.
11. Production judgment and bridge
User relevance: define objectives around decisions/data products. Freshness/history: use partition/as-of semantics. Correctness: reference explicit critical gates. Security: certification must include policy enforcement; a fast but exposed dataset is not a healthy product. Cost: tighter SLOs can require more compute/operational capacity; quantify that tradeoff instead of copying targets. Next: when an SLO fires, Lesson 4 classifies the mechanism before selecting a repair.
12. Verification checklist
- Confirm healthy certification at 08:16 satisfies the fixture deadline.
- Confirm late source and certification miss by 55 and 48 minutes respectively.
- Confirm the late run still reconciles to the same 820 USD revenue.
- Confirm SLO and SLA are not treated as synonyms.
- Confirm the objective explicitly names the partition, measurement, and consumer-visible endpoint.
Knowledge check
Acceptance questions
- Why is a scheduler-runtime SLO insufficient for a data product?
- Can data be correct and still violate reliability?
- Why must completeness define a denominator?
- Why avoid copying another system’s freshness target?
Review the answers
1. It omits source availability, data checks, certification, and consumer access.
2. Yes—correct data that arrives after the business deadline violates freshness/availability expectations.
3. A percentage is meaningless without the population being measured.
4. Consumer needs, source capabilities, cost, and risk differ by product and organization.
Authoritative references
- OpenTelemetry — Metrics specificationAuthoritative metric concepts and separation of instrumentation from collection/export behavior.
- OpenTelemetry — Metrics data modelUseful for reasoning about time-series measurements, aggregation, temporality, and preserving metric semantics.
- Google SRE Book — Service Level ObjectivesDefines SLIs/SLOs and emphasizes selecting indicators from user-relevant service behavior rather than everything that is easy to measure.
- Google SRE Book — Monitoring Distributed SystemsBackground on actionable monitoring signals and avoiding monitoring that produces noise without operator decisions.
- SQLite — TransactionsSupports the local demonstration of atomic repair/backfill state and safe commit boundaries.
- Python — statisticsUsed for small deterministic baseline examples; production anomaly detection needs workload-specific statistical validation.
- Python — hashlibUsed to fingerprint deterministic incident evidence and repaired outputs.
13. Lab cleanup/reset
The SLO scenario uses fixed UTC timestamps; rerunning reproduces the same pass/fail results. Production SLO systems should persist immutable evaluation evidence for the applicable measurement window.