Define and measure end-to-end AtlasMart data SLOs from source availability through certified data products, and prove why task success is not the same as meeting a consumer-facing reliability objective.

End-to-End Data SLOs from Source Availability to Certified Dashboard/Data Product

Build an identity and access model for AtlasMart that separates humans from services, eliminates shared credentials, enforces environment boundaries, and proves least privilege with allow/deny evidence.

Intermediate → Advanced150–190 minutesEnd-to-end SLO labSLIs + SLOs + consumer availabilityLast reviewed: September 2026

Learning outcomes

01

Distinguish service-level indicators, objectives, and agreements in a data context.

02

Define SLOs from consumer-visible data availability rather than scheduler completion.

03

Measure freshness, completeness, and correctness together.

04

Interpret SLO misses without inventing arbitrary universal targets.

05

Use the SLO to drive incident response and communication.

Continuity: observability watches the governed system; it does not redefine it

Chapter 25 begins from the accepted AtlasMart state established through Chapters 01–24: 10 current paid lines, 8 orders, 12 units, 820 USD gross revenue, 495 USD cost, and 325 USD gross profit. The fact grain remains one current accepted paid order line, source progress remains committed through sequence 208, Chapter 20 metric contracts remain authoritative, Chapter 21 certified marts remain dependent on conformed assets, Chapter 22 security/privacy controls remain in force, Chapter 23 lineage/ownership metadata supplies blast-radius context, and Chapter 24 tests remain the correctness gates. Observability adds continuous evidence and incident handling; it must not silently reinterpret business rules merely to make a dashboard look healthy.

Executed local reliability fixture

Runtime: Python 3.13.5 + SQLite 3.46.1. Storage: local in-memory structures plus SQLite semantics where transactions matter. Clock: UTC with explicit timestamps. Security: synthetic identifiers only. Cost: free/local. Healthy run: 10 records in/out, 4,820 bytes, 7.0 minutes, zero rejects/retries, 860 ms fixture CPU, 34 MB fixture peak memory, and certification 12 minutes after source readiness. Important limitation: these CPU/memory numbers are fixture measurements used to teach signal relationships; they are not performance recommendations for any warehouse engine or cloud service.

1. The realistic problem: a successful run that fails the data product

AtlasMart’s late-source scenario completes successfully at 09:14 and certifies at 09:18. If the finance product promises certified data by 08:30, the pipeline is operationally “green” but the product is unreliable for its consumer. An end-to-end data SLO therefore spans producer readiness, data processing, quality/reconciliation, certification, and the consumer-visible product.

2. SLI, SLO, and SLA are not interchangeable

Term Meaning in this chapter Example
SLI Measured indicator of service behavior certified_at minus partition deadline; completeness percentage
SLO Internal target/range for an SLI certified by 08:30 UTC for daily finance partition
SLA External/formal commitment with business consequences, if one exists Not defined by this lab

The mandatory lab defines SLOs only. It does not create a legal/customer SLA or claim that its example times are broadly appropriate.

3. AtlasMart fixture SLO contract

Objective Fixture target Measurement
Source availability manifest ready by 08:10 UTC producer readiness timestamp
Certified freshness certified product by 08:30 UTC certified_at for partition
Completeness 100% of expected accepted rows for this fixed fixture manifest/control comparison
Correctness Chapter 24 blockers pass and gross revenue reconciles to 820 USD test + reconciliation evidence
Availability to consumer certified artifact/semantic result exists after security policy certification registry / consumer query

These values are deliberately fixture-specific. Production teams should choose SLOs from consumer needs, source capabilities, cost, and operational tradeoffs.

4. Compute the end-to-end result, not only task duration

SLO calculation
from datetime import datetimedef ts(s):    return datetime.fromisoformat(s.replace("Z", "+00:00"))certified_deadline = ts("2026-09-21T08:30:00Z")actual_certified   = ts("2026-09-21T09:18:00Z")miss_minutes = (actual_certified - certified_deadline).total_seconds() / 60assert miss_minutes == 48.0

The late run itself lasts only eight minutes, but waiting for the source dominates the consumer-visible delay. Optimizing SQL by one minute would not fix the incident mechanism.

5. Executed SLO evaluation

Healthy versus late scenario
HEALTHYsource ready     08:04 <= 08:10  PASScertified        08:16 <= 08:30  PASScompleteness     100%            PASSrevenue control  820 USD         PASSSOURCE-LATE INCIDENTsource ready     09:05           FAIL by 55 mincertified        09:18           FAIL by 48 mincompleteness     100% after run  PASSrevenue control  820 USD         PASSpipeline status  SUCCESS         (not sufficient)

This is the key distinction: reliability can fail through lateness even when the eventual data is complete and correct.

6. Completeness requires a denominator

“100% complete” must answer 100% of what: source manifest rows, required business entities, expected partitions, or all events through a watermark? AtlasMart’s fixed lab uses the manifest/control set for the partition. In real systems, the source may not expose a trustworthy denominator; then completeness becomes a weaker estimate and that uncertainty must be documented.

7. Correctness is not reducible to one checksum

For AtlasMart, correctness includes schema contracts, declared grain, customer relationships, SCD/event-time semantics, idempotency/restart behavior, and governed metric reconciliation. A matching 820 USD total could still hide wrong historical customer attribution; a perfect schema could still hide the 665 USD dashboard semantic bug. The SLO can reference a bundle of critical test indicators rather than pretending one number proves all correctness.

8. Error budgets and measurement windows are policy choices

An SLO may be evaluated per partition or over a window such as a month. Some teams use an error budget to decide how much unreliability is tolerable before prioritizing reliability work. The lab does not impose a target percentage or budget. The important mechanism is that the objective, window, exclusions, and action are explicit before an incident occurs.

9. Different consumers can need different objectives

Consumer Possible emphasis Why one SLO may not fit
Finance close correctness + reconciliation + fixed deadline late/correct can still miss close process
Marketing exploration freshness may be looser; broad access restricted exploration can tolerate some delay but not policy bypass
Operations nearer-real-time freshness + completeness late order/fulfillment state affects actionability

Do not weaken a shared metric’s semantics for a faster consumer. Instead expose appropriate products/latencies with explicit contracts.

10. Controlled failure: define SLO from task completion

Wrong: “95% of transform tasks finish within 10 minutes.” This can be useful as an internal performance objective but it says nothing about whether data arrived, tests passed, certification happened, or consumers can query it. Repair: include source readiness and certified-product availability; retain task duration as a diagnostic SLI.

11. Production judgment and bridge

User relevance: define objectives around decisions/data products. Freshness/history: use partition/as-of semantics. Correctness: reference explicit critical gates. Security: certification must include policy enforcement; a fast but exposed dataset is not a healthy product. Cost: tighter SLOs can require more compute/operational capacity; quantify that tradeoff instead of copying targets. Next: when an SLO fires, Lesson 4 classifies the mechanism before selecting a repair.

12. Verification checklist

  1. Confirm healthy certification at 08:16 satisfies the fixture deadline.
  2. Confirm late source and certification miss by 55 and 48 minutes respectively.
  3. Confirm the late run still reconciles to the same 820 USD revenue.
  4. Confirm SLO and SLA are not treated as synonyms.
  5. Confirm the objective explicitly names the partition, measurement, and consumer-visible endpoint.

Knowledge check

Acceptance questions

  1. Why is a scheduler-runtime SLO insufficient for a data product?
  2. Can data be correct and still violate reliability?
  3. Why must completeness define a denominator?
  4. Why avoid copying another system’s freshness target?
Review the answers

1. It omits source availability, data checks, certification, and consumer access.

2. Yes—correct data that arrives after the business deadline violates freshness/availability expectations.

3. A percentage is meaningless without the population being measured.

4. Consumer needs, source capabilities, cost, and risk differ by product and organization.

Authoritative references

13. Lab cleanup/reset

The SLO scenario uses fixed UTC timestamps; rerunning reproduces the same pass/fail results. Production SLO systems should persist immutable evaluation evidence for the applicable measurement window.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.