Chapter 02 · Distributed Systems Foundations: Nodes, Networks, Failure, and State

Define SLOs for Reads, Writes, Freshness, Durability, Recovery, and Availability

Turn AtlasMart reliability adjectives into measurable read/write latency, availability, freshness, durability-evidence, RPO/RTO, degraded-mode, and error-budget objectives.

Beginner100–120 minutesSLI/SLO + error-budget evaluatorVendor-neutral · deterministic Python modelLast reviewed: August 2026

Learning outcomes

AtlasMart's Chapter 01 decision notebook contained words such as “fast,” “highly available,” “fresh,” and “recoverable.” Those words cannot choose replication, routing, or failure-domain tradeoffs because they have no denominator, threshold, scope, or measurement method. This lesson turns them into service level indicators (SLIs) and service level objectives (SLOs) for reads, writes, freshness, durability evidence, recovery, and degraded modes.

01

Distinguish an SLI measurement from an SLO target and an SLA contractual commitment.

02

Write request-based availability and percentile-latency objectives with explicit scope, denominator, threshold, and measurement point.

03

Define freshness/staleness objectives for derived views rather than calling them merely “eventually consistent.”

04

Separate normal-operation SLOs from durability/recovery evidence, recovery point objective (RPO), and recovery time objective (RTO).

05

Calculate per-objective error budgets and use a failing freshness objective to drive engineering action instead of hiding it behind a global uptime number.

Important separation

An SLI is a measured indicator, such as successful checkout reads divided by valid checkout read requests. An SLO is the target range for that indicator over a window. An SLA is a contractual/business agreement that may include consequences; it is not automatically identical to the internal SLO.

1. Begin with the user-visible operation

Google SRE guidance recommends starting from what users care about rather than whatever metric happens to be easiest to collect. For a distributed database, “all three nodes have CPU below 70%” is a useful capacity signal but not a user objective. A checkout user cares whether an order read succeeds, whether confirmation arrives before the interaction deadline, whether an acknowledged order disappears, and whether inventory/payment state is correct.

Every objective needs a scope and denominator. Does availability include malformed requests? Load tests? Requests rejected by authentication? Are latency percentiles measured at the database server, service client, gateway, or end-user edge? Does freshness apply to all fields or only the derived search projection? Without those definitions, two dashboards can report different “99.9% availability” values and both be mathematically correct.

Objective type Example AtlasMart definition Evidence source
Read availability ≥99.95% of valid checkout order-read requests succeed over 28 days Client/gateway request outcomes, segmented by tenant/region/path
Read latency 99th percentile of successful checkout reads ≤120 ms over the chosen window Client-observed latency histogram; failed latency tracked separately
Write latency 99th percentile of successful confirmed-order writes ≤250 ms Client-to-ack latency plus database command/replica timing for diagnosis
Search freshness ≥99.9% of catalog changes visible in search within 30 s Source change version/time vs indexed/projected version/time
Durability evidence Recovery/crash test finds all sampled acknowledged order writes Acknowledgement IDs compared with recovered/restored state
RPO Regional recovery loses ≤60 s of eligible order data Backup/replication checkpoint vs recovered state
RTO Required checkout capability restored ≤10 min in declared disaster scenario Timed recovery drill with dependency checklist

2. Percentiles expose tails that averages hide

An average of 40 ms can coexist with a painful 99th percentile if a small fraction of requests take seconds. Interactive systems therefore commonly use percentile objectives: p99 ≤ X ms means the 99th-percentile observation is at or below X under the defined population/window. The exact percentile calculation method must be documented because histogram/quantile implementations differ.

Track failed requests separately. If a system instantly rejects every request, its successful-request latency metric might look excellent while availability is zero. Also segment by region/tenant/partition where a global aggregate can hide one hot or unreachable shard.

3. Freshness is a first-class correctness dimension for derived data

Chapter 01 treated search as a rebuildable derived view, not the source of truth. “Eventually consistent” is not an operational objective. AtlasMart can instead say: “99.9% of eligible catalog changes become visible in free-text search within 30 seconds over 28 days.” Now the system can measure source version/time, projection/index application time, and the number of changes outside the freshness window.

Different data deserves different limits. A recommendation score may tolerate minutes of lag. A session revocation or authorization policy may not safely use a stale derived cache at all. An inventory display may tolerate small staleness while the final checkout reservation must protect the invariant through a stronger authoritative path.

4. Error budgets turn objectives into capacity for failure

If the availability SLO is 99.95%, the allowed bad fraction is 0.05% of the defined valid requests in the window. That error budget is not an excuse to create errors; it is an explicit reliability allowance that helps teams balance engineering/change velocity against user harm. Each SLO can have its own budget. A service may meet availability while exhausting freshness or latency budget.

Request-based and time-based budgets are not interchangeable. “0.05% bad requests” cannot automatically be converted into “21.6 minutes of downtime” unless the SLO itself is defined as time availability with the appropriate weighting. Traffic varies, failures may affect only some users, and partial failure is often scoped by region/partition.

5. Run the SLO evaluator

The deterministic lab evaluates a synthetic observation window. Checkout read availability is 99.960% with eight failed requests against a budget of ten. Read p99 is 110 ms and write p99 is 220 ms, both inside the hypothetical targets. Search freshness is only 99.850%: 15 late updates against a budget of ten, so the freshness objective fails even though availability passes.

The recovery drill records 500 acknowledged writes and restores all 500, plus RPO=45 seconds and RTO=420 seconds against the example bounds. “0 lost in this test” is evidence, not a mathematical durability probability. A real durability claim requires repeated fault testing, implementation guarantees, independent backups, corruption detection, restore validation, and documented failure assumptions.

python · evaluate availability, percentile latency, freshness, durability evidence, RPO, and RTO
from decimal import Decimalfrom math import ceil, floordef nearest_rank(values, percentile):    ordered = sorted(values)    rank = ceil(Decimal(str(percentile)) / Decimal(100) * len(ordered))    return ordered[max(0, rank - 1)]def budget(total, target_percent):    bad_fraction = Decimal(100) - Decimal(str(target_percent))    return floor(Decimal(total) * bad_fraction / Decimal(100))# Request-based availability SLO.reads_total = 20_000reads_failed = 8availability = (Decimal(reads_total - reads_failed) / Decimal(reads_total)) * 100availability_target = Decimal("99.95")availability_budget = budget(reads_total, availability_target)# Latency objectives are based on successful requests only in this lab.read_latencies = [55 + (i % 25) for i in range(19_792)] + [110] * 180 + [118] * 20write_latencies = [90 + (i % 80) for i in range(4_900)] + [220] * 50 + [240] * 50read_p99 = nearest_rank(read_latencies, 99)write_p99 = nearest_rank(write_latencies, 99)# Freshness objective for a derived search view.updates = 10_000late_updates = 15fresh_percent = (Decimal(updates - late_updates) / Decimal(updates)) * 100fresh_target = Decimal("99.9")fresh_budget = budget(updates, fresh_target)# Recovery drill evidence.acked_writes = 500restored_writes = 500rpo_seconds = 45rto_seconds = 420print("objective_evaluation:")print(f"  read_availability actual={availability:.3f}% target>={availability_target}% bad={reads_failed} budget={availability_budget} status={'PASS' if availability >= availability_target else 'FAIL'}")print(f"  read_latency p99={read_p99}ms target<=120ms status={'PASS' if read_p99 <= 120 else 'FAIL'}")print(f"  write_latency p99={write_p99}ms target<=250ms status={'PASS' if write_p99 <= 250 else 'FAIL'}")print(f"  search_freshness actual={fresh_percent:.3f}% target>={fresh_target}% late={late_updates} budget={fresh_budget} status={'PASS' if fresh_percent >= fresh_target else 'FAIL'}")print(f"  crash_restore acknowledged={acked_writes} restored={restored_writes} lost={acked_writes-restored_writes} evidence_only=yes")print(f"  recovery rpo={rpo_seconds}s target<=60s status={'PASS' if rpo_seconds <= 60 else 'FAIL'}")print(f"  recovery rto={rto_seconds}s target<=600s status={'PASS' if rto_seconds <= 600 else 'FAIL'}")print("degraded_mode=if search projection is unavailable, checkout remains authoritative; free-text search may return an explicit partial-outage response")print("note=request-based availability budget is not automatically equivalent to a time-based downtime budget")

Verified deterministic output

text · per-objective status and error-budget evidence
objective_evaluation:  read_availability actual=99.960% target>=99.95% bad=8 budget=10 status=PASS  read_latency p99=110ms target<=120ms status=PASS  write_latency p99=220ms target<=250ms status=PASS  search_freshness actual=99.850% target>=99.9% late=15 budget=10 status=FAIL  crash_restore acknowledged=500 restored=500 lost=0 evidence_only=yes  recovery rpo=45s target<=60s status=PASS  recovery rto=420s target<=600s status=PASSdegraded_mode=if search projection is unavailable, checkout remains authoritative; free-text search may return an explicit partial-outage responsenote=request-based availability budget is not automatically equivalent to a time-based downtime budget

6. Deliberately wrong approach: one uptime number and average latency

AtlasMart could publish “99.99% uptime, 40 ms average database latency” and still fail customers: one region may be unable to write, p99 could be seconds, search could be 20 minutes stale, or recovery could lose an hour of acknowledged data. A global average/uptime number compresses exactly the partial-failure dimensions Chapter 02 has taught us to preserve.

The repair is a small set of user-centered objectives segmented where architecture can fail independently. Use black-box indicators for user-visible outcomes, then correlate with white-box causes such as replica lag, queue depth, disk errors, retry rate, saturation, partition ownership, and recovery checkpoints. Too many SLOs become unmaintainable; choose the minimum set that represents the critical user promises.

7. Define degraded modes as part of the objective

Availability is not always “full feature set or total outage.” During search projection failure, AtlasMart may keep checkout and direct product lookup authoritative while free-text search returns an explicit partial-outage response. During fraud-enrichment slowness, a business rule may decide whether to queue/manual-review rather than block every order. The allowed mode must be designed, secured, instrumented, and tested before the incident.

A degraded mode must never silently weaken tenant isolation, authorization, payment correctness, or durable-order invariants. If a safe degraded path does not exist, failing closed may be the correct availability tradeoff. SLOs should state which mode counts as success for which operation.

8. Update the AtlasMart decision notebook

Add these fields to every critical path: SLI definition; SLO target/window; measurement point; included/excluded traffic; percentile method; freshness/version source; error-budget policy; allowed degraded response; failure-domain scenario; RPO/RTO; backup/restore evidence; and owner/review date. Now the notebook can constrain later CAP, replication, quorum, sharding, and cloud choices.

For example, if checkout requires writes to remain correct during a zone partition but allows them to reject rather than risk conflicting state, that requirement influences the consistency/availability choice in Chapter 03. If search can remain available with 30 seconds of staleness, a different tradeoff may be acceptable. The business invariant and SLO precede the database label.

9. Production judgment and bridge to CAP/PACELC

Good SLOs make architecture falsifiable. They expose when a supposedly “highly available” system fails a real user path, when tail latency grows despite a healthy average, when derived data is too stale, and when a recovery plan exists only on paper. Use error budgets to prioritize reliability work, not to hide incidents; investigate repeated misses by failure domain and user impact.

Chapter 03 now introduces CAP and PACELC. Because Chapter 02 has defined nodes, partitions, partial failure, deadlines, and SLOs, CAP no longer needs to be taught as “pick two.” We can ask a precise question: during a network partition, for a particular replicated operation and invariant, which responses are allowed to preserve consistency, and which availability/latency objective is sacrificed?

Verification checklist

  • Every objective has an explicit measurement population/scope.
  • Read and write latency are expressed as percentiles, not only averages.
  • Search freshness has its own budget and fails in the deterministic dataset.
  • RPO and RTO are separate and measured by a recovery drill.
  • Zero loss in one drill is labeled evidence, not a universal durability proof.
  • The degraded search mode does not make search authoritative for checkout or security.

Check your understanding

  1. What is the difference between an SLI and an SLO?
  2. Why can an average latency metric hide user pain?
  3. How is a freshness SLO stronger than saying “eventually consistent”?
  4. What is an error budget?
  5. Why are RPO and RTO different?
  6. Why is “500 acknowledged writes restored, 0 lost” not a complete durability guarantee?
Review the answers

An SLI is the measured indicator; an SLO is the target range/value for that indicator over a defined window and scope.

A small fraction of very slow requests can be hidden by many fast requests; tail percentiles expose that distributional behavior.

It states a measurable bound and compliance fraction, such as how many changes must become visible within 30 seconds.

It is the allowed fraction/amount of SLO misses over the defined population/window, commonly 1 minus the SLO target for a ratio objective.

RPO bounds tolerable data loss/rollback after recovery; RTO bounds how long restoring the required service capability may take.

It proves only that this test scenario recovered those sampled writes. Other crashes, corruption, correlated failures, software bugs, backup gaps, and untested histories can behave differently.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.