Observe the service from the client boundary before explaining it with server internals.

Golden Signals for Databases: Latency, Throughput, Errors, Saturation, and Freshness

Measure database health from the client outward: latency distributions, throughput, errors, saturation, and freshness must be interpreted together rather than collapsed into averages.

Advanced125–165 minutesObservability labPython 3.13+ · standard libraryVendor-neutral · free/local mandatory pathLast reviewed: August 2026
01

Define client-observed latency distributions, throughput, error classes, saturation, queueing, and freshness as separate service signals.

02

Relate user-visible symptoms to replica lag, queues, storage, CPU, network, and topology without assuming one metric proves causality.

03

Run a deterministic AtlasMart history where the mean looks healthy while p99 latency, queue depth, and freshness violate targets.

04

Build production dashboards and alerts around distributions and SLO-relevant evidence rather than averages alone.

1. Start with the request the customer actually waits for

AtlasMart's checkout API can look healthy on a node dashboard while customers see pauses or stale inventory. Performance engineering therefore begins at the client-observed boundary. Latency is the elapsed time from request start until the client receives a usable response. Throughput is completed work per unit time. Errors must be classified—timeouts, overload/rejection, authentication failures, validation failures, and server faults mean different things. Saturation means some finite resource or concurrency boundary is near or beyond its sustainable working range; queue depth is one common symptom. Freshness measures how far a read model or replica trails the authoritative change, often by source version, log position, or elapsed lag.

Do not reduce a latency distribution to one mean. If 970 requests complete in roughly 6–12 ms but a thin tail takes hundreds of milliseconds, the mean can remain attractive while checkout p99 violates the service level objective (SLO). Percentiles answer a different question: p99 is the latency at or below which roughly 99% of observations fall; it says nothing about the worst 1%, and percentile estimates need enough samples.

2. Golden signals are correlated evidence, not five independent scorecards

A slow request can be caused by client-pool waiting, coordinator queueing, a remote replica, storage stalls, compaction, garbage collection (GC), cache misses, or a network retransmission. A rising queue with flat throughput suggests offered load has exceeded effective service rate or downstream work is blocked. A replica can answer quickly yet be stale. Conversely, a replica can be current but unavailable. Therefore AtlasMart correlates latency, throughput, errors, saturation, and freshness on the same timeline.

Signal Client-visible question Useful correlated evidence
Latency distribution How long did successful and failed requests take? p50/p95/p99/p99.9, queue wait, server/service time, route/replica
Throughput How much useful work completed? offered load vs completed load, read/write mix, batches, retries
Errors Which requests failed and why? timeout vs overload vs validation vs auth; retryability
Saturation What finite resource is backing up? queue depth, pool wait, CPU run queue, disk queue, connection slots
Freshness How current was returned data? source version, replication lag, projection offset, read consistency path
A server-only dashboard can be correct and still incomplete

A coordinator may report 8 ms execution while the client waited 70 ms for a pool slot and 20 ms on the network. Instrument client wait, request route, server timing, and freshness/version metadata so the measurement boundary is explicit.

3. Deliberately wrong approach: “average latency is 14.5 ms, therefore the database is fine”

The lab creates 1,000 client observations. Most are fast, but the tail reaches 450 ms; one replica is 18 seconds behind and one node has a queue depth of 96. The average remains below 20 ms. Treating that mean as the health verdict produces a misleading benchmark/alert: the tail and freshness failure disappear from the headline.

4. AtlasMart lab: measure the distribution, errors, queues, and freshness

Mandatory lab environment

Python 3.13+ standard library only. These numbers are deterministic teaching observations, not product benchmarks or tuning targets.

python · AtlasMart deterministic simulation
from statistics import mean

def pct(xs, p):
    s = sorted(xs)
    rank = max(0, min(len(s)-1, int((p/100) * len(s) + 0.999999) - 1))
    return s[rank]

# 1,000 client-observed requests: most fast, a thin slow tail, and explicit errors.
latency_ms = [6 + (i % 7) for i in range(970)] + [80 + (i % 5) * 20 for i in range(20)] + [250,270,290,310,330,350,370,390,410,450]
errors = {"timeout": 6, "overloaded": 4, "validation": 2}
window_s = 10
freshness_lag_s = {"replica-a": 0.4, "replica-b": 0.7, "replica-c": 18.0}
queue_depth = {"node-a": 3, "node-b": 5, "node-c": 96}

print("CLIENT LATENCY DISTRIBUTION")
for p in (50, 95, 99, 99.9):
    print(f"p{p}: {pct(latency_ms,p)} ms")
print(f"mean: {mean(latency_ms):.1f} ms")
print(f"throughput: {len(latency_ms)/window_s:.1f} req/s")
print("errors:", errors, "error_rate=", f"{sum(errors.values())/len(latency_ms):.2%}")
print("freshness_lag_s:", freshness_lag_s)
print("queue_depth:", queue_depth)

print("\nBROKEN AVERAGE-ONLY JUDGMENT")
print("average under 20 ms?", mean(latency_ms) < 20)
print("p99 under 100 ms?", pct(latency_ms,99) < 100)
print("replica-c freshness under 5 s?", freshness_lag_s["replica-c"] < 5)
print("diagnosis: an acceptable mean coexists with a bad tail, saturation, and stale data")
Expected evidence

The mean is 14.5 ms, yet p99 is 160 ms and p99.9 is 410 ms. Replica C is 18 seconds stale and node C has queue depth 96. The evidence proves the mean alone cannot certify user-visible performance; it does not identify the root cause without correlated internals.

5. Production judgment and instrumentation contract

Define latency at the client boundary, tag operation class and consistency level, separate success from failure latency, and preserve a histogram rather than only averages. Bound label cardinality: raw tenant IDs, user IDs, or full keys can create expensive telemetry and security leakage. Track freshness with a stable source version where possible rather than guessing from wall-clock timestamps. A trace should show routing and waits, but sampling means absence of a trace is not proof an event did not happen.

Alert on service objectives and causal signals together: “checkout p99 above target while queue depth rises” is more actionable than “CPU above 80%.” During incidents, compare offered load with completed throughput; retry storms can make server operations rise even while useful customer throughput falls. The next lesson moves from cluster-level symptoms to the shard and storage-engine internals that explain them.

Check your understanding

  1. Why can mean latency look healthy while users still experience severe stalls?
  2. What is the difference between throughput and offered load?
  3. Why is freshness a performance/observability signal rather than only a consistency label?
  4. What does a high queue depth prove?
  5. Why should telemetry labels avoid unbounded tenant/key identifiers?
Review the answers

1. A small slow tail can have little effect on the arithmetic mean while still affecting many real requests; percentiles preserve that distribution shape.

2. Offered load is work clients attempt to submit; throughput is work the system actually completes per unit time.

3. A read can be fast but operationally wrong for the use case if its source version or replica state is too far behind.

4. It proves work is waiting at that measured boundary; it does not by itself prove whether CPU, disk, network, downstream service, or policy caused the wait.

5. High-cardinality labels can overwhelm monitoring systems, raise cost, and leak sensitive identifiers; use bounded dimensions plus controlled drill-down.

References

Foundational claims use primary research or current official documentation where practical. Product references are implementation anchors only; the mandatory labs are vendor-neutral.

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.