Chapter 10 · Key-Value Databases and Data-Structure Stores

Caching vs System-of-Record Usage: Durability, Recovery, and Source-of-Truth Boundaries

Separate disposable cache state from authoritative system-of-record state, then diagnose lost business data, stale cache entries, warm-up/reconstruction paths, persistence assumptions, and recovery boundaries.

Intermediate90–110 minutesCache/SOR recovery labPython 3.13+ · standard libraryNo external database requiredLast reviewed: August 2026

Learning outcomes

An AtlasMart engineer stores paid orders in the same in-memory key-value tier used for product-page caching because “replication is enabled.” During memory pressure the order key is evicted. There is no relational order row, event log, or backup containing the transaction. The issue is not that key-value databases cannot be durable systems of record—they can. The issue is that the architecture never defined which copy was authoritative and how it would be recovered.

01

Define cache, cache-aside, system of record/source of truth, authoritative state, derived state, invalidation, warm-up, reconstruction, durability, backup, and restore.

02

Explain why persistence/replication do not automatically turn a disposable cache configuration into an authoritative business database.

03

Produce business-data loss when an authoritative order exists only in an evicting cache.

04

Repair the design with a durable authoritative record plus invalidatable/reconstructable cache entries.

05

Diagnose stale cache state and verify rebuild after total cache loss.

Tooling and version snapshot · checked 29 August 2026

The mandatory labs use Python 3.13+ standard library only in one deterministic process with a logical clock; there is no database server, client driver, container, cloud account, network fault injection, or paid feature. Optional implementation references were rechecked against Redis Open Source 8.10.0 (GA July 2026; Redis 8+ is offered under the user's choice of AGPLv3, RSALv2, or SSPLv1) and Valkey 9.1.1 (released 21 July 2026; predominantly BSD-3-Clause). Product commands/defaults are examples only; later Stage 03 product courses teach product-specific operations in depth.

1. Authority is an architectural contract, not a product category

A system of record is the authoritative state from which other representations can be reconstructed or reconciled. A cache is a performance copy whose loss is acceptable because another authoritative source can recreate it. A key-value database can serve either role, depending on durability, replication, eviction settings, backup/restore, invariants, schema, access control, and operational practice.

Therefore “Redis is a cache” and “Redis is a database” are both too coarse. The correct question is: for this key, what is the authority and recovery contract?

2. Cache-aside makes the reconstruction path explicit

In cache-aside, reads check the cache, fetch from the system of record on a miss, then populate the cache. Writes normally commit to the authoritative store and invalidate/update the cache according to a defined consistency policy. The cache can be deleted and rebuilt without losing business truth.

The hard part is invalidation. If the database changes but the cache keeps an old value, clients can read stale state until TTL or explicit invalidation. If invalidation happens before the authoritative write commits, a racing reader can repopulate old data. CDC/outbox/event patterns later in the course provide more robust synchronization paths for important derived views.

3. Failure case: persistence enabled, but recovery is still undefined

Persistence files and replicas can improve durability, but they are not equivalent to backups or a tested recovery plan. A bad application write, accidental delete, eviction policy, logical corruption, or misconfiguration may replicate perfectly. Chapter 22 will cover disaster recovery; for now record the minimum requirement: if state is authoritative, prove how it is restored after the failures you claim to tolerate.

4. Warm-up and cold-cache behavior are part of availability

After cache loss, every miss can hit the origin simultaneously, causing a cache stampede. Production plans may use request coalescing, staggered TTLs/jitter, pre-warming, bounded concurrency, stale-while-revalidate semantics, or load shedding. Cache recovery therefore has both a correctness path (“can data be reconstructed?”) and a capacity path (“can the origin survive reconstruction?”).

5. AtlasMart lab: lose a cache, keep the business truth

The lab first treats an LRU cache as the only home of a paid order and loses it to ordinary cache pressure. It then writes the order to an authoritative dictionary, invalidates/populates a derived cache, demonstrates stale data when invalidation is forgotten, and finally clears the cache completely and reconstructs from the source of truth.

python · cache_vs_system_of_record.py
from collections import OrderedDict

class LRUCache:
    def __init__(self, capacity=2):
        self.capacity = capacity
        self.data = OrderedDict()
    def get(self, key):
        if key not in self.data: return None
        value = self.data.pop(key); self.data[key] = value; return value
    def put(self, key, value):
        if key in self.data: self.data.pop(key)
        self.data[key] = value
        if len(self.data) > self.capacity:
            victim, _ = self.data.popitem(last=False)
            print("cache eviction ->", victim)
    def delete(self, key): self.data.pop(key, None)
    def clear(self): self.data.clear()

db = {}
cache = LRUCache(capacity=2)

print("WRONG: AUTHORITATIVE ORDER ONLY IN CACHE")
cache.put("order:9001", {"status": "PAID", "total": 120})
cache.put("product:1", {"name": "Book"})
cache.put("product:2", {"name": "Cable"})  # evicts order
print("order after pressure", cache.get("order:9001"))
print("authoritative db has order?", "order:9001" in db)

print("\nREPAIR: WRITE SYSTEM OF RECORD, CACHE A DERIVED COPY")
def write_order(order_id, doc):
    db[f"order:{order_id}"] = dict(doc)
    cache.delete(f"order:{order_id}")

def read_order(order_id):
    key = f"order:{order_id}"
    hit = cache.get(key)
    if hit is not None:
        return "cache", dict(hit)
    value = dict(db[key])
    cache.put(key, value)
    return "db", value

write_order("9001", {"status": "PAID", "total": 120, "version": 1})
print("first read", read_order("9001"))
print("second read", read_order("9001"))

print("\nSTALE CACHE WHEN INVALIDATION IS FORGOTTEN")
db["order:9001"] = {"status": "REFUNDED", "total": 120, "version": 2}
print("stale read", read_order("9001"))
cache.delete("order:9001")
print("after invalidation", read_order("9001"))

print("\nCACHE LOSS IS RECOVERABLE")
cache.clear()
print("cache empty; reconstruct from db ->", read_order("9001"))
print("source-of-truth keys", sorted(db))
text · expected output
WRONG: AUTHORITATIVE ORDER ONLY IN CACHE
cache eviction -> order:9001
order after pressure None
authoritative db has order? False

REPAIR: WRITE SYSTEM OF RECORD, CACHE A DERIVED COPY
cache eviction -> product:1
first read ('db', {'status': 'PAID', 'total': 120, 'version': 1})
second read ('cache', {'status': 'PAID', 'total': 120, 'version': 1})

STALE CACHE WHEN INVALIDATION IS FORGOTTEN
stale read ('cache', {'status': 'PAID', 'total': 120, 'version': 1})
after invalidation ('db', {'status': 'REFUNDED', 'total': 120, 'version': 2})

CACHE LOSS IS RECOVERABLE
cache empty; reconstruct from db -> ('db', {'status': 'REFUNDED', 'total': 120, 'version': 2})
source-of-truth keys ['order:9001']

Interpretation

The repair is not “always use a relational database.” A properly configured durable key-value store could be the authoritative db in this model. What matters is explicit authority, non-evictability when required, invariant enforcement, durable acknowledgement, recovery, backup, and reconciliation.

Check your understanding

  1. What makes data cached rather than authoritative?
  2. Why does replication not equal backup?
  3. What is the cache-aside read path?
  4. What caused the stale read in the lab?
  5. What should a cold-cache test measure?
Review the answers

1. It is a reconstructable performance copy; losing it does not destroy the only business truth.

2. Replication can faithfully copy operator mistakes, deletes, corruption, or bad application writes; backup/restore provides a separate recovery history and must be tested.

3. Check cache; on miss read the authoritative store; populate cache; return the value.

4. The authoritative value changed without invalidating/updating the cached copy.

5. Origin load, cache refill rate, tail latency, errors, stampede behavior, and time to restore acceptable hit ratio without violating correctness.

6. Production judgment

Production dimension Questions to record before using the pattern
Correctness / atomicity What is the atomic boundary: one key, one structure, one shard, one transaction, or a broader invariant? Which races remain possible around reads, retries, expiration, failover, and multiple keys?
Consistency / topology Which node owns or coordinates the key, how are replicas acknowledged, what stale reads are allowed, and what changes during partition/failover?
Durability / recovery Is the state disposable, reconstructable, or authoritative? What persistence, replication, backup, restore, and reconciliation evidence supports that claim?
Latency / hot keys Measure p50/p95/p99, queueing, value size, key distribution, per-key operation rate, serialization cost, fan-in/fan-out, and hot-key concentration.
TTL / memory Differentiate logical expiration from physical reclamation and memory-pressure eviction. Define refresh rules, jitter, capacity headroom, and acceptable disappearance.
Security / tenancy Namespace tenant data, authorize keys/operations, protect secrets, prevent cross-tenant scans/collisions, limit abusive large values/structures, and audit privileged coordination actions.
Operations / testing Test ambiguous retries, duplicate requests, eviction, restart, replica lag, failover, clock shifts where relevant, hot keys, schema evolution, restore, and cache rebuild.
Version / license / cost Pin product/client versions when used; verify license/edition/feature status; price memory, replicas, persistence I/O, network egress, backup retention, and operational skill.

Bridge to Lesson 5: With authority and disappearance rules explicit, we can evaluate common key-value use cases one by one—sessions, feature state, rate limits, idempotency, and lightweight coordination.

Authoritative references

Keep knowledge open

Help the academy stay free and grow.

If these tutorials save you time, a small donation supports new lessons, technical review, diagrams, examples, and long-term maintenance.

ETHEthereum / ERC-20 only
0x716c4Ab160C4B66F31a28AE2448BfF68fc3a2ef0

Send only Ethereum or ERC-20 compatible assets to this address.