Trace AtlasMart session expiry through logical TTL, lazy reads, background cleanup, replication delay, failover, persistence, and restore without treating TTL as a precise scheduler.
Expiration Semantics: Lazy vs Active Expiry and What TTL Means Under Replication
Expiration is not merely deletion. AtlasMart must separate logical visibility, physical cleanup, replication behavior, and durable business deadlines.
Differentiate logical expiration from physical reclamation and explain lazy/passive versus active/background expiry.
Reason about expiry under replication, failover, persistence, snapshots, and clock-sensitive behavior without assuming one vendor model.
Show why a nominal TTL deadline does not guarantee a business action executes at that instant.
Choose TTL for lifecycle/visibility cleanup while keeping authoritative scheduling and invariants in durable mechanisms.
1. A TTL is a data-lifecycle promise, not a stopwatch interrupt
AtlasMart stores short-lived sessions, product-page fragments, rate-limit windows, password-reset challenges, and temporary query results. Each entry may carry a time to live (TTL): a duration after which the system should treat the item as expired. Two separate questions immediately appear. First, when does a read stop returning the value? Second, when are the bytes actually reclaimed? Those moments need not be identical.
Lazy (passive) expiration checks expiry when a key is touched. Active expiration is background work that scans or samples candidates and reclaims already-expired items even if no client reads them. This distinction is operationally important: an expired object can consume memory for a while yet already be logically invisible.
2. Make expiration observable
Suppose session:s-7 is written at logical time 0
with TTL 10. A read at t=9 succeeds. At t=11 the key is
logically expired. If background cleanup does not run until
t=14, the memory object can persist for four seconds past the
nominal deadline. That does not imply a correct read
should return it.
| Time | Logical state | Physical state | Client implication |
|---|---|---|---|
| t=0 | live until 10 | entry allocated | readable |
| t=9 | live | allocated | hit |
| t=11 | expired | may still exist before sweep | correct implementation should hide/remove it on access |
| t=14 | expired | active sweep reclaims it | bytes finally disappear |
Seeing an expired key disappear at t=14 proves when this teaching cleaner reclaimed memory. It does not prove that every database runs background expiration on the same cadence or that a TTL event is delivered exactly once.
3. Replication, failover, persistence, and restore complicate the clock story
A replicated cache must decide which node is authoritative for expiry and how expiry metadata is propagated. Redis Open Source, for example, centralizes expiration at the primary and propagates a synthesized deletion to replicas; a replica can still suppress logically expired values while waiting for that message. Other systems may use different rules. Promotion matters because a former replica may become responsible for expiration itself.
Persistence adds another distinction. A snapshot can contain an item whose absolute expiry is already in the past by restore time. A correct restore policy should not silently grant the item a fresh lifetime unless the product explicitly defines that behavior. Likewise, backups of ephemeral stores do not magically make them authoritative. Restore tests must verify both the value and its lifecycle metadata.
4. Deliberately wrong approach: use TTL to schedule a money-moving action
A team stores release:hold-42 with TTL 10 seconds
and assumes “expiration” will release inventory at exactly t=10.
If the implementation performs physical expiry later, or emits
no durable event, the action is late or missing. A failover can
further alter timing semantics.
The safer pattern is to put the business deadline in durable state, execute it with a scheduler/queue/workflow that has retry and idempotency semantics, and use TTL only to clean derived or no-longer-needed state. The deadline can be checked against an authoritative timestamp when the worker runs.
5. AtlasMart lab: separate logical expiry from cleanup
Python 3.13+ standard library only; verified with Python 3.13.5. No Redis server, Docker, cloud account, paid feature, credentials, host-clock change, process killing, or destructive failure injection is required.
The simulator stores absolute expiry metadata, checks it lazily on reads, delays a background sweep, models a lagging replica's in-memory copy, and restores an old snapshot after its original deadline.
from dataclasses import dataclass
@dataclass
class Entry:
value: str
expire_at: int
class Node:
def __init__(self, name):
self.name = name
self.mem = {}
def set(self, key, value, now, ttl):
self.mem[key] = Entry(value, now + ttl)
def get(self, key, now):
e = self.mem.get(key)
if e is None:
return None
if now >= e.expire_at: # lazy/passive expiry on access
del self.mem[key]
return None
return e.value
def active_sweep(self, now):
removed=[]
for k,e in list(self.mem.items()):
if now >= e.expire_at:
removed.append(k); del self.mem[k]
return removed
primary, replica = Node("P"), Node("R")
primary.set("session:s-7", "alice", now=0, ttl=10)
replica.mem["session:s-7"] = Entry("alice", 10) # replicated expiry metadata
print("t=9 primary read:", primary.get("session:s-7", 9))
print("t=11 logical read:", primary.get("session:s-7", 11))
print("t=11 key still in replica memory before delete propagation:", "session:s-7" in replica.mem)
print("t=11 replica logical read suppresses expired value:", replica.get("session:s-7", 11))
# Precise-scheduler mistake: nobody touches the key and background cleanup is delayed.
job = Node("scheduler-cache")
job.set("release:hold-42", "fire-business-action", now=0, ttl=10)
print("t=10.000 nominal deadline; physical entry present:", "release:hold-42" in job.mem)
print("active cleanup actually runs at t=14:", job.active_sweep(14))
print("scheduler lateness:", 14-10, "seconds")
# Restore using the original absolute expiry, not a fresh TTL.
snapshot = {"coupon:c9": Entry("derived", 12)}
restored = Node("restored")
restored.mem.update(snapshot)
print("restore at t=20 yields logical value:", restored.get("coupon:c9", 20))
print("lesson: TTL is a visibility/lifecycle policy, not a precise business scheduler")
The value is readable at t=9, logically absent at t=11, and can remain physically allocated until the simulated sweep. The “scheduler” cleanup occurs four seconds late. Restoring an entry at t=20 whose recorded expiry was t=12 does not resurrect it. These are mechanism demonstrations, not universal vendor timing guarantees.
6. Production judgment
Use TTL when the data has a well-defined freshness/lifetime boundary and delayed physical reclamation is acceptable. Record whether TTL is relative or absolute, what operations refresh it, which node decides expiry, how replicas hide expired values, and what snapshots/backups preserve. Observe expired-key backlog, memory retained by expired objects, active-expire CPU, replica lag, failover events, and restore correctness. Tenant isolation still matters: one tenant must not be able to alter another tenant's expiry metadata.
The next lesson moves from expiry itself to the update paths around a cache: cache-aside, read-through, write-through, write-behind, and refresh-ahead.
Check your understanding
- Why can an expired key still consume memory?
- What does a TTL of 10 seconds guarantee about an external side effect?
- Why does failover matter for expiration?
- What should a restore test verify besides the value?
- When is TTL a good fit?
Review the answers
1. Logical invisibility and physical reclamation can be separate; background cleanup may run later than the nominal deadline.
2. Nothing by itself. It describes data lifetime according to the store's semantics, not a durable exactly-at-t=10 scheduler.
3. Responsibility for deciding/propagating expiry can move to another node, and clock/replication semantics must be documented.
4. Expiry/lifecycle metadata and whether already-expired entries remain logically absent after restore.
5. For lifecycle/visibility cleanup where approximate physical reclamation is acceptable and authoritative business actions live elsewhere.
References
Foundational claims use primary research/specifications where practical. Redis is an optional current implementation example; no product installation is required for the mandatory labs.
- Redis EXPIRE documentation — Current Redis example of passive and active expiry plus replication/AOF handling.
- Redis replication documentation — Current primary/replica expiration behavior and promotion context.
- Redis Open Source 8.10 release notes — Version snapshot: Redis Open Source 8.10.0 GA, July 2026.
- Redis licensing overview — Redis 8+ tri-license options: RSALv2, SSPLv1, or AGPLv3.