Chapter 04 · Time, Ordering, Logical Clocks, and Versioning
Why Wall Clocks Are Dangerous: Drift, Skew, NTP, and Cross-Node Ordering
Show why cross-node wall-clock order can contradict causality, then separate NTP synchronization from the ordering guarantees an application actually needs.
Learning outcomes
AtlasMart has catalog writers in two regions. During an incident review, an engineer sorts audit records by each server's wall-clock timestamp and concludes that an older price overwrote a newer one. The conclusion is wrong: the write that causally happened later carries the numerically smaller wall timestamp because the nodes' clocks disagree. This lesson builds the distinction between physical time, observed wall time, and distributed causality before any logical-clock mechanism is introduced.
Define oscillator drift, clock skew, offset, synchronization, stepping/slewing, and leap-second handling without treating synchronized clocks as perfect.
Construct an event history where a message establishes happens-before even though wall-clock timestamps appear reversed.
Explain why raw cross-node timestamps are unsafe as a universal conflict-resolution or audit-ordering rule.
Use NTP as a clock-synchronization mechanism while separating clock-quality information from causality.
Run a deterministic AtlasMart skew simulation without changing the host system clock.
Chapter 03 showed that correctness is defined by allowed client-visible histories and business invariants. Time metadata is useful only insofar as its assumptions support those histories. A timestamp is evidence about a clock reading; it is not automatically evidence that one distributed event caused another.
1. Physical clocks are measurements made by imperfect oscillators
A machine's clock advances because hardware and software estimate elapsed time from an oscillator. No two oscillators run at exactly the same rate. Drift is the rate error that accumulates over time. Clock skew is the difference between two clocks at a given moment. Offset is the estimated difference between a local clock and a reference. These quantities matter because distributed software often records timestamps on different hosts and later compares them as if they came from one perfect clock.
Network Time Protocol (NTP) can estimate offsets and discipline a local clock, but synchronization is itself performed over a network with variable delay. RFC 5905 describes a clock-discipline process that can slew or step a clock depending on the estimated error and implementation state. That is an operational correction mechanism, not a causal-order oracle. A disciplined clock can still have uncertainty, and a correction can make application-visible wall time advance at a different rate or, in some environments, move discontinuously.
2. Civil time adds another layer: leap seconds and smearing
Distributed applications frequently conflate UTC, local operating-system wall time, and a monotonic elapsed-time source. RFC 8633 discusses leap-second handling and the practice of leap smearing, where some environments spread the adjustment over a window instead of inserting a leap second directly. During a smear, clocks can deliberately disagree with UTC while remaining internally smooth. Mixing smeared and non-smeared time sources can therefore create additional disagreement.
Do not change the learner machine’s real clock for this course. Clock skew, NTP corrections, leap behavior, and backward jumps are modeled with synthetic integers in the lab. Real clock manipulation can break TLS validation, authentication, scheduled jobs, build systems, logs, databases, and other unrelated applications.
3. Causality comes from program order and messages, not from timestamp magnitude
Suppose node A accepts a catalog update and sends it to node B. B receives that message and only then performs another update. The message establishes a causal chain: A's update happened before B's receive, and B's receive happened before B's later update. If B's clock is 300 milliseconds behind A's, B's later event can carry a smaller wall timestamp. Sorting those records numerically reverses a real dependency.
| Event | True sequence used by the simulation | Node wall reading | What is actually known |
|---|---|---|---|
| A writes price=100 | 1 | 1120 ms | A local event |
| B receives A update | 2 | 880 ms | causally after A write because the message arrived |
| B writes price=95 | 3 | 900 ms | causally after the receive and therefore after A write |
The table uses a fictional “true sequence” only for teaching. A real distributed system usually cannot observe a single global true clock with zero uncertainty. What it can observe are local program order, send/receive relationships, version metadata, protocol terms, acknowledgements, and other evidence created by the system itself.
4. Deliberately wrong approach — last-write-wins by raw client or server timestamp
A common last-write-wins (LWW) rule selects the value with the numerically largest timestamp. That rule can be deterministic, but determinism is not correctness. If timestamps come from clocks with unknown skew, a causally later update may lose. If an untrusted client supplies the timestamp, a clock set far into the future can dominate future legitimate writes. If two updates share coarse timestamp resolution, the system needs an additional tie-breaker anyway.
The safer pattern depends on the invariant. A single-owner record can use a monotonic database version or compare-and-set token. Multi-writer data can carry causal context and expose siblings for domain merge. LWW may still be a deliberate policy for fields where silent loss is acceptable, but its clock assumptions and loss semantics must be explicit.
5. AtlasMart lab — make wall-clock order contradict happens-before
from dataclasses import dataclass
@dataclass
class Event:
name: str
node: str
true_ms: int
clock_offset_ms: int
cause: str | None = None
@property
def wall_ms(self):
return self.true_ms + self.clock_offset_ms
events = [
Event("A edits price to 100", "A", 1_000, +120),
Event("message reaches B", "B", 1_060, -180, cause="A edits price to 100"),
Event("B edits price to 95 after seeing A", "B", 1_080, -180, cause="message reaches B"),
]
print("event history")
for e in events:
print(f"{e.name:31} node={e.node} true={e.true_ms:4} wall={e.wall_ms:4}")
print("\ncausal fact:")
print("A edit -> B receives -> B edit")
print("wall-clock comparison says:")
print("B edit wall time", events[2].wall_ms, "< A edit wall time", events[0].wall_ms)
print("so raw timestamp order contradicts happens-before")
versions = [
("price=100", events[0].wall_ms, "A"),
("price=95", events[2].wall_ms, "B"),
]
winner = max(versions, key=lambda x: x[1])
print("\nWRONG timestamp-only LWW winner:", winner)
print("\nNTP-style correction is simulated, not applied to the host clock:")
local_wall = [10_000, 10_050, 9_980, 10_020]
for i, t in enumerate(local_wall, 1):
print(f"sample {i}: wall={t}")
print("wall time can move backward after a correction; monotonic/causal metadata must not assume otherwise")
Expected output
event history
A edits price to 100 node=A true=1000 wall=1120
message reaches B node=B true=1060 wall= 880
B edits price to 95 after seeing A node=B true=1080 wall= 900
causal fact:
A edit -> B receives -> B edit
wall-clock comparison says:
B edit wall time 900 < A edit wall time 1120
so raw timestamp order contradicts happens-before
WRONG timestamp-only LWW winner: ('price=100', 1120, 'A')
NTP-style correction is simulated, not applied to the host clock:
sample 1: wall=10000
sample 2: wall=10050
sample 3: wall=9980
sample 4: wall=10020
wall time can move backward after a correction; monotonic/causal metadata must not assume otherwise
The important evidence is not the numeric timestamp gap. It is the message dependency. The B edit occurs after B has received A's edit, yet its wall clock is smaller. The simulated NTP correction then shows why code that assumes wall time never decreases can also fail.
Verification checklist
- No operating-system clock is changed.
- The causal chain is explicit in the output.
-
The wrong LWW rule selects
price=100even though B wroteprice=95after observing it. - A synthetic wall-clock sequence moves backward after a modeled correction.
- The lab makes no claim about real NTP error bounds on the learner's machine.
Check your understanding
- What is the difference between clock skew and oscillator drift?
- Why can two correctly functioning NTP clients still have different wall-clock readings?
- What fact proves that B’s second edit happened after A’s edit in the AtlasMart history?
- Why is a raw timestamp a weak conflict-resolution token?
- Why does this course simulate clock changes instead of editing the host clock?
Review the answers
Skew is the instantaneous difference between clock readings; drift is a rate error that causes offset to accumulate over time.
NTP estimates and disciplines time over networks with variable delay and local oscillator error; synchronization has uncertainty and implementations may correct clocks differently.
The send/receive and local program-order chain establishes happens-before independently of the wall-clock numbers.
Its ordering can be wrong under skew, it can collide at coarse resolution, and it often says nothing about whether the writer observed the version it is replacing.
Real clock changes can damage unrelated security, scheduling, logging, database, and build behavior; a deterministic simulator gives the teaching evidence without that blast radius.
6. Production judgment and next bridge
Use wall time for what it is good at: human-facing timestamps, retention windows, approximate age, metrics, and protocols that explicitly include uncertainty assumptions. For ordering that protects invariants, prefer metadata whose semantics are tied to the operation: database versions, terms/epochs, compare-and-set tokens, sequence numbers, causal context, or consensus-backed commit order. Audit systems should preserve both physical timestamps and causal/version identifiers when the distinction matters.
Lesson 2 introduces Lamport logical clocks. They do not repair physical time; instead, they assign numbers using local events and messages so every known causal dependency moves forward. That gives a clean ordering invariant without claiming the numbers are UTC.
Authoritative references
- Lamport — Time, Clocks, and the Ordering of Events in a Distributed System — foundational happens-before relation and logical/physical clock discussion
- RFC 5905 — Network Time Protocol Version 4 — NTP protocol and local clock discipline including slew/step behavior
- RFC 8633 — Network Time Protocol Best Current Practices — operational NTP guidance including leap-second handling and leap smearing