Tech
Tech · ◉ Evergreen

The clock on the wall lies. Order events by what caused what.

by · ·5 min·Working Theory

In a distributed system you can't trust timestamps, and you can't always say which of two events came first. Here's the tool that tells causality apart from coincidence — and why 'these two are concurrent' is often the answer you actually needed.

Here is a bug that has humbled a lot of good engineers. Two servers each write a value. You want to keep the newer one, so you compare their timestamps and keep the later. It works in testing. In production it silently keeps the wrong value, sometimes, and you lose a user’s edit with no error anywhere.

The reason is that the two servers’ clocks disagree. They always do. Physical clocks drift, get corrected, leap, and skew by milliseconds that are an eternity to a computer. “Later timestamp” does not mean “happened after.” In a distributed system, a wall clock is a rumor each machine tells itself.

So distributed systems stop asking “what time did this happen?” and start asking a humbler question: “did this event cause that one, or not?” Leslie Lamport framed it in 1978 as the happens-before relation. Event A happens-before B if A could possibly have influenced B — because they’re on the same machine in order, or because A sent a message that B received. If there’s a chain of those links from A to B, A is in B’s past. If there’s no chain either way, the two events are concurrent — not simultaneous, something stranger: genuinely unordered. Neither one is in the other’s history. The universe has no fact about which came first, because nothing connects them.

Lamport’s first tool, a single ever-increasing counter per machine, gets you a consistent total order — handy — but it has a blind spot: if A’s counter is less than B’s, you still can’t tell whether A actually caused B or the two are merely concurrent and the numbers happened to land that way. It can order everything, but it can’t tell causality from coincidence.

The fix, found independently by Colin Fidge and Friedemann Mattern in 1988, is to stop using one counter and use a vector — one slot per node. Each node counts its own events in its own slot. Every message carries the sender’s whole vector; the receiver takes the element-wise maximum (folding in everything the sender knew) and then ticks its own slot. Now the rule is simple: to compare two events, compare their vectors slot by slot. If one is less-than-or-equal in every slot, that event is in the other’s causal past. If each vector beats the other in some slot, neither contains the other — the events are concurrent.

A B C [1,0,0] [1,1,0] [0,0,1] B's event [1,1,0] vs C's [0,0,1] B wins slot B · C wins slot C neither ≤ the other → concurrent: a real conflict
A message folds the sender's knowledge into the receiver. Two vectors that each win on a different slot mean the events are causally unrelated — which is exactly the signal you needed. Original diagram · Working Theory

Why does a builder care? Because “concurrent” is not a failure to answer — it’s the answer. When two writes come back concurrent, you haven’t lost the ability to order them; you’ve learned there’s a genuine conflict that no timestamp could have told you about. Amazon’s Dynamo (2007) leaned on exactly this: it used vector clocks to tag versions, and when two versions turned out concurrent it didn’t silently pick one — it handed both to the application (or a read-repair pass) to reconcile on purpose. The “last timestamp wins” bug from the top of this piece is precisely the thing vector clocks refuse to let you commit by accident.

The honest price: a vector grows with the number of writers, so a system with many clients needs to prune, cap, or switch to refinements like dotted version vectors — otherwise the bookkeeping outgrows the data. And vector clocks tell you that two events conflict; they never tell you how to resolve it. That judgment is yours. But knowing you have a real conflict — instead of a coin flip dressed up as a timestamp — is the difference between a system that loses data quietly and one that loses it loudly, which is the only kind of losing you can fix.

Sources

  • Leslie Lamport, 'Time, Clocks, and the Ordering of Events in a Distributed System' (1978)
  • vector clocks (Colin Fidge 1988
  • Friedemann Mattern 1988)
  • Amazon's Dynamo (DeCandia et al., 2007) and dotted version vectors

Liked this? Get the next one in Working Theory.

Going weekly in August (it's in beta now). One genuinely interesting read on building, the brain, and the science most people missed.

Subscribe →
Got a reaction, a counter-example, or something I missed? Reply by email — I read everything.
◉ join in

Where have you hit this — in a product you use, or one you're building?

Threads open here soon. For now, the conversation lives two clicks away — discuss on GitHub, or just reply by email. I read and answer everything.