Split image of a metal bucket: on the left, golden liquid splashes and overflows over the rim; on the right, the same bucket drips a single steady stream Tech
AI-generated, Working Theory
Tech · ◉ Evergreen

Every request takes a token — or waits for one.

by · ·5 min·Working Theory

Rate limiting in plain English — token bucket vs leaky bucket, and why the choice is really a choice about whether you tolerate bursts or insist on smoothness.

Every system that takes requests eventually meets one that takes too many. A script in a loop with no sleep. A retry storm from a client that thinks you’re down. A genuinely popular launch. A rate limiter is the small, boring component that stands at your door and decides who comes in now, who comes in later, and who gets turned away — so that “too many” never becomes “everything falls over.”

The two classic ways to build that doorman look almost identical from a distance and behave very differently up close. They’re usually called the token bucket and the leaky bucket, and the difference between them is really one question: do you allow bursts, or do you insist on a steady drip?

Token bucket: save up, then spend

Picture a bucket that fills with tokens at a fixed rate — say, ten tokens a second. The bucket has a maximum capacity; once it’s full, new tokens spill over and are lost. Every incoming request has to take one token to be served. If a token’s there, the request goes through immediately and the token’s gone. If the bucket’s empty, the request is refused (or made to wait).

The clever part is what happens when things are quiet. During a lull, tokens accumulate up to the bucket’s capacity. So when a burst arrives, it can spend all that saved-up balance at once — a hundred requests can go through in a blink if a hundred tokens were sitting in the bucket. After that, you’re back to the drip: ten a second, the refill rate.

That’s the signature of the token bucket: it tolerates bursts, up to a limit you set with the bucket size, then settles to a steady average. The refill rate controls your long-run throughput; the bucket capacity controls how big a burst you’ll forgive. This is why it’s the default for most public APIs — real traffic is bursty, users hate being throttled for a brief spike, and a token bucket lets a well-behaved client that’s been quiet spend a little freely.

Leaky bucket: a steady drip, no matter what

Now flip it. Picture a bucket you pour requests into, with a small hole in the bottom that lets them out at a fixed rate. No matter how fast requests pour in, they leave at the same constant pace — the drip never speeds up. If requests arrive faster than the hole drains, the bucket fills; if it overflows, the excess is dropped.

The leaky bucket doesn’t save up anything. A quiet period buys you nothing — there’s no balance to spend, just an empty queue. When a burst hits, it doesn’t rush through; it lines up and drains at the fixed rate. The output is smooth by construction, which is the whole point.

That’s the signature of the leaky bucket: it refuses to pass a burst along, and hands whatever’s downstream a perfectly even stream. You reach for it when the thing you’re protecting genuinely can’t handle spikes — a downstream service with a hard concurrency ceiling, a payment processor that bills you per call, hardware that overheats under bursts. You’re trading responsiveness for predictability.

Same job, opposite instinct: forgive the burst, or flatten it.

token bucket — tolerates bursts tokens drip IN; a request spends one to pass

+10/s balance saved in the quiet burst in: all pass at once (while tokens last) then back to ~10/s

leaky bucket — flattens bursts requests pour IN; a hole drains at a fixed rate

steady drip out → evenly spaced, no matter the input
Token bucket saves credit for a burst; leaky bucket refuses to pass one along. The choice is burst-tolerance vs smoothness. Original diagram · Working Theory

They’re closer than they look — and the choice is a values choice

If you squint, these are two views of the same idea: a bucket, a fixed rate, and a capacity that decides what happens under overload. In fact you can implement behavior that looks like one using the machinery of the other, which is why engineers argue about the definitions. The distinction worth keeping isn’t the plumbing — it’s the promise each one makes to the thing downstream:

Two practical notes that outlast the diagram. First, decide early what “over the limit” means: do you reject (fail fast, tell the client to back off — and if you do, return the header that says when to retry, or you’ll cause the very retry storm you’re guarding against) or do you queue (make them wait, which smooths traffic but adds latency and can hide a backlog until it’s a crisis — a queue is a loan, as an earlier piece put it). Second, in a distributed system the limiter’s own state — the token count — has to live somewhere shared, or each server enforces its own limit and your real ceiling is silently the limit times the number of servers.

Rate limiting looks like a defensive afterthought — a valve you add when something breaks. But the choice between these two buckets is a small, honest statement about your system: whether you’d rather be forgiving and occasionally spiky, or strict and always smooth. Pick the one whose promise matches what’s standing behind the door.

Sources

  • token bucket and leaky bucket rate-limiting algorithms
  • retry-after backoff header pattern
  • distributed rate-limiting and shared limiter state

Liked this? Get the next one in Working Theory.

Going weekly in August (it's in beta now). One genuinely interesting read on building, the brain, and the science most people missed.

Subscribe →
Got a reaction, a counter-example, or something I missed? Reply by email — I read everything.
◉ join in

Where have you hit this — in a product you use, or one you're building?

Threads open here soon. For now, the conversation lives two clicks away — discuss on GitHub, or just reply by email. I read and answer everything.