Rate limiting at scale: an architecture, not an algorithm
A production rate limiter is not a counter bolted onto an endpoint. It is a distributed admission-control system that has to be accurate, fast, synchronized and honest with the clients it throttles.
Amit Shivpratap Singh • • 10 min read
On this page
Why traffic control is an architecture decision
It is tempting to treat rate limiting as a single counter and an if-statement: has this key made too many requests, yes or no. That view survives exactly until the service runs on more than one instance, and then it quietly stops being true. A production limiter is doing three jobs at once — a security job, an economics job and a stability job — and each one shapes a different part of the design.
The security job is to prevent resource starvation, whether it comes from a genuine denial-of-service attempt or from a misbehaving retry loop in someone else's code. The economics job is to keep a lid on anything that costs money per call — a payment gateway, a metered third-party API, a GPU inference endpoint — so a bug three layers upstream cannot turn into an invoice. The stability job is the most familiar one: filter excess load before it reaches the application tier, so response time and queue depth stay inside a range the system can actually sustain.
None of those three jobs is optional in a system that matters, which is why the limiter belongs in the architecture diagram next to the load balancer and the database, not tucked away as a utility function inside one controller.
A baseline before choosing anything
Before comparing algorithms or picking a data store, it helps to fix what "good" means. A useful baseline has six parts: enforcement lives on the server, not the client; accuracy is strict enough that the limiter rarely admits traffic above policy; latency overhead is small enough that the check never becomes the slowest part of the request; memory use per identity stays tiny even at millions of keys; state is shared across every server and microservice instance rather than local to one process; and a throttled request gets an explicit, machine-readable rejection instead of a silent drop.
Those six requirements pull in the same direction more than they conflict. Ultra-low latency argues for in-memory state. Multi-instance scale means that state has to be shared rather than private. Strict accuracy means the read-check-write sequence has to be atomic. And clear rejection means the response has to carry enough information for the caller to behave intelligently instead of hammering the endpoint again immediately. Every later design choice is really just satisfying this list.
Where the limiter should live
Client-side throttling is worth building as a courtesy — it reduces avoidable calls — but it can never be the authoritative control. A client is code you do not run, on a device you do not control, and it can be patched, forked or ignored. Treating it as the enforcement boundary is treating a suggestion as a lock.
Embedding the check directly in each API server is workable, but it ties the implementation to whatever stack that service happens to run, and it means re-solving the same problem in every service that needs it. The placement that scales cleanly is the API gateway or middleware tier, sitting in front of every downstream service. It becomes one policy boundary that also handles authentication, TLS termination and IP allow-listing, and it is the one place a rule change reaches every caller at once.
Five algorithms, one decision
A token bucket holds a capped number of tokens that refill at a steady rate; every admitted request spends one. Because unused tokens accumulate up to the bucket size, it can absorb a short burst without loosening the long-term rate — the shape most APIs actually want from their traffic.
A leaking bucket is the same idea turned around: requests can arrive in any pattern, but they leave through a fixed-rate queue. It is excellent for smoothing bursty input into a predictable outflow for asynchronous workers, and correspondingly bad at ever letting a legitimate burst through faster than that fixed rate.
A fixed window counter is the cheapest option: one counter per identity, reset every interval. Its flaw shows up exactly at the boundary — five requests in the last second of one window plus five more in the first second of the next both pass, and a rolling one-second view now shows ten requests against a limit of five.
A sliding window log fixes that by recording the timestamp of every request and counting how many fall inside the trailing window. It is exact, and it is expensive: the log keeps growing with traffic, including the requests it eventually rejects.
A sliding window counter is the practical middle ground: it blends the count in the current fixed window with a weighted slice of the previous one, using the fraction of the previous window that still overlaps the trailing interval. It gets most of the sliding log's accuracy for a fraction of its memory, which is why it has become the default recommendation for anything that is not either trivially low-stakes or strictly exact-or-nothing.
The hybrid trade-off, in one picture
Put a token bucket and a sliding window counter side by side and the reason production systems combine them becomes obvious: the bucket answers "is this instant allowed?" and the counter answers "has this identity stayed inside its long-term rate?" Neither question is the whole story on its own. A request is admitted only when it clears both checks, which is what lets a hybrid limiter tolerate a legitimate burst while still holding a hard line on sustained abuse.
The counter side of that hybrid is arithmetic, not guesswork, which is exactly why it is cheap to compute per request.
function estimateRequests(
currentWindowCount: number,
previousWindowCount: number,
overlapFraction: number, // portion of the previous window still "in view"
): number {
return currentWindowCount + previousWindowCount * overlapFraction;
}
// 3 requests so far this window, 5 in the previous one, 70% of it still overlaps
estimateRequests(3, 5, 0.7); // => 6.5 — compared directly against the limit
Counters belong in memory, and updates must be atomic
Durable policy — the rules themselves — can live on disk and be loaded into a rules cache by a background worker, because rules change rarely. The counters that get touched on every single request cannot afford that latency; they belong in an in-memory store such as Redis, with INCR to bump a value and EXPIRE to let the key die on its own once the window closes.
The part that is easy to get wrong is concurrency. Two requests can both read a counter at three, both compute four locally, and both write four back — the true count was five, and the limiter just quietly lost an update. Ordinary locks are too slow to sit on this path for every request, so the fix is to make the read, check and write a single atomic step: a Lua script executed inside Redis, or a sorted set keyed by timestamp so concurrent writers cannot collide on the same slot.
async function admit(ruleKey: string, now: number): Promise<AdmissionResult> {
const rule = await rulesCache.get(ruleKey);
// Single round trip: Redis runs the Lua script as one atomic operation, so no
// other request can observe a stale count between the read and the write.
const result = await redis.evalAtomicCheckAndUpdate(rule, now);
return result.allowed
? { allowed: true }
: { allowed: false, retryAfterSeconds: result.retryAfter };
}
The request path end-to-end
Splitting rule distribution from request admission keeps the hot path small. Rules flow, asynchronously, from durable storage through workers into a shared cache. Every request then follows the same short sequence: the gateway loads the applicable rule, checks and updates the shared counter atomically, and either forwards the request or returns a rejection.
What happens on rejection is itself a design choice, not a given. Returning 429 immediately is the simplest option. Dropping the request silently is rarely a good idea, because it looks identical to a network failure from the caller's side. Routing the excess into a queue for delayed processing is the right shape for work that can tolerate being late but not being lost.
One thing worth deciding explicitly, in writing, before it happens in production: what the limiter does when its own shared cache is unreachable. Fail closed and a Redis blip becomes an outage for every client. Fail open and the exact overload the limiter exists to prevent gets through during the one window it is unavailable. Neither answer is universally correct — it depends on whether the rule protects security, budget or stability — but it needs to be a decision, not a default nobody chose.
Give the client something to act on
A 429 with no other information forces the caller to guess. The fix costs nothing at the point of rejection: report how many requests remain in the window, the size of the quota, and — the field that actually changes client behaviour — how long to wait before trying again.
HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
Content-Type: application/json
{ "error": "rate_limited", "retryAfterSeconds": 30 }
Two scaling problems: races and regions
Concurrency is the first scaling problem, and it is solved by the atomic counter path already described: for a given identity and window, every admitted request has to contribute exactly once to state that every competing limiter instance can see.
Distributed synchronization is the second, larger problem. If each stateless limiter instance keeps its own local counter, a client bouncing between instances behind a load balancer effectively gets one quota per instance instead of one quota, period. Sticky sessions look like an easy fix and are not — pinning clients to instances fights the load balancer's entire purpose. The answer is to make every limiter instance stateless and let them all read and write the same shared store.
At global scale that shared store usually cannot be a single machine in a single region without adding latency to every request, so multi-region deployments route each client to its nearest edge and accept eventual consistency between regions. That is a genuine trade-off to name out loud: quotas can be enforced per region, approximately globally, or strictly globally at the cost of cross-region round trips — pick the one the traffic pattern actually needs rather than defaulting to whichever is easiest to stand up first.
Operating it: monitoring, tuning and the client contract
A rate limiter is not a feature you ship once. The useful signal is the volume and shape of 429 responses over time — a spike tells you either that traffic changed or that a rule needs revisiting, and the two require different responses. A service expecting a flash sale, for instance, might deliberately switch to a more burst-tolerant token bucket for the duration rather than eating a wave of rejections from a rule tuned for an ordinary Tuesday.
It is also worth distinguishing hard limits, which block outright once a quota is spent, from soft limits, which allow a bounded, explicit overage — often reserved for premium tiers. And it is worth remembering that Layer 7 rules at the gateway and Layer 3 IP-based rules at the network edge solve different problems and are not substitutes for each other; the first understands endpoints and identities, the second only understands addresses.
None of that operational discipline works without cooperative clients on the other end. Cache what can be cached, treat 429 as a normal, recoverable response rather than an exception to crash on, honour the Retry-After value, and back off exponentially rather than retrying immediately. A well-designed limiter and a well-behaved client are the same system wearing two hats — get either one wrong and the other has to compensate for it.