AmitSingh
All posts

Rate limiting at scale: an architecture, not an algorithm

A production rate limiter is not a counter bolted onto an endpoint. It is a distributed admission-control system that has to be accurate, fast, synchronized and honest with the clients it throttles.

Amit Shivpratap Singh • • 10 min read

Rate limiter policy engine sitting at an API gateway: incoming requests pass through authentication, an IP allow-list and a hybrid rate limiter before a decision engine returns either a 200 OK or a 429 Too Many Requests with a Retry-After value.
On this page

Why traffic control is an architecture decision

It is tempting to treat rate limiting as a single counter and an if-statement: has this key made too many requests, yes or no. That view survives exactly until the service runs on more than one instance, and then it quietly stops being true. A production limiter is doing three jobs at once — a security job, an economics job and a stability job — and each one shapes a different part of the design.

The security job is to prevent resource starvation, whether it comes from a genuine denial-of-service attempt or from a misbehaving retry loop in someone else's code. The economics job is to keep a lid on anything that costs money per call — a payment gateway, a metered third-party API, a GPU inference endpoint — so a bug three layers upstream cannot turn into an invoice. The stability job is the most familiar one: filter excess load before it reaches the application tier, so response time and queue depth stay inside a range the system can actually sustain.

None of those three jobs is optional in a system that matters, which is why the limiter belongs in the architecture diagram next to the load balancer and the database, not tucked away as a utility function inside one controller.

A baseline before choosing anything

Before comparing algorithms or picking a data store, it helps to fix what "good" means. A useful baseline has six parts: enforcement lives on the server, not the client; accuracy is strict enough that the limiter rarely admits traffic above policy; latency overhead is small enough that the check never becomes the slowest part of the request; memory use per identity stays tiny even at millions of keys; state is shared across every server and microservice instance rather than local to one process; and a throttled request gets an explicit, machine-readable rejection instead of a silent drop.

Those six requirements pull in the same direction more than they conflict. Ultra-low latency argues for in-memory state. Multi-instance scale means that state has to be shared rather than private. Strict accuracy means the read-check-write sequence has to be atomic. And clear rejection means the response has to carry enough information for the caller to behave intelligently instead of hammering the endpoint again immediately. Every later design choice is really just satisfying this list.

Where the limiter should live

Client-side throttling is worth building as a courtesy — it reduces avoidable calls — but it can never be the authoritative control. A client is code you do not run, on a device you do not control, and it can be patched, forked or ignored. Treating it as the enforcement boundary is treating a suggestion as a lock.

Embedding the check directly in each API server is workable, but it ties the implementation to whatever stack that service happens to run, and it means re-solving the same problem in every service that needs it. The placement that scales cleanly is the API gateway or middleware tier, sitting in front of every downstream service. It becomes one policy boundary that also handles authentication, TLS termination and IP allow-listing, and it is the one place a rule change reaches every caller at once.

Three placement options for a rate limiter — client-side, server-side and API gateway/middleware — with the gateway tier flagged as the recommended enforcement point.
Client-side throttling is cooperative at best; the gateway tier is where enforcement actually holds.

Five algorithms, one decision

A token bucket holds a capped number of tokens that refill at a steady rate; every admitted request spends one. Because unused tokens accumulate up to the bucket size, it can absorb a short burst without loosening the long-term rate — the shape most APIs actually want from their traffic.

A leaking bucket is the same idea turned around: requests can arrive in any pattern, but they leave through a fixed-rate queue. It is excellent for smoothing bursty input into a predictable outflow for asynchronous workers, and correspondingly bad at ever letting a legitimate burst through faster than that fixed rate.

A fixed window counter is the cheapest option: one counter per identity, reset every interval. Its flaw shows up exactly at the boundary — five requests in the last second of one window plus five more in the first second of the next both pass, and a rolling one-second view now shows ten requests against a limit of five.

A sliding window log fixes that by recording the timestamp of every request and counting how many fall inside the trailing window. It is exact, and it is expensive: the log keeps growing with traffic, including the requests it eventually rejects.

A sliding window counter is the practical middle ground: it blends the count in the current fixed window with a weighted slice of the previous one, using the fraction of the previous window that still overlaps the trailing interval. It gets most of the sliding log's accuracy for a fraction of its memory, which is why it has become the default recommendation for anything that is not either trivially low-stakes or strictly exact-or-nothing.

Four rate limiting algorithms compared side by side — token bucket, leaking bucket, fixed window and sliding window counter — with their best-fit use case.
Capacity algorithms shape bursts; time-window algorithms measure quota consumption.

The hybrid trade-off, in one picture

Put a token bucket and a sliding window counter side by side and the reason production systems combine them becomes obvious: the bucket answers "is this instant allowed?" and the counter answers "has this identity stayed inside its long-term rate?" Neither question is the whole story on its own. A request is admitted only when it clears both checks, which is what lets a hybrid limiter tolerate a legitimate burst while still holding a hard line on sustained abuse.

The counter side of that hybrid is arithmetic, not guesswork, which is exactly why it is cheap to compute per request.

A hybrid rate limiter combining a token bucket for burst capacity with a sliding window counter for long-term rate, feeding a single allow/reject decision engine.
Token bucket for the instant, sliding window counter for the trend — admitted only when both agree.
typescript
function estimateRequests(
  currentWindowCount: number,
  previousWindowCount: number,
  overlapFraction: number, // portion of the previous window still "in view"
): number {
  return currentWindowCount + previousWindowCount * overlapFraction;
}

// 3 requests so far this window, 5 in the previous one, 70% of it still overlaps
estimateRequests(3, 5, 0.7); // => 6.5 — compared directly against the limit
Estimating the rolling count without storing every timestamp.

Counters belong in memory, and updates must be atomic

Durable policy — the rules themselves — can live on disk and be loaded into a rules cache by a background worker, because rules change rarely. The counters that get touched on every single request cannot afford that latency; they belong in an in-memory store such as Redis, with INCR to bump a value and EXPIRE to let the key die on its own once the window closes.

The part that is easy to get wrong is concurrency. Two requests can both read a counter at three, both compute four locally, and both write four back — the true count was five, and the limiter just quietly lost an update. Ordinary locks are too slow to sit on this path for every request, so the fix is to make the read, check and write a single atomic step: a Lua script executed inside Redis, or a sorted set keyed by timestamp so concurrent writers cannot collide on the same slot.

typescript
async function admit(ruleKey: string, now: number): Promise<AdmissionResult> {
  const rule = await rulesCache.get(ruleKey);

  // Single round trip: Redis runs the Lua script as one atomic operation, so no
  // other request can observe a stale count between the read and the write.
  const result = await redis.evalAtomicCheckAndUpdate(rule, now);

  return result.allowed
    ? { allowed: true }
    : { allowed: false, retryAfterSeconds: result.retryAfter };
}
The request path reduces to one atomic decision, not a read-then-write.

The request path end-to-end

Splitting rule distribution from request admission keeps the hot path small. Rules flow, asynchronously, from durable storage through workers into a shared cache. Every request then follows the same short sequence: the gateway loads the applicable rule, checks and updates the shared counter atomically, and either forwards the request or returns a rejection.

What happens on rejection is itself a design choice, not a given. Returning 429 immediately is the simplest option. Dropping the request silently is rarely a good idea, because it looks identical to a network failure from the caller's side. Routing the excess into a queue for delayed processing is the right shape for work that can tolerate being late but not being lost.

One thing worth deciding explicitly, in writing, before it happens in production: what the limiter does when its own shared cache is unreachable. Fail closed and a Redis blip becomes an outage for every client. Fail open and the exact overload the limiter exists to prevent gets through during the one window it is unavailable. Neither answer is universally correct — it depends on whether the rule protects security, budget or stability — but it needs to be a decision, not a default nobody chose.

End-to-end architecture: clients through a load balancer to an API gateway with an embedded rate limiter, backed by a Redis counter store, fanning out to API servers and downstream microservices.
Rule distribution and request admission are two paths — only the second one has to be fast.

Give the client something to act on

A 429 with no other information forces the caller to guess. The fix costs nothing at the point of rejection: report how many requests remain in the window, the size of the quota, and — the field that actually changes client behaviour — how long to wait before trying again.

http
HTTP/1.1 429 Too Many Requests
Retry-After: 30
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
Content-Type: application/json

{ "error": "rate_limited", "retryAfterSeconds": 30 }
What the caller needs is a limit, a remaining count and a wait time.

Two scaling problems: races and regions

Concurrency is the first scaling problem, and it is solved by the atomic counter path already described: for a given identity and window, every admitted request has to contribute exactly once to state that every competing limiter instance can see.

Distributed synchronization is the second, larger problem. If each stateless limiter instance keeps its own local counter, a client bouncing between instances behind a load balancer effectively gets one quota per instance instead of one quota, period. Sticky sessions look like an easy fix and are not — pinning clients to instances fights the load balancer's entire purpose. The answer is to make every limiter instance stateless and let them all read and write the same shared store.

At global scale that shared store usually cannot be a single machine in a single region without adding latency to every request, so multi-region deployments route each client to its nearest edge and accept eventual consistency between regions. That is a genuine trade-off to name out loud: quotas can be enforced per region, approximately globally, or strictly globally at the cost of cross-region round trips — pick the one the traffic pattern actually needs rather than defaulting to whichever is easiest to stand up first.

Multiple stateless API server instances synchronizing through a shared Redis counter store, with allow, reject and delay/queue outcomes feeding monitoring and alerting.
Stateless limiter instances, one shared source of truth — the counter store is the only thing that has to agree.

Operating it: monitoring, tuning and the client contract

A rate limiter is not a feature you ship once. The useful signal is the volume and shape of 429 responses over time — a spike tells you either that traffic changed or that a rule needs revisiting, and the two require different responses. A service expecting a flash sale, for instance, might deliberately switch to a more burst-tolerant token bucket for the duration rather than eating a wave of rejections from a rule tuned for an ordinary Tuesday.

It is also worth distinguishing hard limits, which block outright once a quota is spent, from soft limits, which allow a bounded, explicit overage — often reserved for premium tiers. And it is worth remembering that Layer 7 rules at the gateway and Layer 3 IP-based rules at the network edge solve different problems and are not substitutes for each other; the first understands endpoints and identities, the second only understands addresses.

None of that operational discipline works without cooperative clients on the other end. Cache what can be cached, treat 429 as a normal, recoverable response rather than an exception to crash on, honour the Retry-After value, and back off exponentially rather than retrying immediately. A well-designed limiter and a well-behaved client are the same system wearing two hats — get either one wrong and the other has to compensate for it.

Share LinkedIn X Email

Keep reading