How to Build an API Rate Limiter in Python: Token Bucket and Sliding Window Algorithms
Learn how token bucket, sliding window, and fixed window rate limiting algorithms work by building each one from scratch in Python, then wiring the token bucket into a real FastAPI middleware and...
If you have ever put an API in front of the public internet, you have probably watched one client, sometimes a bug in someone’s retry loop, sometimes a scraper, sometimes just a burst of legitimate traffic, send hundreds of requests in a few seconds and knock over a backend that was fine handling that same load spread across a minute. Rate limiting is the fix: a rule, enforced by your server, that caps how many requests a single client can make in a given amount of time. Get it right and one noisy client gets a polite 429 response instead of taking your database down with it.
Table Of Content
- What You Will Build
- Prerequisites
- Step 1: Set Up Your Project
- Step 2: The Fixed Window Counter, and Its Boundary Problem
- How it works
- The boundary burst problem
- Step 3: The Sliding Window Log, an Exact Fix
- How it works
- Step 4: The Sliding Window Counter, a Cheaper Approximation
- How it works
- Step 5: The Token Bucket, for Controlled Bursts
- How it works
- Step 6: Wire the Token Bucket Into a Real FastAPI App
- Verifying the 429 and Retry-After end to end
- Step 7: The Concurrency Bug Almost Everyone Ships
- Why this is a real risk, not a theoretical one
- Reproducing it for real
- Fixing it with a lock
- Step 8: The Memory Leak Nobody Notices
- Common Mistakes to Watch For
- How to Confirm It All Works End to End
- Next Steps
This tutorial builds four rate limiting algorithms from scratch in Python: fixed window, sliding window log, sliding window counter, and token bucket. You will see exactly why the simplest one (fixed window) lets through twice its stated limit at the worst possible moment, wire the token bucket version into a real FastAPI middleware, and then find and fix two bugs that are easy to ship by accident: a genuine concurrency race condition that lets 30 requests through a hard limit of 10, and a memory leak from tracking every client forever. Every code sample below was written and personally executed against a live local server while writing this post; the output you will see is real, not invented.
What You Will Build
A small Python library, rate_limiters.py, containing four limiter classes that share the same simple interface: an allow(client_id) method that returns True or False. You will then wire the token bucket version into a FastAPI app that:
- Returns HTTP 429 with a correctly formatted
Retry-Afterheader when a client goes over its limit - Tracks each client independently, keyed by an API key or IP address
- Survives real concurrent traffic without letting more requests through than the configured limit allows
- Does not leak memory as thousands of one-off clients pass through it
Prerequisites
- A computer that can run Python 3.10 or newer (this tutorial used Python 3.13.14 on Windows, but nothing here is Windows-specific)
- Comfort with basic Python: classes, dictionaries, and what a decorator is
- Basic familiarity with HTTP status codes (this tutorial explains what you need about 429 specifically)
- No external services, databases, or API keys. Everything runs on your own machine
Step 1: Set Up Your Project
Create an isolated environment and install the two libraries this tutorial needs to run a real HTTP server and send it real requests:
python -m venv venv
venv\Scripts\activate # on Linux/macOS: source venv/bin/activate
pip install fastapi uvicorn requests
This was written against fastapi 0.141.1, uvicorn 0.52.3, and requests 2.34.2. Create a file called rate_limiters.py; every algorithm in this tutorial goes into that one file, each as its own class, so you can compare them side by side.
Every class below accepts a clock argument that defaults to time.monotonic instead of the more familiar time.time(). This is a deliberate choice, not a style preference. Python’s own documentation is explicit about why: time.monotonic() “cannot go backwards” and “is not affected by system clock updates,” while time.time() “can return a lower value than a previous call if the system clock has been set back between the two calls.” A rate limiter that measures elapsed time with a clock that can jump backward can be tricked into never expiring a window, or into expiring one instantly, just by an NTP sync adjusting the system clock. Making the clock injectable also means the tests in this tutorial can advance time deterministically instead of relying on real sleep() calls, which is how the demos below produce exact, reproducible numbers instead of flaky ones.
Step 2: The Fixed Window Counter, and Its Boundary Problem
How it works
The simplest possible rate limiter divides time into fixed blocks (say, every 10 seconds) and counts requests per client within the current block. When a new block starts, the count resets to zero.
class FixedWindowLimiter:
def __init__(self, limit: int, window_seconds: float, clock=time.monotonic):
self.limit = limit
self.window_seconds = window_seconds
self.clock = clock
self._state = {} # client_id -> [window_start, count]
def allow(self, client_id: str) -> bool:
now = self.clock()
window_start, count = self._state.get(client_id, (now, 0))
if now - window_start >= self.window_seconds:
window_start, count = now, 0
if count < self.limit:
self._state[client_id] = (window_start, count + 1)
return True
self._state[client_id] = (window_start, count)
return False
Each client’s state is just two numbers: when its current window started, and how many requests it has made in that window. If the window has expired, both reset. If the count is still under the limit, the request is allowed and the count increments.
The boundary burst problem
This is simple to implement and cheap to store (two numbers per client), which is exactly why it is the first thing people reach for. It also has a well-known flaw: a client can send its full quota right at the end of one window, then its full quota again right at the start of the next, and get double the intended rate through in a short real span of time. Save this as demo_fixed_window.py and run it with a fake, manually advanced clock so the timing is exact:
from rate_limiters import FixedWindowLimiter
class FakeClock:
def __init__(self):
self.t = 0.0
def __call__(self):
return self.t
def advance(self, seconds):
self.t += seconds
clock = FakeClock()
limiter = FixedWindowLimiter(limit=5, window_seconds=10, clock=clock)
allowed = sum(1 for _ in range(5) if limiter.allow("client-a"))
print(f"t=0.0s: 5 requests -> {allowed} allowed")
clock.advance(9.9)
result = limiter.allow("client-a")
print(f"t=9.9s: 1 more request -> allowed={result}")
clock.advance(0.2)
allowed = sum(1 for _ in range(5) if limiter.allow("client-a"))
print(f"t=10.1s: 5 requests -> {allowed} allowed")
Running it prints exactly this:
t=0.0s: 5 requests -> 5 allowed
t=9.9s: 1 more request -> allowed=False
t=10.1s: 5 requests -> 5 allowed
Look at what actually happened on the wall clock: 5 requests landed at t=9.9s (all inside window one) and another 5 landed at t=10.1s (the instant window two opened). That is 10 requests inside a 0.2 second span, against a limit that was supposed to be 5 requests per 10 seconds. The algorithm is not buggy; it is doing exactly what it was told. The problem is that “requests in the current window” is not the same question as “requests in the last 10 seconds,” and for a rate limiter, the second question is almost always the one you actually want answered.
Step 3: The Sliding Window Log, an Exact Fix
How it works
Instead of counting per fixed block, keep a timestamp for every request a client has made, and only count the ones that fall inside a rolling window ending right now. This answers the real question (“how many requests in the last N seconds”) exactly, with no boundary effect.
class SlidingWindowLogLimiter:
def __init__(self, limit: int, window_seconds: float, clock=time.monotonic):
self.limit = limit
self.window_seconds = window_seconds
self.clock = clock
self._log = {} # client_id -> deque[timestamp]
def allow(self, client_id: str) -> bool:
now = self.clock()
log = self._log.setdefault(client_id, deque())
cutoff = now - self.window_seconds
while log and log[0] <= cutoff:
log.popleft()
if len(log) < self.limit:
log.append(now)
return True
return False
On every call, timestamps older than window_seconds are dropped off the front of the deque, then the remaining count is compared against the limit. Run the same boundary scenario from Step 2 against this class instead:
from rate_limiters import SlidingWindowLogLimiter
clock = FakeClock()
limiter = SlidingWindowLogLimiter(limit=5, window_seconds=10, clock=clock)
allowed = sum(1 for _ in range(5) if limiter.allow("client-b"))
print(f"t=0.0s: 5 requests -> {allowed} allowed")
clock.advance(9.9)
allowed_next = sum(1 for _ in range(5) if limiter.allow("client-b"))
print(f"t=9.9s: 5 more requests -> {allowed_next} allowed")
clock.advance(0.2)
result = limiter.allow("client-b")
print(f"t=10.1s: 1 request -> allowed={result}")
The real output:
t=0.0s: 5 requests -> 5 allowed
t=9.9s: 5 more requests -> 0 allowed
t=10.1s: 1 request -> allowed=True
This time the 5 requests at t=9.9s are correctly rejected: the client already used its 5-request budget within the trailing 10-second window, and none of the original requests have aged out yet. Only once the original batch is more than 10 seconds old (at t=10.1s, since the batch happened at t=0.0s) does a new request get through. No double burst. The cost is what you would expect: this class stores every request timestamp for every client, which is exact but grows with traffic volume instead of staying at a fixed size per client.
Step 4: The Sliding Window Counter, a Cheaper Approximation
How it works
The sliding window counter gets most of the accuracy of the log approach while storing only two counters per client (current window and previous window), by blending them with a weight based on how far into the current window you are.
class SlidingWindowCounterLimiter:
def __init__(self, limit: int, window_seconds: float, clock=time.monotonic):
self.limit = limit
self.window_seconds = window_seconds
self.clock = clock
self._state = {} # client_id -> [current_window_start, current_count, previous_count]
def allow(self, client_id: str) -> bool:
now = self.clock()
current_start, current_count, previous_count = self._state.get(
client_id, (now, 0, 0)
)
elapsed_windows = int((now - current_start) // self.window_seconds)
if elapsed_windows >= 2:
current_start, current_count, previous_count = now, 0, 0
elif elapsed_windows == 1:
current_start = current_start + self.window_seconds
previous_count = current_count
current_count = 0
position_in_window = (now - current_start) / self.window_seconds
weight = max(0.0, 1.0 - position_in_window)
weighted_count = previous_count * weight + current_count
if weighted_count < self.limit:
self._state[client_id] = (current_start, current_count + 1, previous_count)
return True
self._state[client_id] = (current_start, current_count, previous_count)
return False
The formula in weighted_count is the whole idea: as the current window ages, less weight is given to the previous window’s count, on the assumption that requests were spread evenly across it. Test it with a slightly larger limit so the weighting is easier to see:
from rate_limiters import SlidingWindowCounterLimiter
clock = FakeClock()
limiter = SlidingWindowCounterLimiter(limit=10, window_seconds=10, clock=clock)
allowed = sum(1 for _ in range(10) if limiter.allow("client-c"))
print(f"t=0.0s (window 1): 10 requests -> {allowed} allowed, window exhausted")
clock.advance(10.0)
results = []
for i in range(10):
results.append(limiter.allow("client-c"))
clock.advance(0.05)
print(f"t=10.0-10.45s (window 2): 10 requests -> {sum(results)} allowed")
Real output:
t=0.0s (window 1): 10 requests -> 10 allowed, window exhausted
t=10.0-10.45s (window 2): 10 requests -> 1 allowed
Only 1 of the 10 requests fired in that first half second of window two gets through, and the reason is worth working through by hand because it shows exactly how the approximation behaves. At the instant window two opens (t=10.0s), the weight on the previous window’s 10 requests is exactly 1.0, so the weighted count is 10, which is not less than the limit of 10, so that very first request is denied. A tiny slice of time later the weight has ticked down just enough (to 0.995) that 10 * 0.995 = 9.95, which is under the limit, so the next request is allowed, immediately pushing the weighted count back over 10 for everything after it. Compare that to the fixed window’s behavior in Step 2, where a full second batch of 5 got through all at once: the sliding window counter still lets a little through right at the seam, but nowhere near double the limit, and it never has to store more than two numbers per client to do it.
Step 5: The Token Bucket, for Controlled Bursts
How it works
The first three algorithms all treat any burst as a problem to prevent. Sometimes a burst is fine, even desirable: a client that has been idle for a while should be able to send a handful of requests immediately instead of being throttled to a strict steady drip. The token bucket models this directly: each client has a bucket holding up to capacity tokens, tokens refill continuously at refill_rate_per_sec, and every request consumes one token. No tokens, no request.
class TokenBucketLimiter:
def __init__(self, capacity: int, refill_rate_per_sec: float, clock=time.monotonic):
self.capacity = capacity
self.refill_rate_per_sec = refill_rate_per_sec
self.clock = clock
self._state = {} # client_id -> [tokens, last_refill]
def allow(self, client_id: str) -> bool:
now = self.clock()
tokens, last_refill = self._state.get(client_id, (float(self.capacity), now))
elapsed = now - last_refill
tokens = min(self.capacity, tokens + elapsed * self.refill_rate_per_sec)
if tokens >= 1:
tokens -= 1
self._state[client_id] = (tokens, now)
return True
self._state[client_id] = (tokens, now)
return False
This is also the algorithm Stripe’s own API documentation recommends to clients that need to control their request rate: “A common technique for controlling API usage is to implement a client-side token bucket rate-limiting algorithm.” Test the burst-then-refill behavior:
from rate_limiters import TokenBucketLimiter
clock = FakeClock()
limiter = TokenBucketLimiter(capacity=5, refill_rate_per_sec=1, clock=clock)
allowed = sum(1 for _ in range(7) if limiter.allow("client-d"))
print(f"t=0.0s: bucket starts full (5 tokens), 7 requests fired instantly -> {allowed} allowed")
clock.advance(1.0)
result = limiter.allow("client-d")
print(f"t=1.0s (+1 token): 1 request -> allowed={result}")
clock.advance(0.5)
result = limiter.allow("client-d")
print(f"t=1.5s (+0.5 token, below 1): 1 request -> allowed={result}")
clock.advance(10.0)
allowed = sum(1 for _ in range(8) if limiter.allow("client-d"))
print(f"t=11.5s (long idle, bucket refilled and capped at 5): 8 requests -> {allowed} allowed")
Real output:
t=0.0s: bucket starts full (5 tokens), 7 requests fired instantly -> 5 allowed
t=1.0s (+1 token): 1 request -> allowed=True
t=1.5s (+0.5 token, below 1): 1 request -> allowed=False
t=11.5s (long idle, bucket refilled and capped at 5): 8 requests -> 5 allowed
The bucket allows an immediate burst of exactly 5 (its capacity), then requires waiting for a fresh token before the next request goes through, and even after 10 idle seconds it caps back out at 5 rather than accumulating forever. This is the algorithm this tutorial wires into a real server, because it is the one most production APIs actually mean when they describe a “requests per second” limit that still tolerates a short burst.
Step 6: Wire the Token Bucket Into a Real FastAPI App
Create app.py in the same folder. Start with the imports, config, and a per-client key function:
import math
import time
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
from rate_limiters import (
TokenBucketLimiter,
UnsafeConcurrentTokenBucketLimiter,
ThreadSafeTokenBucketLimiter,
)
CAPACITY = 10
REFILL_RATE_PER_SEC = 2 # 2 tokens/sec = 10 requests per 5s sustained, burst up to 10
app = FastAPI()
safe_limiter = ThreadSafeTokenBucketLimiter(capacity=CAPACITY, refill_rate_per_sec=REFILL_RATE_PER_SEC)
unsafe_limiter = UnsafeConcurrentTokenBucketLimiter(capacity=CAPACITY, refill_rate_per_sec=REFILL_RATE_PER_SEC)
def client_key(request: Request) -> str:
api_key = request.headers.get("x-api-key")
return api_key or (request.client.host if request.client else "unknown")
Each client is identified by an X-API-Key header if present, falling back to the caller’s IP address. Now add the shared response logic that both protected routes will use:
def respond(limiter, request: Request):
client_id = client_key(request)
allowed = limiter.allow(client_id)
if not allowed:
tokens, _ = limiter._state.get(client_id, (0, time.monotonic()))
seconds_needed = max(0.0, (1 - tokens) / REFILL_RATE_PER_SEC)
# RFC 9110 defines the Retry-After delay-seconds value as a
# non-negative decimal *integer*, so round up rather than send a
# fractional value like "0.37".
retry_after_seconds = math.ceil(seconds_needed)
return JSONResponse(
status_code=429,
content={"detail": "Rate limit exceeded"},
headers={"Retry-After": str(retry_after_seconds)},
)
return {"data": "here is your response", "client_id": client_id}
When a client is over its limit, the response is a 429 with a Retry-After header telling the caller exactly how long to wait. Getting this header’s format right matters more than it looks: MDN’s documentation of the header defines the delay-seconds form as “a non-negative decimal integer indicating the seconds to delay after the response is received,” not a fractional value. An earlier version of this code sent Retry-After: 0.37, which does not match that format even though most HTTP clients will not complain; the fix is math.ceil() before formatting the header, rounding up so the client never retries a moment too early. Finally, the routes:
@app.get("/health")
def health():
return {"status": "ok"}
@app.get("/api/data")
def get_data(request: Request):
# Deliberately synchronous ("def", not "async def") so FastAPI/Starlette
# runs each call in its worker thread pool, giving us real OS-thread
# concurrency, the same concurrency a production deployment gets.
return respond(safe_limiter, request)
@app.get("/api/data-unsafe")
def get_data_unsafe(request: Request):
return respond(unsafe_limiter, request)
@app.get("/api/data-safe")
def get_data_safe(request: Request):
return respond(safe_limiter, request)
Notice these are defined with def, not async def. That is deliberate, and it matters for Step 7. Start the server:
uvicorn app:app --host 127.0.0.1 --port 8123
Confirm it is up:
curl http://127.0.0.1:8123/health
{"status":"ok"}
Verifying the 429 and Retry-After end to end
The bucket in this app has capacity=10 and refill_rate_per_sec=2. Drain it for a fresh client, then check what comes back:
import requests, time
BASE = "http://127.0.0.1:8123"
key = "retry-after-demo-client"
codes = [requests.get(f"{BASE}/api/data-safe", headers={"x-api-key": key}).status_code for _ in range(10)]
print("First 10 requests:", codes)
r = requests.get(f"{BASE}/api/data-safe", headers={"x-api-key": key})
print("11th request:", r.status_code, r.json())
retry_after = int(r.headers["Retry-After"])
print("Retry-After header (seconds):", retry_after)
time.sleep(retry_after + 0.05)
r = requests.get(f"{BASE}/api/data-safe", headers={"x-api-key": key})
print("Request after waiting:", r.status_code, r.json())
Real captured output from running this against the live server:
First 10 requests: [200, 200, 200, 200, 200, 200, 200, 200, 200, 200]
11th request: 429 {'detail': 'Rate limit exceeded'}
Retry-After header (seconds): 1
Request after waiting: 200 {'data': 'here is your response', 'client_id': 'retry-after-demo-client'}
Exactly 10 requests succeed, the 11th is rejected with a valid integer Retry-After, and waiting that long is enough (and only just enough) to get a token back. This is the core loop working end to end. Now break it on purpose.
Step 7: The Concurrency Bug Almost Everyone Ships
Why this is a real risk, not a theoretical one
Starlette (the framework FastAPI is built on) documents this directly. Defining a route with a plain def instead of async def is one of the cases it calls out by name: “Starlette will run your code in a thread pool to avoid blocking the event loop,” using anyio.to_thread.run_sync under the hood. The docs are just as direct about the size of that pool: “The default thread pool size is only 40 tokens. This means that only 40 threads can run at the same time.” That means the /api/data routes in this tutorial are not just conceptually concurrent, they run on real OS threads, at the same time, sharing the same Python dictionary inside the limiter. If reading that shared state and writing it back is not atomic, two threads can both read that 9 tokens are left, both decide to allow the request, and both write back 8, when the correct answer after two requests is 7.
Reproducing it for real
Add this variant to rate_limiters.py, which behaves identically to TokenBucketLimiter except for a deliberate gap between reading the shared state and writing it back:
class UnsafeConcurrentTokenBucketLimiter(TokenBucketLimiter):
"""Same logic as TokenBucketLimiter, but with a deliberate gap between
reading the shared state and writing it back. This stands in for real
checkpoint-then-commit patterns, such as a rate limiter backed by Redis
that does a GET then a separate SET instead of one atomic operation.
Do not use in production; this class exists to reproduce a real race.
"""
def allow(self, client_id: str) -> bool:
now = self.clock()
tokens, last_refill = self._state.get(client_id, (float(self.capacity), now))
elapsed = now - last_refill
tokens = min(self.capacity, tokens + elapsed * self.refill_rate_per_sec)
decision = tokens >= 1
if decision:
tokens -= 1
time.sleep(0.005) # the gap: another thread can read stale state here
self._state[client_id] = (tokens, now)
return decision
That gap stands in for real work that commonly sits between a check and a commit, a database round trip, a log write, or (very commonly in production) a rate limiter backed by Redis that does a GET to read the current count and a separate SET to write it back, instead of one atomic operation. Wire it into the app behind a third route (/api/data-unsafe in the companion code for this tutorial, using this class instead of the safe one) and fire 30 truly concurrent requests at it from 30 worker threads, against a bucket with capacity=10:
from concurrent.futures import ThreadPoolExecutor, as_completed
import requests
def fire_one(i):
r = requests.get("http://127.0.0.1:8123/api/data-unsafe", headers={"x-api-key": "race-test-client"})
return r.status_code
with ThreadPoolExecutor(max_workers=30) as pool:
codes = [f.result() for f in as_completed([pool.submit(fire_one, i) for i in range(30)])]
print("200 OK:", codes.count(200), " 429:", codes.count(429), " (capacity is 10)")
Real output from running this against the live server:
200 OK: 30 429: 0 (capacity is 10)
All 30 requests succeeded against a hard capacity of 10. This is not a rare, hard-to-trigger edge case; under real concurrent load it reproduces every single time, because every one of those 30 threads reads the bucket as full before any of them finish writing their update back. A rate limiter that does this in production is not rate limiting at all under load, which is exactly the condition it exists to handle.
Fixing it with a lock
The fix is to make the read-compute-write sequence atomic, so no other thread can read the state while one thread is in the middle of updating it:
class ThreadSafeTokenBucketLimiter(TokenBucketLimiter):
"""TokenBucketLimiter with the check-and-update wrapped in a lock, so the
read-compute-write sequence is atomic across threads."""
def __init__(self, capacity: int, refill_rate_per_sec: float, clock=time.monotonic):
super().__init__(capacity, refill_rate_per_sec, clock)
self._lock = threading.Lock()
def allow(self, client_id: str) -> bool:
with self._lock:
return super().allow(client_id)
A threading.Lock around the whole allow() call means only one thread can be inside it at a time; every other thread blocks until it releases. Point the exact same 30-worker burst test at the locked version instead:
200 OK: 10 429: 20 (capacity is 10)
Exactly 10 requests succeed, the other 20 are correctly rejected, under the identical concurrent load that let all 30 through a moment ago. The fix is four lines (a lock in __init__, a with self._lock: around the call), and it is the difference between a rate limiter that works and one that only appears to work in single-threaded manual testing.
Step 8: The Memory Leak Nobody Notices
Every limiter class in this tutorial stores one dictionary entry per client it has ever seen, and never removes anything. That is fine in a quick test with a handful of client IDs. It is a real problem in a long-running production process fielding traffic from thousands of distinct IPs or API keys, most of which will only ever be seen once. Demonstrate it directly, reusing the same FakeClock helper class from Step 2:
from rate_limiters import TokenBucketLimiter
clock = FakeClock()
limiter = TokenBucketLimiter(capacity=10, refill_rate_per_sec=2, clock=clock)
for i in range(5000):
limiter.allow(f"client-{i}")
clock.advance(0.01)
print("After 5,000 unique clients:", len(limiter._state))
After 5,000 unique clients: 5000
Every one of those 5,000 entries sits in memory forever, even though almost none of those clients will ever come back. Left running, this dictionary grows without bound. The fix does not need to be complicated: a client that has been idle long enough to have fully refilled its bucket can simply be dropped, since a missing client_id is already treated as a fresh, full bucket the next time allow() sees it.
def sweep_idle(self, idle_seconds: float) -> int:
"""Drop entries that have been idle long enough to have fully
refilled; dropping them is equivalent to keeping them, since a
missing client_id is already treated as a full bucket by allow().
Returns the number of entries removed.
"""
now = self.clock()
to_remove = []
for client_id, (tokens, last_refill) in self._state.items():
idle_for = now - last_refill
if idle_for >= idle_seconds and tokens >= self.capacity - 1e-9:
to_remove.append(client_id)
elif idle_for >= idle_seconds:
# Idle long enough that it *would* be full even though the
# stored value is stale; recompute before deciding.
projected = min(self.capacity, tokens + idle_for * self.refill_rate_per_sec)
if projected >= self.capacity - 1e-9:
to_remove.append(client_id)
for client_id in to_remove:
del self._state[client_id]
return len(to_remove)
Advance the fake clock well past the last client’s activity (standing in for real idle time passing in a running process), then sweep:
clock.advance(120)
removed = limiter.sweep_idle(idle_seconds=60)
print("removed:", removed, " remaining:", len(limiter._state))
removed: 5000 remaining: 0
In production you would call sweep_idle() on a periodic background timer (once a minute is plenty for most APIs) rather than after every request, since the sweep itself has to walk the whole dictionary.
Common Mistakes to Watch For
A short recap of the gotchas this tutorial walked through, since they are easy to reintroduce later even after seeing them fixed once:
- Using
time.time()instead oftime.monotonic(). As Step 1 covered, Python’s own documentation warns that wall-clock time can jump backward if the system clock is adjusted, which can make a rate limiter’s window math briefly wrong right when that adjustment happens. - Choosing fixed window when bursts at the boundary actually matter to you. It is the cheapest option and fine for coarse, forgiving limits; it is the wrong choice anywhere a hard cap needs to be a hard cap.
- Forgetting that synchronous FastAPI routes run on real threads. Any shared, mutable state they touch (a dictionary, a counter, a cache) needs the same thread-safety discipline you would apply to any other multithreaded code.
- Sending a fractional
Retry-After. The header’s delay-seconds form is specified as a non-negative integer; round up, do not truncate or send a float. - Never evicting idle clients. A per-client dictionary with no cleanup path is a slow memory leak that will not show up until the process has been running against real traffic for a while.
How to Confirm It All Works End to End
With the server running, this short checklist confirms every piece from this tutorial is working together correctly:
- Run
curl http://127.0.0.1:8123/healthand confirm you get back{"status":"ok"}, proving the server is reachable. - Send 10 requests to
/api/data-safewith a freshX-API-Keyand confirm all 10 return200. - Send an 11th request with the same key and confirm you get
429with an integerRetry-Afterheader. - Wait that many seconds and confirm the next request succeeds again with
200. - Fire 30 concurrent requests at
/api/data-safewith a fresh key from a thread pool and confirm exactly 10 succeed, never more.
If every step matches, the limiter is correctly enforcing its limit under both normal sequential traffic and real concurrent load, which is the actual bar a production rate limiter has to clear.
Next Steps
Everything in this tutorial runs in a single process, which means the moment you scale to more than one server instance, each instance has its own separate in-memory dictionary and its own separate idea of how many tokens a client has left. A client hitting a load balancer in front of 3 identical instances can get roughly 3 times its intended limit through, simply by getting spread across all 3. Solving that requires moving the state itself somewhere shared, typically Redis, and the exact race condition from Step 7 reappears in a new form: a naive Redis-backed limiter that does a GET then a separate SET has precisely the same read-compute-write gap, just spread across a network round trip instead of a thread scheduler. The standard fix there is the same idea in a different form, an atomic Redis operation (an INCR with an expiry, or a small Lua script executed atomically on the Redis server) instead of two separate commands. For a production Python service, reaching for a maintained library such as limits or, for FastAPI specifically, slowapi is almost always the right call over hand-rolling this in production; building it from scratch, as this tutorial did, is what makes the trade-offs and the failure modes concrete enough to evaluate those libraries (or debug them) with real understanding instead of guesswork.








No Comment! Be the first one.