How to Prevent Duplicate Charges in an API With Idempotency Keys in Python
A hands-on Python and FastAPI tutorial that reproduces a real duplicate-charge bug caused by a client timeout, then fixes it with a spec-compliant idempotency key implementation that safely handles...
Say your API has an endpoint that charges a customer’s card, creates an order, or sends money. A client calls it, the network hiccups, and the client never sees a response. From the client’s point of view, this is completely ambiguous: maybe the request never reached the server, maybe the server processed it and the response got lost on the way back. The only thing the client can safely conclude is “I don’t know.” So it does the only reasonable thing: it retries.
Table Of Content
- What You Will Build
- Prerequisites
- Step 1: See the Problem: Reproduce a Duplicate Charge
- Set Up Your Environment
- Build a Naive Charge API
- Reproduce the Bug
- Step 2: Understand the Idempotency-Key Header
- Step 3: Build an Idempotency Store
- Step 4: Build a Safe, Idempotent Charge Endpoint
- Step 5: Prove the Retry Fix Works
- Step 6: Reject Key Reuse With a Different Payload
- Step 7: Handle True Concurrency Safely
- The Racy Way: Check, Then Insert
- Fire Two Concurrent Requests at Each Version
- Step 8: Expire Old Idempotency Keys
- Common Mistakes and Gotchas
- How to Verify Your Own Implementation End to End
- Next Steps
If your server has no way to recognize “this is the same logical request I already handled,” that retry creates a second charge, a second order, or a second payment. This is not a rare edge case. Timeouts happen constantly in production: a slow database, a garbage collection pause, a flaky load balancer, a mobile client losing signal for two seconds. Any of these can turn one honest retry into a duplicate side effect that a real customer notices on their bank statement.
An idempotency key is the fix. The client generates a unique token for each logical operation (not each HTTP attempt) and sends it on every attempt of that operation, including retries. The server remembers which keys it has already handled and what it returned, so a retry with the same key gets back the exact same result instead of running the operation again. “Idempotent” here comes from mathematics: an idempotent operation produces the same result no matter how many times you apply it. Per RFC 9110, the methods OPTIONS, HEAD, GET, PUT, and DELETE are idempotent by definition in HTTP, while POST and PATCH are not, which is exactly why this technique exists: it is a way to bolt idempotency onto an operation that is not naturally idempotent.
This tutorial builds the whole thing from scratch in Python: a naive charge API that you will watch double-charge a customer for real, then a fixed version that correctly deduplicates retries, rejects a key that gets reused with different data, and survives two genuinely concurrent requests without either duplicating the charge or crashing. Every command below was run for real while writing this post, and the output you see is the actual captured output, not invented numbers.
What You Will Build
- A naive
POST /chargesendpoint with no idempotency handling, and a client script that reproduces a genuine double charge caused by a client-side timeout - An idempotency store (a small SQLite table) that records every key you have seen, its status, and the response you gave it
- A safe
POST /chargesendpoint that deduplicates retries, rejects a key reused with a different payload, and answers a truly concurrent duplicate request with a clean409 Conflictinstead of a race - A side-by-side demonstration of a naive “check, then insert” implementation racing under real concurrent load, compared against the safe version handling the exact same race correctly
Prerequisites
- Python 3.11 or newer (this tutorial used Python 3.13.14 on Windows, but nothing here is platform-specific)
- Comfort with basic Python and the shape of a REST API: what a
POSTrequest and a JSON body are - A terminal and the ability to run
pip install
You do not need any prior experience with FastAPI specifically. Every route handler here is explained as it appears.
Step 1: See the Problem: Reproduce a Duplicate Charge
Set Up Your Environment
Create a fresh project folder and a virtual environment, then install the three libraries this tutorial uses: a web framework, a server to run it, and an HTTP client to test it.
mkdir idempotency-demo
cd idempotency-demo
python -m venv venv
venv\Scripts\activate # on Linux/macOS: source venv/bin/activate
pip install fastapi uvicorn httpx
This tutorial was written against fastapi 0.141.1, uvicorn 0.52.4, and httpx 0.28.1. Newer versions should work the same way; none of the APIs used here are new or unstable.
Build a Naive Charge API
Save this as step1_naive_server.py. It is a small FastAPI app with one endpoint, POST /charges, backed by a SQLite table. To stand in for a real payment processor, the handler sleeps for two seconds before finishing, since a real card-network call is exactly this kind of slow, blocking operation.
import sqlite3
import time
import uuid
from datetime import datetime, timezone
from fastapi import FastAPI
from pydantic import BaseModel
DB_PATH = "naive_charges.db"
app = FastAPI()
def get_db():
conn = sqlite3.connect(DB_PATH)
conn.execute(
"""
CREATE TABLE IF NOT EXISTS charges (
charge_id TEXT PRIMARY KEY,
customer_id TEXT NOT NULL,
amount_cents INTEGER NOT NULL,
currency TEXT NOT NULL,
created_at TEXT NOT NULL
)
"""
)
return conn
class ChargeRequest(BaseModel):
customer_id: str
amount_cents: int
currency: str
@app.post("/charges", status_code=201)
def create_charge(req: ChargeRequest):
# Simulate a slow downstream call, like reaching out to a card network.
time.sleep(2)
charge_id = f"ch_{uuid.uuid4().hex[:16]}"
created_at = datetime.now(timezone.utc).isoformat()
conn = get_db()
conn.execute(
"INSERT INTO charges (charge_id, customer_id, amount_cents, currency, created_at) "
"VALUES (?, ?, ?, ?, ?)",
(charge_id, req.customer_id, req.amount_cents, req.currency, created_at),
)
conn.commit()
conn.close()
return {
"charge_id": charge_id,
"customer_id": req.customer_id,
"amount_cents": req.amount_cents,
"currency": req.currency,
"created_at": created_at,
}
Run it:
uvicorn step1_naive_server:app --port 8001
Reproduce the Bug
Now write a client that does exactly what a real mobile app or browser does when a request hangs: it gives up waiting after a short timeout, and because it genuinely does not know whether the charge went through, it retries. Save this as step1_client_retry_demo.py and run it in a second terminal while the server from the previous step is still running.
import sqlite3
import time
import httpx
DB_PATH = "naive_charges.db"
URL = "http://127.0.0.1:8001/charges"
payload = {"customer_id": "cus_amir_001", "amount_cents": 4999, "currency": "usd"}
print("Attempt 1: sending charge request with a 1-second client timeout...")
try:
with httpx.Client(timeout=1.0) as client:
resp = client.post(URL, json=payload)
print("Attempt 1 response:", resp.status_code, resp.json())
except httpx.ReadTimeout:
print("Attempt 1 timed out on the client side after 1s.")
print("The client has NO IDEA whether the server actually processed this charge.")
print("\nWaiting 3 seconds (simulating the user or app retry logic)...")
time.sleep(3)
print("\nAttempt 2: retrying the exact same charge request, now with a 5-second timeout...")
with httpx.Client(timeout=5.0) as client:
resp = client.post(URL, json=payload)
print("Attempt 2 response:", resp.status_code, resp.json())
print("\nChecking the charges table directly...")
conn = sqlite3.connect(DB_PATH)
rows = conn.execute(
"SELECT charge_id, customer_id, amount_cents, created_at FROM charges"
).fetchall()
conn.close()
print(f"Found {len(rows)} charge(s) in the database for this customer:")
for row in rows:
print(" ", row)
Here is what actually happened when this ran:
Attempt 1: sending charge request with a 1-second client timeout...
Attempt 1 timed out on the client side after 1s.
The client has NO IDEA whether the server actually processed this charge.
Waiting 3 seconds (simulating the user or app retry logic)...
Attempt 2: retrying the exact same charge request, now with a 5-second timeout...
Attempt 2 response: 201 {'charge_id': 'ch_d6328ddb72d24b26', 'customer_id': 'cus_amir_001', 'amount_cents': 4999, 'currency': 'usd', 'created_at': '2026-08-20T19:58:30.568586+00:00'}
Checking the charges table directly...
Found 2 charge(s) in the database for this customer:
('ch_a233078707a84282', 'cus_amir_001', 4999, '2026-08-20T19:58:26.292299+00:00')
('ch_d6328ddb72d24b26', 'cus_amir_001', 4999, '2026-08-20T19:58:30.568586+00:00')
Walk through what just happened. The client’s first attempt gave up after one second because httpx.Client(timeout=1.0) stopped waiting. But the server never received that cancellation. It kept executing the handler in the background: it slept for its full two seconds, inserted a charge, and finished, all while the client had already moved on believing nothing happened. Three seconds later, the client retried with the exact same customer and amount, and the naive server, having no memory of the first request, dutifully processed it as a brand-new charge. The result: two $49.99 charges in the database for what the client experienced as a single logical action.
This is not a bug in httpx or FastAPI. It is the fundamental ambiguity of a network request that does not come back in time: the client cannot distinguish “the server never got this” from “the server got this and finished, but I didn’t hear back.” Any code that reacts to a timeout by retrying a POST has this problem unless something on the server side is explicitly designed to catch it.
Stop the server with Ctrl+C before moving on.
Step 2: Understand the Idempotency-Key Header
The fix is for the client to generate a unique token, an idempotency key, once per logical operation, and send it on every attempt of that operation (the original request and every retry) using an Idempotency-Key HTTP header. The server’s job is to remember which keys it has already seen and reuse the original result instead of repeating the side effect.
This is not a proprietary trick. It is specified in an IETF Internet-Draft, The Idempotency-Key HTTP Header Field (draft-ietf-httpapi-idempotency-key-header-07, published by the HTTPAPI working group). It is worth being precise about what that status means: an Internet-Draft is a work in progress, not a finished standard. This particular draft is marked “Intended Status: Standards Track” but its own header states it expired on 18 April 2026 without being published as an RFC, which is a routine, unremarkable outcome for IETF drafts (they lapse automatically after six months without an update; many are later revived or superseded rather than truly abandoned). Treat it the way the draft itself asks to be treated: as a well-reasoned, widely referenced description of existing practice, not a ratified standard you are contractually bound to.
The draft’s own introduction describes precisely the scenario you just reproduced: “the client sent a POST request to the server, but the request timed out… it doesn’t know if the resource was created or updated, or if the server even completed processing the request… the client does not know if it can safely retry the request.” A few specific rules from the draft are worth calling out because you will implement all of them below:
- The key is a client-generated string, and the draft recommends using a UUID or similarly random identifier so that two unrelated requests never collide by accident.
- A key must never be reused across two requests with a different payload. The draft calls the hash used to detect this an idempotency fingerprint.
- The draft’s enforcement model has exactly three cases, and each gets a different response: a first-time key is processed normally; a retry (the same key and fingerprint, seen again after the original request already finished) gets back the original result; a concurrent request (the same key and fingerprint, seen again before the original request has finished) gets a
409 Conflict, not a second execution. - Reusing a key with a different payload should get a
422 Unprocessable Content. A missing key on an endpoint that requires one should get a400 Bad Request.
You will see this pattern in the wild too. Stripe’s API, one of the most widely used idempotency-key implementations in production, accepts an Idempotency-Key header on write operations, caches the response for 24 hours, and marks a replayed response with an Idempotent-Replayed: true header so the caller can tell it was served from cache. The mechanism you are about to build is the same shape, just self-hosted.
One honest simplification: the draft formally defines Idempotency-Key as an RFC 8941 Structured Header String, which technically means the value should be wrapped in double quotes on the wire, like Idempotency-Key: "8e03978e-40d5-43e8-bc93-6894a57f9324". Most real-world client code, and the example below, sends a bare token without the extra quoting, which is simpler to work with and is what you will see from most HTTP client libraries in practice. The server code below treats the header as an opaque string either way, so it works correctly with both styles; it just is not doing strict RFC 8941 parsing, which would be overkill for what this tutorial is teaching.
Step 3: Build an Idempotency Store
The server needs somewhere to remember keys it has already seen. A single SQLite table does the job: one row per idempotency key, holding the fingerprint of the request it was used with, its current status, and (once finished) the exact response to replay on a retry.
CREATE TABLE IF NOT EXISTS idempotency_keys (
key TEXT PRIMARY KEY,
fingerprint TEXT NOT NULL,
status TEXT NOT NULL,
response_status INTEGER,
response_body TEXT,
created_at TEXT NOT NULL
)
Two design decisions here matter more than they look:
keyis the table’sPRIMARY KEY, not just an indexed column. This is not a minor detail, it is the entire safety mechanism the rest of this tutorial relies on. APRIMARY KEYconstraint is enforced atomically by SQLite itself. That means you can let the database, not your own application logic, be the referee for “who got here first” when two requests arrive at nearly the same instant.statusstarts asin_progressthe moment a key is claimed, and only becomescompletedonce the real work (and the response to replay) is known. This is what lets the server tell the difference between “this is a retry of something that already finished” and “this is a concurrent duplicate of something still running.”
Step 4: Build a Safe, Idempotent Charge Endpoint
Save this as step4_safe_server.py. The key design decision is this: instead of checking whether a key exists and then inserting a new row as two separate steps, the handler tries to insert first and treats a failure of that insert as the signal that someone else got there first. There is deliberately no “check” step for two concurrent requests to race against.
import asyncio
import hashlib
import json
import sqlite3
import uuid
from datetime import datetime, timezone
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
from pydantic import BaseModel, ValidationError
DB_PATH = "safe_charges.db"
app = FastAPI()
def get_db():
conn = sqlite3.connect(DB_PATH, timeout=10)
conn.execute(
"""
CREATE TABLE IF NOT EXISTS charges (
charge_id TEXT PRIMARY KEY,
customer_id TEXT NOT NULL,
amount_cents INTEGER NOT NULL,
currency TEXT NOT NULL,
created_at TEXT NOT NULL
)
"""
)
conn.execute(
"""
CREATE TABLE IF NOT EXISTS idempotency_keys (
key TEXT PRIMARY KEY,
fingerprint TEXT NOT NULL,
status TEXT NOT NULL,
response_status INTEGER,
response_body TEXT,
created_at TEXT NOT NULL
)
"""
)
return conn
class ChargeRequest(BaseModel):
customer_id: str
amount_cents: int
currency: str
def fingerprint_for(method: str, path: str, body: bytes) -> str:
h = hashlib.sha256()
h.update(method.encode())
h.update(b"\n")
h.update(path.encode())
h.update(b"\n")
h.update(body)
return h.hexdigest()
def problem(status_code: int, title: str, detail: str) -> JSONResponse:
return JSONResponse(
status_code=status_code,
media_type="application/problem+json",
content={"type": "https://sxz.io/idempotency-errors", "title": title, "detail": detail},
)
@app.post("/charges")
async def create_charge(request: Request):
idem_key = request.headers.get("Idempotency-Key")
if not idem_key:
return problem(
400,
"Missing Idempotency-Key",
"This endpoint requires an Idempotency-Key header on every POST.",
)
body_bytes = await request.body()
fp = fingerprint_for("POST", "/charges", body_bytes)
conn = get_db()
now = datetime.now(timezone.utc).isoformat()
# Try to atomically claim this key. If someone already claimed it, an
# earlier attempt or a truly concurrent one, this INSERT fails because
# `key` is the table's PRIMARY KEY. That failure IS the coordination
# mechanism: there is no separate "check" step to race against.
try:
conn.execute(
"INSERT INTO idempotency_keys (key, fingerprint, status, created_at) "
"VALUES (?, ?, 'in_progress', ?)",
(idem_key, fp, now),
)
conn.commit()
claimed = True
except sqlite3.IntegrityError:
# Release the lock from our own failed INSERT attempt before we
# read. Skipping this can leave a lock that blocks the OTHER
# request too, turning a harmless conflict into a stall for both.
conn.rollback()
claimed = False
if not claimed:
row = conn.execute(
"SELECT fingerprint, status, response_status, response_body "
"FROM idempotency_keys WHERE key = ?",
(idem_key,),
).fetchone()
conn.close()
existing_fp, status, resp_status, resp_body = row
if existing_fp != fp:
return problem(
422,
"Idempotency-Key already used",
"This Idempotency-Key was already used with a different request "
"payload. Idempotency keys must not be reused across different "
"payloads of the same operation.",
)
if status == "in_progress":
return problem(
409,
"Request in progress",
"A request with this Idempotency-Key is still being processed. "
"Do not retry yet.",
)
# status == "completed": replay the original response verbatim.
return JSONResponse(status_code=resp_status, content=json.loads(resp_body))
# We own this key. Validate the body, then process for real.
try:
req = ChargeRequest.model_validate_json(body_bytes)
except ValidationError as e:
# The operation never ran, so free the key: a corrected retry should
# not be punished with a 422 "reused with a different payload" error.
conn.execute("DELETE FROM idempotency_keys WHERE key = ?", (idem_key,))
conn.commit()
conn.close()
return JSONResponse(status_code=422, content={"detail": e.errors()})
await asyncio.sleep(2) # simulate a slow downstream call, e.g. a card network
charge_id = f"ch_{uuid.uuid4().hex[:16]}"
created_at = datetime.now(timezone.utc).isoformat()
conn.execute(
"INSERT INTO charges (charge_id, customer_id, amount_cents, currency, created_at) "
"VALUES (?, ?, ?, ?, ?)",
(charge_id, req.customer_id, req.amount_cents, req.currency, created_at),
)
response_body = {
"charge_id": charge_id,
"customer_id": req.customer_id,
"amount_cents": req.amount_cents,
"currency": req.currency,
"created_at": created_at,
}
conn.execute(
"UPDATE idempotency_keys SET status = 'completed', response_status = 201, "
"response_body = ? WHERE key = ?",
(json.dumps(response_body), idem_key),
)
conn.commit()
conn.close()
return JSONResponse(status_code=201, content=response_body)
Read through the four things this handler can do with an incoming request:
- No
Idempotency-Keyheader at all: reject immediately with400. This endpoint requires the header; there is nothing safe to do without it. - Brand-new key: the
INSERTsucceeds. The handler now owns this key, validates the body, does the (slow, simulated) real work, and finally updates the row tocompletedwith the response to replay later. - Key already
completed, same fingerprint: this is a genuine retry. The handler never re-runs the charge logic; it reads the stored response straight out of the row and returns it. This is what makes a retry cheap and instant instead of repeating a slow downstream call. - Key already
in_progress: a second request with this exact key showed up while the first one is still running. Per the draft, this gets a409 Conflict, telling the caller to back off rather than assume anything.
The fingerprint_for() function hashes the method, path, and raw request body together with SHA-256. This is what lets the server detect the case the draft calls out explicitly: someone reusing your key with different data. Notice it fingerprints the request body before Pydantic validation, on the raw bytes. This matters: hashing after validation could let two textually different but semantically-equal bodies (different key ordering, different whitespace) produce the same fingerprint by accident, which would be the wrong failure mode to accidentally build in.
Also notice what happens on a validation error: the code deletes the key it just claimed instead of leaving it in_progress forever. The operation never actually ran, so there is nothing to protect the client from repeating. If you left the row behind, a client fixing a typo and retrying with corrected data would hit the exact key-reuse-with-different-payload path you built for a completely different reason, and get a confusing 422 instead of a chance to fix its mistake.
Step 5: Prove the Retry Fix Works
Run the safe server, then reuse the exact same timeout-and-retry scenario from Step 1, this time with an Idempotency-Key attached to both attempts. Save this as step4_client_retry_demo.py.
uvicorn step4_safe_server:app --port 8002
import sqlite3
import time
import uuid
import httpx
DB_PATH = "safe_charges.db"
URL = "http://127.0.0.1:8002/charges"
idem_key = str(uuid.uuid4())
payload = {"customer_id": "cus_amir_002", "amount_cents": 4999, "currency": "usd"}
headers = {"Idempotency-Key": idem_key}
print(f"Using Idempotency-Key: {idem_key}\n")
print("Attempt 1: sending charge request with a 1-second client timeout...")
try:
with httpx.Client(timeout=1.0) as client:
resp = client.post(URL, json=payload, headers=headers)
print("Attempt 1 response:", resp.status_code, resp.json())
except httpx.ReadTimeout:
print("Attempt 1 timed out on the client side after 1s.")
print("Same as before: the client does not know if the server finished.")
print("\nWaiting 3 seconds (simulating the user or app retry logic)...")
time.sleep(3)
print("\nAttempt 2: retrying with the SAME Idempotency-Key and SAME payload...")
start = time.time()
with httpx.Client(timeout=5.0) as client:
resp = client.post(URL, json=payload, headers=headers)
elapsed = time.time() - start
print(f"Attempt 2 response ({elapsed:.2f}s): {resp.status_code} {resp.json()}")
print("\nChecking the charges table directly...")
conn = sqlite3.connect(DB_PATH)
rows = conn.execute(
"SELECT charge_id, customer_id, amount_cents, created_at FROM charges "
"WHERE customer_id = 'cus_amir_002'"
).fetchall()
conn.close()
print(f"Found {len(rows)} charge(s) in the database for this customer:")
for row in rows:
print(" ", row)
Real captured output:
Using Idempotency-Key: 6cfd598d-4cdb-45d5-ac9b-c2ed37b24369
Attempt 1: sending charge request with a 1-second client timeout...
Attempt 1 timed out on the client side after 1s.
Same as before: the client does not know if the server finished.
Waiting 3 seconds (simulating the user or app retry logic)...
Attempt 2: retrying with the SAME Idempotency-Key and SAME payload...
Attempt 2 response (0.29s): 201 {'charge_id': 'ch_13cda48375a444a3', 'customer_id': 'cus_amir_002', 'amount_cents': 4999, 'currency': 'usd', 'created_at': '2026-08-20T19:58:39.954543+00:00'}
Checking the charges table directly...
Found 1 charge(s) in the database for this customer:
('ch_13cda48375a444a3', 'cus_amir_002', 4999, '2026-08-20T19:58:39.954543+00:00')
Compare the timing to Step 1. The first attempt still times out client-side after one second, exactly as before; the server is still silently finishing that same charge in the background. But now the retry comes back in a fraction of a second (well under the two-second processing delay), because the safe server recognized the key as already completed and returned the stored response directly instead of running the charge logic again. The database confirms it: exactly one charge, and it is the same charge_id the client’s second attempt reported, proving the retry got back the result of the first attempt rather than creating a new one.
Step 6: Reject Key Reuse With a Different Payload
An idempotency key is a promise: this key represents one specific operation. Save this as step5_fingerprint_mismatch_demo.py and, with the safe server from Step 5 still running, send two different charge amounts under the same key.
import uuid
import httpx
URL = "http://127.0.0.1:8002/charges"
idem_key = str(uuid.uuid4())
print(f"Using Idempotency-Key: {idem_key}\n")
payload_a = {"customer_id": "cus_amir_003", "amount_cents": 1500, "currency": "usd"}
print("Request A: $15.00 charge with a brand-new Idempotency-Key...")
with httpx.Client(timeout=5.0) as client:
resp_a = client.post(URL, json=payload_a, headers={"Idempotency-Key": idem_key})
print("Response A:", resp_a.status_code, resp_a.json())
payload_b = {"customer_id": "cus_amir_003", "amount_cents": 9999, "currency": "usd"}
print("\nRequest B: REUSING the same key, but now for a $99.99 charge...")
with httpx.Client(timeout=5.0) as client:
resp_b = client.post(URL, json=payload_b, headers={"Idempotency-Key": idem_key})
print("Response B:", resp_b.status_code, resp_b.json())
Real captured output:
Using Idempotency-Key: 31c0adc0-e625-4970-97c1-51e779af4e0e
Request A: $15.00 charge with a brand-new Idempotency-Key...
Response A: 201 {'charge_id': 'ch_b198474fce0549dc', 'customer_id': 'cus_amir_003', 'amount_cents': 1500, 'currency': 'usd', 'created_at': '2026-08-20T19:58:44.537843+00:00'}
Request B: REUSING the same key, but now for a $99.99 charge...
Response B: 422 {'type': 'https://sxz.io/idempotency-errors', 'title': 'Idempotency-Key already used', 'detail': 'This Idempotency-Key was already used with a different request payload. Idempotency keys must not be reused across different payloads of the same operation.'}
The first request succeeds normally and creates a real $15.00 charge. The second, reusing the same key for a $99.99 charge, is rejected outright with 422, exactly the status code the draft specifies for this case. This is an important safety property: without fingerprint checking, a bug that accidentally reused an idempotency key (a common mistake is generating the key once per user session instead of once per operation) would silently return the wrong cached response instead of failing loudly. A loud, immediate 422 is much easier to debug than a customer reporting they got charged the wrong amount.
Stop the server before the next step.
Step 7: Handle True Concurrency Safely
Everything so far has been sequential: one request finishes before the next one starts. But a real “double-click submit” happens when two requests are genuinely in flight at the same time, and that is a different, harder problem. This step builds a version with a subtle, realistic bug on purpose, so you can see exactly what it looks like when it fails, before comparing it against the safe version you already built.
The Racy Way: Check, Then Insert
The most natural-looking way to write this logic is to check whether a key exists, and if it does not, insert it, as two separate statements. Save this as step6_racy_server.py. It reuses the same structure as the safe server, but with that one change.
import asyncio
import hashlib
import json
import sqlite3
import uuid
from datetime import datetime, timezone
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
from pydantic import BaseModel
DB_PATH = "racy_charges.db"
app = FastAPI()
def get_db():
conn = sqlite3.connect(DB_PATH, timeout=10)
conn.execute(
"""
CREATE TABLE IF NOT EXISTS charges (
charge_id TEXT PRIMARY KEY,
customer_id TEXT NOT NULL,
amount_cents INTEGER NOT NULL,
currency TEXT NOT NULL,
created_at TEXT NOT NULL
)
"""
)
conn.execute(
"""
CREATE TABLE IF NOT EXISTS idempotency_keys (
key TEXT PRIMARY KEY,
fingerprint TEXT NOT NULL,
status TEXT NOT NULL,
response_status INTEGER,
response_body TEXT,
created_at TEXT NOT NULL
)
"""
)
return conn
class ChargeRequest(BaseModel):
customer_id: str
amount_cents: int
currency: str
def fingerprint_for(method: str, path: str, body: bytes) -> str:
h = hashlib.sha256()
h.update(method.encode())
h.update(b"\n")
h.update(path.encode())
h.update(b"\n")
h.update(body)
return h.hexdigest()
@app.post("/charges")
async def create_charge(request: Request):
idem_key = request.headers.get("Idempotency-Key")
body_bytes = await request.body()
fp = fingerprint_for("POST", "/charges", body_bytes)
req = ChargeRequest.model_validate_json(body_bytes)
conn = get_db()
# THE BUG: this checks for an existing key, then inserts one, as two
# separate steps. Nothing stops two requests from both passing the
# check before either one inserts.
row = conn.execute(
"SELECT fingerprint, status, response_status, response_body "
"FROM idempotency_keys WHERE key = ?",
(idem_key,),
).fetchone()
# Artificially widen the race window so it reproduces reliably on a
# single test machine. Ordinary request-handling latency does this for
# you in production -- no artificial delay needed there. This must be
# an async sleep: it has to yield control back to the event loop so a
# second in-flight request can actually run its own SELECT here too.
await asyncio.sleep(0.3)
if row is None:
try:
conn.execute(
"INSERT INTO idempotency_keys (key, fingerprint, status, created_at) "
"VALUES (?, ?, 'in_progress', ?)",
(idem_key, fp, datetime.now(timezone.utc).isoformat()),
)
conn.commit()
except Exception:
# Even the buggy version has to close its connection on the
# error path. Otherwise this request's crash leaves a lock
# that blocks the OTHER request too, turning one bad request
# into an outage for both.
conn.close()
raise
else:
existing_fp, status, resp_status, resp_body = row
conn.close()
if status == "completed":
return JSONResponse(status_code=resp_status, content=json.loads(resp_body))
# This naive version has no real handling for "in_progress": it
# just falls through and processes the request a second time.
await asyncio.sleep(2) # simulate a slow downstream call
charge_id = f"ch_{uuid.uuid4().hex[:16]}"
created_at = datetime.now(timezone.utc).isoformat()
conn.execute(
"INSERT INTO charges (charge_id, customer_id, amount_cents, currency, created_at) "
"VALUES (?, ?, ?, ?, ?)",
(charge_id, req.customer_id, req.amount_cents, req.currency, created_at),
)
response_body = {
"charge_id": charge_id,
"customer_id": req.customer_id,
"amount_cents": req.amount_cents,
"currency": req.currency,
"created_at": created_at,
}
conn.execute(
"UPDATE idempotency_keys SET status = 'completed', response_status = 201, "
"response_body = ? WHERE key = ?",
(json.dumps(response_body), idem_key),
)
conn.commit()
conn.close()
return JSONResponse(status_code=201, content=response_body)
This looks reasonable, and it will pass every test from Steps 1 through 6 without complaint, since none of those tests send two requests at once. The bug only shows up under real concurrency: nothing stops two requests from both running the SELECT, both seeing no existing row, and both proceeding to INSERT. The small asyncio.sleep(0.3) between the check and the insert is there deliberately, to widen this window enough to reproduce reliably on a single test machine for this demo. In production you do not need an artificial delay: ordinary request-handling latency, and the fact that a busy server is juggling many requests on the same event loop, creates this exact window on its own.
Fire Two Concurrent Requests at Each Version
Save this as step6_concurrency_demo.py. It fires two requests at the same URL with the same idempotency key from two threads, as close to simultaneously as Python’s threading module allows, then reports both results and checks how many charges actually landed in the database.
import sqlite3
import sys
import threading
import httpx
port = sys.argv[1]
db_path = sys.argv[2]
url = f"http://127.0.0.1:{port}/charges"
payload = {"customer_id": "cus_race_001", "amount_cents": 2500, "currency": "usd"}
headers = {"Idempotency-Key": "race-test-key-001"}
results = [None, None]
def fire(index):
with httpx.Client(timeout=10.0) as client:
resp = client.post(url, json=payload, headers=headers)
results[index] = (resp.status_code, resp.text[:180])
t1 = threading.Thread(target=fire, args=(0,))
t2 = threading.Thread(target=fire, args=(1,))
t1.start()
t2.start()
t1.join()
t2.join()
print("Thread A (double-click #1):", results[0])
print("Thread B (double-click #2):", results[1])
conn = sqlite3.connect(db_path)
rows = conn.execute(
"SELECT charge_id, amount_cents FROM charges WHERE customer_id = 'cus_race_001'"
).fetchall()
conn.close()
print(f"\nCharges actually created for cus_race_001: {len(rows)}")
for r in rows:
print(" ", r)
Run it against the racy server first:
uvicorn step6_racy_server:app --port 8003
python step6_concurrency_demo.py 8003 racy_charges.db
Thread A (double-click #1): (201, '{"charge_id":"ch_6bb4520b70ff466e","customer_id":"cus_race_001","amount_cents":2500,"currency":"usd","created_at":"2026-08-20T19:58:54.843571+00:00"}')
Thread B (double-click #2): (500, 'Internal Server Error')
Charges actually created for cus_race_001: 1
('ch_6bb4520b70ff466e', 2500)
No duplicate charge got created, SQLite’s PRIMARY KEY constraint still refused the second INSERT. But look at what the loser actually got back: a raw 500 Internal Server Error, not a clean, documented 409 Conflict. That is because the racy handler never anticipated this outcome; it let an unhandled sqlite3.IntegrityError propagate straight out of the route function. (Which request wins this race is not deterministic. You might see request A crash and B succeed, or the reverse, run it a few times and you will see both orderings.) For a real customer, the practical difference between these two outcomes is enormous: a 409 with a clear message tells their client “try again shortly,” while an unlabeled 500 looks exactly like your service is broken, and a reasonable client might reasonably retry a 500 in a way that starts this whole problem over again.
Now run the exact same concurrent test against the safe server from Step 4:
uvicorn step4_safe_server:app --port 8002
python step6_concurrency_demo.py 8002 safe_charges.db
Thread A (double-click #1): (201, '{"charge_id":"ch_9c9666f3958b44a3","customer_id":"cus_race_001","amount_cents":2500,"currency":"usd","created_at":"2026-08-20T19:59:05.820858+00:00"}')
Thread B (double-click #2): (409, '{"type":"https://sxz.io/idempotency-errors","title":"Request in progress","detail":"A request with this Idempotency-Key is still being processed. Do not retry yet."}')
Charges actually created for cus_race_001: 1
('ch_9c9666f3958b44a3', 2500)
Same race, same two simultaneous requests, same single charge created, but a completely different failure mode for the loser: a clean, documented 409 Conflict instead of a crash. The only structural difference between the two servers is that the safe version made the INSERT itself the very first thing that touches the key, with no separate check to race against, and it explicitly handles the IntegrityError that means “someone beat you to it” instead of letting it become an unhandled exception.
One more real bug surfaced while building this comparison, and it is worth knowing about even though it is not the headline lesson: the first version of the racy handler let its failed INSERT leave an uncommitted transaction open on that connection, which held a lock long enough to make the other request’s later write fail with database is locked instead of the clean constraint violation shown above. The fix was making sure the connection is always closed on the error path, not just on success. Both handlers above already include this fix. If you are adapting this pattern to your own code and see mysterious lock timeouts under load, check that every code path, including your exception handlers, actually closes or rolls back its connection.
Step 8: Expire Old Idempotency Keys
An idempotency key store that never forgets anything grows forever and eventually slows down every lookup. The draft leaves the exact policy up to you (“The resource MAY require time based idempotency keys to be able to purge or delete a key upon its expiry. The resource SHOULD define such expiration policy and publish it in the documentation.”), but it does not have to be complicated. Stripe’s well-known default is a 24-hour window: long enough to comfortably cover any realistic client retry logic, short enough to keep the table small. A simple cleanup query, run periodically, is enough:
DELETE FROM idempotency_keys
WHERE created_at < datetime('now', '-24 hours');
Document whatever window you choose. A client that assumes keys live forever, and one that assumes they expire in an hour, will disagree about whether a very late retry is safe, and that disagreement is exactly the kind of bug this whole mechanism exists to prevent.
Common Mistakes and Gotchas
- Generating a new key on every retry. This defeats the entire mechanism. The key has to be generated once per logical operation, before the first attempt, and reused unchanged on every retry of that same operation.
- Hashing the parsed, validated body instead of the raw bytes. Validation can normalize data in ways that make two genuinely different requests hash identically, or the reverse. Fingerprint the request as it actually arrived.
- Checking for a key’s existence before inserting it, as two separate steps. Step 7 showed exactly why: it is safe when requests are sequential and can race silently, or not so silently, when they are not.
- Leaving a claimed key
in_progressforever after an error. If the real work fails in a way that means it never actually happened (like a validation error), free the key so a corrected retry is not punished with a confusing “key already used” error. - Not closing a database connection on the error path. An unhandled exception does not automatically release a database lock. Step 7’s postscript is a real example of one broken request stalling an unrelated one because of this.
- Treating idempotency keys as a substitute for authentication or authorization. A key only protects against duplicate execution by the same caller; it says nothing about who is allowed to call the endpoint at all.
How to Verify Your Own Implementation End to End
- Send one request with a fresh key. Confirm you get a
201and the side effect (in this tutorial, a row incharges) happened exactly once. - Immediately resend the identical request with the same key. Confirm you get back the exact same response body, including the same generated ID, and that no new side effect occurred.
- Resend the same key with a different body. Confirm you get a
422, not a silently wrong cached response and not a second side effect. - Send the same key with no body change, but from two clients (or two threads) at nearly the same instant. Confirm exactly one side effect happened, and that the “loser” received a clean, documented error rather than a crash.
- Send a request with no
Idempotency-Keyheader at all, if your endpoint requires one. Confirm you get a clear400instead of the request silently proceeding unprotected.
If your implementation passes all five, you have covered the same enforcement matrix this tutorial built and tested against the IETF draft: first-time request, retry, payload mismatch, genuine concurrency, and the missing-header case.
Next Steps
This tutorial used SQLite because it made the concurrency behavior easy to reason about and required nothing beyond Python’s standard library. The exact same design (an atomic insert-first claim, a status column distinguishing in-progress from completed, a fingerprint check) works with any datastore that gives you an atomic unique-constraint failure, including Postgres, MySQL, or a Redis SET key value NX command, which is a common choice when you need the idempotency store to be shared across multiple application servers rather than living on a single machine’s disk.
If you want to keep building on the reliability patterns in this tutorial, sxz.io has related tutorials on building a circuit breaker in Python to stop cascading failures from a downstream dependency, and on building an API rate limiter with the token bucket and sliding window algorithms. Idempotency keys, circuit breakers, and rate limiting solve three different problems, but they are the same kind of problem: making an API that talks to unreliable networks and impatient clients behave predictably anyway.








No Comment! Be the first one.