How to Build a Circuit Breaker in Python to Stop Cascading Failures
Learn how to build a circuit breaker from scratch in Python that detects a failing dependency, fails fast instead of piling up timeouts, and recovers automatically once it is healthy again.
If you have ever put a real dependency behind an API call, a payment processor, a database, another team’s microservice, you have probably watched what happens when that dependency goes down: every caller keeps calling it anyway. Requests pile up waiting on a timeout that hasn’t fired yet. Threads and connections fill up holding those requests. The one broken dependency starts taking down services that were otherwise perfectly healthy. This is called a cascading failure, and it is one of the most common ways a small outage turns into a big one.
Table Of Content
- What You Will Build
- Prerequisites
- Step 1: Set Up Your Project
- Step 2: Build a Dependency You Can Break On Demand
- Step 3: See the Problem for Real
- Step 4: Build the Circuit Breaker State Machine
- Why the lock, and why the call happens outside it
- Why the half-open transition happens lazily
- Step 5: Do Not Trip the Breaker on Your Own Bugs
- Step 6: Wire the Breaker Around Real Calls, With a Fallback
- Step 7: Watch the Full Recovery Cycle
- Step 8: Fix the Thundering Herd on Half-Open
- Common Mistakes and Gotchas
- How to Verify Your Circuit Breaker Actually Works
- Next Steps
A circuit breaker is the fix. It is a small piece of code that sits between your application and a dependency it calls over the network. It watches for failures, and once it sees enough of them, it stops calling the dependency at all for a while, failing every request immediately instead. This is the same idea as the circuit breaker in your home’s electrical panel: when something downstream draws too much current, the breaker trips and cuts the circuit before the wiring overheats, rather than letting the fault keep drawing power indefinitely.
This tutorial builds a circuit breaker from scratch in Python: no framework, no external service, just the standard library plus a small demo dependency you can break on command. You will reproduce a real cascading failure, measure exactly how much time it wastes, fix it, watch the breaker recover automatically once the dependency comes back, and then find and fix a genuine concurrency bug in the recovery logic itself. Every command below was run for real while writing this post, and the timings and output you will see are the actual captured results, not invented numbers.
What You Will Build
A small Python library, circuit_breaker.py, implementing a three-state circuit breaker (closed, open, half-open) with no third-party dependencies. Around it, you will build:
- A deliberately flaky FastAPI service that you can flip between healthy and unhealthy on demand, so you can reproduce an outage instead of just reading about one
- A naive client that calls the flaky service directly, to measure how bad an unprotected cascading failure actually is
- A protected client that wraps the same calls in your circuit breaker, with a fallback response for graceful degradation
- A script that watches the breaker recover automatically once the dependency comes back healthy
- A concurrency test that proves, with real measured numbers, why the breaker needs to limit itself to exactly one trial call during recovery, and what goes wrong if it does not
Prerequisites
- Python 3.10 or newer (this tutorial used Python 3.13.14 on Windows, but nothing here is platform-specific)
- Comfort with basic Python: classes, exceptions, and what a decorator-free wrapper function does
- Basic familiarity with HTTP status codes and what a client timeout is. This tutorial explains the rest
- No external services, accounts, or API keys. Everything in this tutorial runs on your own machine
- About 40 minutes
Step 1: Set Up Your Project
Create a project folder and an isolated virtual environment, then install the three libraries this tutorial needs: a web framework to build the demo dependency, a server to run it, and an HTTP client to call it.
mkdir circuit-breaker-demo
cd circuit-breaker-demo
python -m venv .venv
.venv\Scripts\activate # on Linux/macOS: source .venv/bin/activate
pip install fastapi uvicorn httpx
This tutorial was written against fastapi 0.141.1, uvicorn 0.52.4, and httpx 0.28.1. Newer versions should work the same way; none of the APIs used here are new or unstable.
Step 2: Build a Dependency You Can Break On Demand
To reproduce a real outage instead of just describing one, you need a dependency whose health you can control. Create flaky_service.py:
"""A fake downstream dependency (e.g. an inventory service) that we can
switch between healthy and unhealthy at runtime, so we can reproduce a
real outage on demand instead of just describing one."""
import time
from fastapi import FastAPI, HTTPException
app = FastAPI()
STATE = {"healthy": True}
@app.post("/admin/set-healthy/{value}")
def set_healthy(value: bool):
STATE["healthy"] = value
return {"healthy": STATE["healthy"]}
@app.get("/inventory/{sku}")
def get_inventory(sku: str):
if not sku.isupper():
# A malformed request. This is a client bug, not evidence that
# the dependency is unhealthy, so it must NOT count toward the
# breaker's failure threshold.
raise HTTPException(status_code=400, detail="sku must be uppercase")
if STATE["healthy"]:
return {"sku": sku, "quantity": 42}
# Simulate a real outage: the dependency does not fail instantly,
# it hangs near its connection/read timeout before giving up. This
# is the case that actually hurts callers, an instant 500 is cheap
# to handle, a slow failure is not.
time.sleep(3.0)
raise HTTPException(status_code=503, detail="inventory-service unavailable")
Two design choices here matter more than they look:
First, when unhealthy, the service does not fail instantly. It calls time.sleep(3.0) before returning a 503. A real dependency rarely fails the instant something goes wrong; a database connection has to time out, a load balancer has to give up on a backend, a TCP connection has to hit its retransmit limit. An instant failure is cheap for a caller to handle. A slow failure is what actually causes cascading failures, because every caller sits there holding a thread, a connection, or an async task waiting for those 3 seconds to pass, over and over, for every single request.
Second, notice the sku.isupper() check at the top, which rejects malformed input with a 400 before the health check ever runs. Keep this in mind, you will use it in Step 5 to prove something important about what should and should not trip a circuit breaker.
Start the service in one terminal and leave it running for the rest of this tutorial:
uvicorn flaky_service:app --host 127.0.0.1 --port 8811
Plain uvicorn does not reload your code automatically. If you edit flaky_service.py later, stop this process (Ctrl+C) and restart it, or your changes will not take effect. This tutorial ran into exactly that: a stale server process kept serving old code after an edit, and produced a confusing false result until it was restarted. If a step’s output does not match what this tutorial shows, restarting the server is the first thing to try.
In a second terminal, confirm it works both ways:
curl http://127.0.0.1:8811/inventory/WIDGET-1
# {"sku":"WIDGET-1","quantity":42}
curl -X POST http://127.0.0.1:8811/admin/set-healthy/false
# {"healthy":false}
curl http://127.0.0.1:8811/inventory/WIDGET-1
# (waits about 3 seconds)
# {"detail":"inventory-service unavailable"}
That 3 second wait before the error is the entire problem this tutorial exists to solve.
Step 3: See the Problem for Real
Before building the fix, measure the actual cost of not having one. Create naive_client.py, which does what most application code does by default: call the dependency, catch whatever error comes back, move on to the next request. Nothing here remembers that the last 5 calls all failed the same way.
"""What most services do by default: call the dependency, catch the
error, move on. No memory of the last call. Every request pays the
full cost of finding out the dependency is still down."""
import time
import httpx
BASE_URL = "http://127.0.0.1:8811"
def check_inventory(sku: str) -> dict:
resp = httpx.get(f"{BASE_URL}/inventory/{sku}", timeout=5.0)
resp.raise_for_status()
return resp.json()
def main():
total_start = time.perf_counter()
successes = 0
failures = 0
for i in range(1, 21):
start = time.perf_counter()
try:
result = check_inventory("WIDGET-1")
elapsed = time.perf_counter() - start
print(f"request {i:2d}: OK in {elapsed:5.2f}s -> {result}")
successes += 1
except httpx.HTTPStatusError as exc:
elapsed = time.perf_counter() - start
print(f"request {i:2d}: FAIL in {elapsed:5.2f}s -> HTTP {exc.response.status_code}")
failures += 1
except httpx.TimeoutException:
elapsed = time.perf_counter() - start
print(f"request {i:2d}: FAIL in {elapsed:5.2f}s -> client timeout")
failures += 1
total_elapsed = time.perf_counter() - total_start
print(f"\n{successes} succeeded, {failures} failed, total wall time {total_elapsed:.2f}s")
if __name__ == "__main__":
main()
With the service still set unhealthy from Step 2, run it:
python naive_client.py
Here is the real output from this run:
request 1: FAIL in 3.18s -> HTTP 503
request 2: FAIL in 3.29s -> HTTP 503
request 3: FAIL in 3.27s -> HTTP 503
...
request 20: FAIL in 3.27s -> HTTP 503
0 succeeded, 20 failed, total wall time 65.30s
Twenty requests, every single one destined to fail, and it still took 65.30 seconds to find that out, because each one had to independently rediscover that the dependency was down. In a real service, those 20 requests are not sequential like this demo, they are concurrent, each one holding a thread or a connection for 3 seconds while it waits. Multiply that by your actual request volume and you can see how one slow dependency turns into an exhausted thread pool for a service that has nothing to do with the original problem.
Step 4: Build the Circuit Breaker State Machine
A circuit breaker is a state machine with three states. Understanding what each one means is more important than any of the code:
- Closed: normal operation. Calls pass through to the real dependency. The breaker counts recent failures.
- Open: the failure count crossed a threshold. Every call fails immediately with an error, without touching the network at all, until a recovery timeout elapses.
- Half-Open: the recovery timeout has elapsed. Exactly one trial call is allowed through to check whether the dependency has recovered. If it succeeds, the breaker closes and resumes normal operation. If it fails, the breaker reopens and the timeout starts again.
This is not a new idea. Michael Nygard popularized it in his book Release It!, and Martin Fowler’s widely read summary of the pattern puts it plainly: You wrap a protected function call in a circuit breaker object, which monitors for failures. Once the failures reach a certain threshold, the circuit breaker trips, and all further calls to the circuit breaker return with an error, without the protected call being made at all.
Microsoft’s Azure Architecture Center describes the same three states and adds the specific reason the half-open state only allows a limited number of requests through: The Half-Open state helps prevent a recovering service from suddenly being flooded with requests. As a service recovers, it might be able to support a limited volume of requests until the recovery is complete. But while recovery is in progress, a flood of work can cause the service to time out or fail again.
You will prove that exact failure mode yourself in Step 8.
Create circuit_breaker.py:
"""A circuit breaker built from scratch: three states, one lock, no
dependencies. Mirrors the state machine used by libraries like pybreaker
and resilience4j, simplified for teaching."""
import threading
import time
from enum import Enum
class CircuitState(Enum):
CLOSED = "closed" # normal operation, calls go through
OPEN = "open" # tripped, calls fail fast without hitting the network
HALF_OPEN = "half_open" # trial period, exactly one call is allowed through
class CircuitOpenError(Exception):
"""Raised by CircuitBreaker.call() instead of invoking the wrapped
function, when the breaker is open or a half-open trial is already
in flight."""
class CircuitBreaker:
def __init__(self, failure_threshold=5, recovery_timeout=10.0, expected_exceptions=(Exception,)):
self.failure_threshold = failure_threshold
self.recovery_timeout = recovery_timeout
self.expected_exceptions = expected_exceptions
self._state = CircuitState.CLOSED
self._failure_count = 0
self._opened_at = None
self._half_open_trial_in_flight = False
self._lock = threading.Lock()
@property
def state(self) -> CircuitState:
"""Read-only view of the state. Does not itself trigger the
OPEN -> HALF_OPEN transition; only call() does that, so that a
transition always happens right before a real trial call."""
with self._lock:
return self._peek_state()
def _peek_state(self) -> CircuitState:
if self._state == CircuitState.OPEN and self._opened_at is not None:
if time.monotonic() - self._opened_at >= self.recovery_timeout:
return CircuitState.HALF_OPEN
return self._state
def call(self, func, *args, **kwargs):
with self._lock:
current = self._peek_state()
if current == CircuitState.OPEN:
raise CircuitOpenError("circuit is open, failing fast")
if current == CircuitState.HALF_OPEN:
if self._half_open_trial_in_flight:
raise CircuitOpenError("circuit is half-open, a trial call is already in flight")
self._state = CircuitState.HALF_OPEN
self._half_open_trial_in_flight = True
# The wrapped call happens OUTSIDE the lock. Holding a lock across
# a network call would serialize every request in the process
# behind the slowest one, exactly the pile-up we are trying to avoid.
try:
result = func(*args, **kwargs)
except self.expected_exceptions:
self._on_failure()
raise
else:
self._on_success()
return result
def _on_success(self):
with self._lock:
self._failure_count = 0
self._state = CircuitState.CLOSED
self._opened_at = None
self._half_open_trial_in_flight = False
def force_open_for_test(self):
"""Test-only seam: jump straight to OPEN without needing real
failures first. Do not ship this in a production breaker, it
exists here purely to set up the concurrency demo quickly."""
with self._lock:
self._state = CircuitState.OPEN
self._opened_at = time.monotonic()
def _on_failure(self):
with self._lock:
self._half_open_trial_in_flight = False
self._failure_count += 1
if self._state == CircuitState.HALF_OPEN:
# The trial call failed: the dependency is still down.
# Re-open and restart the recovery clock.
self._state = CircuitState.OPEN
self._opened_at = time.monotonic()
elif self._failure_count >= self.failure_threshold:
self._state = CircuitState.OPEN
self._opened_at = time.monotonic()
Why the lock, and why the call happens outside it
Every state read and every state transition happens inside self._lock, because multiple threads (or async tasks running on multiple workers) can call .call() at the same time, and two threads racing to both decide independently that they are allowed to open the breaker, or both decide they are the half-open trial call, would corrupt the state. But the line that matters most is this one: func(*args, **kwargs) is called outside the lock. If it were called while holding the lock, every request in your process would queue up behind whichever one happened to be talking to the slow dependency first, turning your circuit breaker into an accidental global mutex. The lock protects the breaker’s bookkeeping, not the call itself.
Why the half-open transition happens lazily
Notice that _peek_state() computes whether enough time has passed to justify moving from open to half-open, but only .call() actually commits that transition (by setting self._state = CircuitState.HALF_OPEN). The read-only .state property can tell you the breaker would be half-open, without that observation itself consuming the one trial slot. You will see why this distinction matters, and a genuine consequence of it, in Step 7.
Step 5: Do Not Trip the Breaker on Your Own Bugs
Fowler’s original description includes a warning that is easy to skip past: Not all errors should trip the circuit, some should reflect normal failures and be dealt with as part of regular logic.
If your circuit breaker treats every exception as evidence the dependency is unhealthy, then a client sending malformed requests, your own validation bug, a typo in a SKU, will look identical to a real outage and trip the breaker for every other caller too, even though retrying the exact same bad request would fail the exact same way forever.
Create inventory_client.py to draw that line explicitly, translating the flaky service’s responses into two distinct exception types:
"""Talks to the inventory service and translates transport-level and
HTTP-level failures into two distinct exception types, so the circuit
breaker can tell "the dependency is unhealthy" apart from "the caller
sent a bad request." Only the first should ever trip a breaker."""
import httpx
BASE_URL = "http://127.0.0.1:8811"
class DependencyError(Exception):
"""The dependency itself is the problem: timeout, connection
refused, or a 5xx response."""
class BadRequestError(Exception):
"""The caller sent something the dependency correctly rejected.
Retrying or tripping a breaker over this would be pointless, the
next identical request would fail the same way."""
def check_inventory(sku: str) -> dict:
try:
resp = httpx.get(f"{BASE_URL}/inventory/{sku}", timeout=5.0)
except httpx.TimeoutException as exc:
raise DependencyError(f"timed out calling inventory service: {exc}") from exc
except httpx.ConnectError as exc:
raise DependencyError(f"could not connect to inventory service: {exc}") from exc
if resp.status_code == 400:
raise BadRequestError(resp.json().get("detail", "bad request"))
if resp.status_code >= 500:
raise DependencyError(f"inventory service returned HTTP {resp.status_code}")
resp.raise_for_status()
return resp.json()
The circuit breaker will be configured to treat only DependencyError as a failure. A BadRequestError will propagate normally without ever reaching _on_failure(), because it will not match the breaker’s expected_exceptions filter. Prove this before moving on. Create bad_request_test.py:
"""Proves that a client error (bad SKU) does not move the breaker
toward OPEN, even when it happens repeatedly, while the service is
otherwise perfectly healthy."""
from circuit_breaker import CircuitBreaker
from inventory_client import check_inventory, DependencyError, BadRequestError
breaker = CircuitBreaker(failure_threshold=3, recovery_timeout=5.0, expected_exceptions=(DependencyError,))
def call(sku):
try:
result = breaker.call(check_inventory, sku)
print(f"sku={sku!r:12} state={breaker.state.value:8} -> OK {result}")
except BadRequestError as exc:
print(f"sku={sku!r:12} state={breaker.state.value:8} -> BadRequestError: {exc}")
except DependencyError as exc:
print(f"sku={sku!r:12} state={breaker.state.value:8} -> DependencyError: {exc}")
for _ in range(5):
call("bad-sku") # lowercase, the service rejects this with a 400 every time
call("WIDGET-1") # a well-formed request right after, to prove the breaker never moved
With the service set back to healthy (curl -X POST http://127.0.0.1:8811/admin/set-healthy/true), run it:
python bad_request_test.py
Real captured output:
sku='bad-sku' state=closed -> BadRequestError: sku must be uppercase
sku='bad-sku' state=closed -> BadRequestError: sku must be uppercase
sku='bad-sku' state=closed -> BadRequestError: sku must be uppercase
sku='bad-sku' state=closed -> BadRequestError: sku must be uppercase
sku='bad-sku' state=closed -> BadRequestError: sku must be uppercase
sku='WIDGET-1' state=closed -> OK {'sku': 'WIDGET-1', 'quantity': 42}
Five consecutive errors, more than enough to trip a breaker with a failure threshold of 3, and the state never moves off closed, because none of them were DependencyError. The well-formed request right after succeeds normally. This is the entire point: the breaker should only ever open because of the dependency, never because of the caller.
Step 6: Wire the Breaker Around Real Calls, With a Fallback
Create protected_client.py. It runs the exact same 20 requests as naive_client.py in Step 3, but every call goes through the circuit breaker, and when the breaker is open, it returns a fallback value instead of letting the caller wait:
"""The same 20 requests as naive_client.py, but every call to the
dependency goes through a CircuitBreaker with a fallback. Compare the
total wall-clock time and the per-request latency to naive_client.py's
output."""
import time
from circuit_breaker import CircuitBreaker, CircuitOpenError
from inventory_client import check_inventory, DependencyError, BadRequestError
breaker = CircuitBreaker(
failure_threshold=3,
recovery_timeout=8.0,
expected_exceptions=(DependencyError,),
)
def get_inventory_with_fallback(sku: str) -> dict:
try:
return breaker.call(check_inventory, sku)
except CircuitOpenError:
# Graceful degradation: tell the truth about staleness instead
# of hanging the caller for 3+ seconds only to fail anyway.
return {"sku": sku, "quantity": None, "source": "fallback (breaker open)"}
except DependencyError:
return {"sku": sku, "quantity": None, "source": "fallback (call failed)"}
def main():
total_start = time.perf_counter()
for i in range(1, 21):
start = time.perf_counter()
result = get_inventory_with_fallback("WIDGET-1")
elapsed = time.perf_counter() - start
print(f"request {i:2d}: state={breaker.state.value:10s} in {elapsed:5.2f}s -> {result}")
total_elapsed = time.perf_counter() - total_start
print(f"\ntotal wall time {total_elapsed:.2f}s")
if __name__ == "__main__":
main()
Graceful degradation is the point of the fallback: instead of hanging for 3 seconds and then failing anyway, a caller with the circuit open gets an honest, immediate answer that the data might be stale. Set the service unhealthy again and run it:
curl -X POST http://127.0.0.1:8811/admin/set-healthy/false
python protected_client.py
Real captured output:
request 1: state=closed in 3.17s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (call failed)'}
request 2: state=closed in 3.27s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (call failed)'}
request 3: state=open in 3.27s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (call failed)'}
request 4: state=open in 0.00s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (breaker open)'}
request 5: state=open in 0.00s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (breaker open)'}
...
request 20: state=open in 0.00s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (breaker open)'}
total wall time 9.71s
9.71 seconds instead of 65.30 seconds, a 6.7x reduction, for the exact same 20 requests against the exact same dead dependency. The first three requests still pay the real cost, that is unavoidable and correct, the breaker needs real evidence before it trips. But requests 4 through 20 fail in 0.00 seconds instead of 3.27 seconds each, because the breaker stopped them before they ever touched the network.
Step 7: Watch the Full Recovery Cycle
So far you have seen closed and open. To see half-open, and the automatic recovery back to closed, you need a longer-running script that flips the dependency back to healthy partway through. Create recovery_demo.py:
"""Runs long enough to observe the full cycle: CLOSED -> OPEN (after 3
real failures) -> repeated fail-fast while OPEN -> HALF_OPEN (once the
recovery timeout elapses) -> CLOSED again, once the dependency is
flipped back to healthy partway through."""
import time
import httpx
from circuit_breaker import CircuitBreaker, CircuitOpenError
from inventory_client import check_inventory, DependencyError
ADMIN_URL = "http://127.0.0.1:8811/admin/set-healthy"
breaker = CircuitBreaker(
failure_threshold=3,
recovery_timeout=4.0,
expected_exceptions=(DependencyError,),
)
def get_inventory_with_fallback(sku: str) -> dict:
try:
return breaker.call(check_inventory, sku)
except CircuitOpenError:
return {"source": "fallback (breaker open)"}
except DependencyError:
return {"source": "fallback (call failed)"}
def main():
start = time.perf_counter()
for i in range(1, 13):
if i == 6:
httpx.post(f"{ADMIN_URL}/true", timeout=5.0)
print(" --- ops fixed the dependency: service is healthy again ---")
pre_state = breaker.state.value
call_start = time.perf_counter()
result = get_inventory_with_fallback("WIDGET-1")
elapsed = time.perf_counter() - call_start
post_state = breaker.state.value
t = time.perf_counter() - start
print(
f"t={t:6.2f}s request {i:2d}: pre={pre_state:10s} post={post_state:10s} "
f"call_took={elapsed:5.2f}s -> {result}"
)
time.sleep(1.5)
if __name__ == "__main__":
main()
This script prints the breaker’s state both immediately before and immediately after each call. That is deliberate, and it exposes something worth understanding before you run it: because .call() always resolves a half-open trial to either closed (success) or open (failure) before returning control to you, the state after a call has completed can never be half_open, only the state before a call can be. If you only log state after each call resolves, as most naive logging does, you will never see half-open in your logs at all, even though the breaker passes through it on every recovery.
Set the service unhealthy and run it:
curl -X POST http://127.0.0.1:8811/admin/set-healthy/false
python recovery_demo.py
Real captured output, unedited:
t= 3.17s request 1: pre=closed post=closed call_took= 3.17s -> {'source': 'fallback (call failed)'}
t= 7.94s request 2: pre=closed post=closed call_took= 3.27s -> {'source': 'fallback (call failed)'}
t= 12.69s request 3: pre=closed post=open call_took= 3.25s -> {'source': 'fallback (call failed)'}
t= 14.19s request 4: pre=open post=open call_took= 0.00s -> {'source': 'fallback (breaker open)'}
t= 15.69s request 5: pre=open post=open call_took= 0.00s -> {'source': 'fallback (breaker open)'}
--- ops fixed the dependency: service is healthy again ---
t= 17.65s request 6: pre=half_open post=closed call_took= 0.19s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 19.40s request 7: pre=closed post=closed call_took= 0.25s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 21.19s request 8: pre=closed post=closed call_took= 0.29s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 22.95s request 9: pre=closed post=closed call_took= 0.26s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 24.70s request 10: pre=closed post=closed call_took= 0.25s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 26.47s request 11: pre=closed post=closed call_took= 0.28s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 28.24s request 12: pre=closed post=closed call_took= 0.26s -> {'sku': 'WIDGET-1', 'quantity': 42}
Read through what actually happened: requests 1 and 2 fail and stay closed (below the threshold of 3). Request 3 is the third consecutive failure, and the breaker flips to open mid-call, note post=open. Requests 4 and 5 fail in 0.00 seconds, correctly rejected without touching the network, because the 4 second recovery timeout has not elapsed yet. Then the dependency is fixed. Request 6 is the first call made after the timeout elapses, its pre-call state genuinely reads half_open, exactly the transient state described above, and because the dependency really is healthy again now, the trial succeeds and the breaker closes before the call even returns. Every request after that is a fast, ordinary success.
Step 8: Fix the Thundering Herd on Half-Open
Look back at circuit_breaker.py and find _half_open_trial_in_flight. It exists to guarantee that when several requests arrive at the exact same moment the breaker becomes eligible for half-open, only one of them is treated as the trial call. Everyone else gets fast-failed until that one trial resolves. Without this guard, every one of those simultaneous requests would see the same half-open state and all charge through as if each were the trial, which is precisely the flood the Microsoft documentation quoted in Step 4 warns against: a service that is barely coming back online gets hit with a burst of traffic at the worst possible moment.
To prove this is a real bug and not a theoretical one, create circuit_breaker_unguarded.py, an intentionally broken copy without that guard:
"""A deliberately broken variant of CircuitBreaker: it transitions to
HALF_OPEN correctly but does not limit how many concurrent callers get
treated as "the trial call." Used only to prove, with real measured
output, why the _half_open_trial_in_flight guard in circuit_breaker.py
is load-bearing and not just defensive style."""
import threading
import time
from enum import Enum
class CircuitState(Enum):
CLOSED = "closed"
OPEN = "open"
HALF_OPEN = "half_open"
class CircuitOpenError(Exception):
pass
class UnguardedCircuitBreaker:
def __init__(self, failure_threshold=5, recovery_timeout=10.0, expected_exceptions=(Exception,)):
self.failure_threshold = failure_threshold
self.recovery_timeout = recovery_timeout
self.expected_exceptions = expected_exceptions
self._state = CircuitState.CLOSED
self._failure_count = 0
self._opened_at = None
self._lock = threading.Lock()
def _peek_state(self):
if self._state == CircuitState.OPEN and self._opened_at is not None:
if time.monotonic() - self._opened_at >= self.recovery_timeout:
return CircuitState.HALF_OPEN
return self._state
def call(self, func, *args, **kwargs):
with self._lock:
current = self._peek_state()
if current == CircuitState.OPEN:
raise CircuitOpenError("circuit is open, failing fast")
# BUG: every caller that observes HALF_OPEN is let through as
# if it were "the" trial call. There is no single-trial gate.
if current == CircuitState.HALF_OPEN:
self._state = CircuitState.HALF_OPEN
try:
result = func(*args, **kwargs)
except self.expected_exceptions:
with self._lock:
self._failure_count += 1
self._state = CircuitState.OPEN
self._opened_at = time.monotonic()
raise
else:
with self._lock:
self._failure_count = 0
self._state = CircuitState.CLOSED
self._opened_at = None
return result
def force_open(self):
with self._lock:
self._state = CircuitState.OPEN
self._opened_at = time.monotonic()
Then create concurrent_test.py, which fires 10 threads at a breaker in the exact instant it becomes eligible for half-open, once against the real guarded breaker and once against the unguarded copy, while the dependency is still down:
"""Fires 10 threads at a breaker in the exact instant it becomes
eligible for HALF_OPEN, once with the guarded CircuitBreaker and once
with the deliberately-unguarded variant, and counts how many threads
actually reached the still-down dependency versus how many were
correctly fast-failed. The dependency is left unhealthy for this test,
so a "real" trial call takes ~3s and fails, exactly the case a
production outage looks like right as someone tries to recover."""
import threading
import time
from circuit_breaker import CircuitBreaker, CircuitOpenError as GuardedOpenError
from circuit_breaker_unguarded import UnguardedCircuitBreaker, CircuitOpenError as UnguardedOpenError
from inventory_client import check_inventory, DependencyError
N_THREADS = 10
def run_batch(breaker, open_error_type, label):
breaker.force_open_for_test() if hasattr(breaker, "force_open_for_test") else breaker.force_open()
time.sleep(1.1) # let the 1.0s recovery_timeout elapse
real_calls = []
fast_fails = []
lock = threading.Lock()
barrier = threading.Barrier(N_THREADS)
def worker():
barrier.wait() # line everyone up so they hit call() at the same instant
start = time.perf_counter()
try:
breaker.call(check_inventory, "WIDGET-1")
except open_error_type:
with lock:
fast_fails.append(time.perf_counter() - start)
except DependencyError:
with lock:
real_calls.append(time.perf_counter() - start)
threads = [threading.Thread(target=worker) for _ in range(N_THREADS)]
batch_start = time.perf_counter()
for t in threads:
t.start()
for t in threads:
t.join()
batch_elapsed = time.perf_counter() - batch_start
print(f"\n{label}")
print(f" {len(real_calls)} of {N_THREADS} threads actually hit the still-down dependency (~3s each)")
print(f" {len(fast_fails)} of {N_THREADS} threads were fast-failed without touching the network")
print(f" batch wall time: {batch_elapsed:.2f}s")
def main():
guarded = CircuitBreaker(failure_threshold=3, recovery_timeout=1.0, expected_exceptions=(DependencyError,))
run_batch(guarded, GuardedOpenError, "GUARDED breaker (circuit_breaker.py)")
unguarded = UnguardedCircuitBreaker(failure_threshold=3, recovery_timeout=1.0, expected_exceptions=(DependencyError,))
run_batch(unguarded, UnguardedOpenError, "UNGUARDED breaker (circuit_breaker_unguarded.py)")
if __name__ == "__main__":
main()
With the service still set unhealthy, run it:
python concurrent_test.py
Real captured output:
GUARDED breaker (circuit_breaker.py)
1 of 10 threads actually hit the still-down dependency (~3s each)
9 of 10 threads were fast-failed without touching the network
batch wall time: 3.18s
UNGUARDED breaker (circuit_breaker_unguarded.py)
10 of 10 threads actually hit the still-down dependency (~3s each)
0 of 10 threads were fast-failed without touching the network
batch wall time: 3.56s
With the guard, 1 request pays the real cost of checking, and 9 are protected. Without it, all 10 threads independently decide they are allowed through, and all 10 hit the dependency at once, exactly the flood a circuit breaker exists to prevent, happening at the one moment (the instant of recovery) when the dependency can least afford it. A single boolean flag, checked and set inside the same lock as the rest of the state machine, is the entire fix.
Common Mistakes and Gotchas
- Counting every exception as a dependency failure. As shown in Step 5, this trips your breaker on client bugs and validation errors, not just real outages. Always scope
expected_exceptionsto failures that indicate the dependency itself is unhealthy: timeouts, connection errors, and 5xx responses. Never 4xx. - Holding a lock across the network call. If your implementation locks before calling the dependency and unlocks after, every concurrent request queues up behind whichever one is currently waiting on a slow dependency, which recreates the exact pile-up problem you were trying to fix, just inside your own process instead of at the dependency.
- Not guarding the half-open trial. Step 8 showed this is not hypothetical: without a single-trial guard, a recovering dependency gets hit with a full burst of traffic at the worst possible moment.
- One breaker instance shared across unrelated dependencies. If your database and your email provider share a breaker, a slow email provider trips protection for your database too. Use one breaker per dependency, not one per process.
- No fallback. An open circuit with no fallback just converts a slow failure into a fast one. That is still an improvement, but the real value of graceful degradation, a cached response, a default value, a queued retry, is a design decision you have to make deliberately, it does not happen automatically.
- Confusing a circuit breaker with a retry. They solve different problems and are often combined, but Microsoft’s own guidance is explicit about the distinction: a retry pattern assumes an operation will eventually succeed and keeps trying it; a circuit breaker assumes an operation is likely to keep failing and stops trying it. If you combine them, your retry logic needs to respect the breaker’s open state instead of retrying straight through it.
- Threshold and timeout picked without thinking about the dependency. A failure threshold of 1 trips on a single transient blip. A threshold of 100 means 100 real users wait through 100 real failures before anyone gets protected. A recovery timeout that is too short reopens the breaker on almost every attempt if the dependency recovers slowly; one that is too long keeps failing fast long after the dependency is actually healthy again.
- Assuming one process’s breaker state is enough. Every example in this tutorial runs in a single process. If your service runs multiple instances, each one gets its own independent breaker unless you deliberately share state (for example, in Redis), which means one instance can be happily closed while another is open against the same dependency. That is often fine, sometimes it is not; know which case you are in.
How to Verify Your Circuit Breaker Actually Works
Do not trust a circuit breaker implementation you have not watched fail and recover for real. Before you consider one done, whether it is this one or your own production version, confirm all four of these, the same way this tutorial did:
- Baseline cost. Call the dependency directly with no protection while it is down, and measure the real wall-clock cost, like Step 3’s 65.30 seconds. Without this number, you cannot tell whether your breaker is actually helping.
- Fail-fast after the threshold. Confirm the breaker opens after exactly your configured number of consecutive failures, not before and not after, and that calls made while open never reach the real dependency (check this by timing them: they should return in single-digit milliseconds, not seconds).
- Full recovery cycle. Fix the dependency while the breaker is open and confirm it automatically returns to closed on its own, without a restart or manual intervention, within one recovery timeout window of the fix.
- Concurrency under real load. Fire multiple simultaneous requests at the breaker right as it becomes eligible for half-open, and confirm only one of them reaches the real dependency. This is the check most homegrown circuit breakers skip, and the one most likely to hide a real bug, as Step 8 demonstrated.
Next Steps
The circuit breaker in this tutorial is deliberately minimal, built to make the state machine visible rather than to ship as-is. A few directions worth exploring next:
- For production Python code, reach for a maintained library instead of hand-rolling this. pybreaker implements the same Nygard pattern with a success threshold (requiring multiple consecutive successes in half-open before fully closing, not just one), pluggable event listeners for monitoring, and optional Redis-backed shared state so multiple instances of your service can share one breaker’s view of a dependency’s health.
- Expose the breaker’s state as a metric (a simple gauge with values for closed, open, and half-open) so you can alert when a dependency trips, instead of only discovering it from a spike in fallback responses.
- Combine a circuit breaker with a rate limiter. A rate limiter protects your service from too many callers; a circuit breaker protects your service from a dependency that cannot handle the callers it already has. This site’s token bucket and sliding window rate limiter tutorial builds the other half of that pair from scratch, the same way this one did.
- If you are routing between multiple upstream options rather than protecting a single dependency, see this site’s multi-tier router with automatic fallback tutorial, which applies a related but distinct pattern: falling back to a different provider entirely, rather than failing fast and waiting.
- Try breaking your own implementation the way Step 8 did here. Write a concurrency test before you trust any resilience code, not after.








No Comment! Be the first one.