TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Build a Circuit Breaker in Python to Stop Cascading Failures
Learning Hub

How to Build a Circuit Breaker in Python to Stop Cascading Failures

Learn how to build a circuit breaker from scratch in Python that detects a failing dependency, fails fast instead of piling up timeouts, and recovers automatically once it is healthy again.

August 20, 2026 22 Min Read
52

If you have ever put a real dependency behind an API call, a payment processor, a database, another team’s microservice, you have probably watched what happens when that dependency goes down: every caller keeps calling it anyway. Requests pile up waiting on a timeout that hasn’t fired yet. Threads and connections fill up holding those requests. The one broken dependency starts taking down services that were otherwise perfectly healthy. This is called a cascading failure, and it is one of the most common ways a small outage turns into a big one.

Table Of Content

  • What You Will Build
  • Prerequisites
  • Step 1: Set Up Your Project
  • Step 2: Build a Dependency You Can Break On Demand
  • Step 3: See the Problem for Real
  • Step 4: Build the Circuit Breaker State Machine
  • Why the lock, and why the call happens outside it
  • Why the half-open transition happens lazily
  • Step 5: Do Not Trip the Breaker on Your Own Bugs
  • Step 6: Wire the Breaker Around Real Calls, With a Fallback
  • Step 7: Watch the Full Recovery Cycle
  • Step 8: Fix the Thundering Herd on Half-Open
  • Common Mistakes and Gotchas
  • How to Verify Your Circuit Breaker Actually Works
  • Next Steps

A circuit breaker is the fix. It is a small piece of code that sits between your application and a dependency it calls over the network. It watches for failures, and once it sees enough of them, it stops calling the dependency at all for a while, failing every request immediately instead. This is the same idea as the circuit breaker in your home’s electrical panel: when something downstream draws too much current, the breaker trips and cuts the circuit before the wiring overheats, rather than letting the fault keep drawing power indefinitely.

This tutorial builds a circuit breaker from scratch in Python: no framework, no external service, just the standard library plus a small demo dependency you can break on command. You will reproduce a real cascading failure, measure exactly how much time it wastes, fix it, watch the breaker recover automatically once the dependency comes back, and then find and fix a genuine concurrency bug in the recovery logic itself. Every command below was run for real while writing this post, and the timings and output you will see are the actual captured results, not invented numbers.

What You Will Build

A small Python library, circuit_breaker.py, implementing a three-state circuit breaker (closed, open, half-open) with no third-party dependencies. Around it, you will build:

  • A deliberately flaky FastAPI service that you can flip between healthy and unhealthy on demand, so you can reproduce an outage instead of just reading about one
  • A naive client that calls the flaky service directly, to measure how bad an unprotected cascading failure actually is
  • A protected client that wraps the same calls in your circuit breaker, with a fallback response for graceful degradation
  • A script that watches the breaker recover automatically once the dependency comes back healthy
  • A concurrency test that proves, with real measured numbers, why the breaker needs to limit itself to exactly one trial call during recovery, and what goes wrong if it does not

Prerequisites

  • Python 3.10 or newer (this tutorial used Python 3.13.14 on Windows, but nothing here is platform-specific)
  • Comfort with basic Python: classes, exceptions, and what a decorator-free wrapper function does
  • Basic familiarity with HTTP status codes and what a client timeout is. This tutorial explains the rest
  • No external services, accounts, or API keys. Everything in this tutorial runs on your own machine
  • About 40 minutes

Step 1: Set Up Your Project

Create a project folder and an isolated virtual environment, then install the three libraries this tutorial needs: a web framework to build the demo dependency, a server to run it, and an HTTP client to call it.

mkdir circuit-breaker-demo
cd circuit-breaker-demo
python -m venv .venv
.venv\Scripts\activate   # on Linux/macOS: source .venv/bin/activate
pip install fastapi uvicorn httpx

This tutorial was written against fastapi 0.141.1, uvicorn 0.52.4, and httpx 0.28.1. Newer versions should work the same way; none of the APIs used here are new or unstable.

Step 2: Build a Dependency You Can Break On Demand

To reproduce a real outage instead of just describing one, you need a dependency whose health you can control. Create flaky_service.py:

"""A fake downstream dependency (e.g. an inventory service) that we can
switch between healthy and unhealthy at runtime, so we can reproduce a
real outage on demand instead of just describing one."""

import time
from fastapi import FastAPI, HTTPException

app = FastAPI()

STATE = {"healthy": True}


@app.post("/admin/set-healthy/{value}")
def set_healthy(value: bool):
    STATE["healthy"] = value
    return {"healthy": STATE["healthy"]}


@app.get("/inventory/{sku}")
def get_inventory(sku: str):
    if not sku.isupper():
        # A malformed request. This is a client bug, not evidence that
        # the dependency is unhealthy, so it must NOT count toward the
        # breaker's failure threshold.
        raise HTTPException(status_code=400, detail="sku must be uppercase")

    if STATE["healthy"]:
        return {"sku": sku, "quantity": 42}

    # Simulate a real outage: the dependency does not fail instantly,
    # it hangs near its connection/read timeout before giving up. This
    # is the case that actually hurts callers, an instant 500 is cheap
    # to handle, a slow failure is not.
    time.sleep(3.0)
    raise HTTPException(status_code=503, detail="inventory-service unavailable")

Two design choices here matter more than they look:

First, when unhealthy, the service does not fail instantly. It calls time.sleep(3.0) before returning a 503. A real dependency rarely fails the instant something goes wrong; a database connection has to time out, a load balancer has to give up on a backend, a TCP connection has to hit its retransmit limit. An instant failure is cheap for a caller to handle. A slow failure is what actually causes cascading failures, because every caller sits there holding a thread, a connection, or an async task waiting for those 3 seconds to pass, over and over, for every single request.

Second, notice the sku.isupper() check at the top, which rejects malformed input with a 400 before the health check ever runs. Keep this in mind, you will use it in Step 5 to prove something important about what should and should not trip a circuit breaker.

Start the service in one terminal and leave it running for the rest of this tutorial:

uvicorn flaky_service:app --host 127.0.0.1 --port 8811

Plain uvicorn does not reload your code automatically. If you edit flaky_service.py later, stop this process (Ctrl+C) and restart it, or your changes will not take effect. This tutorial ran into exactly that: a stale server process kept serving old code after an edit, and produced a confusing false result until it was restarted. If a step’s output does not match what this tutorial shows, restarting the server is the first thing to try.

In a second terminal, confirm it works both ways:

curl http://127.0.0.1:8811/inventory/WIDGET-1
# {"sku":"WIDGET-1","quantity":42}

curl -X POST http://127.0.0.1:8811/admin/set-healthy/false
# {"healthy":false}

curl http://127.0.0.1:8811/inventory/WIDGET-1
# (waits about 3 seconds)
# {"detail":"inventory-service unavailable"}

That 3 second wait before the error is the entire problem this tutorial exists to solve.

Step 3: See the Problem for Real

Before building the fix, measure the actual cost of not having one. Create naive_client.py, which does what most application code does by default: call the dependency, catch whatever error comes back, move on to the next request. Nothing here remembers that the last 5 calls all failed the same way.

"""What most services do by default: call the dependency, catch the
error, move on. No memory of the last call. Every request pays the
full cost of finding out the dependency is still down."""

import time
import httpx

BASE_URL = "http://127.0.0.1:8811"


def check_inventory(sku: str) -> dict:
    resp = httpx.get(f"{BASE_URL}/inventory/{sku}", timeout=5.0)
    resp.raise_for_status()
    return resp.json()


def main():
    total_start = time.perf_counter()
    successes = 0
    failures = 0

    for i in range(1, 21):
        start = time.perf_counter()
        try:
            result = check_inventory("WIDGET-1")
            elapsed = time.perf_counter() - start
            print(f"request {i:2d}: OK   in {elapsed:5.2f}s -> {result}")
            successes += 1
        except httpx.HTTPStatusError as exc:
            elapsed = time.perf_counter() - start
            print(f"request {i:2d}: FAIL in {elapsed:5.2f}s -> HTTP {exc.response.status_code}")
            failures += 1
        except httpx.TimeoutException:
            elapsed = time.perf_counter() - start
            print(f"request {i:2d}: FAIL in {elapsed:5.2f}s -> client timeout")
            failures += 1

    total_elapsed = time.perf_counter() - total_start
    print(f"\n{successes} succeeded, {failures} failed, total wall time {total_elapsed:.2f}s")


if __name__ == "__main__":
    main()

With the service still set unhealthy from Step 2, run it:

python naive_client.py

Here is the real output from this run:

request  1: FAIL in  3.18s -> HTTP 503
request  2: FAIL in  3.29s -> HTTP 503
request  3: FAIL in  3.27s -> HTTP 503
...
request 20: FAIL in  3.27s -> HTTP 503

0 succeeded, 20 failed, total wall time 65.30s

Twenty requests, every single one destined to fail, and it still took 65.30 seconds to find that out, because each one had to independently rediscover that the dependency was down. In a real service, those 20 requests are not sequential like this demo, they are concurrent, each one holding a thread or a connection for 3 seconds while it waits. Multiply that by your actual request volume and you can see how one slow dependency turns into an exhausted thread pool for a service that has nothing to do with the original problem.

Step 4: Build the Circuit Breaker State Machine

A circuit breaker is a state machine with three states. Understanding what each one means is more important than any of the code:

  • Closed: normal operation. Calls pass through to the real dependency. The breaker counts recent failures.
  • Open: the failure count crossed a threshold. Every call fails immediately with an error, without touching the network at all, until a recovery timeout elapses.
  • Half-Open: the recovery timeout has elapsed. Exactly one trial call is allowed through to check whether the dependency has recovered. If it succeeds, the breaker closes and resumes normal operation. If it fails, the breaker reopens and the timeout starts again.

This is not a new idea. Michael Nygard popularized it in his book Release It!, and Martin Fowler’s widely read summary of the pattern puts it plainly: You wrap a protected function call in a circuit breaker object, which monitors for failures. Once the failures reach a certain threshold, the circuit breaker trips, and all further calls to the circuit breaker return with an error, without the protected call being made at all. Microsoft’s Azure Architecture Center describes the same three states and adds the specific reason the half-open state only allows a limited number of requests through: The Half-Open state helps prevent a recovering service from suddenly being flooded with requests. As a service recovers, it might be able to support a limited volume of requests until the recovery is complete. But while recovery is in progress, a flood of work can cause the service to time out or fail again. You will prove that exact failure mode yourself in Step 8.

Create circuit_breaker.py:

"""A circuit breaker built from scratch: three states, one lock, no
dependencies. Mirrors the state machine used by libraries like pybreaker
and resilience4j, simplified for teaching."""

import threading
import time
from enum import Enum


class CircuitState(Enum):
    CLOSED = "closed"        # normal operation, calls go through
    OPEN = "open"            # tripped, calls fail fast without hitting the network
    HALF_OPEN = "half_open"  # trial period, exactly one call is allowed through


class CircuitOpenError(Exception):
    """Raised by CircuitBreaker.call() instead of invoking the wrapped
    function, when the breaker is open or a half-open trial is already
    in flight."""


class CircuitBreaker:
    def __init__(self, failure_threshold=5, recovery_timeout=10.0, expected_exceptions=(Exception,)):
        self.failure_threshold = failure_threshold
        self.recovery_timeout = recovery_timeout
        self.expected_exceptions = expected_exceptions

        self._state = CircuitState.CLOSED
        self._failure_count = 0
        self._opened_at = None
        self._half_open_trial_in_flight = False
        self._lock = threading.Lock()

    @property
    def state(self) -> CircuitState:
        """Read-only view of the state. Does not itself trigger the
        OPEN -> HALF_OPEN transition; only call() does that, so that a
        transition always happens right before a real trial call."""
        with self._lock:
            return self._peek_state()

    def _peek_state(self) -> CircuitState:
        if self._state == CircuitState.OPEN and self._opened_at is not None:
            if time.monotonic() - self._opened_at >= self.recovery_timeout:
                return CircuitState.HALF_OPEN
        return self._state

    def call(self, func, *args, **kwargs):
        with self._lock:
            current = self._peek_state()

            if current == CircuitState.OPEN:
                raise CircuitOpenError("circuit is open, failing fast")

            if current == CircuitState.HALF_OPEN:
                if self._half_open_trial_in_flight:
                    raise CircuitOpenError("circuit is half-open, a trial call is already in flight")
                self._state = CircuitState.HALF_OPEN
                self._half_open_trial_in_flight = True

        # The wrapped call happens OUTSIDE the lock. Holding a lock across
        # a network call would serialize every request in the process
        # behind the slowest one, exactly the pile-up we are trying to avoid.
        try:
            result = func(*args, **kwargs)
        except self.expected_exceptions:
            self._on_failure()
            raise
        else:
            self._on_success()
            return result

    def _on_success(self):
        with self._lock:
            self._failure_count = 0
            self._state = CircuitState.CLOSED
            self._opened_at = None
            self._half_open_trial_in_flight = False

    def force_open_for_test(self):
        """Test-only seam: jump straight to OPEN without needing real
        failures first. Do not ship this in a production breaker, it
        exists here purely to set up the concurrency demo quickly."""
        with self._lock:
            self._state = CircuitState.OPEN
            self._opened_at = time.monotonic()

    def _on_failure(self):
        with self._lock:
            self._half_open_trial_in_flight = False
            self._failure_count += 1

            if self._state == CircuitState.HALF_OPEN:
                # The trial call failed: the dependency is still down.
                # Re-open and restart the recovery clock.
                self._state = CircuitState.OPEN
                self._opened_at = time.monotonic()
            elif self._failure_count >= self.failure_threshold:
                self._state = CircuitState.OPEN
                self._opened_at = time.monotonic()

Why the lock, and why the call happens outside it

Every state read and every state transition happens inside self._lock, because multiple threads (or async tasks running on multiple workers) can call .call() at the same time, and two threads racing to both decide independently that they are allowed to open the breaker, or both decide they are the half-open trial call, would corrupt the state. But the line that matters most is this one: func(*args, **kwargs) is called outside the lock. If it were called while holding the lock, every request in your process would queue up behind whichever one happened to be talking to the slow dependency first, turning your circuit breaker into an accidental global mutex. The lock protects the breaker’s bookkeeping, not the call itself.

Why the half-open transition happens lazily

Notice that _peek_state() computes whether enough time has passed to justify moving from open to half-open, but only .call() actually commits that transition (by setting self._state = CircuitState.HALF_OPEN). The read-only .state property can tell you the breaker would be half-open, without that observation itself consuming the one trial slot. You will see why this distinction matters, and a genuine consequence of it, in Step 7.

Step 5: Do Not Trip the Breaker on Your Own Bugs

Fowler’s original description includes a warning that is easy to skip past: Not all errors should trip the circuit, some should reflect normal failures and be dealt with as part of regular logic. If your circuit breaker treats every exception as evidence the dependency is unhealthy, then a client sending malformed requests, your own validation bug, a typo in a SKU, will look identical to a real outage and trip the breaker for every other caller too, even though retrying the exact same bad request would fail the exact same way forever.

Create inventory_client.py to draw that line explicitly, translating the flaky service’s responses into two distinct exception types:

"""Talks to the inventory service and translates transport-level and
HTTP-level failures into two distinct exception types, so the circuit
breaker can tell "the dependency is unhealthy" apart from "the caller
sent a bad request." Only the first should ever trip a breaker."""

import httpx

BASE_URL = "http://127.0.0.1:8811"


class DependencyError(Exception):
    """The dependency itself is the problem: timeout, connection
    refused, or a 5xx response."""


class BadRequestError(Exception):
    """The caller sent something the dependency correctly rejected.
    Retrying or tripping a breaker over this would be pointless, the
    next identical request would fail the same way."""


def check_inventory(sku: str) -> dict:
    try:
        resp = httpx.get(f"{BASE_URL}/inventory/{sku}", timeout=5.0)
    except httpx.TimeoutException as exc:
        raise DependencyError(f"timed out calling inventory service: {exc}") from exc
    except httpx.ConnectError as exc:
        raise DependencyError(f"could not connect to inventory service: {exc}") from exc

    if resp.status_code == 400:
        raise BadRequestError(resp.json().get("detail", "bad request"))
    if resp.status_code >= 500:
        raise DependencyError(f"inventory service returned HTTP {resp.status_code}")

    resp.raise_for_status()
    return resp.json()

The circuit breaker will be configured to treat only DependencyError as a failure. A BadRequestError will propagate normally without ever reaching _on_failure(), because it will not match the breaker’s expected_exceptions filter. Prove this before moving on. Create bad_request_test.py:

"""Proves that a client error (bad SKU) does not move the breaker
toward OPEN, even when it happens repeatedly, while the service is
otherwise perfectly healthy."""

from circuit_breaker import CircuitBreaker
from inventory_client import check_inventory, DependencyError, BadRequestError

breaker = CircuitBreaker(failure_threshold=3, recovery_timeout=5.0, expected_exceptions=(DependencyError,))


def call(sku):
    try:
        result = breaker.call(check_inventory, sku)
        print(f"sku={sku!r:12} state={breaker.state.value:8} -> OK {result}")
    except BadRequestError as exc:
        print(f"sku={sku!r:12} state={breaker.state.value:8} -> BadRequestError: {exc}")
    except DependencyError as exc:
        print(f"sku={sku!r:12} state={breaker.state.value:8} -> DependencyError: {exc}")


for _ in range(5):
    call("bad-sku")  # lowercase, the service rejects this with a 400 every time

call("WIDGET-1")  # a well-formed request right after, to prove the breaker never moved

With the service set back to healthy (curl -X POST http://127.0.0.1:8811/admin/set-healthy/true), run it:

python bad_request_test.py

Real captured output:

sku='bad-sku'    state=closed   -> BadRequestError: sku must be uppercase
sku='bad-sku'    state=closed   -> BadRequestError: sku must be uppercase
sku='bad-sku'    state=closed   -> BadRequestError: sku must be uppercase
sku='bad-sku'    state=closed   -> BadRequestError: sku must be uppercase
sku='bad-sku'    state=closed   -> BadRequestError: sku must be uppercase
sku='WIDGET-1'   state=closed   -> OK {'sku': 'WIDGET-1', 'quantity': 42}

Five consecutive errors, more than enough to trip a breaker with a failure threshold of 3, and the state never moves off closed, because none of them were DependencyError. The well-formed request right after succeeds normally. This is the entire point: the breaker should only ever open because of the dependency, never because of the caller.

Step 6: Wire the Breaker Around Real Calls, With a Fallback

Create protected_client.py. It runs the exact same 20 requests as naive_client.py in Step 3, but every call goes through the circuit breaker, and when the breaker is open, it returns a fallback value instead of letting the caller wait:

"""The same 20 requests as naive_client.py, but every call to the
dependency goes through a CircuitBreaker with a fallback. Compare the
total wall-clock time and the per-request latency to naive_client.py's
output."""

import time

from circuit_breaker import CircuitBreaker, CircuitOpenError
from inventory_client import check_inventory, DependencyError, BadRequestError

breaker = CircuitBreaker(
    failure_threshold=3,
    recovery_timeout=8.0,
    expected_exceptions=(DependencyError,),
)


def get_inventory_with_fallback(sku: str) -> dict:
    try:
        return breaker.call(check_inventory, sku)
    except CircuitOpenError:
        # Graceful degradation: tell the truth about staleness instead
        # of hanging the caller for 3+ seconds only to fail anyway.
        return {"sku": sku, "quantity": None, "source": "fallback (breaker open)"}
    except DependencyError:
        return {"sku": sku, "quantity": None, "source": "fallback (call failed)"}


def main():
    total_start = time.perf_counter()

    for i in range(1, 21):
        start = time.perf_counter()
        result = get_inventory_with_fallback("WIDGET-1")
        elapsed = time.perf_counter() - start
        print(f"request {i:2d}: state={breaker.state.value:10s} in {elapsed:5.2f}s -> {result}")

    total_elapsed = time.perf_counter() - total_start
    print(f"\ntotal wall time {total_elapsed:.2f}s")


if __name__ == "__main__":
    main()

Graceful degradation is the point of the fallback: instead of hanging for 3 seconds and then failing anyway, a caller with the circuit open gets an honest, immediate answer that the data might be stale. Set the service unhealthy again and run it:

curl -X POST http://127.0.0.1:8811/admin/set-healthy/false
python protected_client.py

Real captured output:

request  1: state=closed     in  3.17s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (call failed)'}
request  2: state=closed     in  3.27s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (call failed)'}
request  3: state=open       in  3.27s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (call failed)'}
request  4: state=open       in  0.00s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (breaker open)'}
request  5: state=open       in  0.00s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (breaker open)'}
...
request 20: state=open       in  0.00s -> {'sku': 'WIDGET-1', 'quantity': None, 'source': 'fallback (breaker open)'}

total wall time 9.71s

9.71 seconds instead of 65.30 seconds, a 6.7x reduction, for the exact same 20 requests against the exact same dead dependency. The first three requests still pay the real cost, that is unavoidable and correct, the breaker needs real evidence before it trips. But requests 4 through 20 fail in 0.00 seconds instead of 3.27 seconds each, because the breaker stopped them before they ever touched the network.

Step 7: Watch the Full Recovery Cycle

So far you have seen closed and open. To see half-open, and the automatic recovery back to closed, you need a longer-running script that flips the dependency back to healthy partway through. Create recovery_demo.py:

"""Runs long enough to observe the full cycle: CLOSED -> OPEN (after 3
real failures) -> repeated fail-fast while OPEN -> HALF_OPEN (once the
recovery timeout elapses) -> CLOSED again, once the dependency is
flipped back to healthy partway through."""

import time

import httpx

from circuit_breaker import CircuitBreaker, CircuitOpenError
from inventory_client import check_inventory, DependencyError

ADMIN_URL = "http://127.0.0.1:8811/admin/set-healthy"

breaker = CircuitBreaker(
    failure_threshold=3,
    recovery_timeout=4.0,
    expected_exceptions=(DependencyError,),
)


def get_inventory_with_fallback(sku: str) -> dict:
    try:
        return breaker.call(check_inventory, sku)
    except CircuitOpenError:
        return {"source": "fallback (breaker open)"}
    except DependencyError:
        return {"source": "fallback (call failed)"}


def main():
    start = time.perf_counter()

    for i in range(1, 13):
        if i == 6:
            httpx.post(f"{ADMIN_URL}/true", timeout=5.0)
            print("          --- ops fixed the dependency: service is healthy again ---")

        pre_state = breaker.state.value
        call_start = time.perf_counter()
        result = get_inventory_with_fallback("WIDGET-1")
        elapsed = time.perf_counter() - call_start
        post_state = breaker.state.value
        t = time.perf_counter() - start
        print(
            f"t={t:6.2f}s  request {i:2d}: pre={pre_state:10s} post={post_state:10s} "
            f"call_took={elapsed:5.2f}s -> {result}"
        )

        time.sleep(1.5)


if __name__ == "__main__":
    main()

This script prints the breaker’s state both immediately before and immediately after each call. That is deliberate, and it exposes something worth understanding before you run it: because .call() always resolves a half-open trial to either closed (success) or open (failure) before returning control to you, the state after a call has completed can never be half_open, only the state before a call can be. If you only log state after each call resolves, as most naive logging does, you will never see half-open in your logs at all, even though the breaker passes through it on every recovery.

Set the service unhealthy and run it:

curl -X POST http://127.0.0.1:8811/admin/set-healthy/false
python recovery_demo.py

Real captured output, unedited:

t=  3.17s  request  1: pre=closed     post=closed     call_took= 3.17s -> {'source': 'fallback (call failed)'}
t=  7.94s  request  2: pre=closed     post=closed     call_took= 3.27s -> {'source': 'fallback (call failed)'}
t= 12.69s  request  3: pre=closed     post=open       call_took= 3.25s -> {'source': 'fallback (call failed)'}
t= 14.19s  request  4: pre=open       post=open       call_took= 0.00s -> {'source': 'fallback (breaker open)'}
t= 15.69s  request  5: pre=open       post=open       call_took= 0.00s -> {'source': 'fallback (breaker open)'}
          --- ops fixed the dependency: service is healthy again ---
t= 17.65s  request  6: pre=half_open  post=closed     call_took= 0.19s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 19.40s  request  7: pre=closed     post=closed     call_took= 0.25s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 21.19s  request  8: pre=closed     post=closed     call_took= 0.29s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 22.95s  request  9: pre=closed     post=closed     call_took= 0.26s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 24.70s  request 10: pre=closed     post=closed     call_took= 0.25s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 26.47s  request 11: pre=closed     post=closed     call_took= 0.28s -> {'sku': 'WIDGET-1', 'quantity': 42}
t= 28.24s  request 12: pre=closed     post=closed     call_took= 0.26s -> {'sku': 'WIDGET-1', 'quantity': 42}

Read through what actually happened: requests 1 and 2 fail and stay closed (below the threshold of 3). Request 3 is the third consecutive failure, and the breaker flips to open mid-call, note post=open. Requests 4 and 5 fail in 0.00 seconds, correctly rejected without touching the network, because the 4 second recovery timeout has not elapsed yet. Then the dependency is fixed. Request 6 is the first call made after the timeout elapses, its pre-call state genuinely reads half_open, exactly the transient state described above, and because the dependency really is healthy again now, the trial succeeds and the breaker closes before the call even returns. Every request after that is a fast, ordinary success.

Step 8: Fix the Thundering Herd on Half-Open

Look back at circuit_breaker.py and find _half_open_trial_in_flight. It exists to guarantee that when several requests arrive at the exact same moment the breaker becomes eligible for half-open, only one of them is treated as the trial call. Everyone else gets fast-failed until that one trial resolves. Without this guard, every one of those simultaneous requests would see the same half-open state and all charge through as if each were the trial, which is precisely the flood the Microsoft documentation quoted in Step 4 warns against: a service that is barely coming back online gets hit with a burst of traffic at the worst possible moment.

To prove this is a real bug and not a theoretical one, create circuit_breaker_unguarded.py, an intentionally broken copy without that guard:

"""A deliberately broken variant of CircuitBreaker: it transitions to
HALF_OPEN correctly but does not limit how many concurrent callers get
treated as "the trial call." Used only to prove, with real measured
output, why the _half_open_trial_in_flight guard in circuit_breaker.py
is load-bearing and not just defensive style."""

import threading
import time
from enum import Enum


class CircuitState(Enum):
    CLOSED = "closed"
    OPEN = "open"
    HALF_OPEN = "half_open"


class CircuitOpenError(Exception):
    pass


class UnguardedCircuitBreaker:
    def __init__(self, failure_threshold=5, recovery_timeout=10.0, expected_exceptions=(Exception,)):
        self.failure_threshold = failure_threshold
        self.recovery_timeout = recovery_timeout
        self.expected_exceptions = expected_exceptions

        self._state = CircuitState.CLOSED
        self._failure_count = 0
        self._opened_at = None
        self._lock = threading.Lock()

    def _peek_state(self):
        if self._state == CircuitState.OPEN and self._opened_at is not None:
            if time.monotonic() - self._opened_at >= self.recovery_timeout:
                return CircuitState.HALF_OPEN
        return self._state

    def call(self, func, *args, **kwargs):
        with self._lock:
            current = self._peek_state()
            if current == CircuitState.OPEN:
                raise CircuitOpenError("circuit is open, failing fast")
            # BUG: every caller that observes HALF_OPEN is let through as
            # if it were "the" trial call. There is no single-trial gate.
            if current == CircuitState.HALF_OPEN:
                self._state = CircuitState.HALF_OPEN

        try:
            result = func(*args, **kwargs)
        except self.expected_exceptions:
            with self._lock:
                self._failure_count += 1
                self._state = CircuitState.OPEN
                self._opened_at = time.monotonic()
            raise
        else:
            with self._lock:
                self._failure_count = 0
                self._state = CircuitState.CLOSED
                self._opened_at = None
            return result

    def force_open(self):
        with self._lock:
            self._state = CircuitState.OPEN
            self._opened_at = time.monotonic()

Then create concurrent_test.py, which fires 10 threads at a breaker in the exact instant it becomes eligible for half-open, once against the real guarded breaker and once against the unguarded copy, while the dependency is still down:

"""Fires 10 threads at a breaker in the exact instant it becomes
eligible for HALF_OPEN, once with the guarded CircuitBreaker and once
with the deliberately-unguarded variant, and counts how many threads
actually reached the still-down dependency versus how many were
correctly fast-failed. The dependency is left unhealthy for this test,
so a "real" trial call takes ~3s and fails, exactly the case a
production outage looks like right as someone tries to recover."""

import threading
import time

from circuit_breaker import CircuitBreaker, CircuitOpenError as GuardedOpenError
from circuit_breaker_unguarded import UnguardedCircuitBreaker, CircuitOpenError as UnguardedOpenError
from inventory_client import check_inventory, DependencyError

N_THREADS = 10


def run_batch(breaker, open_error_type, label):
    breaker.force_open_for_test() if hasattr(breaker, "force_open_for_test") else breaker.force_open()
    time.sleep(1.1)  # let the 1.0s recovery_timeout elapse

    real_calls = []
    fast_fails = []
    lock = threading.Lock()
    barrier = threading.Barrier(N_THREADS)

    def worker():
        barrier.wait()  # line everyone up so they hit call() at the same instant
        start = time.perf_counter()
        try:
            breaker.call(check_inventory, "WIDGET-1")
        except open_error_type:
            with lock:
                fast_fails.append(time.perf_counter() - start)
        except DependencyError:
            with lock:
                real_calls.append(time.perf_counter() - start)

    threads = [threading.Thread(target=worker) for _ in range(N_THREADS)]
    batch_start = time.perf_counter()
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    batch_elapsed = time.perf_counter() - batch_start

    print(f"\n{label}")
    print(f"  {len(real_calls)} of {N_THREADS} threads actually hit the still-down dependency (~3s each)")
    print(f"  {len(fast_fails)} of {N_THREADS} threads were fast-failed without touching the network")
    print(f"  batch wall time: {batch_elapsed:.2f}s")


def main():
    guarded = CircuitBreaker(failure_threshold=3, recovery_timeout=1.0, expected_exceptions=(DependencyError,))
    run_batch(guarded, GuardedOpenError, "GUARDED breaker (circuit_breaker.py)")

    unguarded = UnguardedCircuitBreaker(failure_threshold=3, recovery_timeout=1.0, expected_exceptions=(DependencyError,))
    run_batch(unguarded, UnguardedOpenError, "UNGUARDED breaker (circuit_breaker_unguarded.py)")


if __name__ == "__main__":
    main()

With the service still set unhealthy, run it:

python concurrent_test.py

Real captured output:

GUARDED breaker (circuit_breaker.py)
  1 of 10 threads actually hit the still-down dependency (~3s each)
  9 of 10 threads were fast-failed without touching the network
  batch wall time: 3.18s

UNGUARDED breaker (circuit_breaker_unguarded.py)
  10 of 10 threads actually hit the still-down dependency (~3s each)
  0 of 10 threads were fast-failed without touching the network
  batch wall time: 3.56s

With the guard, 1 request pays the real cost of checking, and 9 are protected. Without it, all 10 threads independently decide they are allowed through, and all 10 hit the dependency at once, exactly the flood a circuit breaker exists to prevent, happening at the one moment (the instant of recovery) when the dependency can least afford it. A single boolean flag, checked and set inside the same lock as the rest of the state machine, is the entire fix.

Common Mistakes and Gotchas

  • Counting every exception as a dependency failure. As shown in Step 5, this trips your breaker on client bugs and validation errors, not just real outages. Always scope expected_exceptions to failures that indicate the dependency itself is unhealthy: timeouts, connection errors, and 5xx responses. Never 4xx.
  • Holding a lock across the network call. If your implementation locks before calling the dependency and unlocks after, every concurrent request queues up behind whichever one is currently waiting on a slow dependency, which recreates the exact pile-up problem you were trying to fix, just inside your own process instead of at the dependency.
  • Not guarding the half-open trial. Step 8 showed this is not hypothetical: without a single-trial guard, a recovering dependency gets hit with a full burst of traffic at the worst possible moment.
  • One breaker instance shared across unrelated dependencies. If your database and your email provider share a breaker, a slow email provider trips protection for your database too. Use one breaker per dependency, not one per process.
  • No fallback. An open circuit with no fallback just converts a slow failure into a fast one. That is still an improvement, but the real value of graceful degradation, a cached response, a default value, a queued retry, is a design decision you have to make deliberately, it does not happen automatically.
  • Confusing a circuit breaker with a retry. They solve different problems and are often combined, but Microsoft’s own guidance is explicit about the distinction: a retry pattern assumes an operation will eventually succeed and keeps trying it; a circuit breaker assumes an operation is likely to keep failing and stops trying it. If you combine them, your retry logic needs to respect the breaker’s open state instead of retrying straight through it.
  • Threshold and timeout picked without thinking about the dependency. A failure threshold of 1 trips on a single transient blip. A threshold of 100 means 100 real users wait through 100 real failures before anyone gets protected. A recovery timeout that is too short reopens the breaker on almost every attempt if the dependency recovers slowly; one that is too long keeps failing fast long after the dependency is actually healthy again.
  • Assuming one process’s breaker state is enough. Every example in this tutorial runs in a single process. If your service runs multiple instances, each one gets its own independent breaker unless you deliberately share state (for example, in Redis), which means one instance can be happily closed while another is open against the same dependency. That is often fine, sometimes it is not; know which case you are in.

How to Verify Your Circuit Breaker Actually Works

Do not trust a circuit breaker implementation you have not watched fail and recover for real. Before you consider one done, whether it is this one or your own production version, confirm all four of these, the same way this tutorial did:

  1. Baseline cost. Call the dependency directly with no protection while it is down, and measure the real wall-clock cost, like Step 3’s 65.30 seconds. Without this number, you cannot tell whether your breaker is actually helping.
  2. Fail-fast after the threshold. Confirm the breaker opens after exactly your configured number of consecutive failures, not before and not after, and that calls made while open never reach the real dependency (check this by timing them: they should return in single-digit milliseconds, not seconds).
  3. Full recovery cycle. Fix the dependency while the breaker is open and confirm it automatically returns to closed on its own, without a restart or manual intervention, within one recovery timeout window of the fix.
  4. Concurrency under real load. Fire multiple simultaneous requests at the breaker right as it becomes eligible for half-open, and confirm only one of them reaches the real dependency. This is the check most homegrown circuit breakers skip, and the one most likely to hide a real bug, as Step 8 demonstrated.

Next Steps

The circuit breaker in this tutorial is deliberately minimal, built to make the state machine visible rather than to ship as-is. A few directions worth exploring next:

  • For production Python code, reach for a maintained library instead of hand-rolling this. pybreaker implements the same Nygard pattern with a success threshold (requiring multiple consecutive successes in half-open before fully closing, not just one), pluggable event listeners for monitoring, and optional Redis-backed shared state so multiple instances of your service can share one breaker’s view of a dependency’s health.
  • Expose the breaker’s state as a metric (a simple gauge with values for closed, open, and half-open) so you can alert when a dependency trips, instead of only discovering it from a spike in fallback responses.
  • Combine a circuit breaker with a rate limiter. A rate limiter protects your service from too many callers; a circuit breaker protects your service from a dependency that cannot handle the callers it already has. This site’s token bucket and sliding window rate limiter tutorial builds the other half of that pair from scratch, the same way this one did.
  • If you are routing between multiple upstream options rather than protecting a single dependency, see this site’s multi-tier router with automatic fallback tutorial, which applies a related but distinct pattern: falling back to a different provider entirely, rather than failing fast and waiting.
  • Try breaking your own implementation the way Step 8 did here. Write a concurrency test before you trust any resilience code, not after.

Tags:

Circuit BreakerFastAPIMicroservicesPythonResilience Engineering

Share

A 1936 photo of a telephone switchboard operator plugging a connection into a manual switchboard, used as a visual metaphor for AI model routing
Previous Post

Stripe’s OpenRouter Deal Turns AI Token Spend Into a Capital-Flows Problem

The Jaquet-Droz Writer automaton, an 18th-century mechanical figure that writes with a quill pen, on display at a museum
Next Post

Pew Research Finds Signs of AI Authorship in a Third of Web Pages Published Since ChatGPT

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026