# How to use circuit breakers, bulkheads and retry budgets when an LLM API or a customer API goes down

> A few seconds of backoff will not clear a 429 caused by a spend cap. If three layers of code all retry, one user request can become 27 calls to a service that is already failing.

Original: https://fdetimes.net/en/guides/circuit-breaker-bulkhead-retry-budget-llm-apis/

Picture 9am at a customer site. The customer-support agent starts getting 429 errors from the LLM API.

The Claude API error documentation says that a 429 caused by hitting a tier's spend cap comes without a `retry-after` header. The error continues until access is restored.

The code retries anyway, and every retry holds one more connection.

By 9:05 the customer's own CRM API is timing out too, even though it has nothing to do with the LLM. The worker pool is tied up with LLM requests that are waiting out their backoff. A failure in one dependency has spread to the whole system.

Forward Deployed Engineers (FDEs) see this scenario often. You integrate an LLM API and the customer's internal systems at the same time, and either one can go down.

This guide builds three layers of protection in about 100 lines of plain Python: a retry budget, a circuit breaker and bulkheads. An error-classification step runs before all three. You need Python 3.10+ and no external libraries.

The code is shortened to make it easier to read. It is not ready for production.

## Why do retries make an outage worse?

Microsoft calls this a retry storm. When a service is busy or not responding, a flood of client retries stops it from recovering and makes the outage worse. The AWS Builders' Library reaches the same conclusion: retrying while a system is overloaded only makes things worse.

LLM APIs add a quota trap. OpenAI's rate-limit documentation notes that failed requests still count towards the per-minute limit, so resending them over and over does not help. Each useless retry uses up quota and delays recovery.

Now multiply. The official Claude SDK retries 2 times by default, so it makes up to 3 calls.

Wrap it in a decorator that also allows up to 3 calls, and put a queue worker outside that runs up to 3 times. One user request can now become 3 × 3 × 3 = 27 calls to a service that is already struggling.

This is why AWS recommends retrying at a single point in the stack.

**Key point:** Retries fix transient errors, but they also add load to a service that is already overloaded.

## Step 1: classify errors before you think about retrying

Some errors are not worth retrying. Microsoft gives the example of a `400 Bad Request`: the server has already said the request is invalid, so sending the same request again achieves nothing. With the Claude API, a 429 caused by a spend cap is a hard error, not a transient one.

```python
def classify(status: int, headers: dict) -> str:
    if status < 400:
        return "ok"
    if status == 400:
        return "client_bug"      # fix the request, do not retry
    if status == 429 and "retry-after" not in headers:
        return "hard_stop"       # e.g. spend cap: open the breaker
    if status == 429 or status >= 500:
        return "transient"       # may retry, within limits
    return "client_bug"
```

This is a simplified version. The rule that "a 429 without retry-after is a hard error" comes from the Claude API's documentation. For another provider's API, or for the customer's API, read their error documentation and adjust the rule.

After this step, write a few unit tests that feed in each status code and check that `client_bug` and `hard_stop` never enter the retry loop.

## Step 2: choose exactly one layer to retry

Microsoft advises putting retry logic only where there is enough context about the failing operation. Lower layers should fail fast, because nested retries cause long delays. Microsoft also notes that retries require the operation to be idempotent.

The Claude SDK lets you set the number of retries through `max_retries`. If you want to control retries in the application layer, turn off the SDK's retries:

```python
import anthropic
client = anthropic.Anthropic(max_retries=0)  # the app layer handles retries
```

The opposite also works: keep the SDK's retries, which use exponential backoff and respect `retry-after`, and add no wrapper outside them. Just do not let both layers retry.

To check, grep the codebase for `retry`, `tenacity`, `max_retries` and the queue configuration. Then write down the maximum number of calls per request. That number should match the call count of a single layer.

## Step 3: a retry budget built on a token bucket

Microsoft's checklist against retry storms includes limiting the number of retries and their total duration, using exponential backoff and respecting `retry-after`. AWS adds another layer: limit retries at the client with a token bucket.

The code below is a short implementation of that idea, not AWS's own formula. Each retry costs one token, and each successful request adds back part of a token. If a dependency stays down for a long time, the bucket empties and the client stops retrying.

```python
import random, time

class RetryBudget:
    def __init__(self, capacity=10, refill=0.1):
        self.tokens, self.capacity, self.refill = capacity, capacity, refill
    def spend(self) -> bool:
        if self.tokens >= 1:
            self.tokens -= 1
            return True
        return False
    def on_success(self):
        self.tokens = min(self.capacity, self.tokens + self.refill)

def backoff(attempt, base=0.5, cap=20.0):
    return random.uniform(0, min(cap, base * 2 ** attempt))  # jitter
```

The values `capacity=10` and `refill=0.1` are only examples. Tune them to your real traffic. Jitter keeps clients from all retrying at the same moment. If the response includes `retry-after`, wait for that value instead of the `backoff` value.

## Step 4: a three-state circuit breaker

Microsoft defines a circuit breaker as a way to block access to a remote service for a while once failures reach a threshold, instead of retrying an operation that will probably fail. In the Closed state, the breaker counts failures.

If failures within a time window go over the threshold, the breaker switches to Open and starts a timer. When the timer runs out, it moves to Half-Open. In the code below, it then lets exactly one trial request through, so a recovering service does not get a sudden rush of requests.

```python
class CircuitBreaker:
    def __init__(self, threshold=5, window=30, open_secs=20):
        self.threshold, self.window, self.open_secs = threshold, window, open_secs
        self.fails, self.state, self.opened_at = [], "closed", 0.0

    def allow(self) -> bool:
        now = time.monotonic()
        if self.state == "open" and now - self.opened_at >= self.open_secs:
            self.state = "half_open"
            return True              # exactly one trial request
        return self.state == "closed"

    def record(self, kind: str):
        now = time.monotonic()
        if kind == "ok":
            self.state, self.fails = "closed", []
            return
        if kind == "hard_stop" or self.state == "half_open":
            self.trip(now)
            return
        self.fails = [t for t in self.fails if now - t < self.window] + [now]
        if len(self.fails) >= self.threshold:
            self.trip(now)

    def trip(self, now):
        self.state, self.opened_at = "open", now
```

Note that `hard_stop` opens the breaker at once, without waiting for the threshold, because retrying a spend-cap error gains nothing. Microsoft also makes clear that retries and breakers do different jobs. Once the breaker signals that a failure is no longer transient, the retry logic must stop. In the call loop, check `breaker.allow()` before every attempt, including retries.

## Step 5: bulkheads, one compartment per dependency

The bulkhead takes its name from the walls that divide a ship's hull. The application is split into pools so that if one part fails, the rest keeps running. Microsoft gives an example that fits FDE work closely: a client that calls several services gets a separate connection pool for each one, so a failing service affects only its own pool.

For AI and inference workloads, Microsoft recommends strict bulkheads, because quotas and concurrency limits are set per deployment. It also recommends isolating by workload or by tenant.

```python
import asyncio

POOLS = {"llm": asyncio.Semaphore(8), "crm_khach": asyncio.Semaphore(4)}

async def with_bulkhead(name, coro_fn):
    sem = POOLS[name]
    if sem.locked():
        raise RuntimeError(f"{name} full: return 503 upstream")
    async with sem:
        return await coro_fn()
```

When a pool is full, the code fails fast and reports the error upstream instead of joining a queue. This follows the advice in Microsoft's throttling pattern: when an external dependency fails, cut the number of requests you are sending, and pass the 429/503 signal up to the next layer instead of swallowing it.

Back to the 9am scenario. With two separate semaphores, a failing LLM API can take at most 8 slots, and the CRM's 4 slots stay free.

## Step 6: test with a fake dependency

Write a fake function that always returns 503. Call it 50 times through the full chain of bulkhead → breaker → classify → retry budget, and count how many times the fake function is actually called. Do the arithmetic first so you have a baseline: with no protection and the three retry layers described above, the maximum is 50 × 27 = 1,350 calls.

With all the layers in place and the parameters used here, the count should stop at about 5, the breaker's threshold, plus one trial call each time the breaker moves to Half-Open.

The retry budget is a second safety net. Even if the breaker is set up wrongly, the bucket allows at most 10 retries. If you count dozens of calls or more, some retry layer is getting past the breaker.

Next, make the fake function return a 429 without `retry-after`, and check that the breaker opens on the very first call.

Three mistakes are common. The first is counting 400 errors towards the breaker, so a single client bug opens the circuit for everyone. The second is sharing one breaker across all tenants, so one tenant that goes over quota takes the others down with it. The third is forgetting that failed requests still count towards the rate limit.

## What will the customer ask on the day the API goes down?

On real projects, customers rarely ask "do you have a circuit breaker?" They ask why their whole support portal died when the LLM API became unstable. A good answer is a diagram: each dependency has its own pool, retries happen in exactly one layer, and the breaker opens on hard errors.

When you read FDE job descriptions, look for phrases such as "production reliability" or "integrate with customer systems". On your CV, do not just write "implemented a circuit breaker". Give the before and after numbers, such as "cut the maximum calls per request from 27 to 3", because interviewers can ask you to explain them in detail.

The next time a customer's API goes down, the goal is not for your system to try harder. The goal is for it to know when to stop.

**Try this week:**

- Grep the codebase for every place that retries (SDK, HTTP client, queue worker, decorators), then work out the maximum number of calls one user request can generate
- Add the classify function to your LLM call layer so that 400 errors and 429 errors without retry-after are never retried
- Run the fake-dependency test from this guide and count the real calls before and after adding the breaker

## Sources

- [Circuit Breaker Pattern - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/patterns/circuit-breaker)

- [Bulkhead Pattern - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/patterns/bulkhead)

- [Retry Storm Antipattern - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/antipatterns/retry-storm/)

- [Retry pattern - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/patterns/retry)

- [Throttling Pattern - Azure Architecture Center | Microsoft Learn](https://learn.microsoft.com/en-us/azure/architecture/patterns/throttling)

- [Timeouts, retries and backoff with jitter](https://builder.aws.com/content/3EumjoZascWd1oZiEgL8ORlv3qE/timeouts-retries-and-backoff-with-jitter)

- [Claude API errors](https://platform.claude.com/docs/en/api/errors)

- [Rate limits | OpenAI API](https://developers.openai.com/api/docs/guides/rate-limits)
