# Making 20,000 LLM requests in Python without 429 failures or hung scripts

> Batch scripts rarely fail because they run slowly. They fail because nothing holds them back: there is no limit on concurrent requests, no pacing, no timeout and no end to the retries.

Original: https://fdetimes.net/en/guides/python-asyncio-llm-batch-rate-limits/

A client sends you 20,000 support tickets and wants them classified by an LLM before Monday morning. A sequential `for` loop would take an estimated full day, so you switch to `asyncio.gather` and fire everything at once. The screen fills with 429 errors, then `PoolTimeout`s, and at 97% the script freezes.

Nearly every FDE has been here, whether the other side is OpenAI's API, Anthropic's, or a client's ageing internal system. Slowness is not the problem: a slow job still finishes eventually.

What breaks things is a script with **no brakes**. This guide covers the four brakes to fit, a complete pipeline, and how to read a 429 correctly so you know when to wait and when to stop.

## Why is firing everything at once slower?

Every API has limits. OpenAI's documentation is explicit: requests that exceed a limit temporarily receive a `429` error. When 20,000 coroutines queue up together, most of them just collect 429s, retry at the same moment and get blocked again together.

The client side jams too. httpx keeps a connection pool of limited size. A request that waits longer than the `pool` timeout for a connection raises `PoolTimeout`, which means you are creating more concurrent requests than the pool can supply connections for.

Hanging at 97% is usually down to a single stuck request with nothing to cut it off, or a retry loop with no exit. Errors like this do not surface. The script simply sits there and you cannot tell where it is stuck.

## A Semaphore is not enough: you also need pacing

The first brake is `asyncio.Semaphore`. The Python documentation describes it as a counter that never drops below zero. When the counter reaches zero, a task calling `acquire()` waits until another task calls `release()`. With `Semaphore(20)`, no more than 20 requests are ever running.

The next part is often missed. Anthropic explains that its API uses a token bucket algorithm: capacity refills continuously rather than resetting at the top of each minute. A 60 RPM limit may therefore be enforced as 1 request per second, and a short burst of requests can exceed the limit.

Picture 20 requests running in parallel, each taking only 300 ms. Within a single second you have sent about 60 requests, the whole minute's quota, without ever exceeding 20 concurrent requests.

**Key point:** Capping concurrent requests is not enough; you must also keep the send rate steady.

The second brake, then, is a pacer, which enforces a minimum gap between sends. Anthropic also advises ramping traffic up gradually and keeping usage steady to avoid acceleration limits. In practice, start with a wide gap and narrow it over the first few minutes.

The two brakes do different jobs, and it pays to do the arithmetic before you start. The pacer caps throughput: at 1 request per second, 20,000 tickets take 20,000 seconds, about 5.5 hours, however many requests you allow in parallel.

The Semaphore limits how many requests are in flight. Roughly, that number is the send rate multiplied by the duration of one request: if each LLM call takes 10 seconds, about 10 requests are in flight, so `Semaphore(20)` is mainly a safety net for when the API slows down unusually.

If 5.5 hours still meets the Monday deadline, let the job run overnight. If it does not, raising the Semaphore will not make the job faster. The thing to do is ask the client for a higher quota.

## A pipeline for classifying 20,000 tickets

Below is a skeleton for calling a classification API with httpx. The two remaining brakes are timeouts and capped retries. Read the numbers carefully: each sits at a different layer.

```python
import asyncio, random, time
import httpx

MAX_CONCURRENT = 20
MIN_INTERVAL = 1.0      # 60 RPM -> 1 request/second -> 20,000 tickets ~ 5.5 hours
RETRY_BUDGET_S = 120    # total time allowed for retries
PER_ITEM_S = 150        # hard ceiling per ticket

# Provider-specific flag:
# Anthropic: a spend-limit 429 has no retry-after, retrying is futile -> True
# OpenAI: if the header is missing, use exponential backoff + jitter -> False
STOP_ON_BARE_429 = True

class SpendLimitError(Exception):
    """Anthropic-style spend-limit 429 without retry-after: retrying is futile."""

class Pacer:
    def __init__(self, interval: float):
        self.interval = interval
        self._lock = asyncio.Lock()
        self._next = 0.0

    async def wait(self):
        async with self._lock:
            now = time.monotonic()
            if self._next > now:
                await asyncio.sleep(self._next - now)
            self._next = max(now, self._next) + self.interval

def retry_after_seconds(resp):
    try:
        return float(resp.headers["retry-after"])
    except (KeyError, ValueError):
        return None

async def call_with_retry(client, pacer, payload, max_attempts=5):
    start = time.monotonic()
    for attempt in range(max_attempts):
        await pacer.wait()
        resp = await client.post("/classify", json=payload)
        if resp.status_code not in (429, 503):
            resp.raise_for_status()
            return resp.json()
        if (resp.status_code == 429 and STOP_ON_BARE_429
                and "retry-after" not in resp.headers):
            raise SpendLimitError("429 without retry-after")  # stop, do not retry
        wait = retry_after_seconds(resp)
        if wait is None:                      # header missing/invalid (flag off) or 503
            wait = min(2 ** attempt, 30)      # exponential backoff
        wait += random.uniform(0, 1)          # jitter
        if time.monotonic() - start + wait > RETRY_BUDGET_S:
            break
        await asyncio.sleep(wait)
    raise RuntimeError(f"retries exhausted, last status {resp.status_code}")

async def worker(client, sem, pacer, ticket):
    async with sem:
        try:
            async with asyncio.timeout(PER_ITEM_S):
                result = await call_with_retry(client, pacer, ticket)
            return {"id": ticket["id"], "ok": True, "result": result}
        except SpendLimitError:
            raise                             # let TaskGroup cancel the whole job
        except Exception as e:
            return {"id": ticket["id"], "ok": False, "error": repr(e)}

async def main(tickets):
    sem = asyncio.Semaphore(MAX_CONCURRENT)
    pacer = Pacer(MIN_INTERVAL)
    limits = httpx.Limits(max_connections=MAX_CONCURRENT)
    timeout = httpx.Timeout(60.0, pool=30.0)
    try:
        async with httpx.AsyncClient(base_url="https://api.khach.example",
                                     limits=limits, timeout=timeout) as client:
            async with asyncio.TaskGroup() as tg:
                tasks = [tg.create_task(worker(client, sem, pacer, t))
                         for t in tickets]
    except* SpendLimitError:
        print("Spend limit reached: stopping job, notify the client's account owner")
        raise
    return [t.result() for t in tasks]
```

The retry logic follows OpenAI's guidance. The `Retry-After` value is only a **minimum**: wait at least that long, then add a small random amount.

If the header is missing or invalid, fall back to exponential backoff with jitter so that clients do not all retry at once. Retries need a ceiling on both the number of attempts and the total time, here 5 attempts and 120 seconds.

The branch that stops outright on a 429 without the header is controlled by the `STOP_ON_BARE_429` flag, which must be set per provider. For Anthropic, the documentation says a spend-limit 429 carries no `retry-after` and retrying is futile, so turning the flag on makes sense.

For OpenAI, the guidance is to back off when the header is missing, so turn the flag off. If you are calling a client's API, ask how they signal "do not retry" and change that one condition accordingly.

Timeouts sit at three layers. By default httpx raises `TimeoutException` after 5 seconds without network activity. That is too short for an LLM call generating a long response, so it is raised to 60 seconds here. The 120-second retry budget sits in the middle, and `asyncio.timeout(150)` wraps everything, turning any stuck case into a `TimeoutError` so that no ticket hangs forever.

## Which errors to contain and which to let escape

The easily missed detail is that the `try/except` sits **inside** the worker. With `TaskGroup`, the first time a task fails with an exception other than `CancelledError`, the remaining tasks in the group are cancelled. One broken ticket must not cancel the other 19,999, so the worker catches ordinary errors itself and returns an `ok: False` record.

That same cancellation is exactly what you want for `SpendLimitError`. The worker re-raises it, TaskGroup cancels all remaining tasks, and `main` catches it via `except*`. The job stops at once instead of letting thousands of tickets hit the same wall one after another.

Why not `gather`? By default, `gather` propagates the first exception, but the other awaitables are **not cancelled** and keep running.

Your `main` function has returned while requests are still in flight, and that is the source of some very hard-to-trace bugs. With TaskGroup, when an error escapes the remaining tasks are cancelled, so no requests are left running out of your control.

## Not every 429 should be retried

If you call through an official SDK, do not rewrite the retry loop above. OpenAI's documentation states that each official SDK already retries eligible `429` and `503` responses according to its retry configuration.

Do the arithmetic: if the SDK retries a few times per call and your loop tries 5 times, the actual number of requests sent is the product of the two numbers, not the sum. Keep the Semaphore, the pacer and the timeouts, and tune retries in the SDK's configuration.

There is another kind of 429. According to Anthropic's documentation, when an organisation hits its spend limit, the 429 response has **no** `retry-after` header, and retries, including the SDK's automatic ones, fail until access is restored. That is why, when running against Anthropic, the code above turns the flag on and raises `SpendLimitError` rather than backing off further.

If you use Anthropic's SDK, check for this header in your logs or in the exception the SDK returns. When you hit exactly this case, stop the job and notify the client's account owner.

## Common mistakes on customer sites

The first is a Semaphore larger than httpx's `max_connections`. The surplus requests sit waiting in the pool and get `PoolTimeout`, which looks like a network error but is really a configuration error. Keep the two numbers in step.

The second is retries without a time budget. Five attempts, each waiting a 60-second `Retry-After`, add up to five minutes for one ticket. The third is logging only the total error count. When the results carry an `id` and an `error` for each ticket, you rerun just the 312 failed tickets rather than all 20,000.

The last is running at full capacity from the first minute on a client's system. Their internal API may not return a polite 429 like OpenAI's; it may simply fall over. Ask about their real limits first, then ramp up gradually.

## How this skill shows up on a CV

When reading FDE job descriptions, look for lines about integrating with client APIs or processing large volumes of data: that is where this skill gets used.

On a CV, rather than writing "proficient in asyncio", describe the outcome: processed N records through a rate-limited API with pacing and capped retries, reported errors per record, and no job ever hung.

In an interview, explaining why `TaskGroup` differs from `gather`, or why with Anthropic a 429 without `retry-after` means stop rather than retry, is far more convincing than listing library names.

A good batch script does not need to be the fastest. It needs to finish, and to tell you exactly which tickets did not complete and why, before the client asks.

**Try this week:**

- Take a batch script that uses unbounded asyncio.gather, add a Semaphore, a pacer and asyncio.timeout, and test it on 500 records.
- Review code that uses OpenAI's or Anthropic's official SDK, remove any hand-written retry loops layered on top of the SDK, and set max_retries in one place only.
- Have the job write a results file with an ok/error column for each record, then count errors by type to decide whether to rerun.

## Sources

- [Rate limits (OpenAI API docs)](https://developers.openai.com/api/docs/guides/rate-limits)

- [Synchronization Primitives — Python 3.15.0 documentation](https://docs.python.org/3/library/asyncio-sync.html)

- [Coroutines and tasks — Python 3.15.0 documentation](https://docs.python.org/3/library/asyncio-task.html)

- [Timeouts - HTTPX](https://www.python-httpx.org/advanced/timeouts/)

- [Rate limits (Claude API docs)](https://platform.claude.com/docs/en/api/rate-limits)
