FDE PulseFDE jobs open 432New in 7 days 28Companies hiring 42Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Making 20,000 LLM requests in Python without 429 failures or hung scripts

Batch scripts rarely fail because they run slowly. They fail because nothing holds them back: there is no limit on concurrent requests, no pacing, no timeout and no end to the retries.

Making 20,000 LLM requests in Python without 429 failures or hung scripts
Photo: Emile Perron / Unsplash

In brief

  • A Semaphore only limits how many requests run at once. An API can still enforce its limit per second, so you also need a pacer, and the pacer is what sets how long the job takes.
  • Each layer of waiting needs its own ceiling: the httpx timeout, a time budget for retries and an outer asyncio.timeout.
  • Catch per-record errors inside the worker. Let errors such as hitting a spend limit escape so that TaskGroup cancels the whole job.
ShareLinkedInFacebookX
GraphicThe path of one request in a batch job
  1. 1SemaphoreAt most 20 requests in flight, matching httpx max_connections
  2. 2Pacer sets the rhythm60 RPM becomes 1 request/second, which also caps throughput for the whole job
  3. 3API call with timeouthttpx timeout per call, outer asyncio.timeout around the whole ticket
  4. 4On 429: wait or stopRetry-After: wait + jitter; no header: backoff; Anthropic spend cap: stop the job
  5. 5Record each ticket's resultWorker catches ordinary errors, returns ok/error so TaskGroup won't cancel others

Every request passes through all four brakes, so the job stays under the rate limit and never hangs forever.

Graphic: FDE Times

A client sends you 20,000 support tickets and wants them classified by an LLM before Monday morning. A sequential for loop would take an estimated full day, so you switch to asyncio.gather and fire everything at once. The screen fills with 429 errors, then PoolTimeouts, and at 97% the script freezes.

Nearly every FDE has been here, whether the other side is OpenAI’s API, Anthropic’s, or a client’s ageing internal system. Slowness is not the problem: a slow job still finishes eventually.

What breaks things is a script with no brakes. This guide covers the four brakes to fit, a complete pipeline, and how to read a 429 correctly so you know when to wait and when to stop.

Why is firing everything at once slower?

Every API has limits. OpenAI’s documentation is explicit: requests that exceed a limit temporarily receive a 429 error. When 20,000 coroutines queue up together, most of them just collect 429s, retry at the same moment and get blocked again together.

The client side jams too. httpx keeps a connection pool of limited size. A request that waits longer than the pool timeout for a connection raises PoolTimeout, which means you are creating more concurrent requests than the pool can supply connections for.

Hanging at 97% is usually down to a single stuck request with nothing to cut it off, or a retry loop with no exit. Errors like this do not surface. The script simply sits there and you cannot tell where it is stuck.

A Semaphore is not enough: you also need pacing

The first brake is asyncio.Semaphore. The Python documentation describes it as a counter that never drops below zero. When the counter reaches zero, a task calling acquire() waits until another task calls release(). With Semaphore(20), no more than 20 requests are ever running.

The next part is often missed. Anthropic explains that its API uses a token bucket algorithm: capacity refills continuously rather than resetting at the top of each minute. A 60 RPM limit may therefore be enforced as 1 request per second, and a short burst of requests can exceed the limit.

Picture 20 requests running in parallel, each taking only 300 ms. Within a single second you have sent about 60 requests, the whole minute’s quota, without ever exceeding 20 concurrent requests.

The second brake, then, is a pacer, which enforces a minimum gap between sends. Anthropic also advises ramping traffic up gradually and keeping usage steady to avoid acceleration limits. In practice, start with a wide gap and narrow it over the first few minutes.

The two brakes do different jobs, and it pays to do the arithmetic before you start. The pacer caps throughput: at 1 request per second, 20,000 tickets take 20,000 seconds, about 5.5 hours, however many requests you allow in parallel.

The Semaphore limits how many requests are in flight. Roughly, that number is the send rate multiplied by the duration of one request: if each LLM call takes 10 seconds, about 10 requests are in flight, so Semaphore(20) is mainly a safety net for when the API slows down unusually.

If 5.5 hours still meets the Monday deadline, let the job run overnight. If it does not, raising the Semaphore will not make the job faster. The thing to do is ask the client for a higher quota.

A pipeline for classifying 20,000 tickets

Below is a skeleton for calling a classification API with httpx. The two remaining brakes are timeouts and capped retries. Read the numbers carefully: each sits at a different layer.

import asyncio, random, time
import httpx

MAX_CONCURRENT = 20
MIN_INTERVAL = 1.0      # 60 RPM -> 1 request/second -> 20,000 tickets ~ 5.5 hours
RETRY_BUDGET_S = 120    # total time allowed for retries
PER_ITEM_S = 150        # hard ceiling per ticket

# Provider-specific flag:
# Anthropic: a spend-limit 429 has no retry-after, retrying is futile -> True
# OpenAI: if the header is missing, use exponential backoff + jitter -> False
STOP_ON_BARE_429 = True

class SpendLimitError(Exception):
    """Anthropic-style spend-limit 429 without retry-after: retrying is futile."""

class Pacer:
    def __init__(self, interval: float):
        self.interval = interval
        self._lock = asyncio.Lock()
        self._next = 0.0

    async def wait(self):
        async with self._lock:
            now = time.monotonic()
            if self._next > now:
                await asyncio.sleep(self._next - now)
            self._next = max(now, self._next) + self.interval

def retry_after_seconds(resp):
    try:
        return float(resp.headers["retry-after"])
    except (KeyError, ValueError):
        return None

async def call_with_retry(client, pacer, payload, max_attempts=5):
    start = time.monotonic()
    for attempt in range(max_attempts):
        await pacer.wait()
        resp = await client.post("/classify", json=payload)
        if resp.status_code not in (429, 503):
            resp.raise_for_status()
            return resp.json()
        if (resp.status_code == 429 and STOP_ON_BARE_429
                and "retry-after" not in resp.headers):
            raise SpendLimitError("429 without retry-after")  # stop, do not retry
        wait = retry_after_seconds(resp)
        if wait is None:                      # header missing/invalid (flag off) or 503
            wait = min(2 ** attempt, 30)      # exponential backoff
        wait += random.uniform(0, 1)          # jitter
        if time.monotonic() - start + wait > RETRY_BUDGET_S:
            break
        await asyncio.sleep(wait)
    raise RuntimeError(f"retries exhausted, last status {resp.status_code}")

async def worker(client, sem, pacer, ticket):
    async with sem:
        try:
            async with asyncio.timeout(PER_ITEM_S):
                result = await call_with_retry(client, pacer, ticket)
            return {"id": ticket["id"], "ok": True, "result": result}
        except SpendLimitError:
            raise                             # let TaskGroup cancel the whole job
        except Exception as e:
            return {"id": ticket["id"], "ok": False, "error": repr(e)}

async def main(tickets):
    sem = asyncio.Semaphore(MAX_CONCURRENT)
    pacer = Pacer(MIN_INTERVAL)
    limits = httpx.Limits(max_connections=MAX_CONCURRENT)
    timeout = httpx.Timeout(60.0, pool=30.0)
    try:
        async with httpx.AsyncClient(base_url="https://api.khach.example",
                                     limits=limits, timeout=timeout) as client:
            async with asyncio.TaskGroup() as tg:
                tasks = [tg.create_task(worker(client, sem, pacer, t))
                         for t in tickets]
    except* SpendLimitError:
        print("Spend limit reached: stopping job, notify the client's account owner")
        raise
    return [t.result() for t in tasks]

The retry logic follows OpenAI’s guidance. The Retry-After value is only a minimum: wait at least that long, then add a small random amount.

If the header is missing or invalid, fall back to exponential backoff with jitter so that clients do not all retry at once. Retries need a ceiling on both the number of attempts and the total time, here 5 attempts and 120 seconds.

The branch that stops outright on a 429 without the header is controlled by the STOP_ON_BARE_429 flag, which must be set per provider. For Anthropic, the documentation says a spend-limit 429 carries no retry-after and retrying is futile, so turning the flag on makes sense.

For OpenAI, the guidance is to back off when the header is missing, so turn the flag off. If you are calling a client’s API, ask how they signal “do not retry” and change that one condition accordingly.

Timeouts sit at three layers. By default httpx raises TimeoutException after 5 seconds without network activity. That is too short for an LLM call generating a long response, so it is raised to 60 seconds here. The 120-second retry budget sits in the middle, and asyncio.timeout(150) wraps everything, turning any stuck case into a TimeoutError so that no ticket hangs forever.

Which errors to contain and which to let escape

The easily missed detail is that the try/except sits inside the worker. With TaskGroup, the first time a task fails with an exception other than CancelledError, the remaining tasks in the group are cancelled. One broken ticket must not cancel the other 19,999, so the worker catches ordinary errors itself and returns an ok: False record.

That same cancellation is exactly what you want for SpendLimitError. The worker re-raises it, TaskGroup cancels all remaining tasks, and main catches it via except*. The job stops at once instead of letting thousands of tickets hit the same wall one after another.

Why not gather? By default, gather propagates the first exception, but the other awaitables are not cancelled and keep running.

Your main function has returned while requests are still in flight, and that is the source of some very hard-to-trace bugs. With TaskGroup, when an error escapes the remaining tasks are cancelled, so no requests are left running out of your control.

Not every 429 should be retried

If you call through an official SDK, do not rewrite the retry loop above. OpenAI’s documentation states that each official SDK already retries eligible 429 and 503 responses according to its retry configuration.

Do the arithmetic: if the SDK retries a few times per call and your loop tries 5 times, the actual number of requests sent is the product of the two numbers, not the sum. Keep the Semaphore, the pacer and the timeouts, and tune retries in the SDK’s configuration.

There is another kind of 429. According to Anthropic’s documentation, when an organisation hits its spend limit, the 429 response has no retry-after header, and retries, including the SDK’s automatic ones, fail until access is restored. That is why, when running against Anthropic, the code above turns the flag on and raises SpendLimitError rather than backing off further.

If you use Anthropic’s SDK, check for this header in your logs or in the exception the SDK returns. When you hit exactly this case, stop the job and notify the client’s account owner.

Common mistakes on customer sites

The first is a Semaphore larger than httpx’s max_connections. The surplus requests sit waiting in the pool and get PoolTimeout, which looks like a network error but is really a configuration error. Keep the two numbers in step.

The second is retries without a time budget. Five attempts, each waiting a 60-second Retry-After, add up to five minutes for one ticket. The third is logging only the total error count. When the results carry an id and an error for each ticket, you rerun just the 312 failed tickets rather than all 20,000.

The last is running at full capacity from the first minute on a client’s system. Their internal API may not return a polite 429 like OpenAI’s; it may simply fall over. Ask about their real limits first, then ramp up gradually.

How this skill shows up on a CV

When reading FDE job descriptions, look for lines about integrating with client APIs or processing large volumes of data: that is where this skill gets used.

On a CV, rather than writing “proficient in asyncio”, describe the outcome: processed N records through a rate-limited API with pacing and capped retries, reported errors per record, and no job ever hung.

In an interview, explaining why TaskGroup differs from gather, or why with Anthropic a 429 without retry-after means stop rather than retry, is far more convincing than listing library names.

A good batch script does not need to be the fastest. It needs to finish, and to tell you exactly which tickets did not complete and why, before the client asks.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
5 sources
Read next on the roadmap · Stage 1: FoundationsHands-on: turn a Python script into a command the client team can install and run themselves with TyperA script that runs on your laptop is not yet something you can hand over. It becomes a tool when the client team can type one command on their own machine and have it work.