Writing the API call layer for a TypeScript agent: timeouts, retries and idempotency keys
If your agent retries a stalled API call, the client's system can issue the same refund twice. About 80 lines of TypeScript prevent it.
In brief
- An API with side effects can only be retried safely if it supports idempotency.
- Generate the idempotency key once per business operation, reuse it on every retry, and attach it only to POST requests.
- Retries need a budget: if every layer retries on its own, load on the database can grow as much as 243-fold.
- 1Generate the idempotency keyPOST only: one UUID V4 per business operation, created outside the loop
- 2Call with a timeoutAbortSignal.timeout() set from the downstream API's p99.9 latency
- 3Classify the resultRetry 408, 409, 429 and network errors; retry 5xx only when the request carries no key
- 4Check the budgetIf the token bucket is empty or attempts are used up, stop and report the error
- 5Wait with jittered backoffWait a random interval, then resend with the same key
The key is generated once, before the loop, so every retry refers to the same operation.
Graphic: FDE Times
Picture your second week on site with a client. The customer support agent your team deployed has a tool called createRefund, which calls the client’s internal refunds API. One afternoon the network is flaky. The first request gets no response, so the agent tries again. The next morning the finance team finds an order that was refunded twice.
The model didn’t cause this. The fault is in the dozen or so lines that wrap fetch, code everyone treats as a chore. For a forward deployed engineer, the API call layer is where the agent touches the client’s money, inventory and production data. You need to get it right the first time.
The layer needs three things: a timeout, disciplined retries and an idempotency key. This guide covers each in turn, then puts them together in one complete TypeScript function you can copy and adapt for your own project.
Why are naive retries dangerous?
The Amazon Builders’ Library is blunt about it: an API with side effects is not safe to retry unless it supports idempotency. When a refund request times out, the client cannot tell whether the server processed it. The packet may never have arrived. Or the money may already have moved and only the response got lost on the way back.
An idempotency key removes exactly that uncertainty. Stripe’s documentation says that if a connection error occurs, you can resend the request without risk of creating a second object or applying the same update twice.
The mechanism is simple. For each key, Stripe stores the status code and body of the first request, whether it succeeded or failed, and returns that same result for every resend.
The words “or failed” matter. As Stripe describes it, if the first attempt returned a 500, a retry with the same key gets that same 500 back. A retry with the old key is therefore mostly useful when the failure happened in transit, such as a timeout or a dropped connection, before the client received any result.
How long should the timeout be?
Defaults are usually far too generous for an agent. OpenAI’s Node SDK sets a default timeout of 10 minutes. That suits a long model call, but an order-lookup API should not be allowed to keep the agent waiting that long.
Amazon offers a principled way to choose. First decide what share of “false” timeouts you can accept, meaning requests that would have succeeded if you had waited a little longer. Then take the matching latency percentile of the downstream service, for example p99.9.
If the refunds API has a p99.9 of 1.2 seconds, a timeout of about 1.2 seconds means you accept cutting off roughly 0.1% of requests that would have succeeded.
In Node you don’t need to build this yourself with setTimeout. AbortSignal.timeout() returns a signal that aborts itself after the given time, and when it fires it aborts with a DOMException named TimeoutError. Your code can then tell a timeout apart from a deliberate cancellation by the user. The two cases need different handling.
How many retries are enough?
OpenAI’s SDK is a good reference design. According to its README, it automatically retries twice with a short exponential backoff. By default it retries 408, 409, 429, every error of 500 and above, and connection errors. It does not retry a 400 or a 401, because resending the same request unchanged would not change the outcome.
Backoff on its own is not enough. You also need jitter. Amazon describes jitter as adding some randomness to the wait time so that retries are spread out over time. Without it, a thousand agent sessions that fail together will all retry at the same moment and knock the service over again.
The biggest danger is retries stacked across layers. Amazon gives the example of a five-layer system where each layer retries independently, multiplying the load on the database by 243.
Working backwards, 243 is 3 to the power of 5, which is what you get when each layer makes up to three attempts. Amazon’s answer is to limit retries locally with a token bucket. When the tokens run out, the layer stops retrying and passes the error up.
The complete TypeScript version
Below is a callApi function that combines all three. Note the line that generates the key: it sits outside the loop.
import { randomUUID } from "node:crypto";
type Method = "GET" | "POST" | "DELETE";
interface CallOptions {
method: Method;
url: string;
body?: unknown;
timeoutMs: number; // taken from the downstream API's p99.9
maxRetries?: number; // defaults to 2
idempotencyKey?: string; // pass one in if the operation already has a key
signal?: AbortSignal; // cancellation signal from the agent or the user
}
// Token bucket: at most 10 retries, refilling 1 token per second
class RetryBudget {
private tokens = 10;
constructor() {
setInterval(() => (this.tokens = Math.min(10, this.tokens + 1)), 1000).unref();
}
take(): boolean {
if (this.tokens <= 0) return false;
this.tokens--;
return true;
}
}
const budget = new RetryBudget();
// Requests with a key: don't retry 5xx, the server has stored it and will return the same error
const isRetryableStatus = (s: number, hasKey: boolean) =>
s === 408 || s === 409 || s === 429 || (!hasKey && s >= 500);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
const backoff = (attempt: number) =>
Math.random() * Math.min(5000, 200 * 2 ** attempt); // jitter
export async function callApi(opts: CallOptions): Promise<Response> {
const maxRetries = opts.maxRetries ?? 2;
// Generated ONCE for the whole operation; only POST carries a key
const key =
opts.method === "POST" ? opts.idempotencyKey ?? randomUUID() : undefined;
for (let attempt = 0; ; attempt++) {
// Each call gets its own timeout, combined with the caller's cancellation signal
const timeout = AbortSignal.timeout(opts.timeoutMs);
const signal = opts.signal ? AbortSignal.any([opts.signal, timeout]) : timeout;
try {
const res = await fetch(opts.url, {
method: opts.method,
headers: {
"Content-Type": "application/json",
...(key ? { "Idempotency-Key": key } : {}), // header name depends on the API
},
body: opts.body ? JSON.stringify(opts.body) : undefined,
signal,
});
const canRetry =
isRetryableStatus(res.status, key !== undefined) &&
attempt < maxRetries &&
budget.take();
if (!canRetry) return res;
await res.body?.cancel(); // discard the body to free the connection before retrying
} catch (err) {
// If the caller cancelled, stop immediately, no retry
if (opts.signal?.aborted || attempt >= maxRetries || !budget.take()) throw err;
// Otherwise it's a TimeoutError or network error: resend with the SAME key
}
await sleep(backoff(attempt));
if (opts.signal?.aborted) throw opts.signal.reason;
}
}
Back to the refund. The first call times out, the catch block receives a TimeoutError, and the function waits a random interval before resending with the same key. If the server had already processed the first request, it returns the stored result and no second refund is issued.
Two small details are worth keeping when you adapt the function. The caller’s cancellation signal is combined with the timeout through AbortSignal.any(), so if the agent is stopped partway through, the function exits at once instead of carrying on with retries. And when a response is discarded in order to retry, its body is cancelled so the connection doesn’t stay stuck in the pool.
The retry condition here is narrower than the OpenAI SDK’s in one respect: a request that carries a key does not retry errors of 500 and above. The reason is the mechanism described earlier. With an API that stores the first result per key, as Stripe does, resending the same key after a 500 only returns the same 500. It uses up budget and gains nothing.
GET and DELETE carry no key, so they still retry 5xx as normal. If the client’s API handles keys differently, ask them before relaxing this condition.
A 409 is still worth retrying when it signals a conflict with another request running concurrently. Stripe states that it does not store the result when parameters fail validation or when the request conflicts with one already in progress, so resending is safe in those cases.
Common mistakes in the field
The most common mistake is generating the key inside the loop. Every retry then carries a new key, the server treats it as a new operation, and idempotency is gone.
Agents add a subtler version of this. The model calls the same tool again at a later step, and your layer generates a fresh key for that call. If your orchestrator can re-run a step, derive the key from that step’s ID and pass it in through idempotencyKey.
The second mistake is reusing a key with different parameters. Stripe compares the new parameters with the original request and returns an error if they differ. It may also delete keys after 24 hours, so a job re-run the next day cannot rely on the old key.
The third mistake is attaching a key to every request. Stripe says plainly that GET and DELETE are already idempotent, so sending a key with them does nothing. Don’t put an email address or account number in the key either. Stripe recommends a UUID V4 or a random string with enough entropy, up to 255 characters, with no sensitive data in it.
The last mistake is the hardest to spot: retries stacked on top of an SDK. If your tool wraps a client that already retries twice, and the agent’s orchestrator retries again on top, the amplification factor has multiplied before you notice. Choose exactly one layer to do the retrying and turn retries off everywhere else.
Putting this skill on your CV
When you read a job description for an FDE role, look for requirements about integrating with client systems or reliability in production. That is where this kind of code belongs in your story.
On your CV, don’t just write “built an agent that calls APIs”. Say which percentile you based your timeouts on, which layer handles retries, and how you used idempotency to prevent duplicate operations.
In an interview, “What does your agent do if a create-order request times out?” is your chance to sketch the loop above. Interviewers remember candidates who know that a timeout does not mean failure. It means the outcome is unknown.
Was this article useful?
Thanks for the feedback!