# APIs and webhooks for FDEs: build every integration as if everything arrives twice

> The costliest bug in an enterprise integration rarely takes the system down. It quietly charges a customer twice, or drops an event that nobody notices is gone.

Original: https://fdetimes.net/en/guides/idempotent-api-webhook-integrations/

At nine on a Monday morning, the customer's accountant sends you a screenshot: the same order has been charged twice. The logs show that the network was unstable for a while over the weekend, the payment-creation request timed out, and your job dutifully retried, exactly as it had been told to.

The code has no syntax errors and it did not crash. What was wrong was the assumption that a request arrives exactly once, in the right order, and always gets a clear answer. For a Forward Deployed Engineer, that assumption breaks almost every day.

Palantir describes the FDE role as embedding engineers alongside customers to solve their most urgent problems. Working inside a customer's systems also means that sooner or later you will have to touch the pipes connecting those systems to others. Build those pipes solidly and you earn trust.

Build them carelessly and you will spend a week explaining why the numbers do not match.

## When you get an error, you do not know what happened

Start with an uncomfortable fact. Stripe's error-handling documentation says it plainly: on a network error, the client does not know whether the server received the request. A 500 response to a write operation must also be treated as an indeterminate result: the operation may have run, or it may not.

So blind retries are dangerous, while not retrying risks losing orders. The way out is an idempotency key: a string generated by the client, up to 255 characters long, for which Stripe suggests a UUID V4. You attach it to the request so the server can recognise resends.

Stripe's mechanism is simple. It stores the status code and body of the first request made with each key, even if that request failed, and returns exactly that result for every retry. If you reuse a key with different parameters, the idempotency layer returns an error to catch the mistake.

Three details often escape newcomers. Keys may be removed after 24 hours, so do not treat them as a permanent record. Every POST accepts a key, while sending one with GET or DELETE has no effect, because those methods are already idempotent.

And Stripe's advice after a network error is to retry with the same key and the same parameters until you get a result from the server.

## A payment job, rewritten correctly

Back to Monday's incident. The most important change is that the key must be generated **once**, when the order is created, and stored with the order, not regenerated on each retry. If it is regenerated, every retry looks like a new transaction and idempotency does nothing.

```python
def create_payment(order):
    key = order.idempotency_key          # UUID v4, generated at order creation, stored in DB
    payload = {
        "amount": order.amount,
        "currency": order.currency,
        "metadata[order_id]": order.id,  # for reconciliation later
    }
    return send_with_retry(payload, key, order)
```

The retry loop is separate, and every attempt uses that same key:

```python
def send_with_retry(payload, key, order):
    for attempt in range(6):
        try:
            r = requests.post(PAYMENTS_URL, data=payload,
                              headers={"Idempotency-Key": key}, timeout=10)
        except (requests.ConnectionError, requests.Timeout):
            time.sleep(2 ** attempt)     # retry with the same key
            continue
        if r.status_code == 429:
            time.sleep(retry_after(r, attempt))
            continue
        if r.status_code >= 500:
            mark_indeterminate(order)    # wait for webhook/reconciliation
            return None
        return r
    mark_indeterminate(order)
    return None
```

The 429 branch deserves attention. Rate limits are routine in integration work, and according to MDN a 429 response may include a Retry-After header saying how long to wait before sending a new request. If the server gives you a number, wait exactly that long rather than guessing:

```python
def retry_after(r, attempt):
    wait = r.headers.get("Retry-After")
    return int(wait) if wait and wait.isdigit() else 2 ** attempt
```

The 500 branch does not retry blindly. The order is marked "indeterminate", and the truth arrives from the other side: the webhook, plus the `order_id` metadata to match the transaction to the order. That is also what Stripe recommends: reconcile through webhooks and metadata.

## Webhooks: late, twice, out of order

The other direction is harder, because you do not control when someone else calls you. Stripe states that an endpoint may occasionally receive the same event more than once, and that events may arrive out of order. Crossmint's documentation says its webhooks guarantee "at least once" delivery, which means the receiver has to be idempotent.

A good handler does four things, in this order.

```python
@app.post("/webhooks/payments")
def receive():
    raw = request.get_data()                       # raw body, not parsed
    sig = request.headers.get("Stripe-Signature")
    try:
        event = stripe.Webhook.construct_event(raw, sig, WEBHOOK_SECRET)
    except Exception:
        return "", 400                             # bad or stale signature
    if not processed_events.insert_if_absent(event["id"]):
        return "", 200                             # already seen: skip
    queue.enqueue(handle_event, event["id"])
    return "", 200                                 # return 2xx immediately
```

The first step is to read the raw body. Stripe signs every webhook with HMAC-SHA256 via the `Stripe-Signature` header, and verification needs the original body exactly. If your framework has parsed and re-serialised it, a single changed whitespace character is enough to break the signature, and you will lose half a day suspecting the secret.

The signature also includes a timestamp to prevent replay attacks, and Stripe's library by default tolerates a gap of up to 5 minutes between that timestamp and the current time. In practice, if the clock on the customer's server drifts significantly, every webhook will be rejected even though the secret is perfectly correct.

Next comes deduplication by event ID, then returning 2xx before any heavy work. Stripe recommends processing events through an asynchronous queue to absorb sudden spikes. In the worker, do not assume a "paid" event always arrives after a "created" event. It is safer to read the object's current state before updating the order.

**Key point:** Treat every request as something that may run twice, and every event as something that may arrive out of order; only code that is correct in both cases belongs in production.

## Every provider promises something different, so read the promise carefully

This is where FDEs earn trust: reading each provider's documentation closely enough to know who is responsible for resending when your system goes down.

| | Stripe | GitHub |
|---|---|---|
| Response deadline | Return 2xx before running heavy logic | Must return 2XX within 10 seconds |
| On failed delivery | Live mode resends automatically for up to three days, with exponential backoff | No automatic redelivery |
| Your job after an outage | Cope with a backlog of old events arriving at once | Redeliver missed webhooks yourself once the server is back |

This difference shapes the architecture. With Stripe, a two-hour outage usually heals itself, as long as the handler can absorb the backlog. With GitHub, without a redelivery script or procedure, missed data is gone for good, and the runbook handed over to the customer must spell out that step.

## The mistakes that keep recurring

The most common mistake is generating a new idempotency key inside the retry loop, which makes the protection useless. Close behind is reusing an old key with different parameters, such as a changed amount, and then being surprised by the error. Changing parameters makes it a new operation, which needs a new key.

The second mistake is treating a 500 as a "definite failure" and recreating the transaction with a different key. The third is doing all the business logic inside the webhook handler: writing to the database, calling the ERP, sending emails. One slow ERP call pushes you past the timeout, the provider treats the delivery as failed, and the resend cycle begins.

The last mistake is rarely discussed: having no reconciliation path. However good the code, you still need a scheduled job that compares state on both sides via metadata, because one day an event will be missed.

## Putting this skill on your CV

When reading job descriptions for FDE or solutions engineer roles, look for phrases such as "integrations", "webhooks", "customer systems" and "data pipelines". When you see them, have a story about idempotency and webhooks ready. Show that you have dealt with real failures, not just called APIs.

So instead of writing "Integrated Stripe", be specific: added idempotency keys and webhook-based reconciliation to eliminate duplicate transactions; moved the webhook handler to queue-based processing with deduplication by event ID. If asked in an interview, walk through the incident, its cause and how you proved it would not happen again.

This week's exercise: take an integration you are running, pull the network cable in the middle of a POST, then send the same webhook three times. If the data is still correct after both tests, you have done the part of the job customers will never see, and that is exactly the part that decides whether they trust you.

**Try this week:**

- Open Stripe test mode, send the same POST twice with the same Idempotency-Key, then a third time with the same key but a different amount, and record all three responses.
- Build a small webhook endpoint with signature verification, a processed_events table and a queue, then resend the same event to it three times to confirm it is processed only once.
- Reopen an integration you have written and check: if a POST times out, what does your code do?

## Sources

- [Idempotent requests (Stripe API Reference)](https://docs.stripe.com/api/idempotent_requests)

- [Advanced error handling (Stripe Docs)](https://docs.stripe.com/error-low-level)

- [Receive Stripe events in your webhook endpoint (Stripe Docs)](https://docs.stripe.com/webhooks)

- [Best practices for using webhooks (GitHub Docs)](https://docs.github.com/en/webhooks/using-webhooks/best-practices-for-using-webhooks)

- [Best Practices (Crossmint Docs, Webhooks)](https://docs.crossmint.com/introduction/platform/webhooks/best-practices)

- [Retry-After header - HTTP | MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Retry-After)

- [Palantir Technologies - Forward Deployed Software Engineer](https://jobs.lever.co/palantir/5168e8fd-fec1-4fea-b7a1-81bdaea65850)
