# Design error responses with RFC 9457 so customer ops teams can fix incidents themselves

> If your API returns only a 500 with "Something went wrong", every incident on the customer's side becomes a ticket for you. A few JSON fields in the right places can change that.

Bản gốc: https://fdetimes.net/en/guides/rfc-9457-api-error-responses/

Picture the first week after you hand an integration over to a customer. At midnight, their operations dashboard turns red. The log holds a single line, `500 Internal Server Error`, and the body reads `{"error": "Something went wrong"}`.

The customer's on-call engineer cannot tell whether the fault is in their system or yours, whether to retry, or where to look in the logs. So they open a ticket. By morning you have ten tickets, and every one begins with "Your API is broken".

For an FDE, error responses are part of the handover, not just a technical detail. Design them well and the customer can handle most incidents on their own. Design them badly and you become the permanent on-call engineer for someone else's system.

## A status code tells only half the story

RFC 9457 starts from a practical observation: an HTTP status code on its own is often not enough for the recipient to understand what happened. A 422 says the request had a problem. It does not say what the problem was, which account it affected, or what the recipient should do next.

To fill that gap, the RFC defines a JSON format with its own media type, `application/problem+json`. You may have heard of RFC 7807 from 2016. RFC 9457 formally replaces it, so cite RFC 9457 when writing documentation or talking to the customer's architects.

The strength of the standard is that each field is written for a specific reader. The table below maps the fields to the questions an on-call engineer typically asks:

| Field | Who reads it | Question it answers |
|---|---|---|
| `type` | The customer's machines | What kind of error is this, so it can be handled automatically? |
| `title` | People | In short, what is this kind of error? |
| `detail` | People | What specifically happened this time? |
| `instance` | Operations team | What ID does this occurrence carry, for searching and cross-checking logs? |
| Extension members | Both | Which account is affected, and which guide should be followed? |

## Example: an API that pushes orders to a warehouse

Suppose you are integrating a retail chain's ordering system with your company's warehouse API. One account has used up its daily order quota. A poor implementation returns this:

```json
{ "error": "Bad request", "code": 4021 }
```

The on-call engineer has no idea what `4021` means, so another ticket gets opened. Here is the RFC 9457 version (the example uses Vietnamese strings; the title means "Daily order limit exceeded"):

```json
{
"type": "https://docs.example.vn/problems/vuot-han-muc-don",
"title": "Vượt hạn mức đơn trong ngày",
"detail": "Tài khoản KH-0192 đã dùng hết hạn mức đơn hôm nay. Hạn mức được đặt lại lúc 00:00.",
"instance": "/incidents/7f3a2c",
"account_id": "KH-0192",
"runbook": "https://docs.example.vn/runbook/han-muc-don",
"retryable": false
}
```

Reading this, the on-call engineer knows at once that the fault is on their side and when it will clear: the detail says account KH-0192 has used up today's quota, which resets at 00:00. They have a link to the runbook, and if they still need to call you, they have the ID `7f3a2c` so you can find the exact log line. `account_id`, `runbook` and `retryable` are extension members.

The RFC allows fields like these and requires clients to ignore any they do not recognise, so you can add them gradually without breaking older clients.

The `detail` field here states only the problem and how to fix it. Postman likewise advises that error messages should contain exactly those two things. Database table names and Java class names are of no use to the customer's on-call engineer.

## Letting the customer's machines handle errors too

The on-call engineer is only half the picture. The other half is the customer's code calling your API, and that code needs to know what to branch on.

RFC 7807 was already explicit: clients must use `type` as the primary identifier for the kind of error and should not parse the `detail` string for information. The reason is visible in the example above. The `detail` string is meant for people, so next week you might reword it, switch it to English or add figures.

If the customer's code is reading that string with a regex, it will break silently. The illustrative code below shows how the client side should branch on the response from the warehouse example:

```python
PROBLEMS = "https://docs.example.vn/problems/"

def xu_ly_loi(resp, method):
if not resp.headers.get("Content-Type", "").startswith("application/problem+json"):
return mo_ticket(None)
problem = resp.json()
if problem.get("type") == PROBLEMS + "vuot-han-muc-don":
return doi_den_ngay_mai(problem.get("account_id"))
if problem.get("retryable") is True and method != "POST":
return thu_lai_sau()
return mo_ticket(problem.get("instance"))
```

Not one line of this function reads `detail`. However often you reword it, the function keeps working. And when it does have to open a ticket, it attaches `instance` so both sides can find the same incident.

**Điểm mấu chốt:** Machines branch on type, people read detail, operations teams search logs by instance. Keep each field to its own job and you can reword messages freely without breaking the customer's code.

The `retryable` flag answers the question every customer asks: should I try again? To be clear, `retryable` is not in the RFC. It is an extension member you define yourself, so you must document its meaning, on the very page that `type` points to.

A Hackernoon article on retries advises retrying only on certain transient error codes, and generally skipping POST to avoid creating duplicate records. In the warehouse example, if the customer's code automatically resends an order-creating POST after a timeout, the warehouse may receive two identical orders, because a timeout does not reveal whether the first order was recorded.

That advice is only complete with an idempotency key. If the customer genuinely needs to retry POSTs, design it so the client generates a unique key for each order and sends that same key on every attempt. When the server sees a key it has already processed, it returns the original result instead of creating a second order.

Until that mechanism exists, `retryable` should be `false` for every error on order-creating POSTs.

## An afternoon's work with Spring

If your API is written in Java/Spring, most of the work is already done. Spring Framework supports RFC 9457 through the `ProblemDetail` class. Spring Boot has a `spring.mvc.problemdetails.enabled` property that makes its built-in exceptions return problem details automatically:

```properties
spring.mvc.problemdetails.enabled=true
```

For your own business errors, build a `ProblemDetail` in an exception handler:

```java
@ExceptionHandler(QuotaExceededException.class)
ProblemDetail handleQuota(QuotaExceededException ex) {
ProblemDetail pd = ProblemDetail.forStatusAndDetail(
HttpStatus.UNPROCESSABLE_ENTITY, ex.getMessage());
pd.setType(URI.create("https://docs.example.vn/problems/vuot-han-muc-don"));
pd.setTitle("Vượt hạn mức đơn trong ngày");
pd.setProperty("account_id", ex.getAccountId());
pd.setProperty("runbook", "https://docs.example.vn/runbook/han-muc-don");
pd.setProperty("retryable", false);
return pd;
}
```

The hard part is not the code. It is sitting down with the customer's operations team to list the errors they actually encounter, naming a `type` for each one and writing a runbook for every error. That is closer to customer discovery than to programming.

## Common traps

The most dangerous trap is leaking implementation details. RFC 9457 specifically warns against exposing things such as stack dumps through the HTTP interface. Write the stack trace to internal logs tied to `instance`, and return only that ID to the customer.

The second trap is every endpoint returning errors in its own shape: one uses `error`, another uses `message`, a third returns HTML. Postman advises that error responses be clearly and consistently structured. A single exception forces the customer's code to add a special-case branch.

The third trap is a `type` that points to an empty URI. The on-call engineer at midnight will click that link. If it leads to a 404, you have lost the chance for them to fix the problem themselves.

## An answer to "How did you reduce tickets after handover?"

In an FDE interview, a question such as "What did you do to reduce tickets after handover?" is a good opening to talk about this skill. Do not say "I handle errors carefully"; describe the actual sequence: standardising errors on RFC 9457, writing a runbook for each `type`, and the customer's operations team closing tickets on their own that previously had to be escalated to you.

The next time a midnight error is resolved by the customer without anyone calling you, that is the sign you designed your error responses correctly.

**Thử ngay tuần này:**

- Take the five most frequent errors in the logs of an API you are working on and rewrite each as RFC 9457 JSON, with type, title, detail, instance and a runbook field.
- For each type you have written, create a short documentation page at that exact URI with three sections: the cause, how to resolve it yourself, and when to call your team.
- Play the customer: write a small client function that branches on type for those five errors without reading the detail string, then reword detail and check the function still works.

## Nguồn

- [RFC 9457 - Problem Details for HTTP APIs](https://www.rfc-editor.org/rfc/rfc9457.html)

- [RFC 7807 - Problem Details for HTTP APIs](https://datatracker.ietf.org/doc/html/rfc7807)

- [Best Practices for API Error Handling](https://blog.postman.com/best-practices-for-api-error-handling/)

- [Spring Framework Reference: Error Responses](https://docs.spring.io/spring-framework/reference/web/webmvc/mvc-ann-rest-exceptions.html)

- [How To Improve Your Backend By Adding Retries to Your API Calls](https://hackernoon.com/how-to-improve-your-backend-by-adding-retries-to-your-api-calls-83r3udx)
