# After go-live, an agent's real eval set lives in customer traces

> A green test suite on launch day only shows that the agent handles the questions you guessed users would ask. The questions they actually ask are in the traces.

Bản gốc: https://fdetimes.net/en/guides/production-traces-agent-evals-after-go-live/

Two weeks after go-live, the agent's eval suite is still 100% green. Then the customer's head of operations sends you a screenshot: the agent answered confidently, in the right tone, and was completely wrong. Nothing in your suite had ever seen a question like that.

This is unsurprising if the eval suite was written only before launch. The suite is not broken. It was simply written for imagined users, and real users ask different things.

The skill needed here is not writing more tests faster. It is turning the customer's production traces into the real eval set, and repeating that steadily until handover.

## Why is the pre-launch test suite always skewed?

Braintrust states the problem plainly: when you write test cases by hand, the eval set leans towards the behaviour you expect from users, not their actual behaviour. You write the questions you can think of.

Real users misspell words, paste in entire emails, ask two things in one sentence, or use the product for a task nobody designed it for.

That is why Anthropic treats production monitoring as a separate layer that starts after launch. Its job is to detect distribution drift and real-world failures that nobody anticipated. A failure nobody anticipated cannot already be sitting in a test suite written before launch day.

With ordinary software, reading the code tells you what the system will do. Not so with agents: LangChain's Harrison Chase argues that the source of truth has shifted from code to traces, the record of what the agent actually did, and that production traces themselves become your evaluation dataset.

## One loop, from start to finish

Take an example: you deploy an agent that handles return requests for a retail chain. The agent reads the customer's message, looks up the order through a tool, then decides to approve, reject or hand over to a staff member. The first week after go-live produces a few thousand traces.

**Step 1: sample deliberately, but always keep a random share.** Hamel Husain and Shreya Shankar advise including some randomly selected traces in every batch. If you only read flagged traces (the user clicked "unsatisfied", a tool returned an error), you only see the failures that were already noisy. The random share is where silent failures show up.

```python
import random

def sample_batch(all_traces, flagged, n=50, random_share=0.3):
# random_share is an example value; tune it for your project
k_random = int(n * random_share)
picked = random.sample(all_traces, k_random)
picked_ids = {t["id"] for t in picked}
rest = [t for t in flagged if t["id"] not in picked_ids]
return picked + rest[: n - k_random]
```

**Step 2: annotate by hand first; don't reach for the machine yet.** Husain and Shankar set a minimum of 30 hand-annotated traces before looking at suggestions from an agent. For each trace, write one short line on the first place the agent went wrong, such as "called the order lookup tool with the phone number instead of the order ID". No need to categorise yet.

This is the step most people want to skip. But the two authors regard error analysis as the most important activity in evals. Everything that follows depends on the quality of those 30 lines of notes.

**Step 3: group into patterns, then rank them.** Braintrust suggests treating each significant pattern in production data as a candidate eval slice, prioritised along three axes: volume, severity and regression risk. The table below is a hypothetical example for the returns agent:

| Pattern (hypothetical) | Volume | Severity | Regression risk |
|---|---|---|---|
| Customer sends several order IDs in one message; agent handles only the first | High | Medium | High, easily recurs when the prompt is changed |
| Approves a return for an order past the policy window | Low | Very high, real money lost | Medium |
| Replies in an overly stiff tone to an annoyed customer | Medium | Low | Low |

The second pattern is rare but ranks at the top, because a single mistake costs real money. The table is also what you bring to the meeting with the customer to agree the order of fixes together.

**Step 4: turn traces into test cases.** Anthropic recommends turning user-reported failures into test cases so the eval set reflects how the product is actually used. Keep the original input exactly as it was, typos included, and state the correct behaviour clearly:

```python
case = {
"id": "return-expired-policy-017",
"source_trace": "trace_8f2c...",
"slice": "expired_policy",
"input": "CUSTOMER_ORIGINAL_MESSAGE (kept verbatim)",
"tool_state": {"order_date": "...", "policy_days": "..."},
"expected": "reject or hand over to staff, do not auto-approve",
}
```

**Step 5: write the evaluator, then check the evaluator itself.** Husain and Shankar's rule is to write evaluators for failures you have found, not failures you imagine. The "approves expired orders" pattern can be checked with plain code, comparing the order date with the policy. The "overly stiff tone" pattern probably needs an LLM as grader.

Graders can be wrong too. Anthropic is explicit that you will not know whether a grader works well unless you read the transcripts and scores from many runs. The simplest approach is to put the grader's scores next to your own annotations from step 2:

```python
def check_grader(rows):
# rows: [{"trace_id": ..., "human": "pass"/"fail", "grader": "pass"/"fail"}]
disagree = [r for r in rows if r["human"] != r["grader"]]
agreement = 1 - len(disagree) / len(rows)
return agreement, disagree  # re-read the transcript of every disagreeing trace
```

The agreement figure only tells you whether there is a problem. The `disagree` list is what needs reading: open each transcript, see where the grader went wrong, and fix the grader before trusting any number it returns.

**Điểm mấu chốt:** Production traces are not logs to archive. They are the real exam users set the agent every day.

## Common traps

The first trap is treating the pre-launch suite as finished and adding cases only when someone complains. Anthropic warns that relying mainly on users to catch failures harms those very users. For an FDE, the cost also includes the trust of the project sponsor on the customer side.

The second trap is reading only flagged traces. It looks efficient, but it means you measure only the kind of failure that makes users click a button, and miss the times the agent was smoothly wrong.

The third trap is letting an LLM categorise failures from the start. If you have not read enough traces yourself, you have no baseline for judging whether the machine's suggestions are right. The fourth trap goes with it: building a dozen evaluators for risks that sounded plausible in a brainstorm, when real data has never shown them.

## Turning this skill into an edge when job hunting

When reading job descriptions for FDE roles or customer-facing AI engineer positions, look for phrases such as "production monitoring", "error analysis" and "eval dataset from traces". When you see them, have a story ready about a loop like the one above, drawn from your own project.

On your CV, don't just write "built an eval system". Describe one loop: how many traces you annotated, which patterns you found, how you prioritised them, how many regression tests you added. In interviews, walking through one specific failure pattern in detail is more convincing than any definition of evals.

Go-live does not mark the end of eval work. From that day on, the customer starts showing you what your eval set is missing.

**Thử ngay tuần này:**

- Take 30 traces from an agent or chatbot you run (or from a personal demo with logs) and write a one-line note for each trace by hand, without LLM suggestions.
- Group those notes into 3-5 failure patterns, give each a rough score for volume, severity and regression risk, then turn the highest-priority pattern into at least one test case whose input comes from a real trace.
- Rewrite one line of your CV using this template: 'Built an eval loop from production traces: annotated X traces, identified Y failure patterns, added Z regression test cases'.

## Nguồn

- [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)

- [Agent observability powers agent evaluation](https://www.langchain.com/conceptual-guides/agent-observability-powers-agent-evaluation)

- [How to analyze AI agent usage patterns to build eval datasets (2026)](https://www.braintrust.dev/articles/analyze-ai-agent-usage-patterns-eval-datasets-2026)

- [AI Evals: Everything You Need to Know – Hamel's Blog](https://hamel.dev/blog/posts/evals-faq/)
