Error analysis from real traces: turning 100 records into a list of what to fix first
Before you write your first evaluator, read your traces, note the first failure in each one, then count. A few focused sessions give you a priority list that no dashboard can generate for you.
In brief
- Hamel Husain calls error analysis the most important activity in evals: read real traces first, measure later.
- Record only the first failure in each trace, label at least 30 traces by hand before asking an LLM to suggest categories, and keep reading until no new failure types appear.
- Rank failure categories by frequency and severity, not by count alone; fix whatever a prompt change can fix first, before building an evaluator.
- 1Sample about 100 tracesPull from production, datasets or test outputs, covering many question types.
- 2Open codingTake free-form notes on the first error per trace; do at least 30 traces by hand.
- 3Axial codingGroup notes into failure types; an LLM can draft them, but a human must review.
- 4Count until saturationCompute each type's error rate; stop when new traces show no new failure types.
- 5Rank and fixPrioritize by frequency and severity; fix what a prompt change can solve first.
Read and name failures by hand first, then count and rank them, and only then decide whether to fix or build an evaluator.
Graphic: FDE Times
An FDE who has just deployed a chatbot for a client usually wants to start with the job that looks most professional: building a suite of automated evaluators. Yet Hamel Husain writes that error analysis is “the most important activity in evals”. The reason is simple: until you know where the system breaks, you do not know what to measure.
The steps below walk through one complete round of error analysis. By the end you will have a label file covering about 100 traces, a list of failure categories you named yourself, a ranking table and a clear decision on what to fix this week.
This is also a skill that separates FDEs from people who only know how to call an API. When a client asks “where is the bot falling short?”, an answer backed by counts from real traces carries far more weight than a general impression.
What do you need?
You need access to the traces of an LLM application: the full record of the input, the intermediate steps such as tool calls or document retrieval, and the output. You also need a spreadsheet or CSV file and Python 3. Langfuse is optional. Its error analysis guide follows exactly this process, but every step can be done with plain files.
To keep things concrete, the article uses one hypothetical case throughout: a customer service bot for a delivery company. The bot answers questions about order status, shipping fees and refund policy. All figures about this bot below are illustrative.
Step 1: Sample a diverse set of traces
Langfuse’s process starts by pulling a representative sample from production traffic, from a dataset or from the outputs of experiments. Husain treats about 100 diverse traces as a practical starting point. Diverse means covering many question types, many kinds of user, and both short and long traces.
Export the sample to a labels.csv file with this structure:
trace_id,note,category
t001,,
t002,,
Check: skim 10 random rows. If all 10 are order-lookup questions, the sample is skewed and you need to resample.
Step 2: Open coding, recording only the first failure
Read each trace and write free-form notes on every problem you see. That is open coding. Do not force yourself into predefined categories; write as if leaving a note for a colleague: “bot promises a 100% refund on an order more than 30 days old”.
The most important rule at this step: record only the first failure that appears in the trace. Husain explains that upstream errors cause downstream ones. If the bot calls the wrong order-lookup tool and then reports the wrong delivery date, the root cause is the tool call; the wrong date is just a consequence.
Husain recommends labelling at least 30 traces by hand before looking at suggestions from an AI agent. The first 30 traces are where you build intuition about the system. Hand the job to a machine too early and you will not understand enough to judge whether its suggestions are right.
Check: the note column for the first 30 rows should contain specific sentences that make sense on first reading. Notes such as “poor answer” are useless; rewrite them.
Step 3: Axial coding, grouping notes into failure categories
Now reread all the notes and group them. That is axial coding. For the hypothetical delivery bot, you might end up with categories such as sai_chinh_sach_hoan_tien (wrong refund policy), bo_qua_ngay_trong_cau_hoi (ignores the date in the question), goi_sai_tool_tra_don (calls the wrong order-lookup tool) and tra_loi_qua_dai (answer too long).
You can ask an LLM to draft the list of categories. But Langfuse’s guide is explicit that you must review the proposed categories yourself, because the model may merge failures with different causes into one group. For example, “wrong shipping fee” because the bot misread the price table is quite different from the case where the bot never asked for the address, even though the two look alike on the surface.
Then fill in the category column for each trace. Leave it blank for traces with no failure. Langfuse handles this step by creating a boolean score for each failure category (Settings → Scores → Create, type: Boolean), so each trace is marked pass or fail per category.
Check: if one category accounts for more than half of all failures, it probably contains several different causes. Split it.
Step 4: Count, and know when to stop reading
The final step of Husain’s process is to count the failures in each category. Langfuse does the same: label every trace against the set of failure categories, then compute a failure rate for each one. The script below is a stripped-down version using only Python’s standard library:
import csv
from collections import Counter
with open("labels.csv", encoding="utf-8") as f:
rows = list(csv.DictReader(f))
total = len(rows)
counts = Counter(r["category"] for r in rows if r["category"])
for cat, n in counts.most_common():
print(f"{cat:30} {n:3} {n/total:.0%}")
When is enough enough? Husain offers a stopping rule: keep reading until you reach theoretical saturation, meaning new traces no longer reveal new failure types or change the existing categories. If you are still creating new categories at trace 100, pull a bigger sample.
Step 5: Rank by frequency and severity
Counting alone is not enough. Product Talk ranks failure categories by both frequency and severity. If the bot rambles 11 times, customers are merely irritated; if it misstates the refund policy 6 times, the company loses real money and customers may complain.
A simple approach is to score each category’s severity from 1 to 3 and multiply by the number of occurrences. This is a simplified calculation, not a standard formula, but it is enough to start a discussion:
SEVERITY = {"sai_chinh_sach_hoan_tien": 3, "goi_sai_tool_tra_don": 3,
"bo_qua_ngay_trong_cau_hoi": 2, "tra_loi_qua_dai": 1}
ranked = sorted(counts, key=lambda c: counts[c] * SEVERITY.get(c, 1),
reverse=True)
With the delivery bot’s hypothetical figures, the ranking might look like this:
| Failure category | Count | Severity | Score | Action |
|---|---|---|---|---|
| goi_sai_tool_tra_don | 8 | 3 | 24 | Fix the tool description in the prompt, re-measure |
| sai_chinh_sach_hoan_tien | 6 | 3 | 18 | Put the policy text into the context |
| bo_qua_ngay_trong_cau_hoi | 7 | 2 | 14 | Consider building an evaluator |
| tra_loi_qua_dai | 11 | 1 | 11 | One line of instruction in the prompt |
The most frequent category ends up at the bottom of the table. That is exactly why you should not rank by frequency alone.
The last column applies Langfuse’s rule: fix first, and do not build an evaluator for a failure a prompt can solve. An evaluator is worth the investment only when a failure persists after you have fixed what you can and needs to be tracked over time.
Common mistakes
In notes from a session given by Husain, Alex Strick van Linschoten lists three common mistakes. The first is skipping the loop: finishing axial coding and never returning to open coding. Once you have failure categories, you read traces with different eyes and often find categories that need splitting or merging.
The second is leaving domain experts out. You may not notice that an answer about refund terms is wrong, but the company’s customer service staff will spot it immediately. The third is automating too early: setting up an LLM-as-judge before you have understood the failures with your own eyes.
Applying it on a client site
On a client engagement, schedule the error analysis session in the first week and invite the person responsible for operations to score severity with you. The finished ranking becomes a plan both sides have agreed on, rather than a list of requests with no order of priority.
For your career, look out for phrases such as “evals”, “trace review” or “failure analysis” in FDE job descriptions.
On your CV, instead of writing “improved chatbot quality”, state that you labelled N traces, identified K failure categories and reduced the failure rate of the top category after a prompt fix.
A before-and-after figure on the same set of traces is evidence a hiring manager understands at once.
Next time someone proposes building a quality dashboard, open 30 traces and read them first. The best dashboards are built on failure categories you named by hand.
Was this article useful?
Thanks for the feedback!
5 sources
- Q: Why is "error analysis" so important in AI evals, and how is it performed? – Hamel's Blog · 2025-06-27
- Error analysis (Langfuse)
- Error Analysis for LLM Applications - Step by step guide
- Error Analysis | Definition and Overview | Product Talk · 2026-09-03
- Error analysis to find failure modes · 2025-05-23