# How to agree with the client on what ‘success’ means before you write any AI code

> The client wants a system that is ‘always right’. Any model will sometimes be wrong. If you want the project signed off, agree on how many errors the client will accept in the very first week.

Bản gốc: https://fdetimes.net/en/guides/define-ai-acceptance-criteria-with-clients/

Picture a second meeting with a logistics company. It wants to use an LLM to classify complaint emails automatically, and the head of operations sums up the requirement in one sentence: "The classification has to be right every time. No mistakes."

If you nod and go back to your editor, the project has probably already failed. Three months later, at sign-off, a single misclassified email is enough for the client to say the system does not meet the requirement, and you will have nothing in writing that says otherwise.

For an FDE, the hardest part of an AI project is often not the prompt or the pipeline. It is agreeing with the client on what "success" means when the system is certain to be wrong some of the time. This guide explains how to do that before you write the first line of code.

## Why "always right" is not a criterion

Software engineers are used to unit tests: a red test means a bug, and every bug gets fixed. Hamel Husain, who advises teams on evals for LLM products, points out that LLM evals are different: an eval suite does not have to pass at 100%. OpenAI's guide to evals likewise says that evals exist precisely because AI systems are nondeterministic.

Husain's more important point is that the pass rate you need is a product decision, and it depends on which errors you are willing to accept. So when a client says "no mistakes", the right answer is not "impossible". The right next question is: which kinds of mistake are unacceptable, and which can you live with?

**Điểm mấu chốt:** How often an AI system may be wrong is a product decision, not a technical one.

Anthropic's documentation on success criteria sets out four requirements: criteria must be specific, measurable, achievable and relevant. Its example is short: don't write "good performance", write "accurate sentiment classification".

"Always right" fails two of those four tests. Nobody can measure it on a finite sample, and no model can achieve it.

## Read real failures before choosing metrics

The usual reflex is to open a ready-made list of metrics and pick accuracy and F1, perhaps adding a "helpfulness" score. Husain warns that this top-down approach tends to lead you to measure what is easy to measure, not what users actually care about.

He suggests going the other way: analyse errors in real data first, then write tests for the specific failure modes you observe. On scale, he suggests starting with about 100 traces to find errors. When you need to validate an LLM judge, label 100 to 200 examples per failure mode.

Applied to the logistics company, that means asking for 100 real complaint emails with personal data removed, running your prototype on them and reading the results alongside a customer service agent. Suppose the two of you find three failure modes.

First, "late delivery" emails are filed as "information request". Second, compensation claims for broken goods are filed as "late delivery". Third, emails that threaten legal action are not escalated to a manager.

These three failures differ widely in severity. Confusing "late delivery" with "information request" costs an agent a few minutes moving the email to the right queue. Missing a threat of legal action can become a legal incident. This is exactly the information that a single accuracy number hides.

## A worked set of acceptance criteria

Anthropic recommends setting criteria across several dimensions at once. One of them is error severity, with the example "90% of errors are inconveniences rather than serious errors".

Latency is on the list too. For safety, Anthropic gives a criterion with a clear number: fewer than 0.1% of outputs flagged as toxic by a content filter, measured over 10,000 runs.

Note how that criterion is written. It has a threshold, a sample size and a measurement tool. The criteria for the email project should follow the same template. The numbers in the table below are illustrative only; you would replace them with the figures you agree with your actual client:

| Dimension | Criterion | Measured on |
|---|---|---|
| Overall classification | At least 92% of emails assigned to the correct category | 500 emails with labels confirmed by the client |
| Error severity | At least 90% of errors are only "wrong queue" errors | The errors in the same 500 emails |
| Legal-threat emails | No legal-threat email missed in the legal-threat test set | 150 labelled legal-threat emails |
| Latency | Classification finishes before the email enters the queue | Staging environment logs |

Now run the numbers on the first row. At 1,000 emails a day, 92% means about 80 misclassified emails. If 90% of those are only queue mix-ups, that leaves about 8 serious errors a day.

The question you put on the table is: "Eight serious emails a day will be misclassified. Can your team live with that, or do we need a human review step?"

That question turns a vague worry into a business choice that can actually be discussed. It also matches how Husain sees pass rates: they are product decisions about which errors are acceptable, so the product owner on the client side has to be part of the decision.

What about the "none missed" row? An absolute threshold only makes sense on a narrow, fixed test set that is large enough for the result not to be down to luck, such as the 150 emails in the illustrative table. If the system cannot meet it, the answer is to add human review, not to quietly lower the bar.

## Five steps for your next project

Step one: rewrite every vague requirement as a sentence with a number in it. OpenAI's guide also starts the eval design process with exactly this question: what does success look like?

Step two: ask for real data and read about 100 traces with a domain expert from the client's side. Record the failure modes in their language, not in ML jargon.

Step three: rank each failure mode by severity, then set criteria across several dimensions: accuracy, severity, latency and safety.

Step four: check that the thresholds are achievable. Anthropic stresses that metrics should not exceed what current frontier models can do, and should be grounded in benchmarks or experiments you have run. Run your prototype on the test set before you make promises, so that in negotiation you have a baseline in hand rather than a hunch.

Step five: get the criteria signed, then automate the evals. OpenAI recommends setting up continuous evaluation so evals run after every change, and calibrating automated grading against human judgement.

## Common traps

The first trap is copying thresholds from one project to the next. Anthropic notes that citation accuracy may be critical for a medical application but far less important for a conversational chatbot. A 92% threshold that suits email classification may not suit summarising medical records.

The second trap is having a single headline number. A good accuracy figure can still hide a serious failure mode. You need at least one criterion on error severity and a separate one for the failure the client fears most.

The third trap is choosing the test set yourself. If you assigned the labels, the client has grounds to distrust the results at sign-off. Have someone on the client side confirm the labels, at least for the test sets covering serious failures.

The last trap is treating the criteria as a document you sign once and file away. Prompts change, models change, and input data changes too. The criteria only mean something if the evals re-run after each of those changes.

When you see phrases such as "define success metrics with customers" or "build evals" in a job description, this is the skill they mean. On your CV, be specific: which thresholds you agreed, how many samples you measured on, and what the client decided based on those numbers.

This week's exercise: pick an AI feature you are working on and write a four-row set of criteria like the table above, each row with a threshold, a sample size and a measurement method. Then try to answer the question "how many serious errors a day are acceptable?" If you cannot answer it, that is the question to ask the client at your next meeting.

**Thử ngay tuần này:**

- Take 100 real outputs from an AI feature you are working on, read each one and tag the errors by failure mode. At the end of the session, count which failure mode comes up most often.
- Rewrite one vague requirement in a current ticket as a criterion with a number, a test dataset and an error threshold.
- On your CV, replace "built a chatbot" with a sentence describing how you agreed acceptance criteria with a client and built evals to check them.

## Nguồn

- [Define success criteria and build evaluations](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests)

- [Your AI Product Needs Evals](https://hamel.dev/blog/posts/evals/)

- [Evals: Doing Error Analysis Before Writing Tests](https://hamel.dev/notes/llm/officehours/erroranalysis.html)

- [Q: How many examples do I need for an eval?](https://hamel.dev/blog/posts/evals-faq/how-many-examples-do-i-need-for-an-eval.html)

- [Evaluation best practices](https://developers.openai.com/api/docs/guides/evaluation-best-practices)
