# LLM-as-judge in practice: calibrating a judge against the client's expert pass/fail labels

> A judge with 90% agreement can still miss every wrong answer. If you want a client to trust your eval numbers, measure the judge against the person on their side who knows the domain best.

Bản gốc: https://fdetimes.net/en/guides/calibrate-llm-judge-expert-labels/

Suppose you have 100 answers from an agent. The client's expert marks 10 of them as fail. Your judge marks all 100 as pass. Agreement is 90%, which sounds respectable, yet the judge has not caught a single wrong answer.

Hamel Husain names this trap plainly: when failures are rare, agreement can be misleading. Anthropic, in its guidance on evals for agents, also says LLM-based judges must be closely calibrated with human experts.

The five steps below take you from collecting the expert's labels to a judge reliable enough for the client to base decisions on.

For anyone aiming to become an FDE, this is a skill worth learning early. There is unlikely to be a public benchmark for one particular insurer's claims-assessment process. The most valuable yardstick is the judgement of the best person on the client's side, and your job is to turn that judgement into a judge that runs automatically.

## What you will build, and what you need

By the end you will have three things: a file of pass/fail labels with critiques, a binary judge prompt, and a script that measures the judge's TPR and TNR against the human labels. Evidently AI describes building an LLM judge as a small ML project. The framing is right: there is labelled data, a model, and a loop of evaluation and iteration.

You need Python, access to an LLM through whatever API you already use and, most importantly, time from a domain expert. The code below is trimmed for illustration. The `call_llm` function is a placeholder for your real client, not the API of any particular library.

## Step 1: Find the one person who gets to say "good enough"

Hamel Husain advises identifying a principal domain expert and bringing that person in as early as possible. Pick one person, not a committee. Picture a chatbot that answers customer questions about insurance terms: the right person might be the head of the underwriting team, the one everyone in the department asks when a case gets hard.

Then sample traces. Braintrust suggests starting with 50 to 100 traces a week, prioritising important samples and edge cases. Do not pull a random batch of easy cases, because the judge will fail precisely where you did not look.

When failures are rare, deliberately oversample traces you suspect will fail, such as answers customers complained about or questions touching complex policy terms. A test set of a few dozen traces with only one or two failures will give a TNR that swings wildly with each case, or one you cannot compute at all.

## Step 2: The expert grades pass/fail and writes a critique

Hamel recommends dropping elaborate scoring scales and keeping a single, clear pass or fail decision. Evidently also observes that binary scores tend to be more stable and consistent, for LLMs and human graders alike. A 1–10 scale sounds more granular, but the expert's "6" and the judge's "6" rarely mean the same thing.

The most valuable part is the critique. According to Hamel, critiques should be detailed enough to drop straight into the judge's few-shot prompt. Store each trace as one JSONL line (the examples here come from a Vietnamese insurance chatbot; the critique says the answer said "yes" but skipped the waiting-period condition, so the customer would misunderstand their benefits):

```json
{"id": "t017", "input": "Hợp đồng có chi trả khi ...?", "output": "Có, ...", "label": "fail", "critique": "Trả lời 'có' nhưng bỏ qua điều kiện thời gian chờ; khách sẽ hiểu sai quyền lợi."}
```

Check after this step: reread 10 random critiques. If you cannot tell why an answer failed, neither will the judge. Go back to the expert straight away, while they still remember.

## Step 3: Write the judge prompt around a rubric

Anthropic recommends building clear, structured rubrics for each dimension of a task. Combined with the binary principle, each dimension becomes a yes/no question, and the final verdict is pass or fail. Save the trimmed prompt below as `judge_prompt.txt`. It tells the model it is grading an insurance assistant, asks two yes/no criteria (does the answer state the clause's conditions correctly; does it avoid promising benefits outside the contract), supplies graded examples, and requests JSON output:

```text
Bạn chấm câu trả lời của trợ lý bảo hiểm.
Tiêu chí (mỗi tiêu chí: CÓ/KHÔNG):
1. Nêu đúng điều kiện áp dụng của điều khoản?
2. Không hứa quyền lợi ngoài hợp đồng?
Ví dụ đã chấm:
{few_shot_critiques}
Câu hỏi: {input}
Câu trả lời: {output}
Trả về JSON: {"critique": "...", "label": "pass" | "fail"}
```

Make the judge write its critique before giving a label. When the judge gets it wrong, its critique tells you which criterion it misread. Any trace used as a few-shot example must be removed from the test set, just as you separate train and test data; otherwise the numbers will look artificially good.

## Step 4: Put the judge's labels next to the human's

Now you measure. Run the judge on the traces not used as few-shot examples, record its label on each line, then compute two separate numbers instead of one overall agreement rate. This is the alignment Evidently describes, comparing judge output against hand-labelled ground truth, and Braintrust likewise treats human scores as the ground truth for calibrating LLM scorers.

```python
import json

def call_llm(prompt):  # thay bằng client thật, trả về chuỗi JSON
raise NotImplementedError

TEMPLATE = open("judge_prompt.txt", encoding="utf-8").read()
FEW_SHOT = open("few_shot.txt", encoding="utf-8").read()

def run_judge(rows):
for r in rows:
prompt = (TEMPLATE.replace("{few_shot_critiques}", FEW_SHOT)
.replace("{input}", r["input"])
.replace("{output}", r["output"]))
result = json.loads(call_llm(prompt))
r["judge"] = result["label"]
r["judge_critique"] = result["critique"]
return rows

def rate(hit, total):
return hit / total if total else None  # tránh chia cho 0

def metrics(rows):
tp = sum(r["label"] == "pass" and r["judge"] == "pass" for r in rows)
tn = sum(r["label"] == "fail" and r["judge"] == "fail" for r in rows)
pos = sum(r["label"] == "pass" for r in rows)
neg = sum(r["label"] == "fail" for r in rows)
return {"TPR": rate(tp, pos), "TNR": rate(tn, neg)}

rows = [json.loads(line) for line in open("test.jsonl", encoding="utf-8")]
print(metrics(run_judge(rows)))
```

(The comments read "replace with a real client, returns a JSON string" and "avoid division by zero".) The script uses `replace` rather than `str.format` because the prompt already contains JSON curly braces. A production version should catch cases where the model returns something that is not JSON; that is omitted here for brevity. If TNR comes back as `None`, the test set has no failing cases: that is a signal to return to Step 1 and oversample failures, not evidence of a perfect judge.

TPR is the share of cases the expert passed that the judge also passed. TNR is the share of cases the expert failed that the judge also failed. Back to the opening example: a judge that passes all 100 answers has a TPR of 100% and a TNR of 0%, and the problem hidden by the 90% agreement rate is exposed at once.

**Điểm mấu chốt:** For the client, TNR is usually the number that matters more: it tells them how many of the errors their expert catches the judge also catches.

## Step 5: Iterate the prompt until nothing surprises you

In each round, filter out the cases where the judge disagrees with the expert and read the `judge_critique` for each. Fix exactly the criterion that was misunderstood, or add a few-shot example of that specific error type, then rerun the whole test set.

Evidently describes this as tuning the judge exactly as you would tune a product prompt, while Anthropic acknowledges that model-based grading often needs careful iteration before its accuracy can be verified.

If endless prompt fixes still leave results fluctuating, look at the model. Confident AI notes that with traditional metrics such as GEval, weaker models struggle to produce reliable results. Another option is DeepEval's DAG metric, described as fully deterministic thanks to its decision-tree structure executed by an LLM.

A rubric made of several yes/no criteria is already close to that approach.

## Three common mistakes when doing this with clients

The most common mistake is reporting only a single agreement number, as shown above. The second is letting engineers label in place of the expert because it is "faster". The judge then ends up calibrated to you, not to the person the client trusts.

The third is treating calibration as a one-off. The client's business changes, new types of question appear, and a judge that once matched well will gradually drift.

Anthropic says LLM-based rubrics for subjective tasks, such as research agents, need frequent calibration against expert judgement; SuperAnnotate describes human reviewers as having the final say on ambiguous cases and continually refining the grading criteria.

So schedule a weekly review of 50 to 100 traces with the expert from the very first meeting.

## What does this skill look like on a CV?

When reading FDE or solutions engineer job descriptions, look for phrases such as "evals", "LLM-as-judge", "human-in-the-loop" and "work with domain experts". On your CV, do not write "experience with evals".

Write something like: built a pass/fail judge for a domain chatbot, calibrated against the head underwriter's labels on N traces, raised TNR from X to Y over K rounds.

A line like that shows a recruiter you can do three things an FDE needs: work with the client's people, measure the right thing, and iterate until the numbers can be trusted. At your first working session with a client, do not open with a judge demo. Ask who everyone in the department turns to when a case gets hard.

**Thử ngay tuần này:**

- Take 50 traces from an LLM application you are working on, ask a colleague who understands the domain to label each pass/fail with a short critique, and save them as a JSONL file. If you can only label them yourself, treat it as practice for the process, not as real labels for calibrating a judge
- Deliberately oversample traces you suspect will fail so the test set has enough failing cases, then write a binary judge prompt using 3 critiques as few-shot examples, run it on the remaining 47 traces, and compute TPR and TNR with the script in this article
- Add a line with numbers to your CV: the TPR/TNR your judge achieved against expert labels, and after how many rounds of prompt tuning

## Nguồn

- [Demystifying evals for AI agents](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents)

- [LLM-as-a-judge: a complete guide to using LLMs for evaluations](https://www.evidentlyai.com/llm-guide/llm-as-a-judge)

- [How to run human-in-the-loop evals for LLM apps](https://www.braintrust.dev/articles/human-in-the-loop-evals-for-llm-apps)

- [Creating a LLM-as-a-Judge That Drives Business Results](https://hamel.dev/blog/posts/llm-judge/)

- [How I Built Deterministic LLM Evaluation Metrics for DeepEval](https://www.confident-ai.com/blog/how-i-built-deterministic-llm-evaluation-metrics-for-deepeval)

- [LLM-as-a-judge vs. human evaluation: Why together is better](https://www.superannotate.com/blog/llm-as-a-judge-vs-human-evaluation)
