# Writing your first eval with Inspect AI: dataset, solver, scorer and reading the log

> With just three questions about a returns policy, a string-matching scorer can grade a model 3/3 when it actually got only 1/3 right. Only the log file will show you that.

Bản gốc: https://fdetimes.net/en/guides/first-eval-inspect-ai-guide/

You set the target to "30 ngày" ("30 days") for a question about the exchange window. The model replies "Bạn được đổi trong 7 ngày, không phải 30 ngày" ("You can exchange within 7 days, not 30 days"), and the scorer still marks it correct. The model is not at fault here. The fault is in how you wrote the eval, and you only find it if you take the trouble to open the log.

Inspect is an AI evaluation framework developed by the UK AI Security Institute and Meridian Labs. For an FDE, its value is concrete. When a client asks "is the new prompt better than the old one?", you need a number you can reproduce and a log file that points to every wrong answer, not a hunch formed after a few manual tries.

This guide works through a hypothetical example: a bot that answers questions about a retailer's returns policy. You will write a dataset, choose a solver, try two kinds of scorer, run the eval and read the log. The sample questions are in Vietnamese, and one of them shows why that matters for string matching.

The code below has been trimmed for readability, so check the import paths and parameter names against the Inspect docs once before running it.

## What do you need?

You need Python, the Inspect package installed as described in the docs, and an API key for a model provider. Create an empty folder, for example `return-policy-eval/`, and a file called `eval.py`. That is all.

## Step 1: the dataset records the answers the client considers correct

A dataset in Inspect is usually a table with two columns, `input` and `target`. Inspect reads CSV, JSON and JSON Lines out of the box. For a small hand-written eval, the quickest option is to build a `MemoryDataset` from a list of `Sample` objects:

```python
from inspect_ai.dataset import Sample, MemoryDataset

dataset = MemoryDataset([
Sample(input="Tôi mua áo được bao lâu thì còn đổi được?", target="30 ngày"),
Sample(input="Hàng giảm giá có được trả lại không?", target="không"),
Sample(input="Đổi hàng có mất phí vận chuyển không?", target="miễn phí"),
])
```

The three questions ask how long after buying a shirt it can still be exchanged (target "30 days"), whether sale items can be returned (target "no") and whether exchanges carry a shipping fee (target "free").

Before moving on, ask yourself: if a customer service agent read each target, would they agree it is exactly the answer they need?

The `target` field can be a literal value such as "30 ngày", or a description that another model uses as the basis for grading. That choice determines which scorer you use in step 3.

The most common mistake is inventing the questions yourself. On a client site, take inputs from real support logs or from the questions the operations team is asked most often. Once the dataset grows beyond a few dozen rows, move it into a CSV or JSON Lines file so people on the client side can edit it themselves.

## Step 2: the solver decides how the model is asked

A solver is a Python function that takes a `TaskState` and a generate function, transforms the state and returns it. The simplest solver is `generate()`: it calls the model and appends the reply to the conversation history.

```python
from inspect_ai import Task, task
from inspect_ai.solver import generate
from inspect_ai.scorer import includes

@task
def return_policy():
return Task(
dataset=dataset,
solver=generate(),
scorer=includes(),
)
```

Starting with a bare `generate()` is deliberate. It gives you a baseline. Later, when you add a system prompt or a document retrieval step, you have a number to compare against.

## Step 3: run it, and do not trust the first number

Run the eval from the command line with `inspect eval`, choosing the model with `--model`:

```bash
inspect eval eval.py --model 
```

The `includes()` scorer checks whether the target appears anywhere in the output: a substring match. It is fast, cheap and good enough when the answer is a number or a specific code. But go back to the example at the top and its limits become obvious.

Suppose the model answers the first question: "Bạn được đổi trong 7 ngày, không phải 30 ngày như nhiều người nghĩ" ("You can exchange within 7 days, not 30 days as many people think"). The string "30 ngày" is in the output, so `includes()` marks it correct.

The second question is just as dangerous. In Vietnamese, "không" means "no" or "not", and also turns a sentence into a question, so it appears in almost every sentence, including "Có, hàng giảm giá vẫn trả lại được, không vấn đề gì" ("Yes, sale items can still be returned, no problem"). Assume further that the model correctly answers the third question, saying exchanges come with free shipping. The scoreboard will report 3/3, when in fact only 1/3 of the answers are right.

**Điểm mấu chốt:** A high eval score proves nothing until you have read every answer that was marked correct.

When the answer is an idea rather than a string, use `model_graded_fact()`. This scorer has another model judge whether the output contains the fact stated in the target. Targets should therefore be written as descriptive sentences rather than short phrases:

```python
from inspect_ai.scorer import model_graded_fact

dataset_graded = MemoryDataset([
Sample(
input="Tôi mua áo được bao lâu thì còn đổi được?",
target="Khách được đổi hàng trong vòng 30 ngày kể từ ngày mua",
),
Sample(
input="Hàng giảm giá có được trả lại không?",
target="Hàng giảm giá không được trả lại",
),
Sample(
input="Đổi hàng có mất phí vận chuyển không?",
target="Đổi hàng được miễn phí vận chuyển",
),
])

@task
def return_policy_graded():
return Task(
dataset=dataset_graded,
solver=generate(),
scorer=model_graded_fact(),
)
```

The new targets read: "Customers may exchange goods within 30 days of purchase", "Sale items cannot be returned" and "Exchanges come with free shipping".

Run it again with the same `inspect eval` command. For the two wrong answers in the example, what you hope for is that the grading model spots that "7 days" contradicts the fact "30 days", and that "can still be returned" contradicts "cannot be returned". Do not assume it will: step 4 is where you check.

| Situation | Scorer to use | Risk to watch |
|---|---|---|
| The answer is an order code, an amount of money, a product name | `includes()` | Marked correct when the answer sits inside a negative sentence |
| The answer is a policy or an explanation | `model_graded_fact()` | The grading model can be wrong, so you still have to read samples |

## Step 4: open the log, because that is the real product

By default Inspect writes logs to a `./logs` subfolder of your current directory. You can change the location with the `INSPECT_LOG_DIR` environment variable, which is useful if you want to collect logs from several client projects in one place.

Since v0.3.46 the default format is `.eval`, a compact and fast binary format. If you need a text file you can read by eye, use `.json`.

Each run produces a structured `EvalLog`. The three fields you will use most are `status` (whether the run completed), `samples` (the input, output, target and score of each sample) and `results` (aggregate figures from the metrics).

Hamel Husain, who has written detailed notes on Inspect, observes that the log keeps enough context to debug, analyse and, most importantly, reproduce the eval.

```python
from inspect_ai.log import read_eval_log

log = read_eval_log("logs/.eval")
print(log.status)
for s in log.samples:
print(s.input, "|", s.output, "|", s.target, "|", s.scores)
```

After this step, check two things. `status` must report completion before you trust `results`. Then filter out every sample marked correct and read its output. This is where you catch the "7 days, not 30 days" answer, and also where you confirm whether the `model_graded_fact()` run has graded correctly.

Log files can run to several GB. `read_eval_log` supports reading only the header, so when you need just the status and aggregate results, do not load every sample into memory.

## What does this skill look like on a client site?

Picture the second week of a project. The client's operations team wants to change the prompt and asks whether the new version is better. Without an eval, the conversation revolves around a handful of screenshots.

With one, you run the same Task on two configurations, open the two logs side by side and show exactly which answers were fixed and which ones broke.

So the first thing to do at a client is sit down with the person who knows the business best and write 20 Samples together, with descriptive targets they agree with. That dataset is usually worth more than any later prompt tuning, because it turns "correct" into something both sides have agreed on.

If you are applying for FDE roles, look for job descriptions that mention "evaluation", "eval harness" or "measuring LLM quality". On your CV, do not just write "experienced with Inspect".

Describe a number instead: how many samples you wrote, which scorer you caught grading wrongly, and how the results changed after you changed the grading method. Recruiters trust a log file more than a line of self-description.

A good eval is not there to prove the model is already good. It is where you find out where the model is going wrong, before the client does.

**Thử ngay tuần này:**

- Write 10 Samples from an internal document you are working with (a policy, an FAQ, an API guide), run them with includes() and note the samples that were graded wrongly.
- Rerun the same Task with model_graded_fact() and descriptive targets, and compare the results across the two logs.
- Write a short script using read_eval_log that prints every sample with a wrong score, along with its input and output.

## Nguồn

- [Inspect (docs home)](https://inspect.aisi.org.uk/)

- [Datasets – Inspect](https://inspect.aisi.org.uk/datasets.html)

- [Solvers – Inspect](https://inspect.aisi.org.uk/solvers.html)

- [Scorers – Inspect](https://inspect.aisi.org.uk/scorers.html)

- [Eval Logs – Inspect](https://inspect.aisi.org.uk/eval-logs.html)

- [Inspect AI, An OSS Python Library For LLM Evals](https://hamel.dev/notes/llm/evals/inspect.html)
