FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Writing your first eval with Inspect AI: dataset, solver, scorer and reading the log

With just three questions about a returns policy, a string-matching scorer can grade a model 3/3 when it actually got only 1/3 right. Only the log file will show you that.

In brief

  • An Inspect eval is a Task with three parts: a dataset (input/target), a solver and a scorer.
  • includes() only checks for a substring, so it can easily mark a wrong answer as correct; model_graded_fact() has another model grade against the fact.
  • Every run produces an EvalLog in ./logs containing the input, output, target and score, enough to debug and rerun.
ShareLinkedInFacebookX

You set the target to “30 ngày” (“30 days”) for a question about the exchange window. The model replies “Bạn được đổi trong 7 ngày, không phải 30 ngày” (“You can exchange within 7 days, not 30 days”), and the scorer still marks it correct. The model is not at fault here. The fault is in how you wrote the eval, and you only find it if you take the trouble to open the log.

Inspect is an AI evaluation framework developed by the UK AI Security Institute and Meridian Labs. For an FDE, its value is concrete. When a client asks “is the new prompt better than the old one?”, you need a number you can reproduce and a log file that points to every wrong answer, not a hunch formed after a few manual tries.

This guide works through a hypothetical example: a bot that answers questions about a retailer’s returns policy. You will write a dataset, choose a solver, try two kinds of scorer, run the eval and read the log. The sample questions are in Vietnamese, and one of them shows why that matters for string matching.

The code below has been trimmed for readability, so check the import paths and parameter names against the Inspect docs once before running it.

What do you need?

You need Python, the Inspect package installed as described in the docs, and an API key for a model provider. Create an empty folder, for example return-policy-eval/, and a file called eval.py. That is all.

Step 1: the dataset records the answers the client considers correct

A dataset in Inspect is usually a table with two columns, input and target. Inspect reads CSV, JSON and JSON Lines out of the box. For a small hand-written eval, the quickest option is to build a MemoryDataset from a list of Sample objects:

from inspect_ai.dataset import Sample, MemoryDataset

dataset = MemoryDataset([
    Sample(input="Tôi mua áo được bao lâu thì còn đổi được?", target="30 ngày"),
    Sample(input="Hàng giảm giá có được trả lại không?", target="không"),
    Sample(input="Đổi hàng có mất phí vận chuyển không?", target="miễn phí"),
])

The three questions ask how long after buying a shirt it can still be exchanged (target “30 days”), whether sale items can be returned (target “no”) and whether exchanges carry a shipping fee (target “free”).

Before moving on, ask yourself: if a customer service agent read each target, would they agree it is exactly the answer they need?

The target field can be a literal value such as “30 ngày”, or a description that another model uses as the basis for grading. That choice determines which scorer you use in step 3.

The most common mistake is inventing the questions yourself. On a client site, take inputs from real support logs or from the questions the operations team is asked most often. Once the dataset grows beyond a few dozen rows, move it into a CSV or JSON Lines file so people on the client side can edit it themselves.

Step 2: the solver decides how the model is asked

A solver is a Python function that takes a TaskState and a generate function, transforms the state and returns it. The simplest solver is generate(): it calls the model and appends the reply to the conversation history.

from inspect_ai import Task, task
from inspect_ai.solver import generate
from inspect_ai.scorer import includes

@task
def return_policy():
    return Task(
        dataset=dataset,
        solver=generate(),
        scorer=includes(),
    )

Starting with a bare generate() is deliberate. It gives you a baseline. Later, when you add a system prompt or a document retrieval step, you have a number to compare against.

Step 3: run it, and do not trust the first number

Run the eval from the command line with inspect eval, choosing the model with --model:

inspect eval eval.py --model <provider/model-name>

The includes() scorer checks whether the target appears anywhere in the output: a substring match. It is fast, cheap and good enough when the answer is a number or a specific code. But go back to the example at the top and its limits become obvious.

Suppose the model answers the first question: “Bạn được đổi trong 7 ngày, không phải 30 ngày như nhiều người nghĩ” (“You can exchange within 7 days, not 30 days as many people think”). The string “30 ngày” is in the output, so includes() marks it correct.

The second question is just as dangerous. In Vietnamese, “không” means “no” or “not”, and also turns a sentence into a question, so it appears in almost every sentence, including “Có, hàng giảm giá vẫn trả lại được, không vấn đề gì” (“Yes, sale items can still be returned, no problem”). Assume further that the model correctly answers the third question, saying exchanges come with free shipping. The scoreboard will report 3/3, when in fact only 1/3 of the answers are right.

When the answer is an idea rather than a string, use model_graded_fact(). This scorer has another model judge whether the output contains the fact stated in the target. Targets should therefore be written as descriptive sentences rather than short phrases:

from inspect_ai.scorer import model_graded_fact

dataset_graded = MemoryDataset([
    Sample(
        input="Tôi mua áo được bao lâu thì còn đổi được?",
        target="Khách được đổi hàng trong vòng 30 ngày kể từ ngày mua",
    ),
    Sample(
        input="Hàng giảm giá có được trả lại không?",
        target="Hàng giảm giá không được trả lại",
    ),
    Sample(
        input="Đổi hàng có mất phí vận chuyển không?",
        target="Đổi hàng được miễn phí vận chuyển",
    ),
])

@task
def return_policy_graded():
    return Task(
        dataset=dataset_graded,
        solver=generate(),
        scorer=model_graded_fact(),
    )

The new targets read: “Customers may exchange goods within 30 days of purchase”, “Sale items cannot be returned” and “Exchanges come with free shipping”.

Run it again with the same inspect eval command. For the two wrong answers in the example, what you hope for is that the grading model spots that “7 days” contradicts the fact “30 days”, and that “can still be returned” contradicts “cannot be returned”. Do not assume it will: step 4 is where you check.

Situation Scorer to use Risk to watch
The answer is an order code, an amount of money, a product name includes() Marked correct when the answer sits inside a negative sentence
The answer is a policy or an explanation model_graded_fact() The grading model can be wrong, so you still have to read samples

Step 4: open the log, because that is the real product

By default Inspect writes logs to a ./logs subfolder of your current directory. You can change the location with the INSPECT_LOG_DIR environment variable, which is useful if you want to collect logs from several client projects in one place.

Since v0.3.46 the default format is .eval, a compact and fast binary format. If you need a text file you can read by eye, use .json.

Each run produces a structured EvalLog. The three fields you will use most are status (whether the run completed), samples (the input, output, target and score of each sample) and results (aggregate figures from the metrics).

Hamel Husain, who has written detailed notes on Inspect, observes that the log keeps enough context to debug, analyse and, most importantly, reproduce the eval.

from inspect_ai.log import read_eval_log

log = read_eval_log("logs/<ten-file>.eval")
print(log.status)
for s in log.samples:
    print(s.input, "|", s.output, "|", s.target, "|", s.scores)

After this step, check two things. status must report completion before you trust results. Then filter out every sample marked correct and read its output. This is where you catch the “7 days, not 30 days” answer, and also where you confirm whether the model_graded_fact() run has graded correctly.

Log files can run to several GB. read_eval_log supports reading only the header, so when you need just the status and aggregate results, do not load every sample into memory.

What does this skill look like on a client site?

Picture the second week of a project. The client’s operations team wants to change the prompt and asks whether the new version is better. Without an eval, the conversation revolves around a handful of screenshots.

With one, you run the same Task on two configurations, open the two logs side by side and show exactly which answers were fixed and which ones broke.

So the first thing to do at a client is sit down with the person who knows the business best and write 20 Samples together, with descriptive targets they agree with. That dataset is usually worth more than any later prompt tuning, because it turns “correct” into something both sides have agreed on.

If you are applying for FDE roles, look for job descriptions that mention “evaluation”, “eval harness” or “measuring LLM quality”. On your CV, do not just write “experienced with Inspect”.

Describe a number instead: how many samples you wrote, which scorer you caught grading wrongly, and how the results changed after you changed the grading method. Recruiters trust a log file more than a line of self-description.

A good eval is not there to prove the model is already good. It is where you find out where the model is going wrong, before the client does.

6 sources
Read next on the roadmap · Stage 6: MeasurementFrom thumbs-down to eval set: turning user complaints into test casesWhen a client sends you a file full of "dissatisfied" clicks, the first job is to read the traces and label them, not to count the clicks.