# Braintrust for FDEs: turning production failures into evals that block merges

> At a customer site, you can only trust an answer to "is the new prompt better?" if a score runs on every pull request. Braintrust is built for exactly that job.

Bản gốc: https://fdetimes.net/en/tools/braintrust-evals-block-merge-regressions/

A prompt is changed at 5pm, looks fine in the customer demo and gets merged. The next morning the customer's support chatbot starts quoting the wrong refund clause, one it used to get right. Anyone who has worked as an FDE long enough has seen this. It doesn't come from bad code. It comes from having nothing that measures quality before the merge.

Braintrust targets exactly that gap. The company describes itself as a proactive observability platform for agents, built around a simple loop: find patterns in production, turn them into evals, and improve quality with every release. For an FDE, it is a way to replace the feeling that something "seems better" with a number that both you and the customer can see.

## What goes into an eval?

Braintrust's docs split every eval into three parts. Data is the set of examples to evaluate: inputs and, optionally, expected outputs or metadata. Task is the function that calls your application to produce an output. Scores are the grading functions.

Scorers are where an FDE's skill shows. Braintrust offers three options: ready-made autoevals, LLM-as-a-judge, or custom code. Autoevals is an open-source library under the MIT licence, and the official example passes the Factuality scorer straight into the scores parameter of Eval().

A custom scorer has a very simple contract. The function takes input, output, expected and metadata, and returns a number from 0 to 1, or an object with a score plus an optional name and metadata. If you leave out the name, Braintrust uses the function name as the score key, so name your functions the way you would name a business metric.

## Example: a refund chatbot for a retail chain

Imagine the customer is a retail chain whose chatbot answers returns questions and must always cite the correct policy code. Factuality helps check whether the content matches the expected answer, but "must cite the policy code" is a hard rule, so checking it in code is cheaper and more stable than asking another model to rule on it. A sketch in Python:

```python
from braintrust import Eval
from autoevals import Factuality

def cites_policy_code(input, output, expected, metadata):
code = metadata["policy_code"]
return 1 if code in output else 0

Eval(
"refund-bot",
data=lambda: load_cases(),   # 40 real questions from the logs
task=lambda input: refund_bot(input),
scores=[Factuality, cites_policy_code],
)
```

Put some numbers on it. Say the dataset has 40 questions and the old prompt cites the right code in 36 of them, so cites_policy_code scores 0.9. A new prompt rewritten to be "friendlier" gets only 28 right, eight fewer, and the score drops to 0.7, even though the answers still read smoothly.

Without evals, those eight wrong answers, a drop of 20 percentage points, only surface when the customer complains. With evals, they show up on the pull request straight away. So the first job at a customer site isn't choosing a tool. It is sitting down with the operations team, pulling out the questions the chatbot has got wrong, and writing the expected answers for them.

## What does putting scores into CI actually block?

The docs recommend running evals on every pull request to catch regressions before they reach production. Braintrust has an official GitHub Action whose stated purpose is to run evals on every PR, block merges based on the results, and report scores to the team automatically.

The Action, in the braintrustdata/eval-action repo, posts the results as a comment on the PR, so the workflow needs the pull-requests: write permission. Two inputs are worth knowing: report_scores, which filters which scores appear in the comment, and terminate_on_failure, which defaults to false.

Don't let the default make the decision for you. Choose deliberately whether a failing eval should stop the workflow, and write down the reason in the README for the customer's team.

If the customer uses GitLab or Jenkins rather than GitHub, the approach is still simple: set BRAINTRUST_API_KEY in the CI environment so the bt eval command can read it. At enterprise customers this is often the biggest snag, because getting a new secret provisioned for a pipeline can take a full week. Ask for it on day one.

**Điểm mấu chốt:** The best eval suites aren't written from imagination. They are written from the failures the customer has already hit.

## The tool doesn't do the hard part for you

Braintrust handles the infrastructure: running, storing, comparing, commenting. It doesn't know which questions matter to your customer. Forty examples picked at random produce a tidy number that means nothing.

LLM-as-a-judge also needs care. It is one model grading another, so before you trust it as a merge gate, grade a small sample by hand and check whether the judge agrees with the humans. Any rule that can be checked in code should be written in code.

One more practical point: scorers can be defined inline in a script, pushed to Braintrust from a file via the CLI, or created in the UI. At a customer site, keep scorers in the repo alongside the code so they are reviewed through PRs, rather than scattered across the web interface.

## What to learn first, and what to put on your CV

A sensible order is to build an Eval() that runs locally with one autoevals scorer, then write a scorer in code, and only then wire it into CI. If a job description for an FDE or applied AI engineer role mentions evals, regression testing for LLM applications or LLM observability, this is the skill you are building.

On your CV, don't write "experienced with Braintrust". Say how many examples your eval suite had and that they came from real logs, what rules your scorers checked, and which regressions your CI gate caught. FDE hiring managers aren't buying tool names. They are buying the ability to turn a customer's worry into a number that can block a merge.

**Thử ngay tuần này:**

- Take 20 real questions the chatbot on your project got wrong, write an expected output for each, and build them into the data for an Eval()
- Write a code-based scorer that checks one hard business rule (a policy code, a currency format) and returns 0 or 1
- Add braintrustdata/eval-action to a test repo, grant the pull-requests: write permission, and open a PR that changes a prompt to see the score comment

## Nguồn

- [Braintrust - The active observability platform for agents](https://www.braintrust.dev/)

- [Evaluate systematically](https://www.braintrust.dev/docs/evaluate)

- [Custom code evaluators](https://www.braintrust.dev/docs/evaluate/custom-code)

- [Run experiments in CI/CD](https://www.braintrust.dev/docs/evaluate/run-in-ci.md)

- [GitHub - braintrustdata/eval-action](https://github.com/braintrustdata/eval-action)

- [GitHub - braintrustdata/autoevals](https://github.com/braintrustdata/autoevals)
