# Evaluating AI Agents: a short course on scoring the router, each skill and the full agent flow

> An overall error rate does not tell you where an agent breaks. A course from DeepLearning.AI and Arize teaches you to score each component on its own so you can find the fault.

Original: https://fdetimes.net/en/books-courses/evaluating-ai-agents-course-router-skill-evals/

A customer-service agent gets 30 out of 100 questions wrong. That number alone does not tell you what to fix: the router may be picking the wrong skill, or it may pick the right one and the skill then botches the job. *Evaluating AI Agents*, a short course from DeepLearning.AI built with Arize AI, teaches you how to find where the error lives.

DeepLearning.AI announced the course on its community site on 19 February 2025. It is taught by John Gilhuly, Head of Developer Relations, and Aman Khan, Director of Product, both at Arize AI.

For a forward deployed engineer, the value of the course is not the tooling. It is the way of thinking you need when an agent breaks at a customer site while everyone waits for you to explain why.

## What does the course teach, and who is it for?

The thread running through the course is that an agent should be scored at two levels: each component on its own, and the full flow from start to finish. For each component, you choose a suitable evaluator, set of test examples and metric, then improve iteratively, both during development and once the agent is in production.

Evaluators come in two main forms: written in code, or using LLM-as-a-Judge. The course also teaches two supporting skills: adding observability (tracing) so you can see every step the agent took and debug it, and organising evaluation into experiments to improve both output quality and trajectory, meaning the sequence of steps the agent takes to reach an answer.

The entry bar is low. You only need basic Python; experience prompting LLMs helps but is not required. The companion GitHub repo, maintained by ksm26 (not an official DeepLearning.AI page), describes it as a hands-on course in which learners test sub-components such as skills and router decisions against real examples.

## Idea one: separate routing errors from execution errors

Go back to the agent that gets 30 of 100 questions wrong. This is a hypothetical example to show how the method works. Suppose the agent has three skills: order lookup, refund handling and policy questions.

The Arize AX documentation describes router evaluation as starting from a simple question: for this input, did the router pick the right skill? Suppose you find the router chose correctly on 85 queries. Then 15 queries went wrong at the routing step.

When scoring skills, the Arize documentation advises assuming the skill was called correctly, so you score skills only on those 85 queries. If 70 produce good results, the remaining 15 are execution errors. The 30% error rate now splits into two groups of 15, and each needs a completely different fix.

For routing errors, the first things to try are clearer skill descriptions or more examples for the router. Execution errors sit inside the skill itself: the prompt, the tools, or the data the skill retrieves. If you look only at the end-to-end rate, you can easily spend a week tuning the router prompt while the rest of the errors stay untouched inside the skill.

**Key point:** An overall error rate tells you the agent is failing, not which step is failing.

## Idea two: pick the evaluator to fit the job

The companion repo lists three kinds of evaluator: code-based metrics, LLM-as-a-judge, and human annotation. The example above shows that each suits a different task. Scoring the router only requires comparing the skill name it chose with the correct one, so a code-based matching function is enough: cheap, fast and consistent across runs.

Scoring an answer about refund policy is much harder, because there is no single correct answer to match against. This is where LLM-as-a-Judge comes in. But an unvalidated judge is not yet trustworthy, so human annotation is the third layer.

Ask a domain expert on the customer side to hand-score a few dozen queries, then compare their scores with the judge's. On FDE projects this step has a further benefit: the customer helps define what "correct" means, so they are more likely to trust the evaluation suite.

## Idea three: tracing first, experiments second

You cannot score the router if you cannot see what it chose. Observability is therefore the precondition for everything above, not an add-on.

Once each step is traced, rerun every prompt change or model swap on the same set of examples as an experiment. That lets you put the new version's output and trajectory side by side with the old one, rather than eyeballing it.

The first thing to do on arriving at a customer site is to turn on tracing before you fix anything, because without traces any conclusion about the cause of a failure is guesswork.

Before taking the course, check whether you are making any of the mistakes below. They follow directly from the course's method:

| Common mistake | What to do instead |
|---|---|
| Tracking only the end-to-end error rate | Score the router and each skill separately, and still measure the full flow |
| Scoring skills on queries the router already got wrong | Score skills only on queries the router got right |
| Using LLM-as-a-Judge to compare skill names | Write a code-based matching function |
| Trusting judge scores without checking them | Compare against a few dozen human-scored queries |
| Changing a prompt and eyeballing a few queries | Run an experiment on the same example set, comparing before and after |

A short exercise: take the 100-query example above and suppose that after you rewrite the skill descriptions, the router picks correctly on 95 queries, and the skill handles 80 of those well. Recalculate the routing and execution errors. The answer is 5 and 15: the overall error rate falls, but the errors inside the skill have not been touched at all.

## When to take it, and how to put it on your CV

The course suits engineers who have already built a simple agent themselves and can see it failing without knowing where. If you have never written an agent with a router, build a small one first, because the exercises mean far more when you have real failures of your own to examine.

Afterwards, read the agent evaluation section of the Arize AX documentation to cement the rule of separating router from skill.

On the career side, when reading FDE job descriptions, look for phrases such as "evals", "observability" or "production monitoring". They signal that the hiring team needs exactly this skill. On your CV, do not just write "completed an evaluation course". Write that you split an agent's error rate into routing errors and execution errors, with figures from before and after the fix.

Customers do not need to hear how good the agent is. They need to know, when it gets something wrong, where it went wrong and when it will be fixed. This course helps you answer the first question with data.

**Try this week:**

- Take 20 real user questions, label the correct skill for each yourself, then write a code-based evaluator that measures how often the router picks the right skill.
- For the queries the router got right, score the skill output separately with LLM-as-a-Judge. Then hand-annotate about 10 of them to see whether the judge agrees with you.
- Add tracing to an agent you are working on, then follow the trajectory of one wrong answer from start to finish.

## Sources

- [Evaluating AI Agents - DeepLearning.AI](https://www.deeplearning.ai/short-courses/evaluating-ai-agents/)

- [✨ New course! Enroll in Evaluating AI Agents](https://community.deeplearning.ai/t/new-course-enroll-in-evaluating-ai-agents/773080)

- [Evaluating AI Agents](https://github.com/ksm26/Evaluating-AI-Agents)

- [Evaluating Agents - Arize AX Docs](https://arize.com/docs/ax/concepts/evaluators/evaluating-agents)
