FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Tools

Pydantic AI: typed tools and evals let FDEs switch models without gambling

In Pydantic AI, switching models takes a one-string edit. To show a client the agent still works afterwards, you also need an eval suite and a pinned library version.

In brief

  • Pydantic AI validates tool arguments against a schema before your code runs, and switching models means changing a single string.
  • Pydantic Evals is a separate package that does not depend on pydantic-ai. It tests agent behaviour much as pytest tests code.
  • v2 changes what the openai: model name means and cuts the no-breaking-changes window between major versions from 6 months to 3, so pin your version.
ShareLinkedInFacebookX

Say you are working on site at a logistics company. An agent that reads complaint emails has run smoothly through two weeks of demos. Then one morning it calls the order-lookup tool with a whole sentence as the order ID, and the backend returns a 500 error. Pydantic AI was built to catch this kind of error early.

Pydantic AI describes itself as a Python SDK for AI, built around a typed, extensible agent loop, where switching to a different model means changing a single string. It is built by the team behind Pydantic and released under the MIT licence. Its GitHub repository sums it up in one line: every model and every interface is typed end to end.

For an FDE, the value is in the word “typed”. On a client site, the hard part is rarely writing a better prompt. It is connecting the agent to the client’s real systems, and those systems reject data in the wrong shape.

Where do types stop errors?

Pydantic AI enforces types in three places: structured output, typed dependency injection and typed tools. Tools are where it matters most. The arguments the model generates are validated against a schema before your code runs.

Go back to the logistics example. You declare a tool, tra_don_hang (look up order), that takes an order ID in a fixed format. When the model passes in the whole sentence “my order is late”, validation catches it before the request reaches the client’s API. The error happens inside the agent, where you can still handle it, rather than in the client’s operations logs.

Output works the same way. You can require the agent to return a ticket with three fields: issue type, order ID and priority. On the receiving end is the client’s ticketing system, which needs exactly those three fields, not a paragraph of prose. The data contract is written in Python, so the client’s engineers can read it without having to infer it from a prompt.

Switching models takes one string; checking the switch takes an eval suite

Switching models with one string is handy when a client asks, “Can we use a cheaper model?” But convenient does not mean safe. To answer that question properly you need a measurable number, and Pydantic Evals is the tool for measuring it.

Pydantic Evals is a separate package, installed on its own with pip or uv, and it does not depend on pydantic-ai. That means you can use it on AI systems not built with Pydantic AI. Logfire is an optional dependency for when you want to observe results.

Evals structures a test simply. A Dataset contains many Cases, each one a test scenario. The Task is what you run, in this case the agent. Evaluators analyse and score the Task’s output on each Case. The Pydantic team compares it to how pytest tests code, and that is how it should be used.

What does a small eval suite look like?

Imagine you take 30 real emails that the client has anonymised. Each email becomes a Case, paired with the correct ticket a support agent assigned to it. The Evaluator checks two things: whether the extracted order ID matches, and whether the priority is right. The figures below are hypothetical, for illustration.

Configuration Cases passed / 30 Rate
Current model 27 90%
Cheaper model proposed by the client 24 80%

The cheaper model fails three more Cases, meaning three more wrong tickets for every 30 emails. Now the client can decide for themselves whether those three wrong tickets are worth the savings. Without this table, all you have to go on is a feeling. With it, the meeting becomes a discussion of trade-offs rather than an argument over impressions.

The first thing to do at a client is to ask for data for the Dataset, before you even write the agent. Thirty correctly labelled Cases are worth more than any demo.

The limits: the API changes faster than you think

v1 shipped on 4 September 2025 with a commitment to no breaking changes for at least 6 months. It introduced durable execution: an agent that crashes partway through a complex workflow can resume exactly where it stopped. For long-running enterprise processes, this feature is worth trying.

But the release cadence is fast. Douwe Maan’s post introducing v2, on 23 June 2026, brought in the concept of a capability, which bundles instructions, tools, hooks and model settings into a single composable unit.

It also came with a breaking change that can fail silently: the openai: model name now uses the Responses API, and to keep Chat Completions you must switch to openai-chat:. The no-breaking-changes window between major versions has also been cut from 6 months to 3.

On PyPI, the latest version is 2.54.0, released on 3 October 2026, and it requires Python 3.10 or later. The lesson for FDEs: always pin the version in client projects, and rerun the eval suite on every upgrade. The same suite serves both to compare models and as a safety net when upgrading.

What to learn first, and what to put on your CV

If you already know Pydantic from FastAPI, you have half the foundation. Learn in this order: structured output, then typed tools, then dependency injection, and only then v2’s capabilities. Alongside that, write your first eval suite in the first week, not at the end of the project.

When reading job descriptions for FDE or AI engineer roles, look for phrases such as “evaluation”, “structured output” and “production agent”. On a CV, “used Pydantic AI” says little. “Built a 30-Case Dataset and found that a replacement model cut the pass rate by 10 points before it reached production” shows a recruiter you work like an FDE.

The library will keep changing. The way of working it encourages is worth keeping: define data in code, measure behaviour with evals, and only then switch models.

6 sources
Read next on the roadmap · Stage 3: Applied AIn8n for FDEs: build an LLM workflow in an afternoon, but know when to move to codeWith drag-and-drop, you can have a customer demo running very quickly. Whether that demo can run in production depends on three things: the limits of the Code node, the scope of the licence and how the system scales.