FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Tools

Instructor: when an LLM breaks the schema, send the error back and ask again

Every client runs a different provider. Forward deployed engineers need schema enforcement that does not depend on any one vendor's API, and Pydantic with Instructor is built for that job.

In brief

  • Instructor uses Pydantic for the schema, validates LLM output and automatically re-asks with the error message when validation fails.
  • The library works across many providers (OpenAI, Anthropic, Google, Mistral, Ollama, Cohere and others), which suits FDEs who must adapt to each client's stack.
  • OpenAI's native Structured Outputs guarantees schema adherence at the API level. Instructor is a provider-independent safety net, and it does not replace good schema design.
ShareLinkedInFacebookX
GraphicInstructor's schema enforcement loop
  1. 1Declare a Pydantic modelFields, types and business validators become the data contract
  2. 2Pass the schema to the LLMPydantic exports JSON Schema; Instructor sends it with the call to any provider
  3. 3Pydantic validates the resultChecks types, constraints and client-specific rules
  4. 4Failure: re-ask with the errorThe error message goes back to the model so it can fix exactly what was wrong
  5. 5Success: a typed objectYour code gets a Pydantic object back, with no json.loads or try/except

When output breaks the schema, Instructor doesn't just retry blindly. It sends the Pydantic error itself back to the model so the model can fix it.

Graphic: FDE Times

OpenAI’s Structured Outputs documentation contains a sentence worth pausing on: JSON mode and Structured Outputs both guarantee valid JSON, but only Structured Outputs guarantees that the output matches the schema. So a JSON string that parses cleanly can still be missing fields, carry the wrong types, or put a meaningless date where an amount of money should be.

In a demo, that hardly matters. For an FDE connecting an LLM to a client’s production systems, it is the failure that surfaces at two in the morning, halfway through a long batch, on a provider you did not choose.

Clients rarely let you choose. One uses OpenAI, another Anthropic, a third runs open models through Ollama on its internal network because the data cannot leave the company. You need schema enforcement that works on all of those stacks. Instructor supports a long list of providers, so it fits that problem well.

Pydantic writes the contract, Instructor enforces it

According to its own documentation, Pydantic is the most widely used data validation library in Python, and its validation core is written in Rust. What matters more here is that a Pydantic model can export JSON Schema, which you can pass straight to an LLM API to say “return exactly this shape”.

Instructor, written by Jason Liu, builds on that foundation. Its homepage describes it as a tool for extracting structured data from any LLM with type safety, validation and automatic retries. One line in the README sums up how it works: failed validations are automatically retried with the error message.

The phrase that matters is “with the error message”. Instructor does not just keep calling until it gets lucky. It sends the error Pydantic raised back to the model, much as a reviewer returns a pull request with specific comments rather than simply rejecting it. The model learns where it went wrong and does not have to guess again from scratch.

Turning the client’s rules into code

Imagine you are working with a logistics company, and your task is to read order emails and push them into the warehouse system. Instead of writing a long prompt that says “remember to return JSON, remember quantity is an integer”, you declare the contract in Pydantic:

from pydantic import BaseModel, Field, field_validator

class OrderLine(BaseModel):
    sku: str = Field(description="Item code from the customer's catalog")
    quantity: int = Field(gt=0)
    unit: str

class Order(BaseModel):
    customer_code: str
    lines: list[OrderLine]
    requested_date: str

    @field_validator("customer_code")
    @classmethod
    def must_be_known(cls, v: str) -> str:
        if not v.startswith("KH-"):
            raise ValueError("customer_code must start with KH-")
        return v

(The field description reads “item code according to the client’s catalogue”, and the error message says “customer_code must start with KH-”, where KH is short for khách hàng, Vietnamese for customer.)

You give this Order model to Instructor along with the email text. If the model returns quantity as the string “hai thùng” (“two cartons”) or a customer_code without the prefix, Pydantic catches the error, Instructor packages it and asks again. What you end up with is a typed Order object that your IDE can autocomplete, with no json.loads and try/except anywhere in your code.

The interesting part is the must_be_known validator. It encodes the client’s business knowledge, not the model’s, and you have put it into the error-correction loop without retraining the model on anything.

This is why Instructor suits FDE work. Every client has its own rules, and Pydantic gives you a place to write them in code rather than in prompt wording.

Retries are insurance, not magic

Every re-ask is another API call, which means more latency and more cost, and validators only catch what you bother to write. A loose schema full of str fields will pass through Instructor without a single re-ask, even if the data inside is nonsense. Output quality still starts with schema quality.

If the client uses only OpenAI, try native Structured Outputs first, because it guarantees schema adherence at the API level instead of fixing errors afterwards. Instructor’s advantage lies elsewhere: it runs on OpenAI, Anthropic, Google, Vertex AI, Mistral, Ollama, llama-cpp-python, Cohere and LiteLLM, so a single codebase can follow you from client to client.

Instructor (validate, then re-ask)

  • Runs on many providers, including open models via Ollama
  • Custom Pydantic validators for each client’s business rules
  • Every retry is an extra API call

OpenAI Structured Outputs (native)

  • Guarantees schema adherence at the API level
  • No error-correction loop needed for anything the schema can describe
  • Only available when the client is on OpenAI

Learn Pydantic first, Instructor second

Installation is just pip install instructor, and the project is MIT-licensed, so the tool itself is not the obstacle. The obstacle is that most developers use Pydantic as a data template, whereas here it is a specification language: Field(description=...) is a prompt, gt=0 is a constraint, field_validator is a business rule.

Practise writing models whose generated JSON Schema would tell a stranger exactly what you want.

When applying for jobs, if the description mentions extracting structured data from LLMs, put this experience near the top of your CV. Describe the work concretely, for example: “moved an extraction pipeline from plain prompts to Pydantic schemas with business validators, running through Instructor on two different providers”.

Language models will keep getting things wrong. Your job is not to hope they get it right, but to write a contract tight enough that every mistake is caught, and corrected, before it touches the client’s data.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
4 sources
Read next on the roadmap · Stage 3: Applied AI10,000 servers in a year: MCP becomes the common socket between LLMs and client systemsThe protocol Anthropic announced in late 2024 now belongs to the Linux Foundation. For FDEs, it turns hand-written integration work into building one server that plugs into Claude, ChatGPT, VS Code and Cursor at once.