FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Context engineering: choosing which customer data the model gets to see

When an agent gives wrong answers at a customer site, the fault usually lies not in the prompt but in what you put into the context window, or forgot to put there.

Kỹ sư ngồi trước laptop đang xem dữ liệu và mã nguồn trong văn phòng làm việc
Photo: Wiki Farazi / CC0

In brief

  • Stuffing more data into the context does not make a model smarter: the longer the context, the worse the model recalls what is in it.
  • The prompt decides what the model does; the context decides what it knows while doing it. With customer data, context is the main lever.
  • Keep references lightweight, load data only when needed, put data in a clear structure and summarise history before you hit the window limit.
ShareLinkedInFacebookX

Picture your second week at a customer site. The order-support agent ran smoothly in the demo, but on real data it starts getting the returns policy wrong. The odd part is that the policy file sits right there in the context, alongside the chat history, three tables of order data and the entire FAQ.

Many engineers’ first reflex is to fix the prompt: add capital letters, add “read the policy carefully”. It rarely works, because the problem is not how you talk to the model but what the model is having to read.

The skill of fixing the right thing has a name, context engineering, and every FDE who works with customer data needs it.

The prompt says “what to do”; the context decides “what it knows”

“Read the policy carefully” changes only how you talk to the model. It does nothing about the fact that the model is reading a policy file buried among chat history, order tables and FAQ entries. Elastic draws the line in exactly this place: prompt engineering handles communication, while context engineering handles what information the model can access when it generates an answer.

Prompt engineering takes the context window as given; context engineering actively curates it. For the returns agent, curation means deciding whether the three order tables go in at all, and whether to include the whole FAQ or only a few passages.

Anthropic describes this as curating and maintaining the optimal set of tokens while the model reasons, while the Prompting Guide stresses design that covers both the instructions and the context that comes with them.

That is why DataHub treats prompt engineering as one component of context engineering, not the other way round: the prompt tells the model what to do, the context decides what it knows while doing it. When an agent goes wrong at a customer, the first question should be “what did the model see at this step?”, not yet “how can the prompt be better written?”.

Why does adding more data make the agent worse?

Intuition says that the more customer data you give a model, the better it understands the situation. Anthropic points to the opposite: as context grows longer, the model’s ability to recall information in it accurately declines, a phenomenon it calls context rot. On its account, an LLM has an “attention budget” that is drawn down as it reads large amounts of context.

Back to the returns agent. The policy file is in the context, but it has to compete for attention with dozens of old chat turns and thousands of order rows that have nothing to do with the question. The model is not broken. It is diluted.

LangChain, quoting Andrej Karpathy, describes context engineering as the art and science of filling the context window with just the right information for the next step. The key words are “the next step”. Context is not a fixed store loaded once at start-up; it is reassembled for every step.

Four questions before every model call

LangChain groups context-engineering strategies into four buckets: write, select, compress and isolate. They work well as a checklist when designing an agent for a customer.

The names are LangChain’s; what follows is how to apply the four buckets to a customer’s agent. Write, in this sense, means recording information outside the window for later reuse. It is close to the memory component the Prompting Guide places within context engineering, covering short-term memory (managing state and history) and long-term memory.

Select means pulling in only the data relevant to the step currently running. Compress means shrinking what is already there, for example with the compaction described below. Isolate, as this checklist reads it, means splitting work apart so each part carries only its own context.

Two concrete techniques from Anthropic help with select and compress. The first is just-in-time retrieval: the agent holds only lightweight identifiers, such as order IDs or document paths, and loads the actual data only when needed, rather than preloading everything.

The second is compaction: when the conversation nears the window limit, summarise its contents and continue with the summary.

Rebuilding the returns agent, line by line

Below is how context is assembled for the returns agent after applying the checklist. The code is an illustrative sketch, but the structure can be used for real.

def build_context(ticket, state):
    blocks = []

    # Chỉ dẫn cố định, ngắn, không lẫn dữ liệu
    blocks.append(("instructions", RETURN_AGENT_RULES))

    # Write: đọc lại ghi chú tiến độ thay vì toàn bộ lịch sử
    blocks.append(("progress_notes", state.notes))

    # Compress: lịch sử dài thì thay bằng bản tóm tắt
    history = state.history
    if count_tokens(history) > HISTORY_LIMIT:
        history = summarize(history)
    blocks.append(("conversation", history))

    # Select: chỉ những đoạn chính sách khớp với ticket
    blocks.append(("policy", search_policy(ticket.text, top_k=3)))

    # Just-in-time: chỉ đưa ID, chi tiết để agent tự gọi tool
    blocks.append(("order_ref", {"order_id": ticket.order_id}))

    return render_with_tags(blocks)

(The comments, in order: fixed, short instructions with no data mixed in; write, re-read progress notes instead of the full history; compress, replace long history with a summary; select, only the policy passages that match the ticket; just-in-time, pass only the ID and let the agent call a tool for details.)

The first change is that the order tables disappear from the context. The agent sees only order_id, and when it genuinely needs a delivery date or payment status, it calls the get_order(order_id) tool to fetch exactly that record. Thousands of irrelevant rows no longer compete with the policy file for attention.

The second change is structure. The Prompting Guide lists structuring input and output, for instance with delimiters or a JSON schema, as a component of context engineering. The render_with_tags function wraps each block in its own named tag, such as policy or conversation, so the model can tell rules from what the customer said from system data.

The output should have a schema too, for example {"decision": ..., "policy_clause": ...}, so you can check which clause the agent relied on.

The third change is isolation. If the agent has to both look up the policy and draft an email to the customer, split those into two steps, each with its own context. The email step needs only the final decision and the brand’s tone of voice, not all ten pages of policy.

At the customer, start from the logs, not the prompt

Before fixing anything, record the exact context the model received on the wrong answer. Reading that log is the quickest way to check two possibilities: the necessary information was buried in a pile of data, or it was missing altogether.

Next, for each step of the agent, write one line describing what that step needs to know to get it right. Every block in the context must answer the question “which description line does this block serve?”. Any block that cannot is the first candidate for removal, or for conversion into a just-in-time reference.

Finally, measure. Each time you change the context, rerun the same eval set and record two numbers: accuracy and average tokens per call. Those two numbers are what persuade the customer, and they are also worth putting on your CV.

A line such as “redesigned the context for an order-support agent, cutting tokens per call while holding eval scores steady” carries far more weight than “experienced in prompt engineering”.

Mistakes that keep recurring

The most common mistake is fixing the prompt when the problem lies in the context. If the model cannot see the right clause, no instruction, however well written, will save it.

The next is letting conversation history grow without limit until the agent starts forgetting what the customer said at the start. That is exactly when compaction is needed: summarise the history before reaching the window limit, rather than waiting for the agent to forget.

The third is letting tools return raw records. A customer API returns JSON with hundreds of fields, and you pass it straight into the context because it is convenient. Write a thin layer that keeps only the fields that step needs.

The fourth is mixing data and instructions without delimiters, which makes it hard for the model to separate system rules from content the customer sent in.

At a customer site, there is always more data than a model can attend to. The FDE’s job is to decide which parts go in, and to own that decision as they would any other piece of code.

5 sources
Read next on the roadmap · Stage 3: Applied AILLM basics for FDEs: tokens, context windows, temperature and why models make things upWhen a client's AI assistant invents a clause that does not exist, these four concepts tell you where the fault lies. Without them, all you can do is edit the prompt and hope.