FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Simulated users: how to test a multi-turn conversational agent before the client sees it

An agent can pass every sample test and still fail on a real user's third question. To catch that before the demo, build a fake customer with its own goal, its own personality and a limited supply of patience.

Một hình nộm thử va chạm màu cam ngồi đối diện một robot qua chiếc bàn, giữa hai bên là ba bong bóng thoại xếp so le theo lượt, bong bóng cuối có tia sét giận dữ, còn đèn báo trên bàn đang sáng đỏ.

In brief

  • Static question–answer pairs cannot test a multi-turn agent, because the user's next message depends on the agent's last reply.
  • A simulated user needs a hidden goal, a consistent persona and a patience budget; grade on database state, not on what the agent says it did.
  • Run each scenario several times and report pass^k: on τ-bench, GPT-4o succeeded on under 50% of tasks and scored pass^8 below 25% in the retail domain.
ShareLinkedInFacebookX
GraphicOne round of agent testing with a simulated user
  1. 1Write the scenarioHidden goal, persona, information-reveal rules, patience budget
  2. 2Simulated user opensAn LLM plays the customer, keeping the same style and expertise from start to finish
  3. 3Agent replies, calls APIsThe agent must draw out the goal through conversation and may write to the database
  4. 4Track the goal, stopStop when the goal is met or the simulated user runs out of patience
  5. 5Grade on the databaseCompare database state with the expected outcome; do not trust the agent's own account
  6. 6Repeat k times, compute pass^kReset the data and rerun the same scenario to measure consistency

The agent passes only when the database reaches the expected state and that result repeats across many runs.

Graphic: FDE Times

Picture the last week before handover. The returns agent you built for a client has passed all two hundred sample question–answer pairs. Then someone on the client’s staff sits down to try it, types “I want to exchange my shoes”, and by the third turn the agent asks for the order number they have just given it.

A bug like this does not live in any single answer. It only shows up as the conversation goes on, when the user answers incompletely, changes their mind or gets impatient. That is why an FDE building conversational agents needs one more skill: building a simulated user to “interview” the agent hundreds of times before a real customer touches it.

Why is a static test set not enough?

The Strands Evals team says plainly that a static dataset of input–output pairs, however large, cannot capture the dynamics of a conversation.

The user’s next message depends on what the agent has just said. If the agent asks the wrong question, a real person answers differently, and every turn after that heads down a branch the sample test set never anticipated.

So you need a conversation partner that reacts. τ-bench, a benchmark published by Sierra, does exactly this: a language model plays the user and talks to an agent that can call business APIs. Sierra describes the simulator as an LLM guided by instructions specific to each scenario.

What does a simulated user need?

The first thing is a goal, and the goal must be hidden. Tian Pan, writing about synthetic users for evaluating multi-turn agents, stresses that the goal must not be revealed to the agent in advance; the agent has to draw it out through conversation. That is precisely the skill you want to test, so do not accidentally put the answer in the agent’s system prompt.

The second is a consistent persona. Strands Evals requires the simulated user to keep the same communication style, level of expertise and personality from start to finish. If an “older customer who rarely uses apps” suddenly writes like an engineer by turn five, the scenario is broken.

The third is knowing when to stop. Strands Evals tracks the simulated user’s goal alongside the conversation to know when it should end. Tian Pan adds something that is often forgotten: real people have finite patience.

An overly patient simulator makes the agent look right on paths that real customers would have abandoned long ago, so you need a “patience budget”.

Example: a returns agent for a retail chain

Suppose your client is a shoe retailer and the agent can look up orders, change sizes and issue refunds. A simulated scenario might look like the one below. The persona is a 50-year-old customer who rarely uses apps, gives short answers and gets impatient easily; the hidden goal is to exchange the shoes in order W1234 for a size 42, or get a refund to the card if size 42 is out of stock; the reveal rules say to give the order number only when asked and never volunteer the old size.

SCENARIO = {
    "persona": "Khách 50 tuổi, ít dùng app, trả lời ngắn, dễ sốt ruột",
    "hidden_goal": "Đổi đôi giày trong đơn W1234 sang size 42; "
                   "nếu hết size 42 thì hoàn tiền về thẻ",
    "reveal_rules": "Chỉ đưa mã đơn khi agent hỏi; không tự nói size cũ",
    "patience_turns": 6,
    "expected_db": {"W1234": {"status": "exchanged", "size": 42}},
}

Only the simulator knows the hidden goal. The reveal rules force the agent to ask the right questions. A six-turn budget means that if the agent beats around the bush, the fake customer will say “forget it, I’ll call the hotline” and end the conversation. The code for a single run takes only a few lines (the inline comments note that the agent may call APIs and write to the database, and that the simulator checks for itself whether the goal has been met):

def run_episode(agent, sim, scenario, db):
    msg = sim.start(scenario)
    for _ in range(scenario["patience_turns"]):
        reply = agent.respond(msg, db)   # agent có thể gọi API, ghi vào db
        msg, done = sim.next(reply)      # sim tự kiểm mục tiêu đã đạt chưa
        if done:
            break
    return db.snapshot() == scenario["expected_db"]

The last line matters most. The agent can say, very politely, “I’ve changed the size for you” while nothing in the database has changed. τ-bench grades by comparing the database state after each task with the expected outcome, and you should do exactly the same.

Passing once is not passing

Running the scenario above once and getting it right tells you little. Sierra uses a metric called pass^k to measure reliability: whether the agent completes the same task across multiple runs.

The results in the τ-bench paper are sobering: even a leading function-calling agent such as GPT-4o succeeded on fewer than 50% of tasks, and its pass^8 in the retail domain was below 25%.

A quick calculation shows why the number falls so fast. If a scenario succeeds 60% of the time at random and runs are independent, the probability of passing all 8 runs is 0.6 to the power of 8, about 1.7%. An agent that is “usually right” is still almost certain to fail someone on the first day of go-live.

How to do this on a client site

Start with the five to ten business flows the client cares about most, taken from discovery sessions or call-centre logs. For each flow, write a scenario with a hidden goal, a persona, reveal rules, a patience budget and an expected database state. Add a few difficult personas: someone who changes their mind halfway, someone who gives wrong information, someone who asks about things out of scope.

Then run each scenario k times against a copy of the database that is reset before every run. Report both the success rate and pass^k to the client, along with a few representative failed transcripts. Transcripts help clients understand faster than any chart, and they double as next week’s fix list.

Three common traps

The first trap is a simulator that is too well-behaved, as discussed above: without a patience budget, every detour looks like success. The second trap is subtler.

Tian Pan warns that when the simulator and the agent use the same kind of model, the evaluation turns into a “hall of mirrors”: the two sides converse fluently but sound nothing like real people.

The third trap is assuming the simulator resembles real people without checking. A paper posted on arXiv in May 2026 proposes the realsim framework for comparing simulated users with real conversations from a distributional perspective rather than sentence by sentence. Before each batch of runs, a quick review with this table helps:

Trap Check before running
Overly patient simulator Does every scenario have a turn budget and a give-up line?
Hall of mirrors Does the user role use a different model, or at least a prompt very different from the agent’s?
Simulator unlike real people Have you compared simulated transcripts with real ones on message length, number of turns and when users give up?

Putting this skill on your CV

On a CV, describe this skill in one concrete line rather than a generic summary, for example: “Built a simulated-user evaluation suite for 10 returns flows, graded on database state, raising pass^8 from X to Y”.

For interviews, have a failed transcript ready (anonymised) and explain how you found the bug, what you fixed and how pass^k changed. A story with before-and-after numbers is more convincing than any list of tools.

Clients will not remember what your agent scored on the sample test set. They will remember the first time it asked again for the order number they had just given it, and ideally you will have seen that bug before they did, in a simulated transcript.

5 sources
Read next on the roadmap · Stage 5: DeploymentBuild a webhook receiver that withstands forged signatures, replays and duplicate eventsStripe can resend the same event for up to three days, with a new signature each time. A time window cannot stop those duplicates, so your receiver needs a separate line of defence.