FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

When an LLM should say "I'm not sure": self-evaluation, confidence and the threshold for handing off to a human

Asking a model "how confident are you?" and trusting the number it gives is the quickest way to put a wrong answer in front of a customer.

In brief

  • LLMs tend to be overconfident when they state how certain they are. Do not use that number as your handoff threshold.
  • Have the model grade its answer as true/false (P(True)), check for bias with an empty input, then choose the threshold on the customer's labelled data.
  • Confidence scores miss questions about values and questions with no answer, so put a rules layer in front of the model.
ShareLinkedInFacebookX
GraphicHow a question passes through the handoff mechanism
  1. 1Rules layerLegal questions, complaints, no matching documents: straight to a person
  2. 2Model answersThe first call generates an answer grounded in the documents
  3. 3Self-grade P(True)A second call asks true/false and reads the probability of the "True" token from logprobs
  4. 4Compare with thresholdThe threshold is chosen with the customer on a labelled question set
  5. 5Answer or hand offAt or above the threshold the bot answers; below it, staff take over

The rules layer filters out the questions a confidence score cannot catch; a threshold measured on real data handles the rest.

Graphic: FDE Times

At the third demo on site, the head of customer service asks the question every FDE deploying an LLM eventually hears: “How does this bot know when it doesn’t know, so it can pass the case to my staff?” Answering “the model will tell you when it isn’t sure” sounds reasonable in a meeting room.

It will not survive the first week in production.

The reason is that a model’s stated confidence cannot be trusted. In 2023, Xiong and colleagues tried a range of prompts to get LLMs to verbalise their confidence, and found that the models were frequently overconfident. A model saying “I am 95% sure” does not mean that 95 out of 100 such answers are correct.

For an FDE, this skill decides whether a demo becomes a system the customer is willing to run in production. The handoff threshold determines how much work is automated and how many wrong answers slip through. It has to be measured, not guessed.

Asking the model a true/false question is more reliable

There is a better approach than asking “how sure are you”. LearnPrompting calls it self-evaluation: using an LLM to check its own output or the output of another LLM. The simplest version is to ask the model a question, then ask it to check the answer it just gave. The loop can be repeated several times.

In 2022, Kadavath and colleagues at Anthropic studied a measurable form of self-grading called P(True). The model produces an answer, then estimates the probability that the answer is correct.

The key is the question format. The research found that large models are fairly well calibrated, with confidence close to actual accuracy, across many multiple-choice and true/false tasks, provided the format is right.

So instead of reading a number the model writes out, turn the self-grading step into a true/false question and read the probability of the “True” token from the logprobs. The same paper also proposes P(IK): training the model to predict whether it knows the answer before it produces any specific answer.

For most FDE projects, P(True) via prompting is the more practical starting point.

One example, from prompt to threshold

Imagine deploying an assistant that answers questions about the refund policy of an e-commerce company. The pipeline calls the model twice: once to answer, once to grade itself.

JUDGE_PROMPT = """Câu hỏi: {q}
Tài liệu: {ctx}
Câu trả lời đề xuất: {a}
Câu trả lời đề xuất có đúng và được tài liệu hỗ trợ không?
(A) Đúng
(B) Sai
Chỉ trả lời A hoặc B."""

def p_true(q, ctx, a):
    out = llm.complete(JUDGE_PROMPT.format(q=q, ctx=ctx, a=a),
                       max_tokens=1, logprobs=True)
    pa, pb = out.prob("A"), out.prob("B")
    return pa / (pa + pb)

def route(q, ctx, threshold):
    a = llm.answer(q, ctx)
    score = p_true(q, ctx, a)
    return ("auto", a) if score >= threshold else ("human", a)

(The prompt is in Vietnamese: it gives the question, the documents and the proposed answer, asks whether the proposed answer is correct and supported by the documents, and offers (A) True or (B) False, answering only A or B.)

The p_true function only works if the API returns per-token probabilities. Many APIs do not. If yours does not, do not fall back to asking the model to write out a percentage.

There are two alternatives: call the grader several times with a temperature above 0 and use the share of “A” responses as the score, or use a self-hosted open-source model that returns logprobs for the grading step alone.

Before trusting p_true, check whether the grader itself is biased. LearnPrompting describes a technique called contextual calibration: given a content-free input, the model should assign roughly 0.5 probability to each label.

If the result is far from that, the model favours one label. Note that this corrects bias between labels; it is not the same as the model verbalising its confidence.

Applied to the example above: put “N/A” into all three slots, question, documents and answer. If the grader still gives “A” around 0.7, it leans towards “True”, and every score you read afterwards is inflated. In that case, fix the prompt, swap the order of the two labels, or subtract the bias when choosing the threshold.

The threshold is a number measured on the customer’s data

A 2024 survey of abstention (a model declining to answer) describes the common approach: estimate a confidence score, and have the model abstain when the score falls below a threshold. The hard part is choosing the threshold, and that cannot be settled in a meeting room.

Continuing the hypothetical example: you take 200 real questions from the ticket history and ask staff to label each of the bot’s answers as correct or incorrect.

Suppose that at a threshold of 0.8 the bot answers 150 questions on its own, gets 140 right (about 93%), and hands 50 to people. Raise the threshold to 0.9 and the bot answers only 110 on its own but gets 107 right (about 97%), handing 90 to people.

No threshold is technically “correct”. The question to put to the head of department is: are 10 wrong answers out of 150 acceptable, or would they rather have staff handle 40 more questions to bring the errors down to 3?

Some questions a confidence score will never catch

Even a well-calibrated grader has limits. The abstention survey looks at declining to answer from three angles: the question, the model, and human values. The authors point out that value-related issues, and whether a question has an answer at all, are hard to capture through model confidence.

In the refund example, when a customer writes “I’m going to sue you”, the problem is not that the model lacks knowledge. The model can answer with great confidence and still handle the situation badly.

Similarly, a question about an order that does not appear in the documents has no answer, yet the model may still invent one that sounds entirely certain.

The system therefore needs a rules layer that runs before the model. Questions containing legal keywords, complaints or sensitive personal data, or questions for which retrieval finds no documents, go straight to a person. The confidence threshold handles only what remains.

Steps on site, and common mistakes

The order of work is as follows. First, sit down with the customer to list the types of question that must always go to a person, and write them as rules. Next, build the true/false self-grader and check its bias with an empty input.

Then collect a labelled set of real questions, build a table of accuracy and automation rate at each threshold, and let the customer pick the level they can accept. Using the hypothetical 200 questions above, with a 0.7 row added so the customer can also see what loosening looks like, the table you bring to the meeting might look like this:

Threshold Bot answers alone Correct Accuracy Wrong answers slipping through Handed to people
0.7 175 158 about 90% 17 25
0.8 150 140 about 93% 10 50
0.9 110 107 about 97% 3 90

The “wrong answers slipping through” and “handed to people” columns are the two things the customer actually pays for: one is risk to their customers, the other is staff hours. Reading across the rows, each increase in the threshold trades a few wrong answers for a few dozen tickets for people to handle.

The most common mistake is trusting the number the model reports about itself, exactly what Xiong’s research warned against. The second is choosing a threshold by feel rather than from labelled data.

The third is using a single threshold for every type of question. In practice, questions about refund amounts may need a higher threshold than questions about delivery times. Many teams also forget to re-measure after changing the model or editing the prompt, which is precisely when the score distribution shifts.

If you are preparing to apply for FDE roles, put exactly this calculation on your CV or portfolio, for example: “designed a human-handoff mechanism based on self-evaluation, chose the threshold on N labelled questions, achieving X% automation at Y% accuracy”.

The topic is also worth preparing before interviews: practise explaining your own threshold table in a few minutes, from how the labelling was done to why you chose a particular row.

The next time a customer asks whether the bot knows when it doesn’t know, do not answer with a promise. Bring the threshold table and let them choose the row that fits.

6 sources
Read next on the roadmap · Stage 6: MeasurementData lineage: tracing a wrong prediction back, step by stepA client sends a screenshot of a wrong number. If you trace that number back to the job and the run that produced it, you can look for the root cause methodically instead of guessing at the model.