FDE PulseFDE jobs open 448New in 7 days 29Companies hiring 52Remote-friendly 25%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Analysis

Reasoning models: when it is worth making a client pay more and wait longer

Reasoning tokens never appear in the response, but they still take up context and still cost money. An FDE who turns reasoning on without measuring it is having the client pay for something nobody has checked.

Reasoning models: when it is worth making a client pay more and wait longer
Photo: Taylor Vick / Unsplash

In brief

  • Reasoning helps models do better on complex problems, but a reasoning chain that sounds plausible can still be wrong.
  • Reasoning tokens are not returned by the API but still occupy context. On Claude they are billed as output tokens, so log them separately.
  • OpenAI and Anthropic both treat effort/budget as a tuning knob with diminishing returns, not as a way to rescue quality.
ShareLinkedInFacebookX
Two horizontal bars compare what the user sees with what is billed: the user sees a 200-token reply, but the billed output is 3,200 tokens, of which 3,000 are hidden thinking tokens that still occupy the context window. Below is a five-step ladder of rising cost: plain prompt, zero-shot CoT, low effort or a budget near 1,024, high effort or a budget from 16k, and adaptive or manual budget.
With Claude, thinking tokens are billed as output: in this example, a 200-token reply is charged as 3,200 tokens. Each step up the reasoning ladder must prove it is worth the money on evals.

OpenAI’s API documentation contains a detail that is easy to skim past: reasoning tokens are not visible through the API, yet they still take up space in the context window. Anthropic, for its part, says plainly that thinking tokens are billed as output tokens. In other words, clients pay for a part of the work they will never get to read.

That is why “should we turn on reasoning?” is rarely a purely technical question. It touches the three things clients care about most: the bill, latency, and how far the answer can be trusted. On a client site, the person expected to answer it is usually you.

This article reaches a specific conclusion. Reasoning is a knob to be measured separately for each workload, not an upgrade switch flipped once for everything. And the answer that persuades a client most does not come from a benchmark leaderboard. It comes from eval results run on the client’s own data.

The benefit is real, but it comes with conditions

The chain-of-thought paper published in 2022 showed that generating intermediate reasoning steps significantly improves large language models’ ability to solve complex problems.

IBM explains the mechanism simply: the problem is broken into smaller logical steps, and the user can see the path the model took to its answer.

The original paper also carries a condition that few people mention: this ability emerges naturally in sufficiently large models. So if a client insists on running a small model on-premise for security reasons, do not assume that adding reasoning steps will rescue quality.

The second condition is more worrying. IBM warns that models can produce reasoning chains that sound plausible but are wrong. When clients see a long, coherent argument, they tend to trust the answer more than it deserves. The transparency CoT provides can turn into false confidence if you have no eval set to check it against.

The bill sits in the part nobody reads

OpenAI describes reasoning models as adding a third type of token alongside input and output. IBM puts it more bluntly: generating and processing many reasoning steps requires more compute and time, which makes adoption more expensive for businesses.

Take Claude as an example, since Anthropic states that thinking tokens are billed as output tokens. Imagine a request with numbers chosen to be easy to add up: 500 input tokens and a 200-token answer. If the model thinks for another 3,000 tokens before answering, the billed output is 3,200 tokens, not 200, even though the user sees only those 200.

Context is eroded in the same way. For an agent that has to fit a long contract, conversation history and tool results into one context window, every thousand tokens of thinking is a thousand tokens of documents that no longer fit. Anthropic returns a figure in the response showing how many billed output tokens were internal reasoning.

The first thing to do on any project that uses reasoning is to log this field from day one.

A knob, not a lifebuoy

Both providers give you a knob, under different names. OpenAI has a reasoning effort parameter, and its documentation advises treating it as a tuning knob, not as the primary way to recover quality. Low effort is faster and uses fewer tokens, which suits latency-sensitive use cases.

Anthropic measures a thinking budget in tokens and says almost the same thing, with more detail on diminishing returns. A higher budget allows more thorough reasoning, but the gains taper off depending on the task, and the price is higher latency.

Anthropic’s concrete advice: for simple tasks, start near the 1,024-token minimum and increase gradually; for complex tasks, start at 16k or above.

For an FDE, the phrase “not the primary way to recover quality” is an operating instruction. When a client complains that answers are poor, the sensible order of checks is: is the context sufficient, is the prompt clear, does the eval reflect what the client actually needs? Only after that does turning the effort knob come into it.

Let the model decide, or keep the decision?

Newer Claude models replace a fixed budget with adaptive thinking: the model decides whether to think, and how much, for each request, and at low effort it can skip thinking altogether on easy inputs. That sounds like a fix for every cost problem, but Anthropic keeps an important exception.

According to its documentation, manual mode remains useful when a workload needs predictable latency or precise control over thinking costs. On client sites those two conditions come up more often than you might expect: a system with a response-time SLA, or a finance department that wants to know monthly costs in advance.

Ordered from cheapest to most expensive, the options give an FDE a five-step ladder to climb in turn, with the final step being a choice between adaptive thinking and a manual budget. At steps 3 and 4, “effort” is OpenAI’s parameter, while the 1,024- and 16k-token budgets are Anthropic’s:

Step Extra cost Latency Suits
1. Plain prompt None Lowest Simple tasks that already pass eval
2. Zero-shot CoT (“Let’s think step by step”) Extra visible output tokens Slight increase A cheap baseline before switching to a reasoning model
3. Low effort (OpenAI) / budget near 1,024 (Claude) Few reasoning tokens Moderate increase Latency-sensitive use cases needing a little more reasoning
4. High effort (OpenAI) / budget from 16k (Claude) Many reasoning tokens (billed as output on Claude) High Complex tasks with a proven benefit on eval
5a. Adaptive thinking (Claude) Varies per request Hard to predict A mix of easy and hard inputs, no tight SLA
5b. Manual budget (Claude) Clear ceiling Predictable An SLA, or a client who needs cost control

A morning on a client site

Picture a logistics company with two workloads. The first is a chatbot that tells customers where their orders are. The second is a tool that reconciles carrier invoices against contracts full of surcharge clauses, run as a nightly batch.

The chatbot is a lookup task, and the user is waiting at the screen. You start at the bottom of the ladder, and if quality is good enough you stop there. If more reasoning is needed, low effort or a small manual budget keeps latency within what the SLA allows.

The reconciliation tool is the opposite. It is the kind of multi-step problem where CoT research shows intermediate reasoning helps, and because it runs overnight, latency matters less. You can start with a large budget, but you should still run evals at several budget levels to find the point where quality starts to plateau.

Diminishing returns mean there is always a threshold beyond which the client is simply paying more.

For both workloads, what you bring into the client meeting is a three-column table: eval quality, latency, and reasoning tokens for each configuration. Clients do not need to understand chain-of-thought. They need to see which step is worth the money, and why.

One more risk should be raised with the client up front: a plausible-sounding reasoning chain can still be wrong. If the client plans to show the reasoning on screen to reconciliation staff, make clear that it is a suggestion to be checked, not evidence.

What developers should practise from this week

The skill here is not memorising API parameters, since the parameter names have already shifted from fixed budgets to adaptive and may change again. What counts is the reflex to measure: build a small eval set, run the same task across several effort levels, read the usage data, and make a recommendation backed by numbers.

When reading job descriptions for FDE or AI engineer roles, look for phrases such as “latency budget”, “cost optimization”, “evals” or “production LLM”. They signal a company that needs someone who understands trade-offs, not just someone who can call an API.

On a CV, a line describing how you compared several reasoning configurations on evals and chose one to fit a client’s SLA is more persuasive than “experienced with GPT and Claude”.

If you do not yet have a real project, take a familiar problem such as bank statement reconciliation or contract clause extraction, run it through each step of the ladder above, and write up the results in a short post.

Attach the quality, latency and token table you measured yourself to your portfolio when applying, because it shows you make decisions from data rather than instinct.

Models will keep getting better at thinking and will decide more for themselves. But whether a client should pay for that thinking remains a decision for whoever stands between the model and the bill, and that position belongs to the FDE.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
5 sources
Read next on the roadmap · Stage 3: Applied AIOpenAI, Anthropic, Gemini or open-weight for a client: where the data goes matters more than the leaderboardAt a client company, the "best" model is rarely the one that gets chosen. The chosen model is the one compliance will sign off on, deployed so that switching it means editing a config file.