# Reasoning models: when it is worth making a client pay more and wait longer

> Reasoning tokens never appear in the response, but they still take up context and still cost money. An FDE who turns reasoning on without measuring it is having the client pay for something nobody has checked.

Original: https://fdetimes.net/en/analysis/when-reasoning-models-worth-the-cost/

OpenAI's API documentation contains a detail that is easy to skim past: reasoning tokens are not visible through the API, yet they still take up space in the context window. Anthropic, for its part, says plainly that thinking tokens are billed as output tokens. In other words, clients pay for a part of the work they will never get to read.

That is why "should we turn on reasoning?" is rarely a purely technical question. It touches the three things clients care about most: the bill, latency, and how far the answer can be trusted. On a client site, the person expected to answer it is usually you.

This article reaches a specific conclusion. Reasoning is a knob to be measured separately for each workload, not an upgrade switch flipped once for everything. And the answer that persuades a client most does not come from a benchmark leaderboard. It comes from eval results run on the client's own data.

## The benefit is real, but it comes with conditions

The chain-of-thought paper published in 2022 showed that generating intermediate reasoning steps significantly improves large language models' ability to solve complex problems.

IBM explains the mechanism simply: the problem is broken into smaller logical steps, and the user can see the path the model took to its answer.

The original paper also carries a condition that few people mention: this ability emerges naturally in sufficiently large models. So if a client insists on running a small model on-premise for security reasons, do not assume that adding reasoning steps will rescue quality.

The second condition is more worrying. IBM warns that models can produce reasoning chains that sound plausible but are wrong. When clients see a long, coherent argument, they tend to trust the answer more than it deserves. The transparency CoT provides can turn into false confidence if you have no eval set to check it against.

## The bill sits in the part nobody reads

OpenAI describes reasoning models as adding a third type of token alongside input and output. IBM puts it more bluntly: generating and processing many reasoning steps requires more compute and time, which makes adoption more expensive for businesses.

Take Claude as an example, since Anthropic states that thinking tokens are billed as output tokens. Imagine a request with numbers chosen to be easy to add up: 500 input tokens and a 200-token answer. If the model thinks for another 3,000 tokens before answering, the billed output is 3,200 tokens, not 200, even though the user sees only those 200.

Context is eroded in the same way. For an agent that has to fit a long contract, conversation history and tool results into one context window, every thousand tokens of thinking is a thousand tokens of documents that no longer fit. Anthropic returns a figure in the response showing how many billed output tokens were internal reasoning.

The first thing to do on any project that uses reasoning is to log this field from day one.

## A knob, not a lifebuoy

Both providers give you a knob, under different names. OpenAI has a reasoning effort parameter, and its documentation advises treating it as a tuning knob, not as the primary way to recover quality. Low effort is faster and uses fewer tokens, which suits latency-sensitive use cases.

Anthropic measures a thinking budget in tokens and says almost the same thing, with more detail on diminishing returns. A higher budget allows more thorough reasoning, but the gains taper off depending on the task, and the price is higher latency.

Anthropic's concrete advice: for simple tasks, start near the 1,024-token minimum and increase gradually; for complex tasks, start at 16k or above.

**Key point:** If an answer is wrong because data is missing or the prompt is vague, raising the thinking budget only makes the model wrong for longer and at greater cost.

For an FDE, the phrase "not the primary way to recover quality" is an operating instruction. When a client complains that answers are poor, the sensible order of checks is: is the context sufficient, is the prompt clear, does the eval reflect what the client actually needs? Only after that does turning the effort knob come into it.

## Let the model decide, or keep the decision?

Newer Claude models replace a fixed budget with adaptive thinking: the model decides whether to think, and how much, for each request, and at low effort it can skip thinking altogether on easy inputs. That sounds like a fix for every cost problem, but Anthropic keeps an important exception.

According to its documentation, manual mode remains useful when a workload needs predictable latency or precise control over thinking costs. On client sites those two conditions come up more often than you might expect: a system with a response-time SLA, or a finance department that wants to know monthly costs in advance.

Ordered from cheapest to most expensive, the options give an FDE a five-step ladder to climb in turn, with the final step being a choice between adaptive thinking and a manual budget. At steps 3 and 4, "effort" is OpenAI's parameter, while the 1,024- and 16k-token budgets are Anthropic's:

| Step | Extra cost | Latency | Suits |
|---|---|---|---|
| 1. Plain prompt | None | Lowest | Simple tasks that already pass eval |
| 2. Zero-shot CoT ("Let's think step by step") | Extra visible output tokens | Slight increase | A cheap baseline before switching to a reasoning model |
| 3. Low effort (OpenAI) / budget near 1,024 (Claude) | Few reasoning tokens | Moderate increase | Latency-sensitive use cases needing a little more reasoning |
| 4. High effort (OpenAI) / budget from 16k (Claude) | Many reasoning tokens (billed as output on Claude) | High | Complex tasks with a proven benefit on eval |
| 5a. Adaptive thinking (Claude) | Varies per request | Hard to predict | A mix of easy and hard inputs, no tight SLA |
| 5b. Manual budget (Claude) | Clear ceiling | Predictable | An SLA, or a client who needs cost control |

## A morning on a client site

Picture a logistics company with two workloads. The first is a chatbot that tells customers where their orders are. The second is a tool that reconciles carrier invoices against contracts full of surcharge clauses, run as a nightly batch.

The chatbot is a lookup task, and the user is waiting at the screen. You start at the bottom of the ladder, and if quality is good enough you stop there. If more reasoning is needed, low effort or a small manual budget keeps latency within what the SLA allows.

The reconciliation tool is the opposite. It is the kind of multi-step problem where CoT research shows intermediate reasoning helps, and because it runs overnight, latency matters less. You can start with a large budget, but you should still run evals at several budget levels to find the point where quality starts to plateau.

Diminishing returns mean there is always a threshold beyond which the client is simply paying more.

For both workloads, what you bring into the client meeting is a three-column table: eval quality, latency, and reasoning tokens for each configuration. Clients do not need to understand chain-of-thought. They need to see which step is worth the money, and why.

One more risk should be raised with the client up front: a plausible-sounding reasoning chain can still be wrong. If the client plans to show the reasoning on screen to reconciliation staff, make clear that it is a suggestion to be checked, not evidence.

## What developers should practise from this week

The skill here is not memorising API parameters, since the parameter names have already shifted from fixed budgets to adaptive and may change again. What counts is the reflex to measure: build a small eval set, run the same task across several effort levels, read the usage data, and make a recommendation backed by numbers.

When reading job descriptions for FDE or AI engineer roles, look for phrases such as "latency budget", "cost optimization", "evals" or "production LLM". They signal a company that needs someone who understands trade-offs, not just someone who can call an API.

On a CV, a line describing how you compared several reasoning configurations on evals and chose one to fit a client's SLA is more persuasive than "experienced with GPT and Claude".

If you do not yet have a real project, take a familiar problem such as bank statement reconciliation or contract clause extraction, run it through each step of the ladder above, and write up the results in a short post.

Attach the quality, latency and token table you measured yourself to your portfolio when applying, because it shows you make decisions from data rather than instinct.

Models will keep getting better at thinking and will decide more for themselves. But whether a client should pay for that thinking remains a decision for whoever stands between the model and the bill, and that position belongs to the FDE.

**Try this week:**

- Pick a real task you are working on, build a small eval set, and run it with a plain prompt, with "Let's think step by step", and with reasoning at two effort levels. Record quality, latency and reasoning tokens in a table.
- Add the field showing how many output tokens were internal reasoning to your project's logging, so the cost is visible every time you change configuration.
- Write a short note for the client (or for your portfolio) explaining why each workload uses adaptive thinking or a manual budget, based on the numbers you just measured.

## Sources

- [Chain-of-Thought Prompting Elicits Reasoning in Large Language Models](https://arxiv.org/abs/2201.11903)

- [What is chain of thought (CoT) prompting? (IBM)](https://www.ibm.com/think/topics/chain-of-thoughts)

- [Chain-of-Thought Prompting - Prompting Guide](https://www.promptingguide.ai/techniques/cot)

- [Reasoning models - OpenAI API docs](https://developers.openai.com/api/docs/guides/reasoning)

- [Extended thinking - Claude API docs](https://platform.claude.com/docs/en/build-with-claude/extended-thinking)
