FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Cutting LLM latency and cost: four levers and the order to pull them

When a system is slow and expensive, the usual first move is a cheaper model. That lever belongs near the end, and only once an eval is in place.

In brief

  • Get the prompt working correctly before you optimise, because optimising early can hide the best quality the task can reach.
  • Trimming output tokens and caching the fixed part of the prompt are cheap, low-risk levers. Pull them first.
  • Move to a smaller model only once you have an eval, and send work to the Batch API only when nobody is waiting for the result.
ShareLinkedInFacebookX

Picture yourself in week three at a client. The internal policy assistant you built is answering correctly and the business team is happy. Then the head of operations sends two lines: staff are complaining about long waits, and this month’s API bill is well over budget.

Many engineers react by switching straight to a smaller model. That may turn out to be right, but it usually comes in the wrong order. There are four levers for cutting latency and cost. Each has its own price, and a good FDE knows which one to pull first.

After the demo stage, this is the skill clients notice most. A system that is correct but slow and expensive will struggle to get through budget review, however good its output.

Why get it right before making it fast?

Anthropic’s documentation puts it plainly: design a prompt that works well without model or prompt constraints first, and only then apply latency-reduction strategies. Optimising too early can hide the best quality the task can reach.

So before optimising, you need a small eval set, such as a few dozen real client questions with the answers you expect. Without one, you cannot tell whether a change that speeds the system up is quietly making it wrong.

Four levers, each moving something different

The first lever is fewer tokens. According to Anthropic, the fewer tokens a model has to process and generate, the faster it responds. OpenAI offers a rule of thumb worth remembering: cutting output tokens by 50% can cut latency by roughly 50%. In practice, ask for shorter, structured answers and set max_tokens as a hard limit.

The second lever is prompt caching. When the start of a prompt repeats across calls, the provider stores the processed portion, which reduces both time and cost. At Anthropic, tokens read from cache cost 0.1 times the base input price. OpenAI says cached input tokens are up to 95% cheaper.

The third lever is choosing the right model. Anthropic calls this one of the most direct ways to reduce latency, and OpenAI notes that smaller models are usually both faster and cheaper. The catch is that this is also the lever most likely to hurt quality if you switch without checking.

The fourth lever is OpenAI’s Batch API. It cuts cost by 50%, and in return each batch completes within 24 hours, often sooner. It suits only work that does not need an immediate response, and it has rate limits separate from the per-model limits.

Streaming is not one of the four levers, because it neither shortens total time nor lowers cost. It works on perception: users see text appear in real time, so the application feels much faster.

A worked example: the policy assistant

Back to the assistant from the opening. Suppose each request contains 10,000 fixed tokens (system instructions plus policy documents) and 200 tokens of question. At peak, 20 questions arrive within the same 5-minute window.

Without caching, you pay 20 × 10,200 = 204,000 units of input price. With caching on Anthropic, the first call writes 10,000 tokens to cache at 1.25 times the price, or 12,500 units.

The next nineteen calls read from cache at 0.1 times, or 19 × 1,000 = 19,000 units. Add 4,000 units for the questions and the total is 35,500 units, less than a fifth of the original figure.

That figure depends on getting the prompt order right. OpenAI requires an exact match on the whole prefix before a cache can be reused, and advises putting stable instructions and shared reference material first. The layout should look like this:

[1] System instructions  (fixed)
[2] Policy documents     (fixed)
---- end of cached part ----
[3] User question        (varies)

Next comes the output. If the average answer runs to 400 tokens but users only need the conclusion plus the cited clause, asking for a compact format of about 200 tokens could nearly halve generation time, by OpenAI’s rule of thumb.

The end-of-day summary of all questions for the business team is a different matter: nobody is waiting for it. Work like this is a good fit for the Batch API. It saves half the cost and does not touch the rate limits of the live assistant.

The order to use them in

The ordering principle is simple: levers that carry little risk to quality come first, and levers that need an eval to justify them come later. The table has six steps because, beyond the four levers, there is a baseline step at the start, and streaming sits in the middle as a near-zero-risk perception lever.

Step What to do What it reduces Condition
1 Get the prompt right, build an eval Nothing yet, it sets the baseline Real client questions available
2 Cut output tokens, set max_tokens Latency and cost Short answers still cover what is needed
3 Cache the fixed part of the prompt Input cost and processing time Stable prefix, steady traffic
4 Turn on streaming (perception lever) Perceived latency, not cost Someone is waiting on screen
5 Switch to a smaller model Latency and cost Eval holds
6 Send non-urgent work to batch Cost A wait of up to 24 hours is acceptable

Mistakes that make optimisation backfire

The most common mistake is breaking the cache without realising. Insert the current date and time, or the user’s name, into the first line of the prompt and the prefix no longer matches, so every request is processed from scratch.

The second is turning on caching for sparse traffic. Anthropic’s default TTL is 5 minutes. If requests arrive further apart than that, every call pays the 1.25x write price and never gets a read, which costs more than not caching at all. The 1-hour TTL costs 2x to write and is only worth it when you are sure of many reads.

The third is treating streaming as a way to cut total time, then using it in a backend pipeline that waits for the full result. Nobody is watching text appear there, so streaming adds nothing.

The fourth is switching models on the strength of a few manual tries. A smaller model may handle the ten questions you thought up and still fail on exactly the kind of question the client asks most. Nor should you push a task with a user waiting into batch just because it costs half as much.

OpenAI also makes a point that is easy to forget: do not default to an LLM. If a step only extracts a contract number in a fixed format, a regular expression is faster and costs nothing.

Clients rarely remember which model you used. They remember that the system answered faster and the bill went down without quality dropping. So when you put this work on your CV, state the result you measured and how you verified it, along the lines of “cut input cost by reordering the prompt to hit the cache, verified with an eval”.

5 sources
Read next on the roadmap · Stage 5: DeploymentDesign error responses with RFC 9457 so customer ops teams can fix incidents themselvesIf your API returns only a 500 with "Something went wrong", every incident on the customer's side becomes a ticket for you. A few JSON fields in the right places can change that.