FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Tools

Langfuse: answering the three hardest questions after an LLM deployment

Clients will ask why the system answered as it did, what it costs and whether it is getting better. One open-source tool puts all three answers on a single platform.

In brief

  • Langfuse traces the full lifecycle of a request, including retrieval and embedding, so FDEs can find where a failure actually sits instead of blaming the model.
  • Cost can be sent by the app or inferred by Langfuse from a price table; values sent by the app always take precedence.
  • After the ClickHouse deal (16 January 2026), core features remain MIT-licensed and self-hostable, the precondition for installing it in client infrastructure.
ShareLinkedInFacebookX
GraphicThe quality loop with Langfuse on a client site
  1. 1Trace every stepPrompt, response, tokens, latency, plus retrieval, embedding and API calls
  2. 2Score productionLLM-as-a-judge automatically scores live traces
  3. 3Dataset + experimentCollect failed traces as test cases, run experiments in CI to catch regressions
  4. 4Deploy prompts by labelPrompts are versioned; switch versions by moving the label

Failed traces become test cases, test cases lock out regressions in CI, and new prompts are deployed under control.

Graphic: FDE Times

Once an LLM application is in production, clients tend to ask three questions: why did it answer like that, how much does it cost each month, and is this week’s version better than last week’s?

Answering them on instinct is the fastest way to lose a client’s trust. Langfuse is designed to answer them with data: tracing for “why”, cost tracking for “how much”, and evals for “is it getting better”, all on one platform.

For an FDE, one condition comes before any discussion of features: the observability tool you install in a client’s infrastructure has to run inside their internal network. Langfuse describes itself as open, self-hostable and extensible, which are exactly the three properties the client’s security team will ask about first.

A trace records more than the LLM call

According to the Langfuse documentation, application tracing captures the full lifecycle of a request as it moves through the system: prompt, response, tokens used and latency. More importantly, a trace also covers the steps that are not LLM calls: retrieval, embedding and API calls. In a RAG application, that is where failures tend to hide.

Imagine you are deploying an internal policy assistant for a bank. Users complain that answers about credit limits are wrong. If you look only at the prompt and response, you will assume the model is making things up and start editing the prompt.

Open the trace in Langfuse and the story may be quite different: the embedding step ran normally, retrieval returned five passages, but none of them contained the new credit-limit table. The problem is a stale index, not the prompt. You give the client an answer in ten minutes instead of losing a day to pointless prompt experiments.

That only works if the SDK is wired into production, and this is where the client’s platform team will worry about performance. The Langfuse documentation is explicit: trace events are queued locally and flushed in batches, so application response times are not affected. That is the sentence you need ready for the architecture review.

The number you send beats the number Langfuse infers

“How much does it cost?” sounds simple, but there are two ways to answer it, and Langfuse supports both. The application can send usage and cost directly; if it does not, Langfuse infers them from model definitions that carry prices per usage type.

The deciding rule: when both exist, the value that was sent takes precedence over the inferred one.

Send cost from the app

  • You control the number, including negotiated pricing or fine-tuned models
  • Always takes precedence in a conflict
  • Requires app-side code to calculate and attach usage

Let Langfuse infer it

  • No app changes needed, only model definitions with prices
  • The number is only as accurate as the price table you load
  • Suits widely used models with public pricing

The advice: if the client uses a model under a private contract, send cost from the app from day one, because the inferred figure will be wrong and you will lose credibility when finance reconciles it against the invoice. If they use a widely available model, check the model definitions before showing the dashboard to anyone.

Score production, then lock out regressions in CI

The third question is the hardest: is the system getting better? Langfuse can automatically score live production traces using LLM-as-a-judge, giving you a continuous measure instead of waiting for users to complain. But a score only matters if it leads to action.

A sensible loop on a client site looks like this. Any trace that scores poorly goes into a dataset of test cases. Every time you change a prompt or the retrieval strategy, you run an experiment against that dataset, and according to the documentation, experiments can run directly in CI to catch regressions before code reaches production.

To make those changes safe, Langfuse also offers prompt management: prompts are versioned and deployed by label.

“Fixing the prompt” then stops being a risky operation and becomes a change with a version number and before-and-after scores. Because production points at a label rather than a fixed prompt, rolling back simply means reassigning the label to the previous version.

Limits to state upfront

Langfuse does not do the judging for you. The judge is an LLM, so its scores need a human reading samples regularly to check them. Treat them as a warning light, not a verdict. Inferred costs also depend entirely on the price table you configure.

On licensing, there is one milestone to know. On 16 January 2026, Langfuse announced on its blog that ClickHouse had acquired the company. Langfuse said it remains committed to open source and self-hosting, with no licence changes planned; ClickHouse stated that core features will continue under the MIT licence.

That commitment applies to “core features”, so read carefully what counts as core before making promises to a client. The tool you install in their infrastructure has to outlive your contract.

On operations, ClickHouse noted that Langfuse was already built on ClickHouse. If you self-host, be ready to run one more storage system inside the client’s infrastructure, and ask their team clearly who will own it.

What to learn first

Start with the concept of a trace: a request is a chain of steps, each with its own input, output, duration and cost. Wire the SDK into a small RAG app of your own, run a few dozen questions and practise reading traces until you can spot a retrieval failure faster than a prompt failure. Only then move on to datasets, experiments and CI.

On your CV, do not write “familiar with Langfuse”. Write that you deployed observability for an LLM application, with full step-level tracing, costs reconciled against invoices and a test suite running in CI. An FDE hiring manager reads that line and sees someone who can stand in front of a client.

When everything works, nobody mentions Langfuse. But on the day the system gets an important number wrong, the person holding the trace is the one who keeps the contract.

6 sources
Read next on the roadmap · Stage 6: MeasurementEscaping the Build Trap: the book that teaches FDEs to measure their work by client outcomes, not feature countsMelissa Perri wrote it for product managers, but the readers who need her 2018 book most are engineers who sit in a client's office and decide every week what to build next.