# Langfuse: answering the three hardest questions after an LLM deployment

> Clients will ask why the system answered as it did, what it costs and whether it is getting better. One open-source tool puts all three answers on a single platform.

Bản gốc: https://fdetimes.net/en/tools/langfuse-llm-observability-for-fdes/

Once an LLM application is in production, clients tend to ask three questions: why did it answer like that, how much does it cost each month, and is this week's version better than last week's?

Answering them on instinct is the fastest way to lose a client's trust. Langfuse is designed to answer them with data: tracing for "why", cost tracking for "how much", and evals for "is it getting better", all on one platform.

For an FDE, one condition comes before any discussion of features: the observability tool you install in a client's infrastructure has to run inside their internal network. Langfuse describes itself as open, self-hostable and extensible, which are exactly the three properties the client's security team will ask about first.

## A trace records more than the LLM call

According to the Langfuse documentation, application tracing captures the full lifecycle of a request as it moves through the system: prompt, response, tokens used and latency. More importantly, a trace also covers the steps that are not LLM calls: retrieval, embedding and API calls. In a RAG application, that is where failures tend to hide.

Imagine you are deploying an internal policy assistant for a bank. Users complain that answers about credit limits are wrong. If you look only at the prompt and response, you will assume the model is making things up and start editing the prompt.

Open the trace in Langfuse and the story may be quite different: the embedding step ran normally, retrieval returned five passages, but none of them contained the new credit-limit table. The problem is a stale index, not the prompt. You give the client an answer in ten minutes instead of losing a day to pointless prompt experiments.

That only works if the SDK is wired into production, and this is where the client's platform team will worry about performance. The Langfuse documentation is explicit: trace events are queued locally and flushed in batches, so application response times are not affected. That is the sentence you need ready for the architecture review.

## The number you send beats the number Langfuse infers

"How much does it cost?" sounds simple, but there are two ways to answer it, and Langfuse supports both. The application can send usage and cost directly; if it does not, Langfuse infers them from model definitions that carry prices per usage type.

The deciding rule: when both exist, the value that was sent takes precedence over the inferred one.

The advice: if the client uses a model under a private contract, send cost from the app from day one, because the inferred figure will be wrong and you will lose credibility when finance reconciles it against the invoice. If they use a widely available model, check the model definitions before showing the dashboard to anyone.

## Score production, then lock out regressions in CI

The third question is the hardest: is the system getting better? Langfuse can automatically score live production traces using LLM-as-a-judge, giving you a continuous measure instead of waiting for users to complain. But a score only matters if it leads to action.

A sensible loop on a client site looks like this. Any trace that scores poorly goes into a dataset of test cases. Every time you change a prompt or the retrieval strategy, you run an experiment against that dataset, and according to the documentation, experiments can run directly in CI to catch regressions before code reaches production.

To make those changes safe, Langfuse also offers prompt management: prompts are versioned and deployed by label.

"Fixing the prompt" then stops being a risky operation and becomes a change with a version number and before-and-after scores. Because production points at a label rather than a fixed prompt, rolling back simply means reassigning the label to the previous version.

**Điểm mấu chốt:** For an FDE, Langfuse earns its place not through its dashboards but because every client question (why, how much, is it getting better) has data behind the answer.

## Limits to state upfront

Langfuse does not do the judging for you. The judge is an LLM, so its scores need a human reading samples regularly to check them. Treat them as a warning light, not a verdict. Inferred costs also depend entirely on the price table you configure.

On licensing, there is one milestone to know. On 16 January 2026, Langfuse announced on its blog that ClickHouse had acquired the company. Langfuse said it remains committed to open source and self-hosting, with no licence changes planned; ClickHouse stated that core features will continue under the MIT licence.

That commitment applies to "core features", so read carefully what counts as core before making promises to a client. The tool you install in their infrastructure has to outlive your contract.

On operations, ClickHouse noted that Langfuse was already built on ClickHouse. If you self-host, be ready to run one more storage system inside the client's infrastructure, and ask their team clearly who will own it.

## What to learn first

Start with the concept of a trace: a request is a chain of steps, each with its own input, output, duration and cost. Wire the SDK into a small RAG app of your own, run a few dozen questions and practise reading traces until you can spot a retrieval failure faster than a prompt failure. Only then move on to datasets, experiments and CI.

On your CV, do not write "familiar with Langfuse". Write that you deployed observability for an LLM application, with full step-level tracing, costs reconciled against invoices and a test suite running in CI. An FDE hiring manager reads that line and sees someone who can stand in front of a client.

When everything works, nobody mentions Langfuse. But on the day the system gets an important number wrong, the person holding the trace is the one who keeps the contract.

**Thử ngay tuần này:**

- Wire Langfuse into a small RAG app of your own, run 10 questions and open each trace to see what the retrieval step returned before you look at the model's answer.
- Collect 20 questions the app answers wrongly into a Langfuse dataset, then run an experiment every time you change the prompt to see whether the score drops.

## Nguồn

- [Langfuse Overview](https://langfuse.com/docs)

- [Overview (Langfuse documentation)](https://langfuse.com/docs/observability/overview)

- [Token & Cost Tracking (Langfuse documentation)](https://langfuse.com/docs/observability/features/token-and-cost-tracking)

- [Evaluation Overview (Langfuse docs)](https://langfuse.com/docs/evaluation/overview)

- [Langfuse joins ClickHouse](https://langfuse.com/blog/joining-clickhouse)

- [ClickHouse welcomes Langfuse: The future of open-source LLM observability](https://clickhouse.com/blog/clickhouse-acquires-langfuse-open-source-llm-observability)
