FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Capstone project: a Q&A app that traces and counts the tokens for every question

Any app can answer questions in a demo. In production, you also have to show how it reached each answer and how many tokens that answer cost.

In brief

  • A full-stack Q&A app has four layers. Traces and token counts must be recorded in the business logic layer, then saved to storage.
  • Tokens map directly to cost, so every request needs a token record carrying its trace_id and question type.
  • OpenTelemetry's GenAI conventions for agents are still in Development status, so pin the spec version before building a dashboard.
ShareLinkedInFacebookX
GraphicOne request through the Q&A app
  1. 1UI receives the questionUser types a question; its type is picked by the user or assigned by a classifier.
  2. 2API forwards the requestEndpoint takes the question, calls business logic, returns the answer with a trace_id.
  3. 3Retrieval logic and LLM callLogs a three-step trace: question received, retrieved chunks, answer.
  4. 4Storage keeps trace and tokensUsage table stores input/output tokens, query_type and timestamp per request.
  5. 5Dashboard computes costSums tokens by request, question type and day, then multiplies by unit price.

Every question must leave a trace and a token count in storage, or the dashboard can't compute cost.

Graphic: FDE Times

Imagine you have just demoed an internal-document Q&A app for a client. The answer is correct. Straight away the head of operations asks two questions: which passage of the documents was that answer based on, and how much will it cost each month if the whole company uses it?

If you cannot answer, what you have built is still only a demo. The project below adds the missing piece: every request leaves a trace and a token count, and a dashboard adds those numbers up into a cost.

If you are aiming for a forward deployed engineer role, this is the kind of exercise that belongs in your portfolio. You do not just make the app work; you prove how it works.

What you will build, and what you need

AWS defines full stack development as building both the frontend and the backend of an application. This Q&A app has all four layers AWS describes. The interface is where the user types a question. The API layer takes interactions from the frontend and passes them down to storage. The business logic layer is the core of the backend. The storage layer manages and stores the application’s data.

You need Python, SQLite (which ships with Python’s standard library), a web framework you already know and access to any LLM. The code below is a simplified version meant to show the structure. The two functions retrieve (retrieve) and call_llm (call the LLM) are where you plug in your own retrieval and LLM SDK.

Step 1: design storage before writing a prompt

Most people building their first project store only the answer. Here, storage needs two tables from the start: one recording each step of a request, and one recording tokens.

CREATE TABLE traces (
  trace_id TEXT, step TEXT, detail TEXT, created_at TEXT
);
CREATE TABLE usage (
  trace_id TEXT, query_type TEXT,
  input_tokens INTEGER, output_tokens INTEGER, created_at TEXT
);

Check: run .schema in sqlite3 and confirm both tables are there. Do not drop the query_type column: in step 4 the dashboard relies on it to break cost down by question type.

So where does the value of query_type come from? The simplest option for this exercise is to let users choose it in the interface, for example a dropdown with “policy lookup” and “contract comparison”. It is easy to build, but the figures are only accurate if users pick correctly.

The second option is to assign it automatically in the business logic layer with a classifier: a few keyword rules, or a short LLM call. If you classify with an LLM, record that call in the trace and the usage table too, because it also consumes tokens. Start with the dropdown, store the real questions, then use that data to test a classifier later.

Step 2: the trace must capture the steps in between

A correct answer did not necessarily go through the right steps. If you store only the final output, you cannot tell which passage the model relied on. IBM lists traces, alongside metrics and logs, among the data to collect from an LLM application; in a Q&A app, the value of a trace lies precisely in these intermediate steps.

JetBrains uses the example of a ReAct loop: when each “thought” is recorded as text, you get a complete trace of the reasoning. That lets you evaluate those steps even when the final answer looks right.

In a Q&A app, “the steps” are the incoming question, the retrieved passages and the answer. Both traces and tokens are recorded in the business logic layer. (In the code, write_trace writes a trace step and handle_question handles a question; the step names mean “question received”, “retrieval” and “answer”.)

# Simplified version: retrieve and call_llm are your own functions
import sqlite3, uuid
from datetime import datetime, timezone

db = sqlite3.connect("qa.db")

def now():
    return datetime.now(timezone.utc).isoformat()

def write_trace(trace_id, step, detail):
    db.execute("INSERT INTO traces VALUES (?,?,?,?)",
               (trace_id, step, detail, now()))

def handle_question(question, query_type):
    trace_id = str(uuid.uuid4())
    write_trace(trace_id, "receive_question", question)
    chunks = retrieve(question)
    write_trace(trace_id, "retrieve", " | ".join(chunks))
    result = call_llm(question, chunks)  # returns the text and token counts
    write_trace(trace_id, "answer", result["text"])
    db.execute("INSERT INTO usage VALUES (?,?,?,?,?)",
               (trace_id, query_type, result["input_tokens"],
                result["output_tokens"], now()))
    db.commit()
    return {"trace_id": trace_id, "answer": result["text"]}

The API endpoint only needs to call handle_question and return JSON to the interface. Remember to return the trace_id as well. When a user reports a wrong answer, you can go straight to the matching trace instead of guessing.

Check: send a question, then run SELECT step FROM traces WHERE trace_id = '...'. The correct result is three rows in order: question received, retrieval, answer.

Step 3: money is counted in tokens

“How much will it cost each month?” can only be answered if you count tokens. IBM recommends tracking the number of tokens processed, especially when tokens are tied to model cost.

On the request side, the prompt’s token count is the main source of cost, according to a guide to OpenTelemetry’s GenAI conventions on the OpenObserve blog. In OpenTelemetry, this figure is the span attribute gen_ai.usage.input_tokens.

So when you move from SQLite to OpenTelemetry, your input_tokens column maps to that attribute. To calculate cost, multiply tokens by the unit price in your model provider’s price list. Keep the unit price in a config file rather than hard-coding it.

Step 4: the dashboard must answer three questions

A single total-cost figure does not tell you where to fix things. Small waste adds up to a lot over time, so JetBrains recommends tracking tokens continuously and breaking cost down by request, by question type and over time. Those three cuts are three queries:

-- By question type
SELECT query_type, COUNT(*) AS request_count,
       AVG(input_tokens) AS avg_input, SUM(input_tokens) AS total_input
FROM usage GROUP BY query_type;

-- By day
SELECT substr(created_at, 1, 10) AS day,
       SUM(input_tokens), SUM(output_tokens)
FROM usage GROUP BY day ORDER BY day;

-- The 10 most expensive requests, with trace_id so you can open the trace
SELECT trace_id, query_type, input_tokens
FROM usage ORDER BY input_tokens DESC LIMIT 10;

The comments read, in order: by question type; by day; the 10 most expensive requests, with trace_id so you can open the trace.

A hypothetical example shows why the breakdown by type matters. Policy lookup questions use 2,000 input tokens on average. Questions comparing two contracts pull in more passages and use 6,000. With 1,000 comparison questions a day, the 4,000-token difference per question alone comes to 4,000,000 tokens a day.

The third query helps you find the cause. Open the trace of the most expensive request and you will see whether retrieval pulled in too many passages or the prompt was duplicated. To cut cost, fix that step.

The dashboard interface can be a single page that runs these three queries and renders tables. For this exercise, that is enough.

Mistakes that make the dashboard wrong

The first mistake is recording only the final answer. The dashboard then has cost figures but cannot explain why costs are high, because the trace is missing the retrieval step.

The second is copying OpenTelemetry attribute names from the web without pinning a version. According to OpenObserve, the conventions for agents and tool orchestration are still in Development status. OpenTelemetry’s official GenAI metrics spec page has also moved to a separate semantic-conventions-genai repository, so record the spec version you use in your README and check current names in the new repository.

The third is confusing evaluation with observability. JetBrains draws the line like this: evaluation determines whether an agent can do the job, while observability determines whether it is doing it. A cost dashboard does not replace a quality test suite. You need both.

What this skill looks like at a client site

When you arrive at a client, the first thing to do is ask which dimension they want to break cost down by: department, document type or user. The answer determines the columns in the usage table, before you write a single line of prompt.

When reading a job description, if it mentions running or monitoring LLM applications, put this project first. On your CV, do not write “built a RAG chatbot”. Write that you built a full-stack Q&A app in which every request has a three-step trace and a cost dashboard by question type, with a link to the repo.

A personal project with all four layers and these two data tables shows you understand that clients are not just buying answers. They also need to know where those answers come from and what they cost.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
5 sources
Read next on the roadmap · Stage 5: DeploymentPrompt injection when your agent can call the client's systems: threat modelling and layered defenceOne hidden line in a ticket is enough for an agent to treat it as a command. What needs careful design is the limit on what the agent can do after it reads that line.