Building an FDE portfolio repo: a simple agent, evals built from real failures, a README and a 3-minute video
A polished demo only proves the agent got it right once. A repo worth sending shows where the agent fails and how often.
In brief
- One deep project with a case study beats many polished demos; pick a real enterprise scenario with dirty data.
- Keep the agent loop simple and readable, invest in tool descriptions, and log every run from the start.
- Build evals from real failures, use cheap assertions first, report pass^k, and read transcripts to check the grader.
- 1Log every runOne JSONL line per run: input, tool calls, final verdict
- 2Read logs, list failuresEach failure links to its log file, e.g. misreading the figure 1.200.000
- 3Turn failures into eval cases20-50 cases, checked with pytest assertions first; read transcripts to audit the grader
- 4Measure pass^kRight 80% of the time per run means pass^3 = 0.8 × 0.8 × 0.8 = 51.2%
- 5Fix the agent, rerunCases still red become the next to-do list, recorded in the README
- ↻ Repeat from step 1
Every failure found in the logs becomes an eval case, and pass^k shows whether the agent can be trusted across many runs.
Graphic: FDE Times
Suppose your agent gives the right answer on 80% of runs. That sounds fine until you work out the probability that it is right three times in a row: 0.8 × 0.8 × 0.8 = 51.2%.
That is why an FDE portfolio repo should not stop at a demo that runs nicely. A demo is a single attempt. A repo worth sending has to show where the agent fails, how often it fails, and how you dealt with each failure. The measurement approach here draws on advice from Anthropic, OpenAI and Hamel Husain on building evals for AI systems.
A guide on the blog of FDE Academy, a training provider, puts it bluntly: one deployed project with a written-up case study carries more weight than eight polished demos.
That is advice, not a law, but the reasoning is easy to follow: a system that works on real data demonstrates more skill than eight prototypes running on clean data.
What follows are five steps for building such a repo. The code in this article is a stripped-down sketch: the function that calls the model is left empty so you can wire in whichever SDK you use.
Step 1: Which problem is “enterprise” enough?
The same FDE Academy guide recommends choosing a realistic enterprise scenario, building only an MVP rather than chasing the full vision, and using dirty data that is real or close to real. That last condition is what separates an FDE portfolio from a course assignment.
Picture a distribution company that receives supplier invoices by email. Some are scanned PDFs printed at an angle; some write the currency as “VND”, others as “đ”; some have tax codes with missing digits. The MVP does one thing: the agent reads the invoice, matches it against the purchase order, and returns “match”, “mismatch” with a reason, or “needs human review”.
invoice-agent/
agent/loop.py # vòng lặp lõi
agent/tools.py # tool + mô tả
data/raw/ # hóa đơn bẩn
logs/ # mỗi lần chạy một file JSONL
evals/cases.jsonl # bộ case
evals/test_basic.py # assertion kiểu pytest
README.md
(The comments read, in order: core loop; tools and descriptions; dirty invoices; one JSONL file per run; case set; pytest-style assertions.)
Check: you can state the MVP scope in one sentence. If that sentence uses “and” more than twice, cut it down.
Step 2: Why must the core loop be readable in five minutes?
In “Building effective agents”, Anthropic advises finding the simplest solution possible and adding complexity only when it is genuinely needed. It also warns that frameworks often add layers of abstraction that obscure the underlying prompts and responses. For a portfolio, that is a serious flaw: a reviewer who opens the repo and cannot find the prompt has no way of judging how you think.
# agent/loop.py — phác thảo rút gọn
def run(invoice_text, po, call_model, tools, log):
messages = [system_prompt(), user_msg(invoice_text, po)]
for step in range(8): # giới hạn số bước
reply = call_model(messages, tools)
log.write(step, messages, reply) # ghi lại mọi thứ
if reply.tool_call is None:
return parse_verdict(reply)
result = tools[reply.tool_call.name](**reply.tool_call.args)
messages += [reply, tool_result(result)]
return {"verdict": "needs_review", "reason": "hết số bước"}
(The comments mark a cap on the number of steps and logging of everything; the fallback reason means “ran out of steps”.)
The part worth investing effort in is the tools. Anthropic calls this the agent-computer interface and recommends designing it carefully: document the tools thoroughly and test them. Applied to this repo, each tool’s docstring should spell out the input format, when it returns empty, and include an example.
def lookup_po(po_number: str) -> dict:
"""Tra đơn đặt hàng theo mã PO.
po_number: dạng 'PO-' + 6 chữ số, ví dụ 'PO-004512'.
Trả về {} nếu không tìm thấy; KHÔNG đoán mã gần đúng."""
(The docstring says: look up a purchase order by PO number; the format is ‘PO-’ plus six digits, for example ‘PO-004512’; return {} if not found, and do NOT guess a near-match.)
Check: hand loop.py to someone who has never seen the repo. After five minutes of reading, they can explain what the agent does.
Step 3: Log before you write evals
OpenAI’s “Evaluation best practices” recommends logging everything during development so the logs can later be mined for eval cases. In “Your AI Product Needs Evals”, Hamel Husain goes further and calls for removing every obstacle to looking at data. If viewing a single run means opening three files and scrolling until your hand aches, you will soon stop looking.
The cheapest approach is one JSONL line per run, containing the input, every tool call and the final verdict. Then write a small script that prints each run to the terminal in readable form.
Run the agent 20-30 times on data/raw/ and read every run. You will very likely find failures you did not anticipate — for instance, whether the agent reads “1.200.000” (Vietnamese notation, with dots as thousands separators) as one million two hundred thousand or as 1.2.
Check: you have a list of real failures that you recorded yourself, each with a path to its log file.
Step 4: Evals start from failures, not theory
Anthropic defines an eval concisely: give an AI system an input, then apply grading logic to its output to measure whether it succeeded. In its view, 20-50 simple tasks drawn from real failures are already a good starting point. The failure list from step 3 is the raw material for the eval suite.
{"id": "c07", "invoice": "data/raw/inv_07.pdf", "po": "PO-004512", "expect": "mismatch", "must_mention": "số lượng"}
Husain suggests that the first layer should be fast, cheap assertions, the kind you already write in pytest. Anything code can check, let code check; ask a model to grade only what code cannot.
# evals/test_basic.py — phác thảo
@pytest.mark.parametrize("case", load_cases("evals/cases.jsonl"))
def test_verdict(case):
out = run_case(case)
assert out["verdict"] == case["expect"]
assert case["must_mention"] in out["reason"]
pytest evals/ -q
According to Anthropic, agent evals usually combine three kinds of grader: code-based, model-based and human. For a portfolio repo, the code layer needs to be solid. Add a model-based grader to judge whether the agent’s stated reason is sound, and state clearly in the README which parts you graded by hand.
OpenAI calls this approach eval-driven development: evaluate early, evaluate often, and write tests for each stage.
The number to put at the top of the README is pass^k. Anthropic distinguishes pass@k from pass^k: pass^k is the probability that all k attempts succeed. The calculation at the start of this article is exactly the pass^3 of an agent that is right 80% of the time per run, assuming runs are independent: 51.2%.
Picture an accountant who uses that agent three times a day. The chance that all three uses on a given day are correct is only just over half, so roughly every other day that person will hit an error.
Be careful reading this number when the case set is still small. With 20 cases, a single case flipping its result moves pass^k by 5 percentage points, and the smaller k is, the more easily the luck of one run flatters the figures.
In practice, run each case at least three times, rerun the whole suite once more to see whether the number jumps around, and always report the number of cases and the value of k alongside the result.
The final step of evaluation is the one most often skipped. Anthropic stresses that you cannot know whether a grader is scoring correctly without reading the transcripts and scores of many attempts. Pick ten runs graded “pass” at random and read them again. An overly lenient grader is more dangerous than a weak agent.
Check: pytest runs the full set of cases, but they do not all need to pass. Because the cases come from real failures, some staying red is normal, and those red cases are your to-do list. The README should state how many cases pass out of the total, with a table of pass^3 by failure category.
Step 5: A README and video for people who won’t run the code
GitHub Docs lists the questions a README should answer, starting with what the project does. For an FDE repo, the README should read like a short case study.
It should cover the hypothetical customer’s problem, the MVP scope and what you deliberately left out, the eval results with pass^k, three failures that remain, and how you would address them given two more weeks.
The “remaining failures” section deserves care. It shows the reader that you can measure your system and are candid about its limits.
The 3-minute video should be made for a manager, not an engineer. One way to split it: 30 seconds stating the problem, 90 seconds running the agent on a real dirty invoice from data/raw/, 45 seconds opening the eval results and reading out the pass^k figure, and a final 15 seconds on one failure you have not yet fixed. Do not film yourself typing code.
Common mistakes
The most common mistake is picking a heavyweight framework before knowing what you need, so that the prompt ends up buried under five layers of classes. The second is using clean, self-generated data, which turns the evals all green while proving nothing. The third is reporting only the best run instead of pass^k.
How does this skill carry over to customer work?
These five steps are not just for a portfolio; they are habits you will need on customer deployments, and your repo is where you practise them first.
On your CV, avoid the generic “built AI agent”. Write one line with numbers instead, for example “built a set of 40 eval cases from real failures, raising pass^3 from X to Y”, where X and Y come from your own repo. If a job description mentions “evaluation” or “reliability”, send the link to this repo with your application.
Let the repo answer in advance the question every customer eventually asks: where does the agent fail, and how often?
6 sources
- Building effective agents · 2024-12-19
- Demystifying evals for AI agents · 2026-01-09
- Your AI Product Needs Evals · 2024-03-29
- Evaluation best practices (OpenAI API docs)
- How to build a forward deployed engineer portfolio · 2026-09-19
- About READMEs - GitHub Docs