Choosing an embedding model for Vietnamese data: build a 50-question test set before trusting the leaderboard
A model fine-tuned on Vietnamese legal text reports that its Accuracy@1 rose from 0.5682 to 0.7274. That number tells you nothing yet about the insurance documents of the client you are working with.

In brief
- Leaderboards such as VN-MTEB only help narrow the list. The best model for a client has to be chosen on the client's data.
- A fair comparison needs a fixed dataset, model revision, preprocessing and hardware. If the model is E5, do not forget the query and passage prefixes.
- A fine-tuned model's self-reported results are a hypothesis. Your 50-question test set is what tests it.
On Hugging Face, the model card for nrl-ai/vietnamese-embedding-bge-m3 reports a sizeable jump. On the Legal Zalo legal-document retrieval task, Accuracy@1 rises from 0.5682 for the original BGE-M3 to 0.7274, and MRR@10 rises from 0.6822 to 0.8181. These figures are published by the authors themselves.
Now imagine that next week you are sitting in the office of an insurance company. It has thousands of pages of policy terms, claims procedures and customer email replies, all written in Vietnamese. Should you adopt the fine-tuned model straight away?
The right answer is: you do not know until you measure it on their data. This guide shows how to build a small test set that answers the question in an afternoon. On a customer site, this is the skill that lets an FDE make a recommendation backed by numbers rather than a gut feeling.
A leaderboard only narrows the list
Supermemory’s comparison of open-source embedding models says plainly that no leaderboard can identify the best model for every corpus. A leaderboard only shortens the list of candidates.
For Vietnamese there is VN-MTEB, an embedding benchmark of 41 datasets covering six task types (retrieval, classification, pair classification, clustering, reranking and STS) that evaluates 18 models.
VN-MTEB was built by using an LLM to translate MTEB into Vietnamese. It is therefore translated data, not the questions real customers type into a search box, complete with abbreviations, typos and internal jargon. The benchmark also reports that larger models using Rotary Positional Embedding outperform those using Absolute Positional Embedding.
That is a useful hint for picking candidates, but it does not replace verification.
What you will build and what you need
The end product is a script that scores three candidate models on about 50 questions and prints Recall@5, MRR@10 and average latency. You need Python, a fixed machine (the same GPU or the same CPU for every run), whatever embedding library your team already uses, and, most importantly, someone on the client side willing to confirm the answers.
The code below is simplified. The scoring is plain Python. The embed() function is left empty: wire it to your own library following the instructions on each model card.
Step 1: write questions in the client’s language
Take questions from real sources such as call-centre logs, tickets or emails; do not invent tidy ones. Then split the documents into passages (chunks) with ids, and attach to each question the id of the passage containing the answer. A hypothetical example for the insurance client (the first query is informal Vietnamese, with abbreviations, asking whether cancelling a contract early earns a premium refund; the second asks how many days in hospital qualify for an allowance):
{"qid": "q01", "query": "huỷ hđ trước hạn có được hoàn phí ko", "relevant": ["dk_12_3"]}
{"qid": "q02", "query": "nằm viện bao nhiêu ngày thì được trợ cấp", "relevant": ["dk_07_1", "dk_07_2"]}
Check: send 10 random lines to the client-side reviewer. If they correct more than two answers, redo the labelling before scoring any model.
Step 2: pick three candidates and pin everything
A sensible shortlist has three models. The first is BAAI/bge-m3 as the baseline, because it supports more than 100 languages and accepts inputs of up to 8192 tokens, which suits long contracts. The second is nrl-ai’s Vietnamese fine-tune.
The third is a model from a different family, to give a point of comparison. How to choose it: open the VN-MTEB retrieval table and take the highest-scoring multilingual E5 model; if the table has no suitable E5, take the top model outside the BGE family. Then copy its exact name, commit and prefixes from its model card into the configuration.
# Đơn giản hóa: thay các chuỗi VIET_HOA bằng giá trị thật từ model card
MODELS = {
"bge-m3": {"name": "BAAI/bge-m3", "revision": "PIN_COMMIT_HASH", "q_prefix": "", "p_prefix": ""},
"vn-bge-m3": {"name": "nrl-ai/vietnamese-embedding-bge-m3", "revision": "PIN_COMMIT_HASH", "q_prefix": "", "p_prefix": ""},
"e5": {"name": "TEN_MODEL_E5_CHON_TU_VN_MTEB", "revision": "PIN_COMMIT_HASH",
"q_prefix": "QUERY_PREFIX_THEO_CARD", "p_prefix": "PASSAGE_PREFIX_THEO_CARD"},
}
(The comment reads: “Simplified: replace the UPPER_CASE strings with real values from the model card.”)
Supermemory advises fixing the dataset, model revision, preprocessing and hardware. Only then do differences in accuracy and latency mean anything. So every model must receive the same set of chunks, the same Vietnamese diacritic normalisation, and run on the same machine.
Step 3: compute scores you can check by hand
def score(ranked, relevant, k=5):
hit = any(d in relevant for d in ranked[:k]) # Recall@k (có ít nhất 1 đoạn đúng)
rr = next((1/(i+1) for i, d in enumerate(ranked[:10]) if d in relevant), 0.0) # cho MRR@10
return hit, rr
def evaluate(results, testset, k=5):
hits, rrs = zip(*(score(results[q["qid"]], q["relevant"], k) for q in testset))
return sum(hits)/len(hits), sum(rrs)/len(rrs)
results is the list of passage ids ranked by cosine similarity between the question vector and the passage vectors, produced by your embed() function. (The comments note that Recall@k counts a hit when at least one correct passage appears, and that rr feeds MRR@10.) Try computing MRR by hand for three questions. Question one has the correct passage at rank 1 (score 1), question two at rank 2 (score 0.5), and question three has no correct passage in the top 10 (score 0).
MRR is (1 + 0.5 + 0) / 3 = 0.5.
Read the number as a position in the result list. An MRR of 0.8181 means the correct passage usually sits at rank 1 or 2. An MRR of 0.5 means users typically have to read past one wrong passage to reach the right one. Measure latency by timing only the embed() call for the question, on the same fixed machine.
Check: run the same model twice. If the scores change, preprocessing or data ordering is not truly fixed.
Step 4: read the misses, not just the scoreboard
Once you have a scoreboard, do not rush to pick the top model. Open the questions that all three models missed. The cause is often a chunk that cuts a clause in half, an abbreviation such as “hđ” (contract) or “ko” (not), or a mislabelled answer. Changing the model will not fix errors of this kind.
BGE-M3 also offers three retrieval modes in a single model: dense, sparse and multi-vector. Once the dense run is done, you can use the same test set to see whether sparse mode rescues questions containing product codes or clause numbers.
Three errors that ruin a comparison without warning
The first is forgetting the prefix. E5 models need separate prefixes for queries and passages. Supermemory’s article warns that missing prefixes can invalidate a comparison. The script still runs, raises no error, and E5 simply looks worse than it is.
The second is trusting self-reported figures. The nrl-ai model card states that the model was trained on roughly 300,000 triplets of query, correct document and incorrect document, and that the Legal Zalo 2021 evaluation data was not in the training set.
Even if that holds, a number measured on legal texts still does not tell you how well the model will handle the insurance terms in the example above. The client’s test set is where that gets verified.
The third is treating open weights as free. Supermemory points out that inference, hosting and maintenance all cost money, and that the licence must be checked for each revision. Put licence and latency columns in the results table you send the client, not just accuracy.
How this skill shows up in FDE work
Picture the question “which model should we use?” being raised in a client meeting. An FDE who brings a 50-question test set already approved by the client is far more persuasive than one who shows a leaderboard screenshot. The test set also outlives the current model: each time a new model appears, you just rerun it.
When reading job descriptions, look for phrases such as “retrieval evaluation”, “RAG quality” or “customer data”. On your CV, do not just write “built RAG”. Say that you built an N-question Vietnamese evaluation set, compared three models and raised MRR@10 from X to Y, along with the reason for your choice.