FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Read the model card before choosing Qwen, Llama, Gemma or Mistral for a Vietnamese project

The Llama 3.1 model card lists Thai but not Vietnamese. You need to catch details like that before you put an open model into a client's system.

Kỹ sư ngồi trước laptop đọc tài liệu kỹ thuật, màn hình hiển thị văn bản và mã, trong không gian văn phòng làm việc.
Photo: Negative Space / CC0

In brief

  • A model card is a README.md with YAML metadata. The license field tells you whether you are using a standard open-source licence or a vendor's own terms.
  • Qwen3-8B and Mistral-7B-Instruct-v0.3 use Apache 2.0. Llama 3.1 and Gemma 3 use custom licences, and Gemma also requires you to accept its terms before downloading files.
  • None of the four cards proves anything about Vietnamese for you. Check VMLU, then run your own tests on text written in regional dialects.
ShareLinkedInFacebookX

The model card for Llama 3.1 8B Instruct lists eight officially supported languages. Thai is on the list. Vietnamese is not.

The detail is on the model page for anyone to read, yet it is easy to miss once a team has decided to “use Llama because it’s popular”. As an FDE, you are usually the person who has to answer the client’s question: which open model works for Vietnamese, and are we legally allowed to use it?

This guide shows how to find that answer in the model card yourself, then check it against your own data.

One evening, one page

By the end of an evening you will have a one-page comparison table for four candidates: Qwen3-8B, Llama-3.1-8B-Instruct, gemma-3-4b-it and Mistral-7B-Instruct-v0.3. For each model it records the licence, what the card promises about languages and the risks you need to test, plus Vietnamese test results you measured yourself.

You need a Hugging Face account, because Gemma requires you to accept its terms before it lets you download files. You also need an editor for notes and, for the last step, access to run at least two of the models through whatever inference environment you already use. Reading the cards needs only a browser.

Step 1: Open README.md and read the YAML block first

The Hugging Face documentation states that the model card of every repo is its README.md file, made of Markdown plus metadata. The metadata is a YAML block at the top of the file with fields such as license, language, datasets and base_model. Read this block first: it is structured data, while the prose below it is the vendor describing its own product.

Here is an illustrative YAML block, shortened and not copied from any real repo:

---
license: apache-2.0
language:
  - en
base_model: some-org/some-base-model
datasets:
  - some-dataset
---

Check: for each model, write the license value into your table. If the card has a base_model, note it too and open the base model’s card to read its licence. Before you answer the client, you need to know the terms of both the version you are using and the model it was built on.

Step 2: Classify the licence before discussing quality

On the Hub, the licence is shown on the model page and can be used as a filter. Crucially, the Hub has its own codes for Llama (llama3.1, llama3.3, llama4) and for Gemma (gemma), separate from standard licences such as apache-2.0. A glance at this field therefore tells you whether you are dealing with a standard open-source licence or a vendor’s own terms.

The hardest case is a custom licence, declared like this (illustrative):

---
license: other
license_name: ten-license-rieng
license_link: LICENSE
---

With license: other, the Hub requires the full licence text to be placed in a LICENSE file in the repo. Do not trust the badge alone: open that file, or the link in license_link, and read it.

Applied to the four models: Qwen3-8B and Mistral-7B-Instruct-v0.3 use Apache 2.0. Llama 3.1 uses its own community licence, which requires products with more than 700 million monthly active users to request a licence from Meta, and comes with an acceptable-use policy governing how it may be used.

Gemma 3 uses the Gemma Terms of Use and is a gated repo, meaning you must agree to the terms before you can download the files.

Check: the licence column in your table should contain only two values: “standard” or “custom, full text read”. If any cell still says “custom, not yet read”, you are not ready for step 3.

Step 3: Read the language section as a promise

The Qwen3-8B card says the model supports more than 100 languages and dialects, but does not name Vietnamese. Gemma 3 4B-it describes multilingual capability across more than 140 languages. Llama 3.1 has eight official languages, and Vietnamese is not among them.

The Mistral-7B-Instruct-v0.3 card makes no multilingual claim at all, and states plainly that the model has no moderation mechanisms.

There are two kinds of signal here, and you need to keep them apart. With Llama, using Vietnamese means operating outside the vendor’s official support, and the client should be told so in advance. With Qwen and Gemma, the figures of 100 or 140 describe breadth, but they do not prove Vietnamese quality for your specific use case.

Check: copy each card’s language statement into the table word for word, without interpretation. Mark any statement that does not mention Vietnamese as “needs testing”.

Step 4: Record context length and operating limits

Qwen3-8B has a native context of 32,768 tokens, extendable to 131,072 tokens with YaRN. Gemma 3 has a 128K context window. These figures matter when a client wants to put an entire contract or a long case file into the prompt.

The extended figure usually requires specific configuration, so record both of Qwen’s values rather than only the larger one. For Llama and Mistral, open the cards and fill in the context length yourself rather than guessing. After this step, the core of your table will look like this:

Model Licence What the card says about languages Risks to test
Qwen3-8B apache-2.0 100+ languages, Vietnamese not named Real-world Vietnamese quality
Llama-3.1-8B-Instruct llama3.1 (custom) 8 languages, no Vietnamese Outside official support, 700 million MAU clause
gemma-3-4b-it gemma (custom, gated) More than 140 languages Gemma terms, Vietnamese quality
Mistral-7B-Instruct-v0.3 apache-2.0 No claims No moderation, Vietnamese unknown

Step 5: Check an independent source, then score with dialects yourself

A card is the vendor talking about itself, so you need an outside source. VMLU is a peer-reviewed Vietnamese benchmark presented at ACL 2025 that compares models including Llama-3, Qwen2.5 and GPT-4. Use it to get a rough sense of which model families are strong in Vietnamese.

Bear in mind that the versions in the benchmark may differ from the version you plan to use.

Next comes the part standard benchmarks do not fully cover. VialectBench, published on arXiv in August 2026, shows that regional dialects degrade LLM performance: an average drop of about 2.82% across ten instruction-tuned models, with no model fully robust. Your client’s real users will not write like a textbook.

So build a small test set and score it with a simple rule: an answer that matches the intent in the expected-answer column scores 1; a wrong or off-topic answer scores 0. The format below is your own, not anyone’s standard, and the two rows of scores are illustrative only. Each row pairs a standard Vietnamese sentence with the same request in dialect (here, Central and Southern forms such as “tui”, “nghen”, “chừng nào” and “rứa”):

id,cau_chuan,cau_phuong_ngu,dap_an_mong_doi,diem_chuan,diem_phuong_ngu
1,"Tôi muốn đổi địa chỉ giao hàng","Tui muốn đổi chỗ giao hàng nghen","Hướng dẫn đổi địa chỉ",1,1
2,"Đơn của tôi bao giờ tới?","Đơn tui chừng nào tới rứa?","Tra cứu trạng thái đơn",1,0

Imagine you run 30 sentence pairs through one model. Suppose the standard column scores 24/30, or 80%, while the dialect column scores only 19/30, about 63.3%. The gap is 5 sentences, or nearly 16.7 percentage points.

That hypothetical gap, not the 80%, is the figure to report. Reread the 5 failed sentences to see whether the model stumbles on regional vocabulary, sentence-final particles or text typed without diacritics. Each type of error points to a different fix, from normalising the input to adding examples to the prompt.

Check: run the same question set on at least two models and record three numbers for each: the standard score, the dialect score and the gap. If the gap is large, that is a finding the client needs to know before launch.

What are the most common mistakes?

The first is trusting the badge. A repo that declares license: other can contain any terms at all, and only the LICENSE file tells you what they are. The second is treating “more than 100 languages” as evidence for Vietnamese, when the Qwen3-8B card does not even name Vietnamese.

The third is forgetting to check the base model’s licence when using a community-built Vietnamese fine-tune. The fourth is choosing Mistral for its Apache 2.0 licence and putting it straight in front of end users, when the card itself says the model has no moderation mechanisms. In that case you need to add your own filtering layer in front of it.

Where does this skill show up in client work?

Picture a bank that wants to run an internal chatbot on its own infrastructure. In the first meeting, the legal team’s question will come before any question about accuracy: “What does this licence allow us to do?”

If you bring the table from step 4, the full texts of the custom licences with the relevant clauses marked, and the three test numbers from step 5, the meeting moves straight on to technical matters.

If you want to put this skill on your CV, do not write “experienced with LLMs”. Write that you evaluated the licences and Vietnamese capability of four open models, built your own dialect test set, and include the score gap you measured.

When reading a job description, phrases such as “model selection”, “evaluation” or “compliance review” describe exactly the work you have just done.

The next time someone asks “Llama or Qwen?”, you can answer with a full table of licences, languages and test results, not a hunch.

8 sources
Read next on the roadmap · Stage 3: Applied AICalling the Claude and OpenAI APIs directly: write your own tool-use loop, then compare with GeminiBefore handing everything to a framework, work through a full cycle of requests, streaming and tool calls with Claude and OpenAI, and learn where Gemini differs. These are the things you will end up debugging at a client's office.