# Estimating token costs before the contract: count Vietnamese with the right tokenizer

> The "4.54 times" multiplier that circulates online can triple a quote. The fix is simple: take the client's real data and count.

Original: https://fdetimes.net/en/guides/estimate-vietnamese-token-costs-right-tokenizer/

Search for "how many tokens does Vietnamese cost" and the answer you will most often find is: 4.54 times as many as English. The author of a post on Viblo, a Vietnamese developer community, did not believe it and re-measured with 13 tokenizers on the same bilingual text set. With newer tokenizers, the gap fell to between 1.05x and 2.14x.

PhoBERT even came in at 0.87x, meaning fewer tokens than the English version. But PhoBERT is not an API model you can quote per token to a client, so that figure only shows how much the tokenizer matters.

The gap between those two numbers decides a quote. If you are an FDE and walk into a client meeting with the 4.54 multiplier, you may quote three times too high; the client finds it expensive and does not sign. Guess a multiplier that is too low, and the project loses money in its first month of production.

The worry is not new. As early as late 2022, a Vietnamese developer complained on OpenAI's forum that the tokenizer counted far too many tokens for Vietnamese, writing that the project had barely started and costs already looked too high. The steps below turn that worry into a number you can verify.

## Preparation: one script, one data file, one pricing page

You will end up with a small script that does two things: counts tokens on the client's sample data, and calculates monthly cost with input and output kept separate. You need Python, a file of the client's real Vietnamese text (or text you write yourself that closely resembles it), and the model provider's pricing page open in a tab.

Two concepts first. Tokens are the small units of data produced by breaking a larger block of information into pieces. Models process text as tokens, and providers charge by the token.

The process of splitting text into tokens is called tokenization, so the same Vietnamese sentence run through two different tokenizers yields two different counts.

## Step 1: get real data, and separate input from output

Ask the client for a sample of exactly the kind of data the system will receive: 50 support tickets, say, or 20 contracts, or 100 user questions. Save the part that will be sent to the model in one file and the expected answers in another.

API pricing usually charges input and output tokens separately, so your data should be split the same way from the start.

Check: open both files and confirm they are Vietnamese with diacritics, in the format the client actually uses. If you use machine-translated text, or Vietnamese without diacritics, the counts will not reflect reality.

## Step 2: count with the model's own tokenizer

This is the most important step. The code below is a shortened version for illustration, using the tiktoken library with the two encodings measured in the Viblo benchmark. Before using it for real, check the installation instructions and function names in the library's documentation.

```python
# Illustration (simplified): count tokens with two tokenizer generations
import tiktoken

new = tiktoken.get_encoding("o200k_base")   # GPT-4o / 4.1
old = tiktoken.get_encoding("cl100k_base")  # GPT-3.5 / 4

text = open("sample_input.txt", encoding="utf-8").read()
print("o200k_base:", len(new.encode(text)))
print("cl100k_base:", len(old.encode(text)))
```

(The variable `new` holds the newer encoding, `old` the older one, and `text` the text.)

Check: the cl100k_base count should be clearly higher. In the Viblo benchmark, o200k_base measured 1.34x and cl100k_base 2.14x. According to the same benchmark, simply switching to the newer tokenizer generation can cut costs by up to 37%, with the prompt unchanged.

These two encodings apply only to OpenAI models. If the client chooses another provider's model, that model has its own tokenizer, so count with that provider's own token-counting tool rather than borrowing o200k_base.

**Key point:** The token count for Vietnamese is set by the tokenizer, not the language, so the only number worth trusting is the one you count on the client's data.

## How one wrong multiplier skews the whole quote

Consider a hypothetical case. A client's traffic would amount to 10 million input tokens a month if written in English. Multiply by the rumoured 4.54 and you quote 45.4 million tokens. Use cl100k_base's 2.14 and you get 21.4 million; o200k_base's 1.34 gives just 13.4 million.

Same client, same workload, and the estimates differ by more than a factor of three purely because of the multiplier. So use these multipliers only to help the client understand the issue. When calculating cost, use the token counts you measured in Step 2 on their own data.

It is also worth knowing why Vietnamese is not the worst case. A 2025 study on arXiv compared many languages and found that those not written in Latin script, or with complex morphology, incur significantly more tokens. Vietnamese uses the Latin alphabet.

The authors also argue that per-token pricing widens the gap between languages further.

## Step 3: build the monthly cost formula

Once you have the average tokens per request, multiply by volume and apply the price. The code below assumes prices are listed per 1 million tokens. If the pricing page uses a different unit, adjust to match.

```python
def monthly_cost(num_requests, tok_in, tok_out,
                 price_in, price_out, surcharge=0.0):
    # price_in, price_out: price per 1 million tokens, checked on the quote date
    base = num_requests * (tok_in * price_in + tok_out * price_out) / 1_000_000
    return base * (1 + surcharge)

# surcharge=0.10 if the customer needs data residency (see Step 4)
```

(Here `monthly_cost` is the monthly cost, `num_requests` the number of requests, `price_in` and `price_out` the input and output prices per million tokens checked on the quote date, and `surcharge` the surcharge.)

Check: run it with tok_out set to 0 and confirm the result matches the input cost calculated by hand. Then raise tok_out and see how much the output portion moves the total. This is the part clients often forget to ask about.

## Step 4: state the pricing tier, the date and any surcharges

OpenAI's pricing changes quickly and is split into several tiers: Standard, Batch, Flex, Fast. Even within 2026 the priority tier was renamed, from Priority processing to Fast mode, on 30 July 2026. A quote without a date and tier will no longer be accurate after a few months.

One item is very easy to miss. For models released from 5 March 2026, regional processing endpoints (data residency) carry a 10% surcharge. So ask the client at the outset whether their data must stay in a particular region; if so, set phu_troi=0.10 in the formula.

## The mistakes that get quotes sent back

The first mistake is using a generic multiplier, whether 4.54 or any other figure found online, instead of counting yourself. The second is merging input and output into a single "tokens per request" number. When the client changes requirements, for instance wanting longer answers, you will not know which part to adjust.

The third is counting with one model's tokenizer and then quoting for a different model. The gap between 1.34x and 2.14x is enough to turn a profitable project into a loss. The last is counting on "clean" sample text you wrote yourself, when the client's real data contains tables, product codes and typos.

## Clients will ask "how much a month?" sooner than you expect

Because API pricing is per token, the question "how much will this cost to run each month?" will come sooner or later, so prepare the answer in advance. A table with three rows (input tokens, output tokens, and pricing tier with the date it was checked) is far more persuasive than "a few hundred dollars or so".

If you are preparing to move into an FDE role, put this skill on your CV in concrete terms. For example: "measured tokens on client Vietnamese data, selected a newer-generation tokenizer, cut estimated cost by X%", where X is a figure you actually measured.

When reading job descriptions, look for lines about cost estimation or solution scoping; that is where you can bring this example up in an interview.

The 4.54 rumour will stay online. But at the negotiating table, whoever brings token counts measured on the client's own data has the stronger case.

**Try this week:**

- Take 20 pieces of real Vietnamese text, such as customer support emails or product descriptions, count them with both o200k_base and cl100k_base, and record the ratio in a spreadsheet.
- Write the chi_phi_thang function from this article, enter today's prices from the pricing page, and note the date checked next to every figure.
- Rewrite an old quote (yours, or an imagined one) so it has all three rows: input tokens, output tokens, and pricing tier with date.

## Sources

- [Explaining Tokens — the Language and Currency of AI](https://blogs.nvidia.com/blog/ai-tokens-explained/)

- [Chi phí token tiếng Việt thực tế không phải đắt gấp 4,5 lần (Viblo)](https://viblo.asia/p/chi-phi-token-tieng-viet-thuc-te-khong-phai-dat-gap-45-lan-y0VGwyn7VPA)

- [Tokenization Disparities as Infrastructure Bias: How Subword Systems Create Inequities in LLM Access and Efficiency](https://arxiv.org/html/2510.12389v1)

- [Tokenizer is so high in Vietnamese](https://community.openai.com/t/tokenizer-is-so-high-in-vietnamese/27692)

- [Pricing | OpenAI API](https://developers.openai.com/api/docs/pricing)
