# Sizing GPUs for a 70B model: from parameters to KV cache and quantisation

> When a client asks how many GPUs they need to run a model in their own datacentre, you should have the answer before you leave the meeting, not after a failed deployment.

Original: https://fdetimes.net/en/guides/estimate-gpus-for-70b-model/

Second meeting with a banking client. Their infrastructure team will not let data leave the building, so the model must run on-prem, and the head of IT asks directly: "To run a 70B model, how many cards do we need to buy?" If you answer "let me check and get back to you", you lose a week.

If you answer wrongly, the client spends money on hardware and the model still will not run.

When a model has to run on a client's infrastructure, you should be able to turn a model name into gigabytes, and gigabytes into a GPU count, right there at the table. The arithmetic is not hard. The hard part is knowing which components to count, and which ones the client will not mention unprompted.

This guide moves from the concept of parameters, through a complete hand calculation for a 70B model, to code you can use to check it. It ends with the mistakes that throw estimates off by a factor of two.

## What are parameters, and why do they determine the GPU count?

IBM defines an LLM's parameters as the settings that control the model's output and behaviour, with two main types: weights and biases. Weights are numerical values that express how much importance the model assigns to a particular input. Taken together, a model can have billions of parameters.

For anyone doing deployment, each parameter is a number that has to sit in GPU memory while the model runs. So the first formula is simple: weight memory = number of parameters × bytes per parameter. The "70B" in the model name is the parameter count; what remains is knowing how many bytes each parameter occupies.

That is where quantisation comes in. IBM describes it as a way of simplifying all the mathematics inside the model, making it smaller and more efficient. Instead of storing each weight in 2 bytes (FP16/BF16), you store it in 1 byte (INT8) or half a byte (INT4).

| Precision | Bytes per parameter | Weights for a 70B model |
|---|---|---|
| FP16 / BF16 | 2 | 140 GB |
| INT8 | 1 | 70 GB |
| INT4 | 0.5 | 35 GB |

Hugging Face's Transformers documentation states that 8-bit quantisation with bitsandbytes halves memory compared with 16-bit, and 4-bit compresses it further. On weights alone, 140 GB at FP16 already exceeds a single 80 GB GPU.

That is why an analysis on Machine Learning at Scale notes that INT4 takes a 70B model from needing two 80 GB GPUs to fitting on one.

## Weights are only half the story

Stop at the table above and you will tell the client "INT4, one 80 GB card is enough", and you may be wrong. When serving requests, the model also holds a KV cache: the key and value vectors for every token already processed, so it does not have to recompute them at each new generation step.

The per-token formula is: kv_bytes_per_token = 2 × n_layers × n_kv_heads × head_dim × bytes_per_element. The factor of 2 accounts for both keys and values. The other values are in the model's config file; read them directly from there.

KV cache grows with the context window, meaning all the text the model can refer to while generating a response. Claude's documentation stresses that the context window includes the response. So a request with a 6,000-token prompt and a 2,000-token answer occupies KV cache for 8,000 tokens, not 6,000.

**Key point:** A GPU holds not just the model but every conversation in progress.

## A worked example

Suppose the bank's 70B model has the following hypothetical configuration: 80 layers, 8 KV heads, head_dim of 128, and KV cache stored at 2 bytes per element. These numbers are for illustration only; in practice, take them from the model's config.json.

KV per token = 2 × 80 × 8 × 128 × 2 = 327,680 bytes, roughly 0.33 MB. The client says typical context is 8,192 tokens and they want to serve 10 users at once. One request costs 327,680 × 8,192 ≈ 2.68 GB; ten requests cost about 26.8 GB.

Now add it up. At INT4: 35 GB of weights + 26.8 GB of KV cache = 61.8 GB. Machine Learning at Scale recommends reserving an extra 10–20% on top of weights and KV cache for activations, CUDA context and the framework; taking 20% to be safe gives about 74.2 GB. It fits on one 80 GB GPU, but only just.

At FP16 the picture changes completely: 140 + 26.8 = 166.8 GB, plus 20% comes to about 200 GB. Two 80 GB GPUs can hold the weights but not this load; you need three cards, or fewer concurrent users. Same model, different precision, a difference of two cards.

Then the client adds: "We want to feed in whole long contracts, around 32 thousand tokens." Say the context is now 32,768 tokens.

One request takes 327,680 × 32,768 ≈ 10.7 GB of KV cache; ten requests take about 107 GB. Add 35 GB of INT4 weights and 20% headroom, and the total rises to about 171 GB: three 80 GB GPUs instead of one. The single-card plan collapsed over one question about document length.

## Long context costs more elsewhere too

Memory is not the only cost. IBM notes that the compute required for attention grows with the square of sequence length, so long requests consume disproportionate resources. Claude's documentation also warns that longer context is not automatically better, because of context rot.

This is the moment to revisit the requirement with the client. Rather than buying more cards to cram an entire 32,768-token contract into the prompt, you can propose splitting documents into chunks and including only the relevant sections. Careful sizing often produces a leaner design, not just a shopping list of cards.

## Don't trust the numbers on paper: measure

Once you have done the arithmetic, verify it on real hardware. The Transformers documentation lets you load a model in 8-bit with device_map="auto" to use whatever GPUs are available, and provides get_memory_footprint to measure the loaded model. If you do not have an 80 GB card yet, test the formula on a small model first:

```python
from transformers import AutoModelForCausalLM, BitsAndBytesConfig

def weight_gb(params, bytes_per_param):
    return params * bytes_per_param / 1e9

def kv_gb(n_layers, n_kv_heads, head_dim, bytes_el, context, concurrency):
    per_token = 2 * n_layers * n_kv_heads * head_dim * bytes_el
    return per_token * context * concurrency / 1e9

# Estimate for the 70B INT4 case in the example above, plus 20% headroom
total = (weight_gb(70e9, 0.5) + kv_gb(80, 8, 128, 2, 8192, 10)) * 1.2
print(f"Estimate for 70B INT4 case: {total:.1f} GB")

# Verify the weight formula on a small model, loaded in 8-bit (1 byte/parameter)
model_id = "your-small-model-name"  # replace with any small model on Hugging Face
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=BitsAndBytesConfig(load_in_8bit=True),
    device_map="auto",
)
print(f"Calculated: {weight_gb(model.num_parameters(), 1):.2f} GB")
print(f"Measured:   {model.get_memory_footprint() / 1e9:.2f} GB")
```

In the code above, the comments estimate the 70B INT4 case with 20% headroom, then check the weight formula on a small 8-bit model (1 byte per parameter); replace `model_id` with any small model on Hugging Face. The last two lines print the hand-calculated figure and the measured one.

get_memory_footprint reports only the loaded model, not the KV cache under load. So the final step is always to run a test at the concurrency and context length the client actually needs, and watch memory.

## Five mistakes that skew the estimate

The most common mistake is counting only the weights, as in the example above: 35 GB looks comfortable until ten users submit long documents at once. The second is forgetting that the response also sits in the context, and so sizing KV cache from prompt length alone.

The third is confusing parameters with bytes, assuming "70B means 70 GB" without asking about precision. The fourth is skipping the 10–20% headroom, and then watching the model crash out of memory in the middle of the demo.

The fifth is subtler: treating quantisation as free. Quantisation simplifies the mathematics inside the model, so you need to run the client's own eval suite on the INT4 version before promising FP16-level quality. Show the client the results and let them decide on the trade-off.

## Bringing this skill to your CV and interviews

When reading FDE job descriptions, look for phrases such as "on-prem deployment", "air-gapped", "GPU sizing" or "inference optimization": that is where this skill gets paid. On a CV, one concrete line carries more weight than any adjective, for instance describing how you estimated VRAM for a model, chose INT4 after measuring quality, and reduced the number of GPUs to buy.

In a case interview, if asked "the client wants to run model X; how many GPUs do they need?", do not rush to a number.

Ask five questions back: what GPUs the client has, how much VRAM each card has, how many concurrent users, how long the context is, and what level of quality is acceptable.

Those five questions show you understand what the final number depends on.

Next time a client asks "how many cards?", the right answer starts with "let me ask you five questions" and ends with a calculation both sides can check.

**Try this week:**

- Write a Python function estimate_vram(params, bytes_per_param, n_layers, n_kv_heads, head_dim, context, concurrency) and run it at FP16, INT8 and INT4.
- Load a small Hugging Face model at 16-bit and then at 8-bit, print get_memory_footprint() both times and compare with your hand calculation.
- Prepare the five sizing questions (which GPUs, how much VRAM, how many concurrent users, how long the context, what quality is acceptable) to take into your next discovery session.

## Sources

- [What Are LLM Parameters? | IBM](https://www.ibm.com/think/topics/llm-parameters)

- [How Much GPU Memory Do You Need to Serve an LLM?](https://www.machinelearningatscale.com/blog/gpu-memory-requirements-llm-serving)

- [bitsandbytes (Transformers v4.44.2 docs)](https://huggingface.co/docs/transformers/v4.44.2/en/quantization/bitsandbytes)

- [Context windows - Claude API Docs](https://platform.claude.com/docs/en/build-with-claude/context-windows)

- [What is a context window? | IBM](https://www.ibm.com/think/topics/context-window)
