# Chunking client documents: measure before you cut, then add context to every chunk

> A passage cut from its source document can lose both the company name and the time period. Adding back a few dozen tokens of context reduced failed retrievals by 49%, and by 67% with reranking added.

Bản gốc: https://fdetimes.net/en/guides/chunking-client-documents-measure-then-contextualise/

Picture your second week on site with a client. The RAG demo breezes through the ten questions you wrote yourself. Then the head of finance types a real one: how much did revenue at the northern subsidiary grow in the second quarter?

The system returns a passage saying "revenue grew 3% on the previous quarter". The passage really is in the documents, but it comes from a different subsidiary's report, for a different quarter. The model invented nothing. It simply received a passage that had been cut away from its name and its dates.

Anthropic describes exactly this problem: on its own, a chunk no longer tells you which company or which period it is about.

For a forward deployed engineer, this is something you will meet on almost every RAG project, because enterprise documents are full of phrases such as "as noted above" and "in this period".

This guide covers three things, in order: measure first, cut second, then add context to every chunk.

## Why not just cut with the defaults?

Microsoft's RAG architecture guidance describes the chunking approach as a near-permanent choice in system design. The reason is practical: changing how you cut means re-embedding, re-indexing and re-measuring everything downstream. A wrong decision in week two will follow you all the way to handover.

Meanwhile, a technical report from Chroma found that the default settings of several popular chunking strategies perform fairly poorly. The library hands you a ready-made chunk size. That number has never seen the client's documents.

**Điểm mấu chốt:** Do not pick a chunk size by default; pick it from measurements on real questions.

## The first question: do you need to cut at all?

Before discussing how to cut, ask whether you need to. Anthropic notes that if a knowledge base is smaller than 200,000 tokens, or about 500 pages, you can put the whole thing in the prompt without RAG.

At many clients, the document set for a first use case is just a few dozen internal procedures. Counting tokens in the first meeting can save you a week of pipeline building. Only if the corpus is larger than that threshold, or will grow quickly, should you move on to the next step.

## Measure with real questions, not intuition

Microsoft recommends doing document analysis first, then testing several chunking approaches against a prepared test set of documents and questions. Document analysis means looking at what the client actually has: plain text, contracts with numbered clauses, or reports with tables and images.

Document structure determines how to cut. According to Microsoft, loosely structured text suits sentence-based splitting or fixed-size chunks with overlap. Semi-structured documents suit layout analysis or custom code written for the specific format.

Pinecone adds a few questions to answer before choosing: what kind of content it is, which embedding model you are using, and how long and complex users' queries are. The last of these can only be answered by sitting with real users. That is why the test set should come from them, not from you.

The simplest metric is hit rate: for each question, does the passage containing the answer appear in the top-k results? Chroma goes further with token-level Intersection over Union (IoU), a way of scoring chunking independently of the rest of the RAG pipeline. IoU penalises the redundant tokens produced by overlapping chunks.

## A worked example from start to finish

Suppose the client is an insurance company with a few thousand policies. You sit with the customer service team and collect 40 questions they genuinely look up often, each paired with the clause that contains the answer.

You then run three configurations and measure hit rate at top-5. The figures below are illustrative only, to show how to read the results.

| Configuration | Correct passage in top-5 | Hit rate |
|---|---|---|
| 200 tokens, no overlap | 24/40 | 60% |
| 200 tokens, 50-token overlap | 31/40 | 77.5% |
| 800 tokens, no overlap | 28/40 | 70% |

The table shows small chunks without overlap clearly trailing the other two, because answers often fall right at a cut. This matches Chroma's observation that with small contexts, overlap is needed to achieve high recall, and Microsoft's warning that relevant context can span several chunks.

The 800-token chunks are not the answer either. Pinecone points out that chunks that are too small or too large both make search results less precise. The reason for large chunks is easy to see: a long passage bundles several different ideas, so its embedding struggles to match a specific question sharply.

Overlap has its own cost: more duplicated tokens, which means lower IoU and a larger index. You choose the second configuration, but you write down the reason and the price paid.

## Adding context to each chunk

Picking the right size still does not fix the error from the opening. The "revenue grew 3%" passage, whether it is 200 or 800 tokens long, still does not say which company it belongs to. Contextual Retrieval, a technique introduced by Anthropic, targets precisely this failure.

The idea is simple: each chunk gets a short piece of context written specifically for it, typically 50-100 tokens, placed at the start of the chunk before embedding and before building the BM25 index. According to Anthropic, this reduces failed retrievals by 49%, and by 67% when combined with reranking.

The sketch below shows one way to do it: ask an LLM to read the document containing the chunk and write that context.

```python
# Sketch: llm() and search() are placeholder functions; swap in the client you use
def contextualize(doc_text, chunk_text, llm):
prompt = (
"DOCUMENT:\n" + doc_text + "\n\n"
"PASSAGE TO CONTEXTUALISE:\n" + chunk_text + "\n\n"
"Write 1-2 short sentences stating which document this passage belongs to, "
"which unit it concerns and which period. Return only the context."
)
context = llm(prompt)  # usually around 50-100 tokens
return context + "\n\n" + chunk_text  # embed and BM25-index this string

def hit_rate(test_set, search, k=5):
hits = 0
for q in test_set:  # each q has "question" and "gold_span"
results = search(q["question"], k=k)
hits += any(q["gold_span"] in r.text for r in results)
return hits / len(test_set)
```

Note that `hit_rate` checks whether the correct passage appears in the results rather than comparing chunk IDs. That way the same test set works for every chunking configuration. Cost is not a serious concern either: with prompt caching, Anthropic estimates about $1.02 per million document tokens, assuming 800-token chunks and 8,000-token documents.

Context is not only prose. Microsoft advises that when describing a table or image split across several chunks, you attach the image URL to each chunk so the metadata travels with every answer that uses it. File names, page numbers and section headings should travel with each chunk in the same way.

## Measuring again: the final round of the example

Back to the insurer. You keep the 200-token, 50-token-overlap configuration, add context to each chunk, and rerun exactly the same 40 questions. Again these are illustrative numbers: suppose the result rises to 35/40, a hit rate of 87.5%, compared with 31/40 before.

What matters is that you compare on the same questions with the same metric. If you change the test set midway, you can no longer tell whether the improvement came from the context or from easier questions. This two-row "before" and "after" table is also what you hand the client when proposing to switch the feature on.

## Common mistakes

The most common mistake is writing the test questions yourself. Engineers' questions tend to use the exact wording of the documents, which inflates hit rate. Real users' questions are shorter, vaguer and full of internal jargon.

The second is measuring once and stopping. Whenever the client adds a new document type, say moving from contracts to meeting minutes, the test set needs questions for that type. The third is chunking tables as if they were plain text: a row of figures separated from its column headers is almost meaningless.

The last is skipping the corpus-size check. More than a few complex pipelines have been built for a document set that would have fit in a single prompt.

## Where do employers look for this skill?

In job descriptions for FDE or AI engineer roles, watch for phrases such as "retrieval quality", "evaluation", "RAG pipeline" or "document ingestion". They signal that the job will need exactly what this guide describes.

On a CV, a line with measurements is far more convincing than "built a RAG system". For example: "Built a 40-question test set from user queries, compared 3 chunking configurations, raised hit rate@5 from X to Y using overlap and per-chunk context".

In interviews, be ready to explain the trade-offs behind your chosen configuration: overlap raises recall but lowers IoU and inflates the index, while 800-token chunks rarely split an answer but are less precise in search. Interviewers want to hear that you know what you paid for that number.

The client will never ask you what chunk size you used. They will only remember the time the system reported the wrong subsidiary's revenue, and the person who understands why that happened is the one they will want to keep.

**Thử ngay tuần này:**

- Take 20 real documents (or public documents of the same type), write 30 questions each paired with the passage containing the answer, then measure hit rate@5 for three chunk configurations.
- Pick the best configuration, add 1-2 sentences of context to each chunk, and measure again on exactly the same questions.
- Record the results as a small table in the README of a personal project, to use as evidence on your CV.

## Nguồn

- [Introducing Contextual Retrieval (Anthropic)](https://www.anthropic.com/news/contextual-retrieval)

- [Evaluating Chunking Strategies for Retrieval (Chroma Technical Report)](https://www.trychroma.com/research/evaluating-chunking)

- [Chunking Strategies for LLM Applications (Pinecone)](https://www.pinecone.io/learn/chunking-strategies/)

- [Develop a RAG Solution on Azure - Chunking Phase - Azure Architecture Center](https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/rag/rag-chunking-phase)
