# RAG on real client documents: handling PDFs, tables and Vietnamese scans

> A RAG demo that runs smoothly on sample files breaks as soon as it meets a client's shared drive, and the failure is usually in reading the documents, not in the model.

Original: https://fdetimes.net/en/guides/rag-messy-pdfs-tables-vietnamese-scans/

Picture your first week at a client. You are given access to a shared folder with a few hundred files: contracts exported from Word as PDFs, quarterly reports dense with tables, and a stack of skewed scanned invoices with red company stamps over the text.

The RAG demo ran beautifully on sample files, but here it gets the very first question wrong: "What is the payment due date for contract no. 0457?"

Many engineers' first reflex is to swap the embedding model or raise top-k. But that is rarely where the fault lies. The answer is not in the top 20 because the system never read that page correctly: the table was cut in half, OCR dropped the tone marks, and the embedding treated the code "0457" as noise.

The right approach is to trace the data from raw file to retrieval result, one stage at a time. For an FDE, the most valuable skill is pinpointing which stage is failing before starting to fix anything.

**Key point:** RAG on real documents usually breaks at the document-reading stage before it breaks at retrieval.

## Step one: classify files before parsing

Do not push the whole folder through a single parser. Open a few dozen files at random and sort them into three groups, because each group fails in a different way and needs its own processing branch.

| Document type | Main risk | Suggested handling |
|---|---|---|
| PDF with a text layer (exported from Word) | Heading structure lost, repeated headers/footers | Layout-aware parser, chunk by heading |
| Table-heavy PDF | Table rows split, columns merged into a single line of text | Extract tables separately, keep each table with its heading |
| Scans, photos | OCR errors on tone marks, noise, stamps | Specialised Vietnamese OCR, or retrieve directly from page images |

The test is simple: try to highlight the text in the file. If you can, the file has a text layer. If you cannot, it is an image and you need OCR.

This step alone gives you a number to report to the client in the first week, such as what percentage of the document store is scanned. That number determines much of the effort that follows.

## Tables: do not let the parser split a row

A price list or payment schedule read as ordinary text becomes a string of numbers with no indication of which column each belongs to. Worse, a chunker that splits by token count can put the column headers in one chunk and the data rows in the next.

Cohere's documentation on RAG over mixed data recommends Unstructured for parsing PDFs, because it can separate tables from text. It also chunks tables and text by heading during parsing, so related elements stay together.

IBM's Docling takes a similar approach, using specialised models for layout analysis (trained on DocLayNet) and table structure recognition. A minimal snippet to get started:

```python
from docling.document_converter import DocumentConverter


converter = DocumentConverter()
result = converter.convert("q3_report.pdf")
markdown = result.document.export_to_markdown()
# Open the markdown and read it yourself: does the table still have every row and column?
```

The final comment, which asks you to open the markdown and check by eye whether every row and column survived, is the part that matters. Before indexing, reread a few of the hardest tables, such as those with merged cells or those spanning two pages. If a table is already broken here, no reranker will save it.

## Vietnamese scans: tone marks break first

Vietnamese uses the Latin alphabet but carries many diacritics to mark tones and distinguish vowels. For OCR, each diacritic is a small detail that noise, faded strokes or a red stamp can erase.

A 2025 survey of Vietnamese document recognition notes that Tesseract supports Vietnamese, but its accuracy remains limited because of how it handles diacritics and document noise.

If "Hạn thanh toán" (payment due date) becomes "Han thanh toan", BM25 no longer matches the words in the user's question. And if "bảo hành" (warranty) is read as "bào hành", the embedding may pull it towards an entirely different topic.

So measure diacritic errors before indexing. Retype one page by hand as a reference, run Tesseract with `lang="vie"`, then count words whose letters are right but whose diacritics are wrong:

```python
import difflib, unicodedata
import pytesseract
from PIL import Image

def strip_diacritics(s):
    s = s.replace("đ", "d").replace("Đ", "D")
    return "".join(c for c in unicodedata.normalize("NFD", s)
                   if unicodedata.category(c) != "Mn")

ocr = pytesseract.image_to_string(Image.open("invoice_01.png"), lang="vie")
reference = open("invoice_01_reference.txt", encoding="utf-8").read()  # hand-typed transcription

ocr_w = unicodedata.normalize("NFC", ocr).split()
ref_w = unicodedata.normalize("NFC", reference).split()

diacritic_errors = 0
sm = difflib.SequenceMatcher(a=ref_w, b=ocr_w, autojunk=False)
for tag, i1, i2, j1, j2 in sm.get_opcodes():
    if tag == "replace":
        for r, o in zip(ref_w[i1:i2], ocr_w[j1:j2]):
            if strip_diacritics(r) == strip_diacritics(o):  # same word, different diacritics
                diacritic_errors += 1

print(f"Words with wrong diacritics: {diacritic_errors}/{len(ref_w)}")
```

The NFC normalisation step matters. The same character "ệ" can be stored as two different Unicode sequences, and without normalising you will also count words that are in fact correct.

The same survey describes a typical pipeline from the MC-OCR 2021 competition on receipts: CRAFT detects text regions, VietOCR recognises the text, and information is then extracted with rules. The lesson is to split OCR into two stages, detection and recognition, so you know which stage is producing errors.

There is another option: skip OCR altogether. ColPali uses a Vision Language Model to create multi-vector embeddings directly from images of document pages, avoiding the fragile text-extraction step. For a store of poor-quality scans, it is worth running ColPali alongside OCR and comparing results on the same question set.

## A chunk needs to know where it belongs

Suppose parsing and OCR are clean. One problem remains: the chunk "Party B shall pay within 30 days of acceptance" does not say which contract it comes from, or which client.

Anthropic's Contextual Retrieval addresses exactly this. Before indexing, a model writes a short piece of context, typically 50 to 100 tokens, which is prepended to each chunk. For example: "Excerpt from Article 5 of contract no. 0457 between company A and company B, payment terms section."

This context does not only go into the embedding. With Contextual BM25, the contextualised chunk is also added to the BM25 index, so a question containing "0457" can match the payment clause even though the clause itself never mentions that code.

That is why BM25 is worth adding. It is a ranking function based on exact word matching, so it catches precisely what embeddings tend to miss: contract numbers, reference codes and proper names, which are everywhere in client documents.

According to Anthropic's own measurements, Contextual Embeddings reduce the failed-retrieval rate in the top 20 chunks by 35%. Adding Contextual BM25 brings the reduction to 49%, and adding reranking takes it to 67%. Translated into a hypothetical test set with 100 failed retrievals, that figure falls to 65, then 51, then 33 after each step.

Do not promise these numbers to the client. They were measured on Anthropic's data, not on your client's invoice archive. What to take away is the method: measure the failure rate on that specific document store after each single change.

## Rebuilding it in order

1. Collect 20 to 30 real questions from the client's users, each with the page containing the answer. This question set is the yardstick for every later decision.
2. Classify files into three groups: text PDFs, table-heavy PDFs, scans.
3. Choose a branch for each group: a layout-aware parser (Unstructured, Docling) for the table group, Vietnamese OCR or ColPali for the scans.
4. Inspect a small sample of each group by eye, and run the diacritic-error script on the scans.
5. Chunk by heading, then prepend a context passage to each chunk.
6. Index the contextualised chunks in both embeddings and BM25, and add a reranker.
7. Re-measure the share of correct chunks in the top 20 on the question set. Change only one variable at a time, or you will not know where the improvement came from.

## Common mistakes

The most common mistake is judging by the final answer without looking at the top 20 chunks. When an answer is wrong, you need to know whether it is a retrieval error or a generation error. The two are fixed in completely different ways.

The second is stripping Vietnamese diacritics from both the index and the questions for the sake of "consistency". This erases the distinction between words with different meanings. If you need to normalise, keep an accented version alongside.

The third is testing only on clean files. The question set must include questions answered in tables with merged cells, on blurry scanned pages and in text containing reference codes. Real users will ask exactly those questions.

When reading FDE job descriptions, look for phrases such as "document ingestion", "OCR" and "unstructured data". On a CV, a line like "raised correct retrieval on a store of scanned Vietnamese invoices from X to Y" carries far more weight than "built a RAG system".

Clients never hand you clean data. The people who deliver are the ones who open the PDF and read it before opening a notebook.

**Try this week:**

- Take 10 pages of real Vietnamese PDFs (invoices, contracts, reports with tables), run them through Tesseract with lang='vie' and through Docling, then use the script in this article to count diacritic errors on each page
- Write 20 questions whose answers sit in tables or contain reference codes, then measure the share of correct chunks in the top 20 with embeddings alone and with BM25 added
- Add a line to your CV describing that pipeline with your own measurements, for example correct retrieval rates before and after the fixes

## Sources

- [Introducing Contextual Retrieval](https://www.anthropic.com/news/contextual-retrieval)

- [Agentic RAG with mixed data (Cohere docs)](https://docs.cohere.com/page/agentic-rag-mixed-data)

- [Docling Technical Report](https://arxiv.org/abs/2408.09869)

- [A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions](https://arxiv.org/html/2506.05061v1)

- [ColPali: Efficient Document Retrieval with Vision Language Models](https://arxiv.org/abs/2407.01449)
