RAG on real client documents: handling PDFs, tables and Vietnamese scans
A RAG demo that runs smoothly on sample files breaks as soon as it meets a client's shared drive, and the failure is usually in reading the documents, not in the model.
In brief
- RAG failures on real documents usually start at parsing and OCR, before embeddings come into play.
- Tables must be extracted separately and chunked by heading. Vietnamese scans need a diacritics check.
- Add context to chunks, combine BM25 with embeddings and rerank. Anthropic's own measurements show up to a 67% reduction in failed retrievals.
- 1Classify the filesSplit into text PDFs, table-heavy PDFs and scans, each with its own processing branch
- 2Parse each group its own wayTables: Unstructured or Docling. Scans: Vietnamese OCR with diacritic error counts, or ColPali
- 3Add context to each chunkAttach 50-100 tokens saying which document and section the chunk belongs to
- 4Hybrid retrievalIndex contextualized chunks in both embeddings and BM25, then rerank
- 5Measure with real questionsTrack how often the right chunk lands in the top-20 after changing one variable
Failures usually start in the first two steps, so check those before tuning retrieval.
Graphic: FDE Times
Picture your first week at a client. You are given access to a shared folder with a few hundred files: contracts exported from Word as PDFs, quarterly reports dense with tables, and a stack of skewed scanned invoices with red company stamps over the text.
The RAG demo ran beautifully on sample files, but here it gets the very first question wrong: “What is the payment due date for contract no. 0457?”
Many engineers’ first reflex is to swap the embedding model or raise top-k. But that is rarely where the fault lies. The answer is not in the top 20 because the system never read that page correctly: the table was cut in half, OCR dropped the tone marks, and the embedding treated the code “0457” as noise.
The right approach is to trace the data from raw file to retrieval result, one stage at a time. For an FDE, the most valuable skill is pinpointing which stage is failing before starting to fix anything.
Step one: classify files before parsing
Do not push the whole folder through a single parser. Open a few dozen files at random and sort them into three groups, because each group fails in a different way and needs its own processing branch.
| Document type | Main risk | Suggested handling |
|---|---|---|
| PDF with a text layer (exported from Word) | Heading structure lost, repeated headers/footers | Layout-aware parser, chunk by heading |
| Table-heavy PDF | Table rows split, columns merged into a single line of text | Extract tables separately, keep each table with its heading |
| Scans, photos | OCR errors on tone marks, noise, stamps | Specialised Vietnamese OCR, or retrieve directly from page images |
The test is simple: try to highlight the text in the file. If you can, the file has a text layer. If you cannot, it is an image and you need OCR.
This step alone gives you a number to report to the client in the first week, such as what percentage of the document store is scanned. That number determines much of the effort that follows.
Tables: do not let the parser split a row
A price list or payment schedule read as ordinary text becomes a string of numbers with no indication of which column each belongs to. Worse, a chunker that splits by token count can put the column headers in one chunk and the data rows in the next.
Cohere’s documentation on RAG over mixed data recommends Unstructured for parsing PDFs, because it can separate tables from text. It also chunks tables and text by heading during parsing, so related elements stay together.
IBM’s Docling takes a similar approach, using specialised models for layout analysis (trained on DocLayNet) and table structure recognition. A minimal snippet to get started:
from docling.document_converter import DocumentConverter
converter = DocumentConverter()
result = converter.convert("q3_report.pdf")
markdown = result.document.export_to_markdown()
# Open the markdown and read it yourself: does the table still have every row and column?
The final comment, which asks you to open the markdown and check by eye whether every row and column survived, is the part that matters. Before indexing, reread a few of the hardest tables, such as those with merged cells or those spanning two pages. If a table is already broken here, no reranker will save it.
Vietnamese scans: tone marks break first
Vietnamese uses the Latin alphabet but carries many diacritics to mark tones and distinguish vowels. For OCR, each diacritic is a small detail that noise, faded strokes or a red stamp can erase.
A 2025 survey of Vietnamese document recognition notes that Tesseract supports Vietnamese, but its accuracy remains limited because of how it handles diacritics and document noise.
If “Hạn thanh toán” (payment due date) becomes “Han thanh toan”, BM25 no longer matches the words in the user’s question. And if “bảo hành” (warranty) is read as “bào hành”, the embedding may pull it towards an entirely different topic.
So measure diacritic errors before indexing. Retype one page by hand as a reference, run Tesseract with lang="vie", then count words whose letters are right but whose diacritics are wrong:
import difflib, unicodedata
import pytesseract
from PIL import Image
def strip_diacritics(s):
s = s.replace("đ", "d").replace("Đ", "D")
return "".join(c for c in unicodedata.normalize("NFD", s)
if unicodedata.category(c) != "Mn")
ocr = pytesseract.image_to_string(Image.open("invoice_01.png"), lang="vie")
reference = open("invoice_01_reference.txt", encoding="utf-8").read() # hand-typed transcription
ocr_w = unicodedata.normalize("NFC", ocr).split()
ref_w = unicodedata.normalize("NFC", reference).split()
diacritic_errors = 0
sm = difflib.SequenceMatcher(a=ref_w, b=ocr_w, autojunk=False)
for tag, i1, i2, j1, j2 in sm.get_opcodes():
if tag == "replace":
for r, o in zip(ref_w[i1:i2], ocr_w[j1:j2]):
if strip_diacritics(r) == strip_diacritics(o): # same word, different diacritics
diacritic_errors += 1
print(f"Words with wrong diacritics: {diacritic_errors}/{len(ref_w)}")
The NFC normalisation step matters. The same character “ệ” can be stored as two different Unicode sequences, and without normalising you will also count words that are in fact correct.
The same survey describes a typical pipeline from the MC-OCR 2021 competition on receipts: CRAFT detects text regions, VietOCR recognises the text, and information is then extracted with rules. The lesson is to split OCR into two stages, detection and recognition, so you know which stage is producing errors.
There is another option: skip OCR altogether. ColPali uses a Vision Language Model to create multi-vector embeddings directly from images of document pages, avoiding the fragile text-extraction step. For a store of poor-quality scans, it is worth running ColPali alongside OCR and comparing results on the same question set.
A chunk needs to know where it belongs
Suppose parsing and OCR are clean. One problem remains: the chunk “Party B shall pay within 30 days of acceptance” does not say which contract it comes from, or which client.
Anthropic’s Contextual Retrieval addresses exactly this. Before indexing, a model writes a short piece of context, typically 50 to 100 tokens, which is prepended to each chunk. For example: “Excerpt from Article 5 of contract no. 0457 between company A and company B, payment terms section.”
This context does not only go into the embedding. With Contextual BM25, the contextualised chunk is also added to the BM25 index, so a question containing “0457” can match the payment clause even though the clause itself never mentions that code.
That is why BM25 is worth adding. It is a ranking function based on exact word matching, so it catches precisely what embeddings tend to miss: contract numbers, reference codes and proper names, which are everywhere in client documents.
According to Anthropic’s own measurements, Contextual Embeddings reduce the failed-retrieval rate in the top 20 chunks by 35%. Adding Contextual BM25 brings the reduction to 49%, and adding reranking takes it to 67%. Translated into a hypothetical test set with 100 failed retrievals, that figure falls to 65, then 51, then 33 after each step.
Do not promise these numbers to the client. They were measured on Anthropic’s data, not on your client’s invoice archive. What to take away is the method: measure the failure rate on that specific document store after each single change.
Rebuilding it in order
- Collect 20 to 30 real questions from the client’s users, each with the page containing the answer. This question set is the yardstick for every later decision.
- Classify files into three groups: text PDFs, table-heavy PDFs, scans.
- Choose a branch for each group: a layout-aware parser (Unstructured, Docling) for the table group, Vietnamese OCR or ColPali for the scans.
- Inspect a small sample of each group by eye, and run the diacritic-error script on the scans.
- Chunk by heading, then prepend a context passage to each chunk.
- Index the contextualised chunks in both embeddings and BM25, and add a reranker.
- Re-measure the share of correct chunks in the top 20 on the question set. Change only one variable at a time, or you will not know where the improvement came from.
Common mistakes
The most common mistake is judging by the final answer without looking at the top 20 chunks. When an answer is wrong, you need to know whether it is a retrieval error or a generation error. The two are fixed in completely different ways.
The second is stripping Vietnamese diacritics from both the index and the questions for the sake of “consistency”. This erases the distinction between words with different meanings. If you need to normalise, keep an accented version alongside.
The third is testing only on clean files. The question set must include questions answered in tables with merged cells, on blurry scanned pages and in text containing reference codes. Real users will ask exactly those questions.
When reading FDE job descriptions, look for phrases such as “document ingestion”, “OCR” and “unstructured data”. On a CV, a line like “raised correct retrieval on a store of scanned Vietnamese invoices from X to Y” carries far more weight than “built a RAG system”.
Clients never hand you clean data. The people who deliver are the ones who open the PDF and read it before opening a notebook.
Was this article useful?
Thanks for the feedback!
5 sources
- Introducing Contextual Retrieval · 2024-09-19
- Agentic RAG with mixed data (Cohere docs)
- Docling Technical Report · 2024-08-19
- A Survey on Vietnamese Document Analysis and Recognition: Challenges and Future Directions · 2025-06-05
- ColPali: Efficient Document Retrieval with Vision Language Models · 2024-06-27