# Fine-tuning or RAG: diagnose the problem before you pick the technique

> Two studies tested the same pair of techniques. In one, the gains added up. In the other, the model got worse. For FDEs, that means classifying failures and measuring a baseline before choosing anything.

Bản gốc: https://fdetimes.net/en/guides/fine-tuning-vs-rag-diagnose-first/

In 2024 a study called "Fine-Tuning or Fine-Failing?" reported a result many people did not expect. In the RAG setup it tested, the fine-tuned model performed worse than the original.

The same year, a Microsoft team fine-tuned a model on agricultural data and gained more than 6 percentage points of accuracy. Adding RAG brought another 5 points. The same pair of techniques produced opposite results. The two teams worked in different setups, though, so neither result directly refutes the other.

If you work as an FDE, you will hear this question in your first week: "Should you fine-tune for us or build RAG?" A good answer is not the name of a technique. It is a diagnosis.

The two opposite results are a reminder that no technique wins by default. Only measurements taken under the customer's actual conditions tell you which case you are in.

## Does the model not know, or does it know and get it wrong?

OpenAI's guide to optimising accuracy splits model failures into two types. In the first, the model lacks knowledge because the information was not in its training data. OpenAI calls this a context optimisation problem, and RAG is the tool for it.

In the second, the model produces inconsistent output or the wrong format. That is a problem with the model itself, and the answer is fine-tuning.

According to OpenAI, RAG is useful but can only fix what the model needs to take from its context. It does not teach the model a new habit. Fine-tuning, for its part, cannot give the model the contract the customer signed last week.

Picture an internal assistant deployed for an insurance company with a few thousand pages of claims rules. On demo day the model invents a clause that does not exist. That is a knowledge failure, and retrieval is what needs fixing.

Now suppose the model cites the right clause but sometimes answers in a table, sometimes in a paragraph, and sometimes forgets the case number. That is a behaviour failure. No change to chunking will fix it.

So before choosing a technique, start by collecting the customer's real failures and sorting them. Whichever column is longer decides what you work on the following week.

## Don't fine-tune while retrieval is still weak

If the left-hand column is longer, do not rush into training. The retrieval layer may have plenty of room left. In its post introducing Contextual Retrieval, Anthropic showed that combining contextual embeddings with contextual BM25 alone cut the top-20-chunk retrieval failure rate by 49%.

Adding a reranking step pushed the reduction to 67%.

Not a single model parameter changed to get those results. So make the RAG pipeline as good as it can be before anyone mentions training. Check whether chunks carry their context, whether keyword search is in the mix, and whether there is a reranker.

Sometimes you do not even need RAG. Anthropic notes that if the knowledge base is smaller than 200,000 tokens, about 500 pages, you can put all of it in the prompt. Before building a pipeline, ask the customer how much material actually needs to be used.

If it fits within that limit, putting it straight into the prompt is the option to try first.

## When fine-tuning is worth the cost

Fine-tuning makes sense when the right-hand column is long: the model has the information but does not work the way the customer wants. Here OpenAI's guidance has a principle worth remembering: the quality of training data matters more than its quantity.

In practice, a small set of examples that the customer's domain experts have reviewed line by line deserves priority over a huge pile of raw chat logs.

The cost is not just GPU time either. Microsoft's agriculture paper stresses that the upfront cost is high because fine-tuning on new data takes a lot of work, and inference costs later on have to be counted too.

Each time the customer updates its rules, a RAG pipeline only needs re-indexing. A fine-tuned model may need retraining.

Then there is the risk the "Fine-Tuning or Fine-Failing?" paper pointed out: after fine-tuning, the model may be worse than when you started. Without an eval set and a baseline number from before, you will not know that you have just spent the customer's money on a worse system.

## What the two conflicting results teach

Go back to the two studies at the start. Microsoft's results show that the two techniques can stack: more than 6 points from fine-tuning, another 5 from RAG. The other study is a reminder that the stacking is not guaranteed. The lesson for anyone deploying these systems is to measure each layer's contribution separately, so you know which one is actually helping.

There are also deliberate hybrids. RAFT trains a model to ignore irrelevant retrieved documents, and it improved results consistently across three datasets: PubMed, HotpotQA and Gorilla. Here fine-tuning does not teach new knowledge. It teaches the model to use RAG better.

In other words, it teaches a behaviour, exactly as OpenAI's classification would predict.

**Điểm mấu chốt:** Fine-tuning and RAG are not mutually exclusive. They answer two different questions, and you only know which question you are asking once you have an eval set in hand.

## What belongs on your CV is not a technique's name

For a developer aiming at an FDE role, the thing to learn is the order of work, not the technology. Build the habit of creating an eval set before writing the first line of pipeline code: 30-50 real customer questions, the expected answers and a baseline number.

From then on, every decision, whether adding a reranker, changing the chunking or fine-tuning, should come with a number compared against that baseline.

When reading job descriptions, check whether the role mentions evaluation or retrieval quality. If it does, put your experience building eval sets at the top of your application.

On your CV, instead of writing "experience with fine-tuning", record what you measured: how much you cut retrieval failures, how many points you gained on the customer's eval set, and why you chose one approach over another.

Customers will keep asking "fine-tune or RAG?" for years. A good FDE answers with an eval set and an error classification, not with the name of a technique.

**Thử ngay tuần này:**

- Take an internal chatbot project you are working on, collect 30-50 real questions with expected answers into an eval set, run a baseline and record the accuracy before fixing anything.
- Put each wrong answer in the eval set into one of two columns: 'the model does not have the information' or 'it has the information but gets the format or tone wrong'. Invest in whichever column is longer.
- Add a reranker or contextual chunking to an existing RAG pipeline, then measure the top-20-chunk retrieval failure rate again and compare it with the original number.

## Nguồn

- [Optimizing LLM Accuracy](https://developers.openai.com/api/docs/guides/optimizing-llm-accuracy)

- [RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture](https://arxiv.org/html/2401.08406v2)

- [Fine-Tuning or Fine-Failing? Debunking Performance Myths in Large Language Models](https://arxiv.org/abs/2406.11201)

- [Introducing Contextual Retrieval](https://www.anthropic.com/news/contextual-retrieval)

- [RAFT: Adapting Language Model to Domain Specific RAG](https://arxiv.org/abs/2403.10131)
