FDE PulseFDE jobs open 448New in 7 days 29Companies hiring 52Remote-friendly 25%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Analysis

A context window holds 5,000 pages, but the point where you can drop RAG is 500: FDEs need numbers to answer clients

Context windows can hold ten times more than the safe threshold for skipping RAG. In that gap sit lost-in-the-middle failures, KV cache memory and a bill that grows non-linearly, and clients rarely see any of them coming.

A context window holds 5,000 pages, but the point where you can drop RAG is 500: FDEs need numbers to answer clients
Photo: The National Archives (United Kingdom) / CC BY 3.0

In brief

  • A 2-million-token window holds about 5,000 pages. According to Anthropic, a knowledge base under 200,000 tokens (about 500 pages) can go straight into the prompt without RAG.
  • Long context costs disproportionately more, the first token can take over 2 minutes to arrive, and models use information in the middle of the context far less well.
  • RAG and long context draw on the same token budget. For large corpora, the answer is to combine them, not to pick one.
ShareLinkedInFacebookX
Horizontal bar chart comparing two thresholds measured in pages. Gemini 1.5 Pro's 2 million-token context window equals about 5,000 pages. The 200,000-token level at which Anthropic says a whole knowledge base can go into the prompt without RAG is only about 500 pages. A bracket marks the tenfold gap between the two. The article calls this the zone for combining retrieval with long context.
The context window holds about 5,000 pages, but the safe threshold for dropping RAG is only about 500 pages. Source: Google Cloud (Gemini 1.5 Pro), Anthropic.

According to Google Cloud, Gemini 1.5 Pro has a context window of up to 2 million tokens, roughly 5,000 pages of text. Anthropic gives a far smaller figure: a knowledge base under 200,000 tokens, about 500 pages, can be placed in the prompt whole, without RAG.

The two numbers differ by a factor of ten. Most of the work of designing a document Q&A system for a client, and most of the arguments in the meeting room, happen inside that gap.

On a client site, the question comes up early. An IT manager who has just read about million-token context will ask why your team is spending two weeks building chunking, embeddings and a vector store. If you answer “RAG is best practice”, you have already lost. You need to answer in three terms the client understands: accuracy, latency and the bill.

A large window does not mean the model reads carefully

The strongest evidence against stuffing a whole corpus into the prompt comes from “Lost in the Middle” by Liu et al., published in TACL in 2024. The authors showed that performance drops markedly when the needed information sits in the middle of a long context, compared with when it sits at the beginning or the end.

The result holds even for models built specifically for long context.

Introl, an AI infrastructure company, draws the practical conclusion: enterprises cannot assume that “more context is better”. Say this plainly to the client, because their intuition points the other way.

Picture a 400-page insurance rulebook whose most important exclusion clause is on page 200. Put the whole book in the prompt and that clause lands exactly where the model is most likely to miss it. With RAG, the passage is retrieved and placed right next to the question, where the model uses it well.

So the first job on a client site is to measure, not to choose an architecture. Insert a distinctive factual sentence at the start, middle and end of the client’s real documents, ask about it, and count how often the model answers correctly at each position.

The bill grows faster than the page count

Even when accuracy is acceptable, cost remains a problem. IBM explains that compute requirements grow quadratically with sequence length. Introl makes the same point in terms budget holders will follow more easily: the relationship is non-linear, and longer contexts become disproportionately expensive.

A quick calculation makes this concrete. A lean RAG prompt of about 10,000 tokens and a whole-corpus prompt of 1 million tokens differ by a factor of 100 in token count. Apply the quadratic rule and the corresponding compute could differ by as much as 10,000 times for the same question.

That calculation applies the quadratic rule to the entire prompt, so treat 10,000 times as a ceiling for illustration, not the actual difference on an invoice. The figure is still useful because it shows the client that the bill does not grow with the page count. It grows faster.

Even Google, which sells the 2-million-token window, acknowledges that long-context queries tend to increase processing time and require more compute. Introl gives a more specific figure: at maximum context length, users may wait more than 2 minutes before the model starts generating text.

For clients that host their own models, such as a bank required to run on-prem, there is another barrier. According to Introl, a 70B-parameter model with a 128K context needs about 40GB of KV cache per user. Imagine 50 employees asking questions at once: the KV cache alone would need about 2,000GB of memory, before counting the model weights.

So in customer discovery, do not stop at asking how large the corpus is. Ask how many concurrent users there will be, how long they will tolerate waiting for an answer, and whether the system runs in the cloud or in the client’s own data centre.

Under 500 pages, put the whole corpus in

None of this means long context is useless. Anthropic is fairly direct: a knowledge base smaller than 200,000 tokens can go entirely into the prompt, with no RAG. What makes this viable is prompt caching, which Anthropic says cuts latency by more than 2x and cost by up to 90%.

IBM describes the general mechanism: prompt caching reduces the number of tokens processed, API costs and latency for repeated or similar requests. “Repeated” is the key condition. Caching works best when the documents at the top of the prompt rarely change while the questions change constantly.

Picture a company’s 300-page HR policy handbook, revised a few times a year and fielding hundreds of questions a day. That is the ideal case for putting the whole corpus in the prompt and caching it. Building RAG here means spending two weeks solving a problem the client does not have.

Conversely, if documents change hourly, as price lists or inventory do, the cache keeps getting invalidated and the cost advantage shrinks. Test this argument with real measurements before promising anything to the client.

Above that threshold, RAG is not dead but upgraded

Now take a corpus of 20,000 pages of contracts, four times the roughly 5,000 pages that Google’s 2-million-token window can hold. At this scale there is nothing to debate: retrieval is mandatory.

The question becomes how to retrieve better. Anthropic reports that Contextual Embeddings, a technique that attaches context from the whole document to each chunk before embedding, reduce the retrieval miss rate in the top 20 chunks by 35%.

The most overlooked detail in the whole debate is a point IBM makes: content retrieved by RAG also sits in the context window at inference time. RAG and long context are not rival camps. They share the same token budget.

That is why Introl recommends a hybrid approach: combine long context with retrieval so that critical information is reliably surfaced. A large window lets you retrieve more generously, bringing in full sections rather than fragments. Retrieval, in turn, ensures the decisive passage is not buried in the middle.

Corpus size Sensible approach Main risk First task on the client site
Under 200,000 tokens (~500 pages) Put the whole corpus in the prompt, enable prompt caching Frequently changing documents make the cache ineffective Measure how often documents are updated and how many questions arrive per day
200,000 to 2 million tokens (~500 to 5,000 pages) Hybrid: retrieve broadly, use the long window to hold full sections Lost in the middle, non-linear cost, long waits at the top of the range Run fact-insertion tests at the start, middle and end, and measure latency
Over 2 million tokens (over ~5,000 pages) RAG is mandatory; improve retrieval (e.g. Contextual Embeddings) Retrieval misses the decisive passage Build a benchmark question set to measure the top-20 miss rate

The table is only a starting point. The real boundaries depend on the number of concurrent users, latency limits and whether the client runs in the cloud or on-prem. You only learn those things after sitting down with them.

What should developers practise from this debate?

The most valuable skill here is not knowing the API of a particular vector store. It is the ability to turn the “context or RAG” question into three numbers measured on the client’s data: accuracy by position of the information, time to first token, and cost per thousand questions.

Get into the habit of counting tokens with a real tokenizer rather than estimating from page counts. The figure of roughly 400 tokens per page that both Google and Anthropic implicitly use is only an approximation, and a client’s documents may be full of tables, contract codes or scanned text. Getting this estimate wrong by an order of magnitude means choosing the wrong architecture altogether.

When reading job descriptions for FDE or AI engineer roles, look for phrases such as retrieval evaluation, latency budget and on-prem deployment. They signal that the company is stuck on exactly the problem described here.

On your CV, instead of writing “built a RAG chatbot”, say that you compared full-context with RAG on a corpus of a given number of pages, what you measured, which approach you chose and why, with your own numbers attached.

Context windows will keep growing. But the question clients pay you to answer stays the same: which tokens deserve a place in the window, and where they should sit.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
6 sources
Read next on the roadmap · Stage 3: Applied AIReasoning models: when it is worth making a client pay more and wait longerReasoning tokens never appear in the response, but they still take up context and still cost money. An FDE who turns reasoning on without measuring it is having the client pay for something nobody has checked.