RAG vs Long Context
A bigger context window — the amount of text a model can take in for one request — shrinks the case for RAG on small knowledge bases. It doesn't remove the case for RAG on large ones, ones that change often, or ones where you need to show exactly which source an answer came from. If your material is small enough to paste directly into a request and stays roughly the same, long context can genuinely replace retrieval. If it isn't, retrieval is still doing real work that a bigger window doesn't do for you.
What RAG does
RAG searches your documents for the passages relevant to the current question and hands the model only those, alongside the question itself. The model never sees the whole corpus — just a small, chosen slice of it. The RAG page covers the pipeline stage by stage.
What long context does
Long context skips the search step entirely. Instead of retrieving a subset, you hand the model a much larger piece of the material directly — sometimes the whole document, sometimes the whole corpus if it's small enough — and let the model read all of it itself. There's no separate page for this because it isn't really a technique of its own; it's the absence of one, made possible by a bigger context window.
Side by side
| RAG | Long context | |
|---|---|---|
| What the model sees | A small, retrieved subset | Most or all of the source material |
| Per-query cost | Lower — only the relevant passages are sent | Higher — the same large block gets sent on every query |
| Latency | Retrieval step adds a little, but the model reads less | No retrieval step, but the model reads far more |
| Citations | Yes — each answer can point at the passage it came from | Weaker — with no retrieval step marking what mattered, tracing an answer back to one spot is harder |
| Freshness | Update the index; the next query sees the change | Repaste or re-supply the updated material every time |
| Ceiling as the corpus grows | Scales to any size — you're always sending the same small slice | Hits the window's limit eventually, no matter how large that limit is |
| Reliability across the input | Not applicable — the model only sees the chosen slice | Uneven — models use text near the start or end of a long input more reliably than text buried in the middle |
Which one your problem calls for
Long context fits when the material is genuinely small enough to paste in comfortably, doesn't change often, and you're not running enough queries for the extra tokens on every single one to add up. A single support ticket answered against one 40-page manual is a reasonable case: paste the manual, ask the question, done — building a retrieval pipeline for that is real machinery solving a problem you don't have yet.
RAG fits once the corpus is larger than comfortably fits, updates regularly, or the answer needs a traceable source — a legal research tool over millions of pages of case law that's updated weekly has no context window large enough to hold it all, and "which case does this come from" matters enough that a citation isn't optional. It also fits any high-volume system, because paying to re-read a large block of text on every query adds up in a way a one-off question never will.
Using both
They aren't strictly either/or. A common real setup still retrieves — because most corpora are too large or too dynamic for any window to hold the whole thing — but retrieves a somewhat larger, less aggressively filtered set of passages than a small-window system would need, since there's now room to be a little less precise about the cutoff. The window got bigger; the reason to be selective about what fills it didn't go away, it just moved.
In this guide
FAQ
If a model's context window is large enough to hold your entire knowledge base, is RAG definitely unnecessary?
Not automatically. Fitting is not the only cost — every query re-reading that entire knowledge base is slower and more expensive than retrieving a few relevant passages, and it's true on every single query, not just the first one. If you're answering more than a handful of questions against the same material, that cost compounds in a way a one-off use case never hits.