AI glossary · Data and training
What is RAG (retrieval-augmented generation)?
Also called: retrieval-augmented generation, retrieval augmented generation
Definition
RAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
Explained
How it works
The name comes from a 2020 research paper (arXiv 2005.11401) that paired a language model’s built-in knowledge, which it called parametric memory, with a searchable “dense vector index of Wikipedia”. Today the term covers any setup that fetches text at question time and hands it to a model.
A typical pipeline has two halves. Indexing runs once, and again when documents change: split them into chunks, turn each chunk into an embedding, and store the vectors with their text in a vector database. Answering runs on every question: embed the question, fetch the closest few chunks (the top-k), paste them into the prompt with an instruction to answer from them, and generate.
Retrieval doesn’t have to be vector search. Keyword search, a SQL query or a web search count too, and many systems combine keyword and vector results. The model only sees what retrieval returns, so a wrong or missing chunk leads to a wrong or missing answer.
Example
Five 512-token chunks per question
A help-centre bot retrieves the top 5 chunks of 512 tokens for each question. The question “Can I cancel my annual plan and get a refund for the unused months?” is 15 tokens, and with a 200-token instruction each prompt carries 2,775 input tokens, of which 2,560 (92%) are retrieved text. Each answer is about 300 tokens.
The table prices 10,000 questions on three models. Retrieved chunks dominate the input bill, so chunk size and top-k move cost as much as the choice of model. Halve either and that part of the bill halves, as long as answer quality holds.
| Model | Retrieved chunks | Total with question and answer |
|---|---|---|
| GPT-6 Luna | $2.56 | $4.28 |
| Gemini 3.8 Flash | $19.20 | $32.06 |
| Claude Sonnet 5.5 | $51.20 | $85.50 |
Prices from our daily data, 2026-10-11. Question measured with o200k_base; the 200-token instruction and 300-token answer are assumptions. Embedding the documents is a separate, mostly one-off cost.
Cost and quality
Why it matters
RAG lets a model answer about private, recent or fast-changing information without retraining, and lets you show which passage an answer came from. It usually reduces hallucination about your own content but doesn’t remove it: the model can still misread a passage, and if retrieval misses the right chunk, it answers from the wrong one.
Most RAG problems are retrieval problems, so measure retrieval on its own: collect real questions, check whether the right passage is in the top-k, and only then tune the prompt or swap the model.
Don’t mix up
Common confusions
- RAG vs fine-tuning
- Fine-tuning changes the model’s weights to teach a task, format or style. RAG leaves the model alone and changes what it reads. For facts that change, RAG is the usual choice, because editing a document is cheaper than a training run.
- RAG vs a long context window
- With a large context window you can paste a whole document in, but you pay for every token on every question. RAG sends only the passages that matter. For a one-off question about one document, skipping RAG can be the simpler choice.
Go deeper
Try it and read more
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolContext Window CheckerSee whether your text fits each model’s context window.
- Guide · 12 min readRAG chunk size and overlap: how to choose the right settingsHow to choose a RAG chunk size and overlap: token vs character splitters, embedding limits, what research shows, overlap cost maths and a test you can run.
Related
Related terms
- EmbeddingsEmbeddings are lists of numbers (vectors) that an embedding model produces from text, images or other data, placed so that inputs with similar meaning end up close together and can be compared mathematically.
- ChunkingChunking is splitting documents into smaller pieces, usually a few hundred tokens each, before embedding them, so that a RAG system can search, retrieve and paste in only the parts relevant to a question.
- Vector databaseA vector database is a data store that keeps embeddings alongside their source data and quickly finds the stored vectors closest to a query vector, which is how semantic search and RAG fetch relevant text.
- Fine-tuningFine-tuning is training an existing model further on a set of your own example inputs and ideal outputs, so its weights change and it follows a task, format or style without long instructions in every prompt.
- HallucinationA hallucination is a confident, plausible-sounding statement from a language model that is false or unsupported by its input, such as an invented citation, date, function or quote.