Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Data and training

What is RAG (retrieval-augmented generation)?

Also called: retrieval-augmented generation, retrieval augmented generation

Definition

RAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.

Explained

How it works

The name comes from a 2020 research paper (arXiv 2005.11401) that paired a language model’s built-in knowledge, which it called parametric memory, with a searchable “dense vector index of Wikipedia”. Today the term covers any setup that fetches text at question time and hands it to a model.

A typical pipeline has two halves. Indexing runs once, and again when documents change: split them into chunks, turn each chunk into an embedding, and store the vectors with their text in a vector database. Answering runs on every question: embed the question, fetch the closest few chunks (the top-k), paste them into the prompt with an instruction to answer from them, and generate.

Retrieval doesn’t have to be vector search. Keyword search, a SQL query or a web search count too, and many systems combine keyword and vector results. The model only sees what retrieval returns, so a wrong or missing chunk leads to a wrong or missing answer.

Example

Five 512-token chunks per question

A help-centre bot retrieves the top 5 chunks of 512 tokens for each question. The question “Can I cancel my annual plan and get a refund for the unused months?” is 15 tokens, and with a 200-token instruction each prompt carries 2,775 input tokens, of which 2,560 (92%) are retrieved text. Each answer is about 300 tokens.

The table prices 10,000 questions on three models. Retrieved chunks dominate the input bill, so chunk size and top-k move cost as much as the choice of model. Halve either and that part of the bill halves, as long as answer quality holds.

10,000 questions, top-5 chunks of 512 tokens
ModelRetrieved chunksTotal with question and answer
GPT-6 Luna$2.56$4.28
Gemini 3.8 Flash$19.20$32.06
Claude Sonnet 5.5$51.20$85.50

Prices from our daily data, 2026-10-11. Question measured with o200k_base; the 200-token instruction and 300-token answer are assumptions. Embedding the documents is a separate, mostly one-off cost.

Cost and quality

Why it matters

RAG lets a model answer about private, recent or fast-changing information without retraining, and lets you show which passage an answer came from. It usually reduces hallucination about your own content but doesn’t remove it: the model can still misread a passage, and if retrieval misses the right chunk, it answers from the wrong one.

Most RAG problems are retrieval problems, so measure retrieval on its own: collect real questions, check whether the right passage is in the top-k, and only then tune the prompt or swap the model.

Don’t mix up

Common confusions

RAG vs fine-tuning
Fine-tuning changes the model’s weights to teach a task, format or style. RAG leaves the model alone and changes what it reads. For facts that change, RAG is the usual choice, because editing a document is cheaper than a training run.
RAG vs a long context window
With a large context window you can paste a whole document in, but you pay for every token on every question. RAG sends only the passages that matter. For a one-off question about one document, skipping RAG can be the simpler choice.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary