AI glossary · Data and training
What is chunking in RAG?
Also called: text chunking, document chunking, text splitting
Definition
Chunking is splitting documents into smaller pieces, usually a few hundred tokens each, before embedding them, so that a RAG system can search, retrieve and paste in only the parts relevant to a question.
Explained
How it works
Each chunk gets its own embedding and is retrieved on its own, so the chunk is both the unit you search and the unit you pay for when it lands in a prompt. Small chunks match precise questions well but can lose the sentence that makes them useful. Large chunks keep context but blend several topics into one vector.
A fixed-size splitter has two settings: chunk size, and overlap, the number of tokens each chunk repeats from the end of the one before, so a sentence cut at a boundary still appears whole somewhere. Better splitters cut on headings and paragraphs first and fall back to a token window only for pieces that are still too long.
A 2025 study of fixed-size chunking across several datasets and embedding models (arXiv 2505.21700) found 64 to 128 tokens best for short, fact-based answers and 512 to 1,024 for questions needing wider context, and the embedding models differed too, so there is no single right size.
Example
A 10,000-token document in 512-token chunks
With 64 tokens of overlap, each new chunk starts 448 tokens after the last, so a 10,000-token handbook becomes 23 chunks and 11,408 embedded tokens: 14.1% more than the document, because 22 chunks each repeat 64 tokens. The last chunk holds only 144 tokens.
The general formula is chunks = ⌈(document − overlap) ÷ (size − overlap)⌉. We checked every row below by running a token-window chunker (gpt-tokenizer, o200k_base) over a 10,000-token text.
| Chunk size | Overlap | Chunks | Tokens embedded | Last chunk |
|---|---|---|---|---|
| 512 tokens | 0 tokens | 20 | 10,000 (+0.0%) | 272 tokens |
| 512 tokens | 64 tokens | 23 | 11,408 (+14.1%) | 144 tokens |
| 512 tokens | 128 tokens | 26 | 13,200 (+32.0%) | 400 tokens |
| 1,024 tokens | 128 tokens | 12 | 11,408 (+14.1%) | 144 tokens |
Fixed-size token windows. Structure-aware splitters produce uneven chunks, so count yours.
Cost and quality
Why it matters
Chunk size times the number of chunks you retrieve is the context you add to every question, billed as input tokens each time: five 1,024-token chunks are 5,120 tokens per question, five 256-token chunks are 1,280. Overlap matters less for cost; mostly it adds vectors to store.
Changing chunk size means re-embedding everything, so test two or three settings on real questions before you index a large collection. The RAG chunk size guide walks through that test.
Don’t mix up
Common confusions
- Chunk size vs context window
- The context window limits the whole prompt. Chunk size is your choice about how to cut documents, and the retrieved chunks are only part of what fills that window.
- Characters vs tokens
- Some splitters count characters, not tokens, so the same
chunk_sizesetting can produce very different chunks. Embedding models limit and bill input in tokens, so check which unit your library uses.
Go deeper
Try it and read more
- Free toolAI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.
- Free toolTokens to Words ConverterConvert between tokens, words and pages.
- Free toolLLM API Cost CalculatorEstimate per-request, daily and monthly API costs.
- Guide · 12 min readRAG chunk size and overlap: how to choose the right settingsHow to choose a RAG chunk size and overlap: token vs character splitters, embedding limits, what research shows, overlap cost maths and a test you can run.
Related
Related terms
- RAGRAG (retrieval-augmented generation) is a technique where an application searches your own documents for passages relevant to a question and adds them to the prompt, so the model answers from that text rather than from memory alone.
- EmbeddingsEmbeddings are lists of numbers (vectors) that an embedding model produces from text, images or other data, placed so that inputs with similar meaning end up close together and can be compared mathematically.
- Vector databaseA vector database is a data store that keeps embeddings alongside their source data and quickly finds the stored vectors closest to a query vector, which is how semantic search and RAG fetch relevant text.
- TokenA token in AI is the unit of text a language model reads and writes, usually a whole word, part of a word or a punctuation mark, which the model sees only as a number from its vocabulary.
- Context windowA context window is the maximum number of tokens a language model can work with in one request, counting the system prompt, tool definitions, conversation history, documents and the reply it writes.