Guide · Tokens & costs
RAG chunk size and overlap: how to choose the right settings
There is no single best RAG chunk size. A sensible start for prose is about 512 tokens with 10% overlap, split on paragraphs and headings, then tested against real questions from your own documents. Research shows short factual answers favour small chunks (64 to 128 tokens) and questions that need context favour larger ones (512 to 1,024).
By Tahir NazirUpdated 12 min read
On this page
- What is chunking in RAG, and why does chunk size matter?
- Token, character, sentence or semantic splitting?
- What is the best chunk size for RAG?
- How much chunk overlap should you use?
- How chunk size changes what each answer costs
- Add headings and metadata to every chunk
- How to test chunk sizes on your own data
- Questions people ask
What is chunking in RAG, and why does chunk size matter?
Chunking is cutting your documents into pieces before you embed them. In retrieval-augmented generation (RAG), each chunk gets its own embedding (a vector that represents its meaning), a question is matched against those vectors, and the best few chunks are pasted into the prompt. The chunk is the unit you search and the unit you pay for, so its size changes three things:
- What a hit contains. A small chunk is one idea, so its vector matches precise questions well, but it may lack the sentence that makes the answer usable. A large chunk carries context, but its single vector blends several topics.
- What the model sees. Retrieve five chunks of 1,024 tokens and you send about 5,000 tokens of context with every question. Five chunks of 128 tokens is about 640.
- What you pay. Embedding is charged once per token indexed. The retrieved chunks are charged as input tokens on every single question, which usually costs far more over a month.
LlamaIndex’s documentation states the trade-off plainly: smaller chunks make embeddings more precise, larger ones more general, and if you shrink chunks you may want to retrieve more of them. Its own example halves the chunk size and doubles the number retrieved.
Token, character, sentence or semantic splitting?
Split on the document’s own structure first (headings, paragraphs, sentences) and measure the result in tokens. Embedding models limit and bill input in tokens, so a size in characters only roughly predicts what you’ll pay or whether a chunk fits. The common methods:
| Method | How it cuts | Good for | Watch out for |
|---|---|---|---|
| Fixed tokens | Every N tokens, optionally with overlap | Predictable size and cost | Cuts mid-sentence and mid-table |
| Fixed characters | Every N characters | Quick prototypes | Token count varies with language and content |
| Recursive | Paragraphs first, then lines, then words, until pieces fit | General prose; the usual default | Size limit still needs to be in tokens |
| Document structure | Markdown or HTML headings, code functions, JSON keys | Docs, wikis, code | Very long sections still need a second split |
| Semantic | Embeds sentences and cuts where the meaning shifts | Text without clear structure | Every sentence is embedded at index time, on top of the chunks |
Sentence and paragraph splitters are recursive splitters with different separators. “Semantic” describes the usual approach; implementations differ.
The unit catches people out. In LangChain, chunk_size is measured by len(), which counts characters: the base splitter defaults to 4,000 characters with 200 overlap, and RecursiveCharacterTextSplitter tries the separators "\n\n", "\n", " " and "" in that order. LlamaIndex’s SentenceSplitter counts tokens and defaults to 1,024 tokens with 200 overlap. We split the 2,270-token Apache License 2.0 text both ways in LangChain:
# pip install langchain-text-splitters tiktoken
import tiktoken
from langchain_text_splitters import RecursiveCharacterTextSplitter
text = open("LICENSE-2.0.txt", encoding="utf-8").read()
enc = tiktoken.get_encoding("cl100k_base")
by_chars = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=50)
by_tokens = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
encoding_name="cl100k_base", chunk_size=512, chunk_overlap=50
)
for name, splitter in [("characters", by_chars), ("tokens", by_tokens)]:
chunks = splitter.split_text(text)
sizes = [len(enc.encode(c)) for c in chunks]
print(f"chunk_size=512 {name}: {len(chunks)} chunks, "
f"{min(sizes)}-{max(sizes)} tokens each, {sum(sizes)} tokens embedded")
# Output:
# chunk_size=512 characters: 33 chunks, 6-109 tokens each, 2217 tokens embedded
# chunk_size=512 tokens: 6 chunks, 251-477 tokens each, 2281 tokens embeddedSame number, five times as many chunks. If a tutorial says “use 512” without a unit, check which one its code uses. To turn a token budget into words for your own writers, the tokens to words converter uses measured ratios per model.
What is the best chunk size for RAG?
The best size depends on the questions, not on the embedding model’s limit. A model’s maximum input is a ceiling for one chunk, and modern limits are far above what retrieval usually needs:
| Model | Max input tokens |
|---|---|
| OpenAI text-embedding-3-small, -3-large | 8,192 |
| Google gemini-embedding-2 | 8,192 |
| Google gemini-embedding-001 | 2,048 |
| Voyage voyage-4, voyage-4-large, voyage-4-lite | 32,000 |
| Cohere embed-v4.0, embed-v5.0-pro, embed-v5.0-fast | 128k |
| Cohere embed-english-v3.0, embed-multilingual-v3.0 | 512 |
From each provider’s embeddings documentation, checked 2026-10-11. OpenAI says to count tokens for its third-generation embedding models with cl100k_base.
Squeezing 8,000 tokens into one vector is allowed, but that vector then has to represent everything in the chunk at once, and every retrieved chunk lands in your prompt. What the research says:
- It depends on the answer type. Bhat et al. (2025) tested chunks of 64 to 1,024 tokens, with no overlap, on six question-answering datasets and two embedding models. On SQuAD, with short factual answers, recall at rank 1 for the Stella model fell from about 64% with 64-token chunks to about 50% at 512. On NarrativeQA, where answers need story context, it rose from 4.2% at 64 tokens to 10.7% at 1,024. The authors conclude that 64 to 128 tokens suits concise fact-based answers and 512 to 1,024 suits questions needing broader context.
- It depends on the embedding model. In the same study the two models peaked at different sizes on some datasets (on Natural Questions, Stella at 512 tokens and Snowflake at 1,024), so a size tuned for one model may not carry over when you switch.
- It depends on the text. A 2026 study of two textbooks (Garrido-Lestache Belinchon and Garrido-Lestache Belinchon) found paragraph chunks retrieved best in a maths book and sentence chunks in a narrative one, and concluded that no single chunk size wins across kinds of text.
For scale, the influential Dense Passage Retrieval work (Karpukhin et al., 2020) split Wikipedia into disjoint 100-word passages, roughly 125 tokens of English prose at the 1.25 tokens per word we measured. Our practical reading: start near 512 tokens for mixed documentation, go smaller (128 to 256) for FAQs, glossaries and other one-fact-per-paragraph text, and larger (up to about 1,024) for narrative or argument where meaning spans paragraphs. Then measure.
How much chunk overlap should you use?
Use 10% to 20% overlap with fixed-size splitting, and little or none when you already split on paragraph or section boundaries. Overlap means each chunk repeats the last few tokens of the one before it, so a sentence cut at a boundary still appears whole in one of the two chunks:
Chunk 1tokens 1–40 · 40 tokens
Refunds. You can cancel a monthly plan at any time from the Billing page. The cancellation takes effect at the end of the current billing period, and you keep access until then. Annual plans can
Chunk 2tokens 33–72 · 40 tokens
keep access until then. Annual plans can be refunded in full within 14 days of purchase. After 14 days, we refund the unused months minus a 10% processing fee. Refunds go
Chunk 3tokens 65–90 · 26 tokens
10% processing fee. Refunds go back to the original payment method and usually arrive within 5 to 10 working days.
- Document
- 90 tokens
- Embedded
- 106 tokens in 3 chunks
- Overlap overhead
- +16 tokens (+17.8%)
o200k_base · 40-token chunks · 8-token overlap (20%) · scaled down to fit
The cost is easy to underestimate. Each chunk after the first adds the overlap again, so on long documents the extra tokens approach overlap ÷ (chunk size − overlap), which is more than the percentage you set:
| Overlap | Chunks per document | Tokens embedded | Extra | OpenAI text-embedding-3-small | OpenAI text-embedding-3-large | Google gemini-embedding-2 (text, paid tier) |
|---|---|---|---|---|---|---|
| None | 4 | 20,000,000 | – | $0.40 | $2.60 | $4.00 |
| 51 tokens (10%) | 5 | 22,040,000 | +10.2% | $0.44 | $2.87 | $4.41 |
| 102 tokens (20%) | 5 | 24,080,000 | +20.4% | $0.48 | $3.13 | $4.82 |
Embedding prices per million tokens: OpenAI text-embedding-3-small $0.02, OpenAI text-embedding-3-large $0.13, Google gemini-embedding-2 (text, paid tier) $0.20, from OpenAI’s and Google’s pricing pages on 2026-10-11. On very long documents 10% overlap tends to +11.1% and 20% to +24.9%.
So overlap barely moves the embedding bill. Its real costs are more vectors to store and search, and duplicated text in the prompt when two neighbouring chunks are both retrieved. A token-window chunker with overlap is a few lines; this one produced the figure above:
// npm install gpt-tokenizer
import { decode, encode } from "gpt-tokenizer/encoding/o200k_base";
/** Split text into windows of `size` tokens; neighbours share `overlap` tokens. */
export function chunkByTokens(text: string, size: number, overlap: number): string[] {
if (overlap >= size) throw new Error("overlap must be smaller than size");
const ids = encode(text);
const chunks: string[] = [];
for (let start = 0; start < ids.length; start += size - overlap) {
const end = Math.min(start + size, ids.length);
chunks.push(decode(ids.slice(start, end)));
if (end === ids.length) break;
}
return chunks;
}
const chunks = chunkByTokens(policy, 40, 8); // the 90-token refund policy
console.log(chunks.map((c) => encode(c).length)); // [ 40, 40, 26 ]In production, split on paragraphs and headings first and use a token window only for pieces that are still too long. Cutting at arbitrary token boundaries can also split a character in non-Latin scripts, because these tokenizers work on bytes.
How chunk size changes what each answer costs
Indexing is paid once; retrieval is paid on every question. Chunk size times the number of chunks you retrieve (top-k) is the context you add to each prompt, billed at the chat model’s input price:
| Chunk size | Context per question | GPT-5 Mini | Gemini 3.8 Flash | Claude Sonnet 5.5 |
|---|---|---|---|---|
| 256 tokens | 1,280 tokens | $32.00 | $96.00 | $256.00 |
| 512 tokens | 2,560 tokens | $64.00 | $192.00 | $512.00 |
| 1,024 tokens | 5,120 tokens | $128.00 | $384.00 | $1,024 |
Input prices per million tokens: GPT-5 Mini $0.25, Gemini 3.8 Flash $0.75, Claude Sonnet 5.5 $2 (our daily data, 2026-10-11). Retrieved chunks only: the question, system prompt and answer are extra.
Compare that with the one-off indexing bill in the overlap table: embedding the whole 20-million-token corpus with OpenAI text-embedding-3-small costs $0.40. Halving the chunk size halves the monthly context bill only if you keep top-k the same, and you may need to raise top-k to keep the same recall. That trade-off is exactly what an evaluation measures. Put your own numbers into the LLM cost calculator, and check long contexts against limits with the context windows guide.
Free toolLLM cost calculatorEnter input and output tokens per request and requests per day to see the monthly bill on every model, with today’s prices.Add headings and metadata to every chunk
A chunk cut from the middle of a document often doesn’t say what it’s about. Anthropic’s example is a chunk that reads “The company’s revenue grew by 3% over the previous quarter”, which names neither the company nor the quarter. Two fixes:
- Prefix the context before embedding. Put the document title and heading path at the top of each chunk, for example “Billing help › Refunds › Annual plans”. Dense Passage Retrieval prepended each passage’s Wikipedia title. Anthropic’s Contextual Retrieval goes further and uses a model to write 50 to 100 tokens of chunk-specific context; in its tests that cut the top-20 retrieval failure rate by 35%, by 49% when combined with contextual keyword (BM25) search, and by 67% with reranking added.
- Store metadata next to the vector. Keep the source URL, section, position and last-updated date as fields. You can then filter by product or date before ranking, cite sources in answers, and fetch the neighbouring chunks when one hit isn’t enough.
Prefixes count towards the embedding model’s input limit and your token bill: a 100-token prefix on a 512-token chunk is 19.5% more tokens to embed. Count a sample with the token counter or in code (see counting tokens in Python and JavaScript).
How to test chunk sizes on your own data
Build a small test set and compare two or three settings, so the choice rests on your own numbers rather than guesswork:
- Collect 30 to 100 real questions (from search logs, support tickets or users) and, for each, a short phrase the correct passage must contain.
- Index your documents at each candidate setting, for example 256, 512 and 1,024 tokens at 10% overlap. Changing the size means re-embedding everything.
- For each question, retrieve the top k chunks and record a hit if the answer phrase is in them. Judge by text, not chunk IDs, because chunk boundaries differ between settings.
- Record the tokens retrieved per question next to the hit rate, then run your best two settings through the full pipeline and read the answers.
# pip install tiktoken scikit-learn
import tiktoken
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
enc = tiktoken.get_encoding("cl100k_base")
def chunk(text, size, overlap):
ids = enc.encode(text)
step = size - overlap
return [enc.decode(ids[i:i + size]) for i in range(0, max(len(ids) - overlap, 1), step)]
def hit_rate(chunks, questions, k=3):
# Stand-in retriever so the harness runs offline: swap in your embedding model.
vec = TfidfVectorizer().fit(chunks)
scores = cosine_similarity(vec.transform([q for q, _ in questions]), vec.transform(chunks))
hits = 0
for (q, answer), row in zip(questions, scores):
top = row.argsort()[::-1][:k]
hits += any(answer in " ".join(chunks[i].split()) for i in top) # judge by answer text
return hits / len(questions)
text = open("LICENSE-2.0.txt", encoding="utf-8").read()
questions = [ # (question, a phrase the retrieved text must contain): six in our run
("When does a patent license terminate?", "shall terminate as of the date such litigation is filed"),
("Is the work provided with a warranty?", "WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND"),
# ...
]
for size in (64, 128, 256, 512):
chunks = chunk(text, size, overlap=size // 10)
sent = 3 * max(len(enc.encode(c)) for c in chunks) # worst case tokens sent per query
print(f"{size:>4}-token chunks: {len(chunks):>2} chunks, hit rate@3 {hit_rate(chunks, questions):.2f}, up to {sent} tokens per query")
# Output:
# 64-token chunks: 40 chunks, hit rate@3 0.33, up to 192 tokens per query
# 128-token chunks: 20 chunks, hit rate@3 0.67, up to 384 tokens per query
# 256-token chunks: 10 chunks, hit rate@3 0.67, up to 768 tokens per query
# 512-token chunks: 5 chunks, hit rate@3 0.83, up to 1536 tokens per queryThe toy run shows the trap to avoid. The 512-token setting “wins”, but its top 3 chunks are 1,536 of the document’s 2,270 tokens, so it is close to sending everything. Always read hit rate next to tokens per question, and prefer the smallest setting that reaches the hit rate you need. On a real corpus with your real embedding model the curve will look different, which is the point of measuring.
FAQ
Questions people ask
What is a good chunk size for RAG?
For mixed documentation, about 512 tokens with 10% overlap is a reasonable first setting. Research by Bhat et al. found 64 to 128 tokens best for short factual answers and 512 to 1,024 for questions that need context. Test two or three sizes on real questions from your own documents before you settle.
What chunk overlap should I use?
Between 10% and 20% of the chunk size for fixed-size splitting, for example 50 to 100 tokens on 512-token chunks. If you split on paragraphs or headings, you need little or no overlap. Remember that overlap repeats tokens: 20% overlap adds about 25% to the tokens you embed and store.
Is chunk_size measured in tokens or characters?
It depends on the library. LangChain’s text splitters count characters by default because they measure with len(); build them with from_tiktoken_encoder to count tokens. LlamaIndex’s SentenceSplitter and TokenTextSplitter count tokens. In our test, 512-character chunks of English held at most 109 tokens.
Should chunk size match the embedding model’s maximum input?
No. The maximum (8,192 tokens for OpenAI’s text-embedding-3 models and Google’s gemini-embedding-2) is a hard ceiling, not a target. One vector has to stand for everything in its chunk, so very long chunks blur together different topics, and every retrieved chunk is pasted into the prompt, so long chunks also cost more per question.
How many chunks should I retrieve?
Start with 3 to 5 and adjust with chunk size: smaller chunks usually need a larger top-k to bring back the same information. LlamaIndex’s docs give the same advice, doubling top-k when halving chunk size. Check that top-k times chunk size still fits comfortably in your model’s context window and budget.
Do long context windows make chunking unnecessary?
Not for most apps. You can paste a whole document into a model with a million-token window, but you pay for every one of those tokens on every question. Chunked retrieval sends a few thousand relevant tokens instead. For one-off analysis of a single long document, skipping RAG can be the simpler choice.
Try it
Tools from this guide
Keep reading
Related guides
- Tokens & costsHow to count tokens in Python and JavaScript (OpenAI, Claude, Gemini, Qwen)13 min read
- Tokens & costsContext windows explained: what counts, what happens at the limit, and how to check fit10 min read
- Tokens & costsHow to estimate LLM API costs: the formula, worked examples and the traps11 min read