Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

Guide · Tokens & costs

RAG chunk size and overlap: how to choose the right settings

There is no single best RAG chunk size. A sensible start for prose is about 512 tokens with 10% overlap, split on paragraphs and headings, then tested against real questions from your own documents. Research shows short factual answers favour small chunks (64 to 128 tokens) and questions that need context favour larger ones (512 to 1,024).

By Tahir NazirUpdated 12 min read

On this page
  1. What is chunking in RAG, and why does chunk size matter?
  2. Token, character, sentence or semantic splitting?
  3. What is the best chunk size for RAG?
  4. How much chunk overlap should you use?
  5. How chunk size changes what each answer costs
  6. Add headings and metadata to every chunk
  7. How to test chunk sizes on your own data
  8. Questions people ask

What is chunking in RAG, and why does chunk size matter?

Chunking is cutting your documents into pieces before you embed them. In retrieval-augmented generation (RAG), each chunk gets its own embedding (a vector that represents its meaning), a question is matched against those vectors, and the best few chunks are pasted into the prompt. The chunk is the unit you search and the unit you pay for, so its size changes three things:

  • What a hit contains. A small chunk is one idea, so its vector matches precise questions well, but it may lack the sentence that makes the answer usable. A large chunk carries context, but its single vector blends several topics.
  • What the model sees. Retrieve five chunks of 1,024 tokens and you send about 5,000 tokens of context with every question. Five chunks of 128 tokens is about 640.
  • What you pay. Embedding is charged once per token indexed. The retrieved chunks are charged as input tokens on every single question, which usually costs far more over a month.

LlamaIndex’s documentation states the trade-off plainly: smaller chunks make embeddings more precise, larger ones more general, and if you shrink chunks you may want to retrieve more of them. Its own example halves the chunk size and doubles the number retrieved.

Token, character, sentence or semantic splitting?

Split on the document’s own structure first (headings, paragraphs, sentences) and measure the result in tokens. Embedding models limit and bill input in tokens, so a size in characters only roughly predicts what you’ll pay or whether a chunk fits. The common methods:

Ways to split text for RAG
MethodHow it cutsGood forWatch out for
Fixed tokensEvery N tokens, optionally with overlapPredictable size and costCuts mid-sentence and mid-table
Fixed charactersEvery N charactersQuick prototypesToken count varies with language and content
RecursiveParagraphs first, then lines, then words, until pieces fitGeneral prose; the usual defaultSize limit still needs to be in tokens
Document structureMarkdown or HTML headings, code functions, JSON keysDocs, wikis, codeVery long sections still need a second split
SemanticEmbeds sentences and cuts where the meaning shiftsText without clear structureEvery sentence is embedded at index time, on top of the chunks

Sentence and paragraph splitters are recursive splitters with different separators. “Semantic” describes the usual approach; implementations differ.

The unit catches people out. In LangChain, chunk_size is measured by len(), which counts characters: the base splitter defaults to 4,000 characters with 200 overlap, and RecursiveCharacterTextSplitter tries the separators "\n\n", "\n", " " and "" in that order. LlamaIndex’s SentenceSplitter counts tokens and defaults to 1,024 tokens with 200 overlap. We split the 2,270-token Apache License 2.0 text both ways in LangChain:

split_demo.py
# pip install langchain-text-splitters tiktoken
import tiktoken
from langchain_text_splitters import RecursiveCharacterTextSplitter

text = open("LICENSE-2.0.txt", encoding="utf-8").read()
enc = tiktoken.get_encoding("cl100k_base")

by_chars = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=50)
by_tokens = RecursiveCharacterTextSplitter.from_tiktoken_encoder(
    encoding_name="cl100k_base", chunk_size=512, chunk_overlap=50
)
for name, splitter in [("characters", by_chars), ("tokens", by_tokens)]:
    chunks = splitter.split_text(text)
    sizes = [len(enc.encode(c)) for c in chunks]
    print(f"chunk_size=512 {name}: {len(chunks)} chunks, "
          f"{min(sizes)}-{max(sizes)} tokens each, {sum(sizes)} tokens embedded")

# Output:
# chunk_size=512 characters: 33 chunks, 6-109 tokens each, 2217 tokens embedded
# chunk_size=512 tokens: 6 chunks, 251-477 tokens each, 2281 tokens embedded

Same number, five times as many chunks. If a tutorial says “use 512” without a unit, check which one its code uses. To turn a token budget into words for your own writers, the tokens to words converter uses measured ratios per model.

What is the best chunk size for RAG?

The best size depends on the questions, not on the embedding model’s limit. A model’s maximum input is a ceiling for one chunk, and modern limits are far above what retrieval usually needs:

Maximum input per text for popular embedding models
ModelMax input tokens
OpenAI text-embedding-3-small, -3-large8,192
Google gemini-embedding-28,192
Google gemini-embedding-0012,048
Voyage voyage-4, voyage-4-large, voyage-4-lite32,000
Cohere embed-v4.0, embed-v5.0-pro, embed-v5.0-fast128k
Cohere embed-english-v3.0, embed-multilingual-v3.0512

From each provider’s embeddings documentation, checked 2026-10-11. OpenAI says to count tokens for its third-generation embedding models with cl100k_base.

Squeezing 8,000 tokens into one vector is allowed, but that vector then has to represent everything in the chunk at once, and every retrieved chunk lands in your prompt. What the research says:

  • It depends on the answer type. Bhat et al. (2025) tested chunks of 64 to 1,024 tokens, with no overlap, on six question-answering datasets and two embedding models. On SQuAD, with short factual answers, recall at rank 1 for the Stella model fell from about 64% with 64-token chunks to about 50% at 512. On NarrativeQA, where answers need story context, it rose from 4.2% at 64 tokens to 10.7% at 1,024. The authors conclude that 64 to 128 tokens suits concise fact-based answers and 512 to 1,024 suits questions needing broader context.
  • It depends on the embedding model. In the same study the two models peaked at different sizes on some datasets (on Natural Questions, Stella at 512 tokens and Snowflake at 1,024), so a size tuned for one model may not carry over when you switch.
  • It depends on the text. A 2026 study of two textbooks (Garrido-Lestache Belinchon and Garrido-Lestache Belinchon) found paragraph chunks retrieved best in a maths book and sentence chunks in a narrative one, and concluded that no single chunk size wins across kinds of text.

For scale, the influential Dense Passage Retrieval work (Karpukhin et al., 2020) split Wikipedia into disjoint 100-word passages, roughly 125 tokens of English prose at the 1.25 tokens per word we measured. Our practical reading: start near 512 tokens for mixed documentation, go smaller (128 to 256) for FAQs, glossaries and other one-fact-per-paragraph text, and larger (up to about 1,024) for narrative or argument where meaning spans paragraphs. Then measure.

How much chunk overlap should you use?

Use 10% to 20% overlap with fixed-size splitting, and little or none when you already split on paragraph or section boundaries. Overlap means each chunk repeats the last few tokens of the one before it, so a sentence cut at a boundary still appears whole in one of the two chunks:

The first chunk ends mid-sentence at “Annual plans can”. Because the second chunk repeats those 8 tokens, the full refund rule appears in one chunk. The price is 16 duplicated tokens.

The cost is easy to underestimate. Each chunk after the first adds the overlap again, so on long documents the extra tokens approach overlap ÷ (chunk size − overlap), which is more than the percentage you set:

What overlap adds when you index 10,000 documents of 2,000 tokens with 512-token chunks
OverlapChunks per documentTokens embeddedExtraOpenAI text-embedding-3-smallOpenAI text-embedding-3-largeGoogle gemini-embedding-2 (text, paid tier)
None420,000,000–$0.40$2.60$4.00
51 tokens (10%)522,040,000+10.2%$0.44$2.87$4.41
102 tokens (20%)524,080,000+20.4%$0.48$3.13$4.82

Embedding prices per million tokens: OpenAI text-embedding-3-small $0.02, OpenAI text-embedding-3-large $0.13, Google gemini-embedding-2 (text, paid tier) $0.20, from OpenAI’s and Google’s pricing pages on 2026-10-11. On very long documents 10% overlap tends to +11.1% and 20% to +24.9%.

So overlap barely moves the embedding bill. Its real costs are more vectors to store and search, and duplicated text in the prompt when two neighbouring chunks are both retrieved. A token-window chunker with overlap is a few lines; this one produced the figure above:

chunk.ts
// npm install gpt-tokenizer
import { decode, encode } from "gpt-tokenizer/encoding/o200k_base";

/** Split text into windows of `size` tokens; neighbours share `overlap` tokens. */
export function chunkByTokens(text: string, size: number, overlap: number): string[] {
  if (overlap >= size) throw new Error("overlap must be smaller than size");
  const ids = encode(text);
  const chunks: string[] = [];
  for (let start = 0; start < ids.length; start += size - overlap) {
    const end = Math.min(start + size, ids.length);
    chunks.push(decode(ids.slice(start, end)));
    if (end === ids.length) break;
  }
  return chunks;
}

const chunks = chunkByTokens(policy, 40, 8); // the 90-token refund policy
console.log(chunks.map((c) => encode(c).length)); // [ 40, 40, 26 ]

In production, split on paragraphs and headings first and use a token window only for pieces that are still too long. Cutting at arbitrary token boundaries can also split a character in non-Latin scripts, because these tokenizers work on bytes.

How chunk size changes what each answer costs

Indexing is paid once; retrieval is paid on every question. Chunk size times the number of chunks you retrieve (top-k) is the context you add to each prompt, billed at the chat model’s input price:

Retrieved context for 100,000 questions a month, top-5 chunks
Chunk sizeContext per questionGPT-5 MiniGemini 3.8 FlashClaude Sonnet 5.5
256 tokens1,280 tokens$32.00$96.00$256.00
512 tokens2,560 tokens$64.00$192.00$512.00
1,024 tokens5,120 tokens$128.00$384.00$1,024

Input prices per million tokens: GPT-5 Mini $0.25, Gemini 3.8 Flash $0.75, Claude Sonnet 5.5 $2 (our daily data, 2026-10-11). Retrieved chunks only: the question, system prompt and answer are extra.

Compare that with the one-off indexing bill in the overlap table: embedding the whole 20-million-token corpus with OpenAI text-embedding-3-small costs $0.40. Halving the chunk size halves the monthly context bill only if you keep top-k the same, and you may need to raise top-k to keep the same recall. That trade-off is exactly what an evaluation measures. Put your own numbers into the LLM cost calculator, and check long contexts against limits with the context windows guide.

Free toolLLM cost calculatorEnter input and output tokens per request and requests per day to see the monthly bill on every model, with today’s prices.

Add headings and metadata to every chunk

A chunk cut from the middle of a document often doesn’t say what it’s about. Anthropic’s example is a chunk that reads “The company’s revenue grew by 3% over the previous quarter”, which names neither the company nor the quarter. Two fixes:

  1. Prefix the context before embedding. Put the document title and heading path at the top of each chunk, for example “Billing help › Refunds › Annual plans”. Dense Passage Retrieval prepended each passage’s Wikipedia title. Anthropic’s Contextual Retrieval goes further and uses a model to write 50 to 100 tokens of chunk-specific context; in its tests that cut the top-20 retrieval failure rate by 35%, by 49% when combined with contextual keyword (BM25) search, and by 67% with reranking added.
  2. Store metadata next to the vector. Keep the source URL, section, position and last-updated date as fields. You can then filter by product or date before ranking, cite sources in answers, and fetch the neighbouring chunks when one hit isn’t enough.

Prefixes count towards the embedding model’s input limit and your token bill: a 100-token prefix on a 512-token chunk is 19.5% more tokens to embed. Count a sample with the token counter or in code (see counting tokens in Python and JavaScript).

How to test chunk sizes on your own data

Build a small test set and compare two or three settings, so the choice rests on your own numbers rather than guesswork:

  1. Collect 30 to 100 real questions (from search logs, support tickets or users) and, for each, a short phrase the correct passage must contain.
  2. Index your documents at each candidate setting, for example 256, 512 and 1,024 tokens at 10% overlap. Changing the size means re-embedding everything.
  3. For each question, retrieve the top k chunks and record a hit if the answer phrase is in them. Judge by text, not chunk IDs, because chunk boundaries differ between settings.
  4. Record the tokens retrieved per question next to the hit rate, then run your best two settings through the full pipeline and read the answers.
eval_chunks.py
# pip install tiktoken scikit-learn
import tiktoken
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity

enc = tiktoken.get_encoding("cl100k_base")

def chunk(text, size, overlap):
    ids = enc.encode(text)
    step = size - overlap
    return [enc.decode(ids[i:i + size]) for i in range(0, max(len(ids) - overlap, 1), step)]

def hit_rate(chunks, questions, k=3):
    # Stand-in retriever so the harness runs offline: swap in your embedding model.
    vec = TfidfVectorizer().fit(chunks)
    scores = cosine_similarity(vec.transform([q for q, _ in questions]), vec.transform(chunks))
    hits = 0
    for (q, answer), row in zip(questions, scores):
        top = row.argsort()[::-1][:k]
        hits += any(answer in " ".join(chunks[i].split()) for i in top)  # judge by answer text
    return hits / len(questions)

text = open("LICENSE-2.0.txt", encoding="utf-8").read()
questions = [  # (question, a phrase the retrieved text must contain): six in our run
    ("When does a patent license terminate?", "shall terminate as of the date such litigation is filed"),
    ("Is the work provided with a warranty?", "WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND"),
    # ...
]
for size in (64, 128, 256, 512):
    chunks = chunk(text, size, overlap=size // 10)
    sent = 3 * max(len(enc.encode(c)) for c in chunks)  # worst case tokens sent per query
    print(f"{size:>4}-token chunks: {len(chunks):>2} chunks, hit rate@3 {hit_rate(chunks, questions):.2f}, up to {sent} tokens per query")

# Output:
#   64-token chunks: 40 chunks, hit rate@3 0.33, up to 192 tokens per query
#  128-token chunks: 20 chunks, hit rate@3 0.67, up to 384 tokens per query
#  256-token chunks: 10 chunks, hit rate@3 0.67, up to 768 tokens per query
#  512-token chunks:  5 chunks, hit rate@3 0.83, up to 1536 tokens per query

The toy run shows the trap to avoid. The 512-token setting “wins”, but its top 3 chunks are 1,536 of the document’s 2,270 tokens, so it is close to sending everything. Always read hit rate next to tokens per question, and prefer the smallest setting that reaches the hit rate you need. On a real corpus with your real embedding model the curve will look different, which is the point of measuring.

FAQ

Questions people ask

What is a good chunk size for RAG?

For mixed documentation, about 512 tokens with 10% overlap is a reasonable first setting. Research by Bhat et al. found 64 to 128 tokens best for short factual answers and 512 to 1,024 for questions that need context. Test two or three sizes on real questions from your own documents before you settle.

What chunk overlap should I use?

Between 10% and 20% of the chunk size for fixed-size splitting, for example 50 to 100 tokens on 512-token chunks. If you split on paragraphs or headings, you need little or no overlap. Remember that overlap repeats tokens: 20% overlap adds about 25% to the tokens you embed and store.

Is chunk_size measured in tokens or characters?

It depends on the library. LangChain’s text splitters count characters by default because they measure with len(); build them with from_tiktoken_encoder to count tokens. LlamaIndex’s SentenceSplitter and TokenTextSplitter count tokens. In our test, 512-character chunks of English held at most 109 tokens.

Should chunk size match the embedding model’s maximum input?

No. The maximum (8,192 tokens for OpenAI’s text-embedding-3 models and Google’s gemini-embedding-2) is a hard ceiling, not a target. One vector has to stand for everything in its chunk, so very long chunks blur together different topics, and every retrieved chunk is pasted into the prompt, so long chunks also cost more per question.

How many chunks should I retrieve?

Start with 3 to 5 and adjust with chunk size: smaller chunks usually need a larger top-k to bring back the same information. LlamaIndex’s docs give the same advice, doubling top-k when halving chunk size. Check that top-k times chunk size still fits comfortably in your model’s context window and budget.

Do long context windows make chunking unnecessary?

Not for most apps. You can paste a whole document into a model with a million-token window, but you pay for every one of those tokens on every question. Chunked retrieval sends a few thousand relevant tokens instead. For one-off analysis of a single long document, skipping RAG can be the simpler choice.

Try it

Tools from this guide

Keep reading