Skip to content
AI Dev Toolkit.
Esc
  • AI Token CounterCount tokens for GPT, Claude, Gemini, DeepSeek, Qwen and more.Tool
  • LLM API Cost CalculatorEstimate per-request, daily and monthly API costs.Tool
  • AI Model ComparisonCompare prices, context windows and features across models.Tool
  • AI Model Pricing PagesSpecs, real costs and cheaper alternatives for popular models.Tool
  • Context Window CheckerSee whether your text fits each model’s context window.Tool
  • Subscription vs API CalculatorFind out whether a chat plan or the API is cheaper for you.Tool
  • GPU / VRAM CalculatorCheck how much VRAM a local model needs and which GPUs fit.Tool
  • Prompt Caching CalculatorEstimate savings from prompt caching.Tool

AI glossary · Data and training

What is chunking in RAG?

Also called: text chunking, document chunking, text splitting

Definition

Chunking is splitting documents into smaller pieces, usually a few hundred tokens each, before embedding them, so that a RAG system can search, retrieve and paste in only the parts relevant to a question.

Explained

How it works

Each chunk gets its own embedding and is retrieved on its own, so the chunk is both the unit you search and the unit you pay for when it lands in a prompt. Small chunks match precise questions well but can lose the sentence that makes them useful. Large chunks keep context but blend several topics into one vector.

A fixed-size splitter has two settings: chunk size, and overlap, the number of tokens each chunk repeats from the end of the one before, so a sentence cut at a boundary still appears whole somewhere. Better splitters cut on headings and paragraphs first and fall back to a token window only for pieces that are still too long.

A 2025 study of fixed-size chunking across several datasets and embedding models (arXiv 2505.21700) found 64 to 128 tokens best for short, fact-based answers and 512 to 1,024 for questions needing wider context, and the embedding models differed too, so there is no single right size.

Example

A 10,000-token document in 512-token chunks

With 64 tokens of overlap, each new chunk starts 448 tokens after the last, so a 10,000-token handbook becomes 23 chunks and 11,408 embedded tokens: 14.1% more than the document, because 22 chunks each repeat 64 tokens. The last chunk holds only 144 tokens.

The general formula is chunks = ⌈(document − overlap) ÷ (size − overlap)⌉. We checked every row below by running a token-window chunker (gpt-tokenizer, o200k_base) over a 10,000-token text.

Splitting a 10,000-token document
Chunk sizeOverlapChunksTokens embeddedLast chunk
512 tokens0 tokens2010,000 (+0.0%)272 tokens
512 tokens64 tokens2311,408 (+14.1%)144 tokens
512 tokens128 tokens2613,200 (+32.0%)400 tokens
1,024 tokens128 tokens1211,408 (+14.1%)144 tokens

Fixed-size token windows. Structure-aware splitters produce uneven chunks, so count yours.

Cost and quality

Why it matters

Chunk size times the number of chunks you retrieve is the context you add to every question, billed as input tokens each time: five 1,024-token chunks are 5,120 tokens per question, five 256-token chunks are 1,280. Overlap matters less for cost; mostly it adds vectors to store.

Changing chunk size means re-embedding everything, so test two or three settings on real questions before you index a large collection. The RAG chunk size guide walks through that test.

Don’t mix up

Common confusions

Chunk size vs context window
The context window limits the whole prompt. Chunk size is your choice about how to cut documents, and the retrieved chunks are only part of what fills that window.
Characters vs tokens
Some splitters count characters, not tokens, so the same chunk_size setting can produce very different chunks. Embedding models limit and bill input in tokens, so check which unit your library uses.

Go deeper

Try it and read more

Related

All 40 terms in the AI glossary