What is chunking in RAG?
In a retrieval-augmented generation pipeline, chunking is the step where a source document gets split into smaller pieces before each piece is embedded and stored in a vector database. It happens before retrieval even starts, but it silently caps what retrieval can ever return: the retriever can only hand back whatever unit you indexed, so a badly chosen chunk boundary is a ceiling nothing downstream can fix (Redis — Chunking for RAG; IBM Developer — Enhancing RAG performance with chunking).
The core trade-off is size. Chunks that are too small are incomplete fragments — a sentence like "the transformer architecture was" carries no useful meaning on its own, and the model can't construct a good answer from disconnected pieces. Chunks that are too large add noise: irrelevant sentences ride along with the one that matters, diluting the match and sometimes pulling the generator off-topic.
The four common chunking strategies
Most RAG pipelines pick from four approaches, in roughly increasing order of sophistication and ingestion cost (IBM — Chunking strategies for RAG with LangChain; Adaptive Chunking: Optimizing Chunking-Method Selection for RAG):
| Strategy | How it works | Pros | Cons |
|---|---|---|---|
| Fixed-size | Split every N tokens/characters with a fixed overlap. | Simple, fast, predictable cost | Ignores structure — can cut a sentence or table row in half |
| Recursive (LangChain default) | Try paragraph breaks first, fall back to sentences, then words, then raw characters. | Keeps natural boundaries most of the time; no extra API calls | Chunk size still varies; can still split mid-idea occasionally |
| Semantic | Embed consecutive sentences and start a new chunk where similarity drops. | Chunks track topic shifts, not arbitrary length | Needs an embedding call per sentence at ingest time — slower, costs more |
| Document-aware (Markdown/HTML/code) | Split on structural units the format already provides — headings, functions, table rows. | Best fit for structured docs and code | Needs a format-specific splitter per source type |
Recursive chunking is the practical default for most teams: LangChain's RecursiveCharacterTextSplitter tries paragraph breaks first, then sentences, then words, then raw characters as a last resort, so it usually respects natural structure without needing an embedding call at ingest time (IBM Developer; Chunking Strategies for RAG — worked example). Semantic chunking scores better on topic coherence in benchmarks but adds an embedding call per sentence during ingestion, which is a real latency and cost trade-off worth measuring before defaulting to it everywhere.
Get the weekly AI engineering brief
RAG, agents, evals, and the tools worth using — one practical email a week. Plus the free roadmap PDF.
By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.
What chunk size should I use?
There's no universal number — the right size depends on your document type and query patterns — but practitioner write-ups converge on 200–500 tokens as a reasonable starting point for general prose, with roughly 10–20% overlap between adjacent chunks (Chunking Strategies for RAG; IBM Developer — chunking guidelines).
| Chunk size | Effect | Risk |
|---|---|---|
| ~50–100 tokens | Fragments — each chunk is often an incomplete thought | Retrieval returns pieces the generator can't answer from |
| ~200–500 tokens | Sweet spot for most prose — a self-contained idea per chunk | Still document-dependent; test against your own corpus |
| ~1000+ tokens | More context per chunk, but more irrelevant text riding along | Noisy retrieval dilutes precision and can push the answer off-topic |
Overlap exists because a boundary can land in the middle of a fact — without it, that fact can be split across two chunks and effectively lost to retrieval, since neither chunk alone contains it intact. Too much overlap just duplicates content across chunks, wasting storage and retrieval budget without adding coverage (Chunking Strategies for RAG — overlap section).
Chunking different document types
- • Prose / documentation. Recursive chunking on paragraph and sentence boundaries works well; semantic chunking is worth the extra cost for high-value, low-volume corpora.
- • Code. Split on function/class boundaries where possible, not fixed character counts — a function cut in half is worse than a slightly oversized chunk.
- • Tables. Fixed-size or naive splitting routinely breaks rows apart; keep each row (or a coherent group of rows) intact and consider serializing tables to a row-per-chunk format.
- • Markdown/HTML. Use the format's own structure — headings, list items — as chunk boundaries before falling back to generic splitting.
This is the case for document-aware chunking: a splitter that understands the source format's structure (Markdown headers, code AST, table rows) generally out-performs a size-only splitter on the same corpus, because it aligns chunk boundaries with the boundaries the format already gives you for free (IBM — Chunking strategies tutorial).
How to know if your chunking strategy is actually working
Chunk size and strategy are hyperparameters — treat them like any other pipeline change and measure the effect on retrieval, not just eyeball a few examples. Run the same eval set through each candidate configuration and compare context precision and context recall — the two retrieval-side metrics from RAGAS and similar frameworks — since chunking only affects the retrieval half of the pipeline, not generation. See our guide to RAG evaluation metrics for how those scores are computed and what thresholds are commonly used.
A pattern worth adopting: rebuild the vector index with two or three chunk-size candidates (e.g. 200, 400, 800 tokens) against the exact same source documents and the exact same eval questions, then pick the configuration with the best context recall at the retrieval depth you actually query with. A chunking change that looks good on paper but wasn't measured against your own corpus is a guess, not a decision.
Frequently Asked Questions
What is chunking in RAG?
Chunking is splitting a source document into smaller pieces before embedding and storing them in a vector database. It matters because retrieval can only return whatever unit you indexed — if a chunk is too small it's an incomplete fragment, and if it's too large it buries the relevant sentence in irrelevant text, so how you chunk directly caps how good retrieval can ever be.
What is the best chunk size for RAG?
There's no universal number, but 200–500 tokens with roughly 10–20% overlap is a commonly cited starting point for prose. The right size depends on document type: code, tables, and long-form prose behave differently, so treat it as a hyperparameter to tune against your own eval set rather than a fixed rule.
What is recursive chunking?
Recursive chunking (LangChain's default text splitter) tries to split on the most meaningful boundary first — paragraph breaks — and only falls back to sentences, then words, then raw characters if a chunk is still too big. It keeps natural structure most of the time without requiring extra embedding calls at ingest time.
What is semantic chunking?
Semantic chunking embeds consecutive sentences and starts a new chunk when the similarity between them drops below a threshold, so boundaries track actual topic shifts instead of a fixed length. It typically retrieves more coherent chunks than fixed-size splitting, but costs an embedding call per sentence at ingestion time, which makes it slower and more expensive to index.
Does chunk overlap matter?
Yes — without overlap, information that falls exactly on a chunk boundary can be split across two chunks and effectively lost to retrieval, since neither chunk alone contains the full idea. A common starting point is overlap equal to about 10% of the chunk size; too much overlap just duplicates content and wastes storage and retrieval budget.
How do I know if my chunking strategy is working?
Measure it, don't guess — run retrieval evaluation (context precision and context recall) against a labeled eval set for each candidate chunking configuration and compare scores, the same way you'd evaluate any other RAG pipeline change. See our guide to RAG evaluation metrics for how to set that up.
References
- • Redis — Chunking for RAG: Strategies, Tradeoffs & Common Mistakes
- • IBM Developer — Enhancing RAG Performance with Intelligent Chunking Strategies
- • IBM — Chunking Strategies for RAG Tutorial (LangChain + Granite)
- • Adaptive Chunking: Optimizing Chunking-Method Selection for RAG (arXiv)
- • Chunking Strategies for RAG — Why How You Split Your Documents Changes Everything