RAG & Retrieval

RAG Chunking Strategies: Fixed-Size vs Recursive vs Semantic (And How to Pick a Chunk Size)

Chunking caps what your retriever can ever find — split a document too small and you get incomplete fragments; too large and irrelevant text dilutes the match. This guide covers the four common chunking strategies (fixed-size, recursive, semantic, document-aware), what chunk size and overlap to start with, how each choice shows up in retrieval evaluation scores, and the mistakes that quietly cap a RAG pipeline's ceiling before generation even runs.

Gurram Poorna Prudhvi

Lead AI Engineer

Technical Guide
Sep 28, 2026
10 min read
RAG CHUNKING STRATEGIEShow you split a document decides what your retriever can ever findSourcedocumentFixed-sizeN tokens/chars + overlap — fast, ignores structureRecursiveparagraph → sentence → word fallback — LangChain defaultSemanticsplits where embedding similarity drops — costs embedding callsChunks →Vector storeembedded & indexedRetrievertop-k searchChunk size trade-offToo small: fragmentsSweet spot: 200–500 tokToo large: noisy contextOptimal size and overlap depend on document type — prose, code, and tables behave differently.

What is chunking in RAG?

In a retrieval-augmented generation pipeline, chunking is the step where a source document gets split into smaller pieces before each piece is embedded and stored in a vector database. It happens before retrieval even starts, but it silently caps what retrieval can ever return: the retriever can only hand back whatever unit you indexed, so a badly chosen chunk boundary is a ceiling nothing downstream can fix (Redis — Chunking for RAG; IBM Developer — Enhancing RAG performance with chunking).

The core trade-off is size. Chunks that are too small are incomplete fragments — a sentence like "the transformer architecture was" carries no useful meaning on its own, and the model can't construct a good answer from disconnected pieces. Chunks that are too large add noise: irrelevant sentences ride along with the one that matters, diluting the match and sometimes pulling the generator off-topic.

The four common chunking strategies

Most RAG pipelines pick from four approaches, in roughly increasing order of sophistication and ingestion cost (IBM — Chunking strategies for RAG with LangChain; Adaptive Chunking: Optimizing Chunking-Method Selection for RAG):

StrategyHow it worksProsCons
Fixed-sizeSplit every N tokens/characters with a fixed overlap.Simple, fast, predictable costIgnores structure — can cut a sentence or table row in half
Recursive (LangChain default)Try paragraph breaks first, fall back to sentences, then words, then raw characters.Keeps natural boundaries most of the time; no extra API callsChunk size still varies; can still split mid-idea occasionally
SemanticEmbed consecutive sentences and start a new chunk where similarity drops.Chunks track topic shifts, not arbitrary lengthNeeds an embedding call per sentence at ingest time — slower, costs more
Document-aware (Markdown/HTML/code)Split on structural units the format already provides — headings, functions, table rows.Best fit for structured docs and codeNeeds a format-specific splitter per source type

Recursive chunking is the practical default for most teams: LangChain's RecursiveCharacterTextSplitter tries paragraph breaks first, then sentences, then words, then raw characters as a last resort, so it usually respects natural structure without needing an embedding call at ingest time (IBM Developer; Chunking Strategies for RAG — worked example). Semantic chunking scores better on topic coherence in benchmarks but adds an embedding call per sentence during ingestion, which is a real latency and cost trade-off worth measuring before defaulting to it everywhere.

Get the weekly AI engineering brief

RAG, agents, evals, and the tools worth using — one practical email a week. Plus the free roadmap PDF.

By subscribing you agree to receive emails from AI Engineer Insights. Unsubscribe anytime. See our Privacy Policy.

What chunk size should I use?

There's no universal number — the right size depends on your document type and query patterns — but practitioner write-ups converge on 200–500 tokens as a reasonable starting point for general prose, with roughly 10–20% overlap between adjacent chunks (Chunking Strategies for RAG; IBM Developer — chunking guidelines).

Chunk sizeEffectRisk
~50–100 tokensFragments — each chunk is often an incomplete thoughtRetrieval returns pieces the generator can't answer from
~200–500 tokensSweet spot for most prose — a self-contained idea per chunkStill document-dependent; test against your own corpus
~1000+ tokensMore context per chunk, but more irrelevant text riding alongNoisy retrieval dilutes precision and can push the answer off-topic

Overlap exists because a boundary can land in the middle of a fact — without it, that fact can be split across two chunks and effectively lost to retrieval, since neither chunk alone contains it intact. Too much overlap just duplicates content across chunks, wasting storage and retrieval budget without adding coverage (Chunking Strategies for RAG — overlap section).

Chunking different document types

  • • Prose / documentation. Recursive chunking on paragraph and sentence boundaries works well; semantic chunking is worth the extra cost for high-value, low-volume corpora.
  • • Code. Split on function/class boundaries where possible, not fixed character counts — a function cut in half is worse than a slightly oversized chunk.
  • • Tables. Fixed-size or naive splitting routinely breaks rows apart; keep each row (or a coherent group of rows) intact and consider serializing tables to a row-per-chunk format.
  • • Markdown/HTML. Use the format's own structure — headings, list items — as chunk boundaries before falling back to generic splitting.

This is the case for document-aware chunking: a splitter that understands the source format's structure (Markdown headers, code AST, table rows) generally out-performs a size-only splitter on the same corpus, because it aligns chunk boundaries with the boundaries the format already gives you for free (IBM — Chunking strategies tutorial).

How to know if your chunking strategy is actually working

Chunk size and strategy are hyperparameters — treat them like any other pipeline change and measure the effect on retrieval, not just eyeball a few examples. Run the same eval set through each candidate configuration and compare context precision and context recall — the two retrieval-side metrics from RAGAS and similar frameworks — since chunking only affects the retrieval half of the pipeline, not generation. See our guide to RAG evaluation metrics for how those scores are computed and what thresholds are commonly used.

A pattern worth adopting: rebuild the vector index with two or three chunk-size candidates (e.g. 200, 400, 800 tokens) against the exact same source documents and the exact same eval questions, then pick the configuration with the best context recall at the retrieval depth you actually query with. A chunking change that looks good on paper but wasn't measured against your own corpus is a guess, not a decision.

Frequently Asked Questions

What is chunking in RAG?

Chunking is splitting a source document into smaller pieces before embedding and storing them in a vector database. It matters because retrieval can only return whatever unit you indexed — if a chunk is too small it's an incomplete fragment, and if it's too large it buries the relevant sentence in irrelevant text, so how you chunk directly caps how good retrieval can ever be.

What is the best chunk size for RAG?

There's no universal number, but 200–500 tokens with roughly 10–20% overlap is a commonly cited starting point for prose. The right size depends on document type: code, tables, and long-form prose behave differently, so treat it as a hyperparameter to tune against your own eval set rather than a fixed rule.

What is recursive chunking?

Recursive chunking (LangChain's default text splitter) tries to split on the most meaningful boundary first — paragraph breaks — and only falls back to sentences, then words, then raw characters if a chunk is still too big. It keeps natural structure most of the time without requiring extra embedding calls at ingest time.

What is semantic chunking?

Semantic chunking embeds consecutive sentences and starts a new chunk when the similarity between them drops below a threshold, so boundaries track actual topic shifts instead of a fixed length. It typically retrieves more coherent chunks than fixed-size splitting, but costs an embedding call per sentence at ingestion time, which makes it slower and more expensive to index.

Does chunk overlap matter?

Yes — without overlap, information that falls exactly on a chunk boundary can be split across two chunks and effectively lost to retrieval, since neither chunk alone contains the full idea. A common starting point is overlap equal to about 10% of the chunk size; too much overlap just duplicates content and wastes storage and retrieval budget.

How do I know if my chunking strategy is working?

Measure it, don't guess — run retrieval evaluation (context precision and context recall) against a labeled eval set for each candidate chunking configuration and compare scores, the same way you'd evaluate any other RAG pipeline change. See our guide to RAG evaluation metrics for how to set that up.

Want 1:1 help? Book a session

5.0 · 18 reviews

Career guidance, resume & interview prep, or tech consulting — 1:1 with a Lead AI Engineer. Start with a free 30-min quick chat.

References

Found this useful? Share it.

Share:

Related Articles