
5 RAG Chunking Strategies in 2026: Fidelity, Cost, and Complexity
TLDR: Chunking decides what a retrieval system can find, because a chunk is the unit of retrieval. The five strategies in real pipelines are fixed-size splitting, recursive structure-aware splitting, semantic chunking, late chunking, and contextual retrieval. Each trades implementation cost against retrieval fidelity. This guide compares all five, maps each to the corpus it fits, and shows how to detect the failures each one introduces.
A RAG system is only as retrievable as its smallest unit. Whatever you split your documents into is what the index can match and what the model sees as context, so a bad splitting rule quietly caps answer quality no matter how good the embedding model or the generator is. The original RAG paper from Lewis and colleagues at Facebook AI, Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, indexed passages rather than whole documents for exactly this reason. This guide compares the five chunking strategies you will actually meet in production, from the trivial to the expensive. For the surrounding architecture, see our RAG pipeline guide, and for the deepest treatment of one strategy, our semantic chunking guide covers that technique end to end. If your corpus is the live web rather than files you own, You.com covers that case with the Web Search API, which does the retrieval for you, and chunking then applies only to the pages you extract.
What Are the Five Chunking Strategies?
Each strategy answers the same question, where to cut, with a different amount of information about the document.
- Fixed-size splitting: cut every N tokens, with optional overlap. No document understanding at all.
- Recursive structure-aware splitting: cut on headings, then paragraphs, then sentences, down to a size budget. Uses the document's own structure.
- Semantic chunking: embed sentences, then cut where embedding similarity drops, so each chunk holds one coherent idea.
- Late chunking: embed the whole document first, then chunk the long contextualized embedding vectors. Described by Jina AI researchers in Late Chunking: Contextual Chunk Embeddings for Long Texts.
- Contextual retrieval: prepend a model-generated context blurb to each chunk before embedding it, so a chunk that starts mid-thought carries its provenance. Announced by Anthropic in September 2024.
Which Strategy Fits Which Corpus?
The corpus decides the strategy, because the failure modes of each strategy are failures against a specific document shape.
- Uniform prose such as articles and reports: recursive splitting is the default. Headings already mark topic boundaries, and the method costs nothing beyond a splitter configuration.
- Heterogeneous reference material such as wiki pages and mixed-format docs: semantic chunking pays off, because topic shifts do not align with headings and similarity drops find them anyway.
- Long, coherent technical documents where entity definitions span pages: late chunking or contextual retrieval, because both preserve document-level context that per-chunk embedding destroys.
- Low-stakes prototypes: fixed-size splitting with overlap. It is wrong in known ways, it is fast, and you will replace it once you know what your retrieval misses.
One corpus does not need any of them: the live web. When the source is current web pages rather than a private corpus, the retrieval unit is the search result and the extraction is page content, so the chunking problem reduces to how much of each page you keep. Our RAG with web search guide covers that path.
What Is Late Chunking?
Late chunking inverts the usual order. Instead of splitting the document and embedding each piece, you run the embedding model over the whole document first, so every token's vector is computed in full document context, and then you pool the per-token vectors into chunk embeddings by slicing afterward.
The property this buys is context permanence: a pronoun or a bare product name in the middle of a page keeps the meaning it had in the full document, because its embedding was computed while the document was still whole. The Jina AI paper on the method reports improved retrieval on long documents relative to chunk-then-embed baselines, particularly where entities repeat and reference each other across sections.
The cost is an implementation constraint: the method requires per-token embeddings from a model you can run that way, which rules out opaque embedding APIs and raises the compute per document. Treat it as a fit for long, densely cross-referenced documents, not a default.
What Is Contextual Retrieval?
Contextual retrieval keeps the ordinary chunk-then-embed order but adds a step in the middle: a language model writes a short context blurb for each chunk, stating what the chunk is about and where it sits in the document, and that blurb is prepended to the chunk before embedding and indexing.
The numbers Anthropic published with its September 2024 announcement are the most cited evidence in this space: contextual embeddings reduced the rate at which the correct chunk was missing from the top 20 retrieved chunks by 49 percent, and adding contextual BM25 retrieval on top of the embeddings brought the reduction to 67 percent, on Anthropic's own corpus and benchmarks. Those are vendor-run evaluations on their data, so treat them as a strong directional signal rather than a guarantee, but the mechanism is sound: a chunk that carries its own provenance matches queries that never mention the chunk's literal words.
The cost is twofold. Every chunk pays a generation call before indexing, which multiplies ingestion cost, and the quality of the blurbs depends on the summarizing model. It is the strategy to reach for when retrieval misses are expensive, as in support and research corpora, and the corpus is stable enough that the indexing cost amortizes.
What Goes Wrong and How Do You Detect It?
Every chunking strategy fails by manufacturing units that cannot answer some class of question, and the detection method is the same for all of them: build a small labeled set of questions where you know which chunk should be retrieved, and measure retrieval, not just answer quality. Our RAG evaluation guide covers the metrics.
The split through a table or list. Fixed-size and naive recursive splitters cut tables in half, and both halves retrieve badly because neither carries the header row. Detection: log chunk contents whose first or last lines are table rows, and inspect a sample at ingestion time. Fix by treating tables as atomic or splitting them by row with the header repeated.
The orphaned entity. A chunk that says "the 2024 model supports it" after the split lost the sentence that named the model. This is the exact failure late chunking and contextual retrieval exist to fix, and its signature in evals is a query that names the entity explicitly and retrieves nothing. Detection: track per-query retrieval misses against a keyword baseline. A miss on an exact entity name that a raw keyword search finds means context loss, not an embedding quality problem.
The oversized chunk. A chunk longer than the embedding model's input window is silently truncated by some pipelines, and the tail is never indexed at all. Detection: log chunk token counts against the embedding model's documented limit at ingestion, and alert on any chunk near the cap. Fix by splitting at the limit with overlap, never by truncating.
The drift after a strategy change. Re-chunking the corpus with a new strategy changes every chunk ID, and any stored citations or eval baselines break silently. Detection: version the chunking configuration alongside the index, and re-run the labeled retrieval set after any change. The delta between the two runs is the real cost of the switch.
Related Guides
- Semantic Chunking: A Developer's Guide to Smarter Data
- How to Build a RAG Pipeline: A Step-by-Step Guide for Developers
- What Is RAG Evaluation? A Developer Guide to Measuring Retrieval-Augmented Generation Quality
- What Is an API for RAG? Retrieval-Augmented Generation for Modern AI Applications
- Web Search API guide
FAQ
What chunk size should I start with? There is no universal answer, because the right size depends on the density of the corpus and the embedding model's window. Start at roughly 512 tokens with 10 to 20 percent overlap, then tune against a labeled retrieval set rather than by intuition.
Does chunking matter if I use a search API instead of a vector index? Less, but not zero. With live web retrieval the search engine is the retriever, and chunking applies only to the pages you extract and re-index, where the same failure modes, such as split tables and orphaned entities, still apply.
Is contextual retrieval worth the cost? It depends on how expensive a retrieval miss is. The generation call per chunk raises ingestion cost, so it pays off on stable, high-value corpora where misses are costly, and it is overkill for prototypes and frequently re-ingested sources.
How do I know my chunking change actually helped? Re-run the same labeled retrieval set before and after the change with the index version pinned, and compare hit rates at the same cutoff, such as top 5. Answer quality alone is too noisy to attribute to chunking.
LI Test
LI Test
Share Article:
Related resources.

What Is a Legal Research API? Building Cited Legal Research Into Applications
September 16, 2026
Blog

What Is a Price Monitoring API? How to Build One With the You.com Contents API
September 2, 2026
Blog
.png)
What Is the You.com Contents API? Clean Page Content From Any URL
September 2, 2026
Blog

What Is a Product Data API? A Practical Guide for Commerce Pipelines
September 1, 2026
Blog

