When I shipped the Chunking Visualizer, I wrote that chunk size should be tuned against your own eval set, not a blog post default. Fair to ask what the defaults are even worth, then. So I measured them. This is the first of what I intend to be quarterly data reports: a real experiment, real numbers, and a script you can run yourself. No vendor sponsored it and no result was removed for being inconvenient.
The Setup
Three chunking strategies, identical budgets, head to head:
- Fixed-size: cut every 200 tokens regardless of structure.
- Sentence-aware: pack whole sentences until the 200-token budget runs out.
- Recursive: paragraph boundaries first, sentences when a paragraph is too big, raw splits as a last resort. The strategy most frameworks ship as a default.
Corpus and questions: 20 articles from the SQuAD v1.1 dev set, each article's paragraphs concatenated into one document, and 800 of their questions sampled evenly (seed pinned). Embeddings: all-MiniLM-L6-v2, cosine similarity, retrieval scoped per document. The metric is answer recall@k: does any gold answer string appear in the top-k retrieved chunks, for k of 1, 3, and 5. Two passes: one with zero overlap, one with the common 30-token overlap request. The full script is published here; it runs on a laptop CPU in a few minutes.
Result One: Structure Is Free
The zero-overlap pass is the clean comparison, because all three strategies embed exactly the same corpus: 161,426 tokens, zero duplication, identical cost.
| Strategy (no overlap) | recall@1 | recall@3 | recall@5 | chunks | embedded tokens |
|---|---|---|---|---|---|
| Fixed-size | 61.1% | 80.1% | 87.1% | 818 | 161,426 |
| Sentence-aware | 66.9% | 85.6% | 90.5% | 925 | 161,426 |
| Recursive | 70.8% | 88.1% | 91.4% | 1,010 | 161,426 |
Recursive beats fixed-size by 9.6 points of recall@1 while embedding the same tokens for the same money. There is no tradeoff in this table. Respecting document structure is a pure win on prose, and if you are shipping fixed-size chunking because a tutorial did, this is the cost: roughly one in ten top-answer retrievals, free.
The mechanism is exactly the one you would guess. Fixed-size cuts slice facts in half, the half-fact embeds as mush, and the right chunk loses a similarity contest it should have won. Sentence awareness fixes the worst of it; paragraph awareness fixes most of the rest.
Result Two: Overlap Lies About Its Size
The second pass requested a standard 30-token overlap, about 15% of chunk size. Here is what each strategy actually did with that request:
| Strategy (30-tok overlap) | recall@1 | gain vs no overlap | duplication | embedded tokens | cost vs no overlap |
|---|---|---|---|---|---|
| Fixed-size | 65.8% | +4.6 pts | 17.5% | 189,596 | +17% |
| Sentence-aware | 70.3% | +3.4 pts | 44.6% | 233,473 | +45% |
| Recursive | 73.3% | +2.5 pts | 78.3% | 287,882 | +78% |
Look at the duplication column. Fixed-size overlap behaves as advertised: ask for 15%, pay about 17%. The structure-aware strategies do not, because their overlap is measured in whole spans. To honor a 30-token overlap request, the packer carries back the last complete sentence or paragraph, and a complete paragraph is rarely 30 tokens. The recursive strategy duplicated 78% of the corpus to deliver a requested 15% overlap. You asked for a safety margin and bought a second copy of your index.
And the recall it bought shrinks as the strategy gets smarter: 4.6 points for fixed, 2.5 for recursive. Which makes sense. Overlap exists to paper over bad cuts, and structure-aware strategies make fewer bad cuts, so there is less to paper over. The strategy that most needs overlap is the one you should not be using, and the strategy you should be using barely needs it.
Result Three: By k=5, Nobody Cares
At recall@5 the three strategies land at 87.1%, 90.5%, and 91.4% without overlap, and within about 1.3 points of each other with it. If your pipeline retrieves five or more chunks and reranks, chunking strategy is not your bottleneck; your reranker is. At k=1, the regime every token-constrained and latency-constrained system actually lives in, chunking is one of the biggest levers on the board.
What I Would Do With These Numbers
- Default to recursive, skip overlap, spend the savings on k. Recursive at k=3 with no overlap (88.1%) beats fixed at k=3 with overlap (85.0%) and embeds 15% fewer tokens.
- If you use overlap with a structure-aware splitter, cap it in characters, not spans. Trim the carried-back span to the budget. Nobody's overlap intent is “one full paragraph, whatever that costs.”
- Audit what your framework's splitter actually does. The 78% number hides inside a default most people have never printed. The visualizer now shows a duplication stat for exactly this reason.
Caveats, Because Data Without Caveats Is Marketing
This is one corpus of Wikipedia-style prose with a small embedding model and a substring-match recall metric. Tables, code, and legal numbering will behave differently and generally favor format-aware splitters. Answer-string recall is a generous metric that ignores whether the chunk gives enough context to actually answer. Token counts are the same 4-characters-per-token estimate the visualizer uses. And a single benchmark on public data tells you where to start tuning, not where to stop: the point of publishing the script is that you can swap in your own corpus in about ten lines.
FAQ
What is the best chunking strategy for RAG?
On prose, recursive splitting won here by 9.6 points of recall@1 over fixed-size at identical embedding cost. Structure-aware is the right default; format-aware splitters take over for tables and code.
Does chunk overlap improve retrieval?
Modestly: 2.5 to 4.6 points of recall@1 in this test. The catch is span-granularity overlap overshooting its budget, up to 78% duplication for a 15% request. Cap overlap in characters or skip it and retrieve one more chunk instead.
Does the strategy matter at higher k?
By recall@5 the spread collapses to about 2 points. Retrieval depth and reranking dominate past k=3; chunking dominates at k=1.
Can I reproduce this?
Yes. One Python file, public dataset, pinned seed, laptop CPU, a few minutes. If you get different numbers on your own corpus, that is the benchmark working as intended.