When I shipped the Chunking Visualizer, I wrote that chunk size should be tuned against your own eval set, not a blog post default. Fair to ask what the defaults are even worth, then. So I measured them. This is the first of what I intend to be quarterly data reports: a real experiment, real numbers, and a script you can run yourself. No vendor sponsored it and no result was removed for being inconvenient.

The Setup

Three chunking strategies, identical budgets, head to head:

Corpus and questions: 20 articles from the SQuAD v1.1 dev set, each article's paragraphs concatenated into one document, and 800 of their questions sampled evenly (seed pinned). Embeddings: all-MiniLM-L6-v2, cosine similarity, retrieval scoped per document. The metric is answer recall@k: does any gold answer string appear in the top-k retrieved chunks, for k of 1, 3, and 5. Two passes: one with zero overlap, one with the common 30-token overlap request. The full script is published here; it runs on a laptop CPU in a few minutes.

Result One: Structure Is Free

The zero-overlap pass is the clean comparison, because all three strategies embed exactly the same corpus: 161,426 tokens, zero duplication, identical cost.

Strategy (no overlap)recall@1recall@3recall@5chunksembedded tokens
Fixed-size61.1%80.1%87.1%818161,426
Sentence-aware66.9%85.6%90.5%925161,426
Recursive70.8%88.1%91.4%1,010161,426

Recursive beats fixed-size by 9.6 points of recall@1 while embedding the same tokens for the same money. There is no tradeoff in this table. Respecting document structure is a pure win on prose, and if you are shipping fixed-size chunking because a tutorial did, this is the cost: roughly one in ten top-answer retrievals, free.

The mechanism is exactly the one you would guess. Fixed-size cuts slice facts in half, the half-fact embeds as mush, and the right chunk loses a similarity contest it should have won. Sentence awareness fixes the worst of it; paragraph awareness fixes most of the rest.

Result Two: Overlap Lies About Its Size

The second pass requested a standard 30-token overlap, about 15% of chunk size. Here is what each strategy actually did with that request:

Strategy (30-tok overlap)recall@1gain vs no overlapduplicationembedded tokenscost vs no overlap
Fixed-size65.8%+4.6 pts17.5%189,596+17%
Sentence-aware70.3%+3.4 pts44.6%233,473+45%
Recursive73.3%+2.5 pts78.3%287,882+78%

Look at the duplication column. Fixed-size overlap behaves as advertised: ask for 15%, pay about 17%. The structure-aware strategies do not, because their overlap is measured in whole spans. To honor a 30-token overlap request, the packer carries back the last complete sentence or paragraph, and a complete paragraph is rarely 30 tokens. The recursive strategy duplicated 78% of the corpus to deliver a requested 15% overlap. You asked for a safety margin and bought a second copy of your index.

And the recall it bought shrinks as the strategy gets smarter: 4.6 points for fixed, 2.5 for recursive. Which makes sense. Overlap exists to paper over bad cuts, and structure-aware strategies make fewer bad cuts, so there is less to paper over. The strategy that most needs overlap is the one you should not be using, and the strategy you should be using barely needs it.

Result Three: By k=5, Nobody Cares

At recall@5 the three strategies land at 87.1%, 90.5%, and 91.4% without overlap, and within about 1.3 points of each other with it. If your pipeline retrieves five or more chunks and reranks, chunking strategy is not your bottleneck; your reranker is. At k=1, the regime every token-constrained and latency-constrained system actually lives in, chunking is one of the biggest levers on the board.

What I Would Do With These Numbers

Caveats, Because Data Without Caveats Is Marketing

This is one corpus of Wikipedia-style prose with a small embedding model and a substring-match recall metric. Tables, code, and legal numbering will behave differently and generally favor format-aware splitters. Answer-string recall is a generous metric that ignores whether the chunk gives enough context to actually answer. Token counts are the same 4-characters-per-token estimate the visualizer uses. And a single benchmark on public data tells you where to start tuning, not where to stop: the point of publishing the script is that you can swap in your own corpus in about ten lines.

FAQ

What is the best chunking strategy for RAG?
On prose, recursive splitting won here by 9.6 points of recall@1 over fixed-size at identical embedding cost. Structure-aware is the right default; format-aware splitters take over for tables and code.

Does chunk overlap improve retrieval?
Modestly: 2.5 to 4.6 points of recall@1 in this test. The catch is span-granularity overlap overshooting its budget, up to 78% duplication for a 15% request. Cap overlap in characters or skip it and retrieve one more chunk instead.

Does the strategy matter at higher k?
By recall@5 the spread collapses to about 2 points. Retrieval depth and reranking dominate past k=3; chunking dominates at k=1.

Can I reproduce this?
Yes. One Python file, public dataset, pinned seed, laptop CPU, a few minutes. If you get different numbers on your own corpus, that is the benchmark working as intended.