Chunking Benchmark Lab
Most chunking advice is vibes. This is measured: 800 questions from SQuAD v1.1 against three splitters, embedded with all-MiniLM-L6-v2, scored on answer recall@k. The data below is the real output of a published, seeded script anyone can re-run. Explore it, then put your own corpus through the same chunkers and metrics, entirely in your browser.
# Loading benchmark data...
The measured results
Your corpus, same metrics
Paste a representative document. The three chunkers below are the same algorithms the benchmark ran (fixed windows, sentence-aware packing, recursive paragraph → sentence → word splitting), so the structural metrics transfer: chunk counts, sizes, duplication from overlap, and what indexing it will cost in embedded tokens. Recall does not transfer without your own labeled questions, and anyone who tells you otherwise is selling something.
Methodology: 20 SQuAD v1.1 dev articles, paragraphs concatenated per article; 800 questions (40 per article, seed 7); chunks approximated at 4 characters per token; retrieval is cosine similarity over all-MiniLM-L6-v2 embeddings; recall@k scores whether any gold answer string appears in a retrieved chunk. The lab's browser chunkers mirror the script's splitting logic; embedding cost uses your editable price (default is a common small-embedding rate; check your provider). One corpus, one embedder, honest limits: treat the findings as a strong prior to test on your own data, not a law of nature. Everything on this page runs client-side. The engine is open source, tested, with a CLI: github.com/AlexRyan92/rag-chunking-benchmark.