RAG Chunking Visualizer
Chunking decisions are invisible until retrieval fails. Paste a document, pick a strategy, and see every boundary and every overlapping token before your embedding model does. Companion to RAG vs. CAG vs. KAG. Runs entirely in your browser.
Token counts are estimated at 4 characters per token, the standard rule of thumb. Real tokenizers vary by about 10 to 20 percent either way; treat these as planning numbers.
What to look for
- Mid-sentence cuts. If fixed-size splitting slices a fact in half, neither chunk embeds it well and retrieval misses both. That is the whole argument for structure-aware strategies.
- Overlap is a tax, not a fix. The highlighted tokens are stored and embedded twice. If you need heavy overlap to keep answers intact, your chunk size is wrong or your documents need better structure.
- Watch the small chunks. A 15-token chunk of header text embeds as noise and pollutes nearest-neighbor results. Filter or merge them before indexing.
- Chunk size is a retrieval decision, not a storage one. Smaller chunks sharpen precision but lose context; bigger chunks blur embeddings. Tune against your own eval set, not a blog post default.
Where chunked retrieval fits against cache-augmented and knowledge-graph approaches: the essay. If your answer spans need reasoning across chunks, read RAG Is Dead. Long Live Agentic Retrieval.