Chunk size tuning stopped mattering once retrieval was the bottleneck
- Goal
- Improve answer quality on a 40k-document internal knowledge base.
- What was tried
- Six weeks of chunking experiments — sizes from 256 to 2048 tokens, overlap sweeps, semantic vs. fixed splitting.
- What happened
- Every configuration landed within a few points of the others. Quality stayed flat.
- Why
- The retriever was returning plausible-but-wrong passages regardless of how they were cut. Chunking was never the constraint.
- What changed
- Measured retrieval recall separately from answer quality, then added a reranker. Recall was the actual gap.
- Lesson: measure each stage of the pipeline before tuning any single one. A flat sweep means you're tuning the wrong stage.