Chunking is how you split documents into retrievable units. The strategy you choose determines what the search system can find. Poor chunking — chunks too large, too small, or split at wrong boundaries — is one of the most common root causes of bad RAG performance.
Larger chunks preserve more context but dilute relevance — a chunk covering multiple topics scores weakly for all of them. Smaller chunks are more precise but may be too isolated to be meaningful without surrounding context.
Fixed-size — Split every N tokens. Simple, fast. Risk: may split mid-sentence. Mitigated by overlap.
Sentence-based — Split at sentence boundaries. Better coherence, no mid-sentence splits. Variable chunk sizes.
Paragraph-based — Each paragraph is a chunk. Natural semantic unit. Works well for structured documents; breaks down for poorly structured ones.
Recursive — Split at paragraphs first, then sentences, then fixed-size as fallback. Respects structure where it exists. Used by LangChain's RecursiveCharacterTextSplitter.
Semantic — Detect topic shifts by embedding sentences and finding similarity drops. Most coherent chunks. Requires two-pass processing and more compute.
Include the last N tokens of the previous chunk at the start of the next. Prevents key content from being split across boundaries and lost. Typical overlap: 10–20% of chunk size. For 512-token chunks, use 50–100 tokens of overlap.
Precise Q&A over dense technical text. Each chunk covers a narrow topic. Best when users ask specific factual questions.
Broader context needed. Answers require surrounding paragraphs. Better when queries are conceptual rather than fact-seeking.
Always evaluate empirically on real queries. No universal optimal chunk size exists.
Ask the AI assistant about chunking strategies, how to choose chunk size, or how overlap prevents boundary problems.