RAG Chunking Strategies: How to Prepare Documents for AI Retrieval
How to chunk documents for RAG: fixed-size, recursive, structure-aware and semantic chunking, chunk size and overlap, tables, metadata, parent-document retrieval and how to test chunking choices.
Quick answer
Chunking splits documents into passages for retrieval. For most business documents, structure-aware chunking works best: split at headings and sections, keep headings attached to their text, keep tables and lists intact, and add metadata such as source, section path, date and permissions to every chunk. Use recursive splitting as a general fallback, modest overlap where splits cut ideas, and parent-document retrieval when small chunks lack context. There is no universal chunk size; test strategies against real questions and measure retrieval recall.
Where This Fits
Chunking is part of the indexing phase in RAG, before embeddings are created. Ranking retrieved chunks is covered in reranking, and document parsing at scale in enterprise RAG architecture.
Why Chunking Matters
Retrieval returns chunks, not documents. If a chunk cuts a definition in half, mixes two unrelated topics or loses the heading that gives it meaning, the right answer may never be retrieved, or may be retrieved without the context the model needs. Many RAG quality problems blamed on the model are chunking problems.
Chunking Strategies Compared
| Strategy | How it splits | Strengths | Weaknesses |
|---|---|---|---|
| Fixed size | Every N tokens, optional overlap | Simple, predictable | Cuts mid-sentence and mid-idea |
| Recursive | Tries paragraphs, then sentences, then words | Good general default | Ignores document semantics |
| Structure-aware | Headings, sections, lists, tables | Coherent, citable chunks | Needs good parsing |
| Semantic | Where topic shifts | Coherent for narrative text | Extra cost, variable sizes |
| Document-specific | Custom rules per format (FAQs, contracts, code) | Best fit for known formats | More engineering |
Structure-Aware Chunking in Practice
Parse the document into its structure first: headings, paragraphs, lists, tables, code blocks. Split at section boundaries, merge very small sections with neighbours, and split very long sections recursively. Prepend the heading path ('HR Policy > Leave > Parental leave') to each chunk's text or metadata so it carries context. FAQs become one chunk per question and answer; contracts often split by clause.
Chunk Size and Overlap
Smaller chunks match questions precisely but may lack context; larger chunks carry context but dilute the embedding and cost more tokens when retrieved. Start around a paragraph to a short section, then test smaller and larger variants. Overlap of a sentence or two helps fixed-size chunking; structure-aware chunking usually needs little. Check your embedding model's input limit and avoid silently truncated chunks.
RAG answers missing information that is clearly in your documents?
ZSpace Labs can audit parsing and chunking on your corpus and test alternatives against real questions.
Tables, Lists and Special Content
Parsing, OCR and transcription before chunking are covered in unstructured data processing.
- Keep small tables whole; split large tables by row groups and repeat headers
- Add a caption or summary sentence so tables are findable by meaning
- Keep numbered steps together where possible
- Treat code blocks as units
- Use OCR and layout parsing for scanned PDFs before chunking
- Drop boilerplate such as repeated headers, footers and navigation
Metadata and Parent-Document Retrieval
Every chunk should carry metadata: source ID, title, heading path, page, date, document type, owner and access groups. Metadata supports filtering, citations and freshness. Parent-document retrieval indexes small chunks but returns the surrounding section to the model, combining precise matching with enough context. Some systems also index a short summary per document to help with broad questions.
Testing Chunking Choices
- 1. Build a question set with the passages that answer each question
- 2. Index the same corpus with two or three strategies
- 3. Measure recall at k and where the right chunk ranks
- 4. Check answer quality end to end on a sample
- 5. Inspect failures to see whether splits, headings or tables caused them
- 6. Pick the strategy and document it with the index version
Advantages and Limitations of Each Approach
Simple strategies are fast to build and good enough for uniform text. Structure-aware chunking gives the best results on manuals, policies and documentation but depends on parsing quality. Semantic chunking helps with long narrative text but adds cost and makes results less predictable. Whatever you choose, re-chunking means re-embedding, so test before indexing everything.
Chunking by Document Type
| Document type | Recommended approach |
|---|---|
| Policies and handbooks | Section-based with heading paths |
| FAQs | One question and answer per chunk |
| Contracts | Clause-based, keeping definitions retrievable |
| Product documentation | Section-based; keep code blocks and steps together |
| Support tickets | Summarize or chunk resolution separately from conversation |
| Spreadsheets and tables | Row groups with headers, plus a table summary |
| Transcripts | Time or topic windows with speaker labels |
Adding Context to Each Chunk
A chunk that says 'the limit is 30 days' is useless without knowing which policy it belongs to. Besides heading paths, some teams prepend a short, generated description of where the chunk sits in its document before embedding it, an approach Anthropic has described as contextual retrieval. It can improve retrieval for chunks that are ambiguous on their own, at the cost of an extra model call per chunk during indexing. Test whether it helps on your corpus before applying it everywhere.
Worked Example
An illustrative scenario, not a client case: a company indexes its employee handbook with fixed 500-token chunks. Questions about parental leave return chunks that start mid-policy without the heading. Switching to section-based chunks with heading paths and keeping eligibility tables intact makes the correct section the top result for most leave questions in the test set.
Common Mistakes
- One chunk size for every document type
- Losing headings and table structure during parsing
- Chunks larger than the embedding model's input limit
- No metadata for filtering and citations
- Changing chunking without re-running evaluations
Preparing a document corpus for AI retrieval?
Talk to ZSpace Labs about RAG development and document pipelines.
Conclusion
Good chunking keeps meaning intact: split by structure, keep context with each chunk, handle tables carefully and test against real questions. Related: RAG guide, embeddings and reranking.
Common questions
Splitting documents into smaller passages before embedding and indexing them, so retrieval can return the specific parts that answer a question rather than whole documents.