Skip to content
AI & Automation

RAG Chunking Strategies: How to Prepare Documents for AI Retrieval

How to chunk documents for RAG: fixed-size, recursive, structure-aware and semantic chunking, chunk size and overlap, tables, metadata, parent-document retrieval and how to test chunking choices.

Quick answer

Chunking splits documents into passages for retrieval. For most business documents, structure-aware chunking works best: split at headings and sections, keep headings attached to their text, keep tables and lists intact, and add metadata such as source, section path, date and permissions to every chunk. Use recursive splitting as a general fallback, modest overlap where splits cut ideas, and parent-document retrieval when small chunks lack context. There is no universal chunk size; test strategies against real questions and measure retrieval recall.

Where This Fits

Chunking is part of the indexing phase in RAG, before embeddings are created. Ranking retrieved chunks is covered in reranking, and document parsing at scale in enterprise RAG architecture.

Why Chunking Matters

Retrieval returns chunks, not documents. If a chunk cuts a definition in half, mixes two unrelated topics or loses the heading that gives it meaning, the right answer may never be retrieved, or may be retrieved without the context the model needs. Many RAG quality problems blamed on the model are chunking problems.

Chunking Strategies Compared

StrategyHow it splitsStrengthsWeaknesses
Fixed sizeEvery N tokens, optional overlapSimple, predictableCuts mid-sentence and mid-idea
RecursiveTries paragraphs, then sentences, then wordsGood general defaultIgnores document semantics
Structure-awareHeadings, sections, lists, tablesCoherent, citable chunksNeeds good parsing
SemanticWhere topic shiftsCoherent for narrative textExtra cost, variable sizes
Document-specificCustom rules per format (FAQs, contracts, code)Best fit for known formatsMore engineering

Structure-Aware Chunking in Practice

Parse the document into its structure first: headings, paragraphs, lists, tables, code blocks. Split at section boundaries, merge very small sections with neighbours, and split very long sections recursively. Prepend the heading path ('HR Policy > Leave > Parental leave') to each chunk's text or metadata so it carries context. FAQs become one chunk per question and answer; contracts often split by clause.

Keeping structure is the step that most improves retrieval for business documents.

Chunk Size and Overlap

Smaller chunks match questions precisely but may lack context; larger chunks carry context but dilute the embedding and cost more tokens when retrieved. Start around a paragraph to a short section, then test smaller and larger variants. Overlap of a sentence or two helps fixed-size chunking; structure-aware chunking usually needs little. Check your embedding model's input limit and avoid silently truncated chunks.

RAG answers missing information that is clearly in your documents?

ZSpace Labs can audit parsing and chunking on your corpus and test alternatives against real questions.

Start a Project

Tables, Lists and Special Content

Parsing, OCR and transcription before chunking are covered in unstructured data processing.

  • Keep small tables whole; split large tables by row groups and repeat headers
  • Add a caption or summary sentence so tables are findable by meaning
  • Keep numbered steps together where possible
  • Treat code blocks as units
  • Use OCR and layout parsing for scanned PDFs before chunking
  • Drop boilerplate such as repeated headers, footers and navigation

Metadata and Parent-Document Retrieval

Every chunk should carry metadata: source ID, title, heading path, page, date, document type, owner and access groups. Metadata supports filtering, citations and freshness. Parent-document retrieval indexes small chunks but returns the surrounding section to the model, combining precise matching with enough context. Some systems also index a short summary per document to help with broad questions.

Testing Chunking Choices

  • 1. Build a question set with the passages that answer each question
  • 2. Index the same corpus with two or three strategies
  • 3. Measure recall at k and where the right chunk ranks
  • 4. Check answer quality end to end on a sample
  • 5. Inspect failures to see whether splits, headings or tables caused them
  • 6. Pick the strategy and document it with the index version

Advantages and Limitations of Each Approach

Simple strategies are fast to build and good enough for uniform text. Structure-aware chunking gives the best results on manuals, policies and documentation but depends on parsing quality. Semantic chunking helps with long narrative text but adds cost and makes results less predictable. Whatever you choose, re-chunking means re-embedding, so test before indexing everything.

Chunking by Document Type

Document typeRecommended approach
Policies and handbooksSection-based with heading paths
FAQsOne question and answer per chunk
ContractsClause-based, keeping definitions retrievable
Product documentationSection-based; keep code blocks and steps together
Support ticketsSummarize or chunk resolution separately from conversation
Spreadsheets and tablesRow groups with headers, plus a table summary
TranscriptsTime or topic windows with speaker labels

Adding Context to Each Chunk

A chunk that says 'the limit is 30 days' is useless without knowing which policy it belongs to. Besides heading paths, some teams prepend a short, generated description of where the chunk sits in its document before embedding it, an approach Anthropic has described as contextual retrieval. It can improve retrieval for chunks that are ambiguous on their own, at the cost of an extra model call per chunk during indexing. Test whether it helps on your corpus before applying it everywhere.

Worked Example

An illustrative scenario, not a client case: a company indexes its employee handbook with fixed 500-token chunks. Questions about parental leave return chunks that start mid-policy without the heading. Switching to section-based chunks with heading paths and keeping eligibility tables intact makes the correct section the top result for most leave questions in the test set.

Common Mistakes

  • One chunk size for every document type
  • Losing headings and table structure during parsing
  • Chunks larger than the embedding model's input limit
  • No metadata for filtering and citations
  • Changing chunking without re-running evaluations

Preparing a document corpus for AI retrieval?

Talk to ZSpace Labs about RAG development and document pipelines.

Start a Project

Conclusion

Good chunking keeps meaning intact: split by structure, keep context with each chunk, handle tables carefully and test against real questions. Related: RAG guide, embeddings and reranking.

FAQ

Common questions

Splitting documents into smaller passages before embedding and indexing them, so retrieval can return the specific parts that answer a question rather than whole documents.

Related services
Relevant industries
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.