Lesson 10.7 · 31 min
Document Chunking Strategies for RAG
Your RAG bot has the right document, yet it answers "I don't know". Very often the culprit is not the model or the search, but where we cut the text.
In short: RAG retrieves pieces of documents, called chunks, and hands them to an LLM. How we cut those chunks decides what can be found and whether the retrieved text makes sense on its own. Strategies range from simple fixed-size windows and sentence packing to recursive, structure-aware, semantic, contextual, small-to-big and LLM-driven (agentic) chunking; the right choice depends on the documents, the questions and the budget.
What is RAG, what is a chunk, and why chunk at all?
RAG (Retrieval-Augmented Generation) answers a question in two moves: first retrieve the passages most relevant to the question from our own documents, then generate an answer with an LLM that reads those passages. It lets an LLM answer from private or fresh data without retraining it.
A chunk is one piece of a document that we store and retrieve as a unit: a paragraph, a few sentences, a section, or a fixed number of tokens. Our running example is an online shop's policy handbook, 80 pages covering refunds, shipping, warranties and accounts.
Why not store the whole handbook as one item? Three reasons:
- Embedding models have input limits (often 512 to 8,192 tokens) and a single vector for 80 pages would be a blurry average of every topic in it.
- LLM context is limited and costly. We want to give the model the two paragraphs that matter, not 80 pages. Irrelevant text also distracts the model.
- Precision. Small, focused chunks let retrieval point at exactly the part that answers the question.
Think of it like index cards Imagine copying a textbook onto index cards for revision. Cards that are too small ("30 days.") are useless alone. Cards that are too big (a whole chapter) are hard to search and full of noise. Cards cut mid-sentence lose meaning. Good chunking is writing good index cards: one clear idea each, readable on its own.
How retrieval actually works, and what bad chunking breaks
Retrieval compares the question's vector with each chunk's vector. The retriever never sees the document as a whole, only chunks. So a chunk has two jobs: its vector must clearly represent one topic (so it can be found), and its text must contain enough to answer (so it is useful once found).
What goes wrong with bad chunks:
- Split facts: "Items can be returned within" ends one chunk and "30 days" starts the next. Neither chunk answers "how long is the return window?".
- Lost context: a chunk says "This fee is waived for members", but the chunk does not say which fee. Its vector does not mention shipping at all.
- Mixed topics: a big chunk covering refunds, shipping and passwords gets a muddy vector that matches nothing strongly.
- Too tiny: "30 days." matches poorly and gives the LLM no context.
Fixed-size, sentence and recursive chunking
Fixed-size chunking cuts every N characters or tokens (say 500 tokens), regardless of content. It is simple, fast and predictable, and every chunk fits the embedding model. But it cuts mid-sentence and mid-word, splitting facts across chunks. Overlap (repeating the last part of one chunk at the start of the next) softens this.
Chunking by sentence first splits text into sentences, then packs whole sentences into a chunk until a size limit is reached. No sentence is ever cut. It can still glue the end of one topic to the start of another, because it does not know where paragraphs or sections end.
Recursive chunking tries a list of separators from biggest to smallest: first split on blank lines (paragraphs); any piece still too big is split on sentence ends; any piece still too big on spaces; and so on. Small neighbouring pieces are packed back together up to the limit. The result respects the natural structure as much as the size limit allows. It is the default in popular frameworks (for example LangChain's RecursiveCharacterTextSplitter) and a strong baseline.
three_chunkers.py
import re
doc = ("Refunds. Items can be returned within 30 days. The refund goes to the "
"original card.\n\nShipping. Orders ship in 2 days. Express costs $9 and "
"arrives next day.")
def fixed(text, size=60, overlap=15): # slide a window of characters
step = size - overlap
return [text[i:i + size] for i in range(0, len(text) - overlap, step)]
def by_sentence(text, max_chars=80): # pack whole sentences
sents = re.split(r"(?<=[.!?])\s+", text.replace("\n\n", " "))
chunks, cur = [], ""
for s in sents:
if cur and len(cur) + len(s) + 1 > max_chars:
chunks.append(cur); cur = ""
cur = (cur + " " + s).strip()
return chunks + [cur]
def recursive(text, max_chars=80, seps=("\n\n", ". ", " ")):
if len(text) <= max_chars or not seps: # small enough (or give up)
return [text]
sep, out, cur = seps[0], [], ""
for part in text.split(sep): # biggest separator first
joined = cur + sep + part if cur else part
if len(joined) <= max_chars:
cur = joined # keep packing pieces together
continue
if cur:
out.append(cur)
if len(part) > max_chars: # piece still too big: recurse
out += recursive(part, max_chars, seps[1:]); cur = ""
else:
cur = part
return out + ([cur] if cur else [])
for name, fn in [("fixed", fixed), ("sentence", by_sentence), ("recursive", recursive)]:
print(f"--- {name}")
for c in fn(doc):
print(repr(c))Output:
--- fixed 'Refunds. Items can be returned within 30 days. The refund go' '. The refund goes to the original card.\n\nShipping. Orders sh' 'ping. Orders ship in 2 days. Express costs $9 and arrives ne' ' and arrives next day.' --- sentence 'Refunds. Items can be returned within 30 days.' 'The refund goes to the original card. Shipping. Orders ship in 2 days.' 'Express costs $9 and arrives next day.' --- recursive 'Refunds. Items can be returned within 30 days' 'The refund goes to the original card.' 'Shipping. Orders ship in 2 days. Express costs $9 and arrives next day.'
Compare the outputs. Fixed-size produces fragments like '. The refund goes to the original card.\n\nShipping. Orders sh', mixing two topics and cutting words. Sentence chunking never cuts a sentence, but its middle chunk glues the refund rule to the shipping rule. Recursive chunking keeps the shipping paragraph whole and only splits the refund paragraph (which was slightly too long) at a sentence boundary.
Document-structure and semantic chunking
Document-structure based chunking uses the document's own layout: Markdown or HTML headings, PDF sections, table boundaries, code functions, slides, FAQ question/answer pairs. Each section (or subsection) becomes a chunk, and the heading path such as "Handbook > Refunds > Damaged items" is usually stored as metadata or prepended to the text. Authors already grouped related content, so this often gives the most coherent chunks. It needs reliable parsing; messy PDFs and scanned documents make it hard.
Semantic chunking looks at meaning. Split the text into sentences, embed each one, and measure the similarity between neighbouring sentences. Where similarity drops sharply, the topic probably changed, so we cut there. For example, if consecutive-sentence similarities are 0.82, 0.79, 0.31, 0.85, we cut at the 0.31 gap. It adapts to content without needing headings, but it costs an embedding per sentence at indexing time, the cut threshold needs tuning, and results can be uneven in size.
Pause and think: A handbook converted from Markdown has clear headings for every policy. Which strategy is likely to give the most coherent chunks with the least effort?
Document-structure chunking: split on headings, keep each policy section as a chunk (splitting long ones recursively), and attach the heading path as context. The author already grouped related content.
Contextual, small-to-big and agentic chunking
Contextual chunking fixes the "lost context" problem. Before embedding a chunk, we add a short note that situates it in the whole document, for example: "From the Shipping section of the 2026 handbook; describes the express delivery fee." That note can be the title and heading path, or a sentence an LLM writes after reading the full document and the chunk. Anthropic described the LLM-written version as "Contextual Retrieval" in 2024 and reported substantially fewer retrieval failures, especially when combined with BM25 and reranking. The cost is one LLM call per chunk at indexing time (prompt caching of the shared document makes this cheaper).
Small-to-big chunking (also called parent-document or sentence-window retrieval) separates what we search from what we return. We embed small units, such as single sentences or 100-token pieces, because small units give sharp, precise vectors. Each small unit points to its bigger parent (the full paragraph or section). When a small unit matches, we hand the LLM the parent, so it gets full context. Precision of small chunks, context of big ones; the price is more bookkeeping and more tokens sent to the LLM.
Agentic chunking lets an LLM decide the boundaries. It reads the document and groups content into self-contained ideas, sometimes rewriting text into standalone statements called propositions ("The return window for unused items is 30 days."). It can produce excellent chunks for messy text, but it is slow, expensive, harder to reproduce, and the LLM might drop or alter information. It is usually reserved for small, high-value collections.
Small-to-big retrieval, step by step
- Split into parents: Cut documents into large parent chunks, e.g. sections of ~1,000 tokens.
- Split parents into children: Cut each parent into small child chunks, e.g. ~100 tokens or single sentences.
- Embed children only: Store child vectors, each with a pointer to its parent ID.
- Search children: The question is matched against the precise child vectors.
- Return parents: Swap each matching child for its parent (removing duplicates) and give those to the LLM.
Chunk overlap, and how to choose the chunk size
Chunk overlap repeats some text at the boundary: with size 500 and overlap 50 tokens, chunk 2 starts 450 tokens after chunk 1 starts. A fact sitting on a boundary then appears whole in at least one chunk. The cost is duplicated storage and embeddings (about size / (size − overlap) times as many chunks) and sometimes near-duplicate results. Overlap of roughly 10–20% is a common starting point; structure-aware methods need less, because they rarely cut mid-thought.
Choosing the chunk size is a trade-off. Small chunks (100–256 tokens) give precise vectors and suit short factual questions, but may lack context. Large chunks (512–1,024+ tokens) carry context and suit "explain" or summary questions, but their vectors are blurrier and they use more of the LLM's context. Things to consider: the embedding model's max input length, the typical answer length in your documents, how many chunks you will pass to the LLM, and its context budget. Many teams start around 300–800 tokens with modest overlap, then test.
The only reliable way to choose is to measure: build 30–100 real questions with the passage that answers each, try two or three chunking settings, and check how often the right passage appears in the top-k (recall@k) and how good the final answers are.
Comparison of all the strategies
| Strategy | How it cuts | Strength | Weakness | Indexing cost |
|---|---|---|---|---|
| Fixed-size | Every N tokens/characters | Simple, predictable | Cuts sentences and topics | Very low |
| Sentence | Pack whole sentences | Never cuts a sentence | Ignores topic/section ends | Low |
| Recursive | Paragraph → sentence → word | Respects structure, good default | Still rule-based | Low |
| Document structure | Headings, sections, tables | Coherent, author-defined units | Needs clean parsing | Low–medium |
| Semantic | Where meaning shifts | Adapts to content | Embedding per sentence; tuning | Medium |
| Contextual | Any method + added context | Fixes lost context | LLM call per chunk | Medium–high |
| Small-to-big | Search small, return big | Precision plus context | More bookkeeping and tokens | Low–medium |
| Agentic | LLM picks units | Best for messy text | Slow, costly, less reproducible | High |
Worked example, step by step
Let us size chunks for the 80-page handbook with real arithmetic. Assume about 500 tokens per page, so 40,000 tokens in total (an illustrative figure). We compare a small setting and a large one, and we pass the top 5 chunks to the LLM.
Two settings, side by side
- Option A: size 400, overlap 60: Each chunk advances 340 tokens. Chunks = ⌈(40,000 − 60) / 340⌉ = ⌈117.5⌉ = 118.
- Option B: size 1,000, overlap 100: Each chunk advances 900 tokens. Chunks = ⌈(40,000 − 100) / 900⌉ = ⌈44.3⌉ = 45.
- Context cost per question: A sends 5 × 400 = 2,000 tokens to the LLM. B sends 5 × 1,000 = 5,000 tokens, 2.5 times as much, on every single question.
- How much of the handbook the LLM sees: A shows it at most 2,000 of 40,000 tokens, or 5%. B shows 12.5%. B forgives sloppier retrieval; A needs retrieval to be right.
- Decide by test: Run both on the evaluation questions. If A finds the right passage in the top 5 as often as B does, take A: same quality, fewer tokens.
When a test question fails, the first move is always the same: open the retrieved chunks and read them. What we see usually matches one of these patterns.
| What we see | Likely cause | First thing to try |
|---|---|---|
| Right chunk retrieved, answer still incomplete | The fact is split across a boundary | Add overlap, or cut on sentences and paragraphs |
| Right chunk retrieved, LLM asks "which fee?" | The chunk lost its context | Prepend the heading path, or return the parent |
| Right chunk sits around rank 20 | Chunk too big, topics mixed | Smaller or structure-based chunks |
| Top results are near-identical text | Too much overlap, or repeated boilerplate | Lower the overlap; remove duplicates |
| Chunk matches only on its first part | Longer than the embedding model limit, so cut off | Count tokens; lower the chunk size |
Practice: try it yourself
The lesson coded fixed-size, sentence and recursive chunking. Now we build semantic chunking: compare each sentence with the next one and cut where the similarity drops. Our "embedding" is a simple stand-in, a bag of crude word stems, so two sentences are similar when they share words. A real embedding model would compare meaning, but the cutting logic is identical.
practice_semantic_chunking.py
import math
import re
from collections import Counter
sentences = [
"Refunds are available within 30 days of purchase.",
"A refund goes back to the card used for the purchase.",
"Refunds for sale items are given as store credit.",
"Shipping takes 2 days for standard orders.",
"Express shipping costs 9 dollars and orders arrive next day.",
"Passwords must be at least 12 characters long.",
"Reset a forgotten password from the login page.",
]
STOP = {"a", "the", "are", "of", "to", "for", "as", "and", "be", "at", "must", "from"}
def embed(sentence): # stand-in embedding: bag of crude word stems
words = re.findall(r"[a-z0-9]+", sentence.lower())
return Counter(w.rstrip("s") for w in words if w not in STOP)
def cosine(a, b):
dot = sum(a[w] * b[w] for w in a)
return dot / math.sqrt(sum(v * v for v in a.values()) * sum(v * v for v in b.values()))
THRESHOLD = 0.10 # cut where neighbour similarity falls below this
chunks, current = [], [sentences[0]]
for prev, nxt in zip(sentences, sentences[1:]):
sim = cosine(embed(prev), embed(nxt))
cut = sim < THRESHOLD
print(f"sim = {sim:.2f} {'CUT ' if cut else 'keep'} before: {nxt[:34]}")
if cut:
chunks.append(current)
current = []
current.append(nxt)
chunks.append(current)
for i, c in enumerate(chunks, 1):
print(f"chunk {i}: {len(c)} sentences, starts with {c[0][:28]!r}")Output:
sim = 0.33 keep before: A refund goes back to the card use sim = 0.17 keep before: Refunds for sale items are given a sim = 0.00 CUT before: Shipping takes 2 days for standard sim = 0.41 keep before: Express shipping costs 9 dollars a sim = 0.00 CUT before: Passwords must be at least 12 char sim = 0.20 keep before: Reset a forgotten password from th chunk 1: 3 sentences, starts with 'Refunds are available within' chunk 2: 2 sentences, starts with 'Shipping takes 2 days for st' chunk 3: 2 sentences, starts with 'Passwords must be at least 1'
Now change it:
- Set
THRESHOLD = 0.25. Look at the six similarities in the output and predict how many chunks we get before you run it. - Remove
"for"from theSTOPset. Sentences 3 and 4 both contain "for". Predict whether the cut between refunds and shipping survives. - Move the "Express shipping…" sentence to second place, between the two refund sentences. Predict where the cuts fall now, and what that says about text that jumps between topics.
Pause and think: The similarity between the second and third refund sentences is 0.17, only just above the threshold of 0.10. What would a slightly higher threshold do, and is that better or worse?
It would cut there, so the rule about sale items would become its own chunk, separated from the general refund rules. We would get more and smaller chunks. Whether that is better depends on the questions, which is why the threshold has to be tuned on real data and not picked by feel.
Pause and think: Take "Refunds take 5 days." followed by "The money returns to your card." They share no words. What does our stand-in do with this pair, and what would a real embedding model do?
Our stand-in gives a similarity of 0 and cuts between them, splitting one topic in two. A real embedding model would place the two sentences close together, because they mean related things, and would keep them in one chunk. Semantic chunking is only as good as the embeddings behind it.
Common mistakes and conclusion
Common mistakes Choosing a chunk size without testing. Exceeding the embedding model's input limit (text gets silently truncated). Splitting tables row by row so headers are lost. Throwing away titles and headings that would give context. Using one strategy for everything (code, tables and prose need different handling). Forgetting metadata such as source, section and date. Re-chunking documents without re-embedding them, or mixing old and new chunks.
Real-world patterns Documentation assistants often split Markdown by headings and prepend the heading path. Legal and policy RAG uses structure (articles, clauses) plus small-to-big. Code assistants chunk by function or class rather than by line count. Support bots chunk FAQs as one question-and-answer pair per chunk.
Pause and think: Our bot retrieves the chunk "This fee is waived for Plus members." for the question "Do Plus members pay for express shipping?" but the LLM says it cannot tell which fee is meant. Which strategy fixes this most directly?
Contextual chunking (or small-to-big): add the section heading or an LLM-written context such as "Shipping section: express delivery fee" to the chunk, or return the parent paragraph, so the chunk says which fee it is about.
Chunking is the quiet foundation of RAG. Retrieval can only find what our chunks express, and the LLM can only answer from what the chunks contain. Choose units that are small enough to be precise, big enough to be meaningful, aligned with the document's structure, and carrying enough context to stand alone. Then measure and iterate.
Key takeaways
- A chunk must be findable (a focused vector) and useful (enough context to answer).
- Fixed-size is simple but cuts ideas; sentence and recursive chunking respect boundaries; structure-based uses headings.
- Semantic chunking cuts where meaning shifts; contextual chunking adds situating context; small-to-big searches small and returns big.
- Overlap of about 10–20% protects boundary facts; chunk sizes of a few hundred tokens are a common start.
- There is no universal best: build a small evaluation set and measure recall and answer quality.
Key terms
- Chunk: A piece of a document stored, embedded and retrieved as one unit.
- Chunk overlap: Text repeated at the boundary between neighbouring chunks.
- Recursive chunking: Splitting by a hierarchy of separators (paragraph, sentence, word) until pieces fit a size limit.
- Semantic chunking: Cutting where embedding similarity between neighbouring sentences drops.
- Contextual chunking: Adding document-level context to each chunk before embedding it.
- Small-to-big: Searching small child chunks but returning their larger parent chunks to the LLM.
- Proposition: A short, self-contained statement of a single fact, used in agentic chunking.
← 10.6 ColBERT: Token-Level Late Interaction for Retrieval · 10.8 HyDE: Generating Hypothetical Documents to Improve RAG →