Lesson 10.13 · 25 min
Vectorless RAG: Retrieval Without Embeddings or a Vector Store
When you look something up in a 300-page manual, you don't compute embeddings: you open the table of contents, pick a chapter, then a section. Can an LLM retrieve the same way?
In short: Vectorless RAG retrieves context for an LLM without embeddings or a vector database. The best-known form turns a document into a tree, like a table of contents with section summaries, and lets an LLM reason its way down the tree to the right section; other forms use keyword search, agentic file search, SQL or simply long context. It avoids chunking and similarity problems and gives explainable retrieval paths, but costs more LLM calls per query and does not scale to huge unstructured collections as easily as vector search.
What is an LLM, and what is RAG?
A Large Language Model (LLM) is a neural network trained on huge amounts of text to predict the next token. That simple skill lets it answer questions, summarise, write code and reason step by step. But it only knows what was in its training data, up to a cutoff date, and nothing about our private documents.
RAG (Retrieval-Augmented Generation) fixes that: before the LLM answers, we retrieve relevant passages from our own documents and put them in the prompt, so the model generates an answer grounded in them. The "retrieve" part does not have to use vectors. That is the whole idea of this lesson.
Our running example: a company's employee handbook, a long, well-structured document with chapters on leave, expenses, security and so on. Employees ask things like "How many vacation days can I carry over?"
How normal vector RAG works, and its problems
This works very well in many cases, but it has known weak spots:
- Similarity is not relevance: the chunk most similar to the question is not always the one that answers it. "Vacation days carry over" might match an intro paragraph that mentions vacation many times.
- Chunking breaks context: a rule can be split across chunks, or a chunk can say "this limit" without saying which limit.
- Lost structure: chunks forget which chapter and section they came from, and cross-references ("see section 4.2") are not followed.
- Opaque: it is hard to explain why a chunk was chosen beyond "its vector was close".
- Extra infrastructure: an embedding model, a vector database, re-embedding when the model changes.
Think of it like two ways to find a recipe Vector RAG is like tearing every cookbook into single pages, shuffling them, and picking the pages that "feel" closest to "chocolate cake". Vectorless tree RAG is like opening the cookbook's contents page, going to Desserts, then Cakes, then Chocolate cake. The second keeps the book intact and you can explain every step.
What is vectorless RAG?
Vectorless RAG is any RAG system that finds the context for the LLM without embeddings and without a vector database. Instead of measuring vector similarity, it uses structure, keywords, queries, or the LLM's own reasoning to decide what to read.
The approach that made the term popular is reasoning-based tree retrieval: build a hierarchical index of a document, like a table of contents where each node has a title and a short summary, and let an LLM navigate it the way a person would. Open-source projects such as PageIndex describe themselves this way. The core claim is that for long, structured, professional documents (financial reports, contracts, manuals, regulations), retrieval by reasoning over structure finds the relevant section more reliably than retrieval by similarity of chunks. How much better it is depends on the documents and questions, so test on your own data.
Pause and think: Is a plain BM25 keyword search over the handbook "vectorless RAG"?
Yes, by definition: it retrieves context without embeddings or a vector database. In practice the term is most often used for LLM-guided navigation of a document tree, but keyword search, SQL and file-search agents are all vectorless retrieval.
How vectorless (tree-based) RAG works
Indexing and querying
- Parse the structure: Read the document's headings, sections and page ranges (from Markdown, HTML, or PDF layout) to build a tree: document → chapters → sections → subsections.
- Summarise each node: Store a title and a short summary for every node (often written by an LLM once, at indexing time). This is the "table of contents with notes".
- Show the LLM the top level: At question time, give the LLM the question plus the titles and summaries of the root's children.
- Reason and descend: The LLM picks the most promising child (and can explain why), then sees that node's children, and so on. It may backtrack or open several branches.
- Read the leaf text: When it reaches the right section, the full text of that section goes into the prompt.
- Answer with a path: The LLM answers and can cite the exact path, e.g. Handbook > Leave > Vacation.
Notice what changed: the "index" is the document's own structure, not a set of vectors; the "search algorithm" is the LLM's reasoning, not a nearest-neighbour lookup. Sections stay whole, so no chunk boundary cuts a rule in half.
An example of vectorless RAG, in code
The script below builds a tiny handbook tree and navigates it. To keep it runnable without an API key, the choose function is a stand-in for the LLM: it picks the child whose title and summary share the most words with the question. A real system sends the question and the child summaries to an LLM, which can also understand synonyms and reason about the question.
tree_navigation.py
import re
# A document turned into a tree: every node has a title and a short summary.
tree = {"title": "Employee Handbook", "summary": "all company policies", "children": [
{"title": "1 Leave", "summary": "vacation sick parental leave days", "children": [
{"title": "1.1 Vacation", "summary": "25 vacation days per year, carry over up to 5",
"text": "Staff get 25 vacation days a year. Up to 5 unused days carry over."},
{"title": "1.2 Parental leave", "summary": "parental leave weeks for new parents",
"text": "New parents get 16 weeks of paid parental leave."}]},
{"title": "2 Expenses", "summary": "travel meals expense claims receipts", "children": [
{"title": "2.1 Travel", "summary": "flights hotels travel booking rules",
"text": "Book economy for flights under 6 hours."},
{"title": "2.2 Meals", "summary": "meal allowance per day when travelling",
"text": "The meal allowance is 50 EUR per travel day."}]}]}
def words(t):
return set(re.findall(r"[a-z]+", t.lower()))
def choose(question, children):
# Stand-in for the LLM step "which section should I open next, and why?"
# Here: count shared words with each title + summary.
q = words(question)
return max(children, key=lambda c: len(q & words(c["title"] + " " + c["summary"])))
def navigate(question, node, path=()):
path = path + (node["title"],)
if "children" not in node: # reached a leaf: read it
return path, node["text"]
return navigate(question, choose(question, node["children"]), path)
for q in ["How many vacation days can I carry over?",
"What is the meal allowance when I travel?"]:
path, text = navigate(q, tree)
print(q)
print(" path:", " > ".join(path))
print(" read:", text)Output:
How many vacation days can I carry over? path: Employee Handbook > 1 Leave > 1.1 Vacation read: Staff get 25 vacation days a year. Up to 5 unused days carry over. What is the meal allowance when I travel? path: Employee Handbook > 2 Expenses > 2.2 Meals read: The meal allowance is 50 EUR per travel day.
Each question needed only two decisions (chapter, then section) and retrieved one complete, self-contained section. The output also shows why that text was chosen: the path. With a real LLM doing choose, a question like "Can I bring unused days into next year?" would also work, even though it shares no word with "carry over", because the model reasons about meaning.
Other vectorless approaches
| Approach | How it finds context | Good for |
|---|---|---|
| Keyword search (BM25) | Inverted index of words, ranked by term statistics | Exact terms, codes, names; cheap baseline |
| Tree / table-of-contents navigation | LLM reasons down a hierarchy of section summaries | Long structured documents: reports, contracts, manuals |
| Agentic file search | An agent uses tools like grep, find and open-file, iterating like a developer | Codebases, folders of text files |
| Text-to-SQL / structured queries | LLM writes a database query and reads the rows | Tables, metrics, records |
| Knowledge-graph queries | Traverse explicit entity relationships | Connection and multi-hop questions |
| Long-context stuffing | Put the whole document in the prompt | Small or medium documents that fit the context window |
These are often combined. A coding assistant might grep for a function name, open the file, follow an import, and read another file: retrieval through reasoning and simple tools, no embeddings at all. As LLM context windows have grown to hundreds of thousands of tokens or more, simply including a whole document has also become practical for many cases, though it costs more per call and models can still overlook details buried in very long inputs.
Advantages, disadvantages and the comparison
Advantages. No embedding model or vector database to run. No arbitrary chunking: sections stay intact. Retrieval follows document structure and can follow cross-references. Every decision is explainable as a path ("Leave > Vacation"). Often better on long, professional, well-structured documents where similar-sounding passages are common.
Disadvantages. Several LLM calls per query (one per level, more with backtracking), so higher latency and cost than one vector lookup. It depends on good structure: messy PDFs, chat logs or millions of short unrelated documents have no useful tree. LLM navigation can take a wrong branch early and miss the answer. Building good node summaries at indexing time also costs LLM calls. And for "find anything similar to X across a huge corpus", vector search remains far more scalable.
When to use which one
- Choose vectorless tree RAG for a modest number of long, structured, high-stakes documents (annual reports, legal contracts, technical manuals, regulations) where precision and explainability matter more than milliseconds.
- Choose vector (or hybrid) RAG for large, diverse, unstructured collections (support tickets, chat logs, web pages, millions of short documents), high query volume, or tight latency budgets.
- Choose long-context stuffing when the document is small enough to fit comfortably in the context window and query volume is low.
- Choose SQL or graph queries when the knowledge is already structured as tables or relationships.
- Combine them when needed: vectors to pick candidate documents, tree navigation to find the exact section.
Common mistakes Treating "vectorless" as automatically better: it trades cheap lookups for LLM reasoning, so measure accuracy, latency and cost on your own questions. Applying tree navigation to documents with no real structure (the tree becomes arbitrary). Writing vague node summaries, which give the LLM nothing to reason about. Forgetting a fallback (such as keyword search) when navigation reaches a dead end.
Pause and think: We have 2 million customer-support chat transcripts and need answers in under a second. Is tree-based vectorless RAG a good fit?
No. Chat transcripts have little hierarchical structure, the collection is huge, and multiple LLM navigation calls would break the latency budget. Vector or hybrid search, perhaps with a reranker, fits much better.
Worked example, step by step
What does one tree-navigation query really cost, and how often does it reach the right section? Let us count for a manual with 1,000 sections arranged as a tree with 10 children per node, so 3 levels. All token counts and success rates are illustrative.
One query, counted
- One navigation call: The LLM reads the question (about 30 tokens) and 10 child entries of about 40 tokens each (title plus summary): roughly 430 input tokens.
- Three levels: 3 calls × 430 ≈ 1,290 tokens spent on reading summaries.
- The answer call: The chosen section, say 1,500 tokens, goes into the final prompt. Total: about 2,800 input tokens and 4 LLM calls.
- Vector RAG for comparison: 5 chunks × 400 tokens = 2,000 tokens and 1 LLM call. The token totals are close. The real difference is 4 calls in a row instead of 1, which shows up as latency.
- Risk of a wrong turn: If each choice is right 95% of the time, all three are right 0.95³ ≈ 0.86 of the time. One wrong turn at the top loses the answer.
- Widen the search: Keeping the best 2 children at every level opens 1 + 2 + 4 = 7 nodes instead of 3. More calls, but a wrong first choice is no longer fatal.
| Strategy | Nodes opened (LLM calls) | Survives one wrong turn? |
|---|---|---|
| Greedy: best child only | 3 | No |
| Keep the best 2 at each level | 7 | Yes, if the right branch was the second choice |
| Greedy, then backtrack on a bad leaf | 3, more only when needed | Yes, but it needs a check that the leaf answers the question |
To find wrong turns in a running system, log the path for every question. Take the questions that were answered badly and see where their paths left the correct route. If most of them go wrong at level 1, the chapter summaries are too vague. The fix is then to rewrite those summaries so each one lists what its children cover, not to change the model.
Practice: try it yourself
The earlier code always followed the single best child. Now we add a beam: at each level we keep the best beam children instead of one, then pick the best leaf at the end. We ask a question that sends the greedy walk into the wrong chapter and watch a beam of 2 recover. As before, word overlap stands in for the judgement of the LLM.
practice_beam_navigation.py
import re
tree = {"title": "Handbook", "summary": "all policies", "children": [
{"title": "1 Leave", "summary": "vacation days sick days parental leave", "children": [
{"title": "1.1 Vacation", "summary": "vacation days per year and carry over"},
{"title": "1.2 Sick leave", "summary": "sick days and doctor notes"}]},
{"title": "2 Expenses", "summary": "claims receipts travel meals", "children": [
{"title": "2.1 Travel", "summary": "flights hotels booking"},
{"title": "2.2 Training", "summary": "paid training days and conference fees"}]}]}
def score(question, node): # stand-in for the LLM judging one node
words = lambda t: set(re.findall(r"[a-z]+", t.lower()))
return len(words(question) & words(node["title"] + " " + node["summary"]))
def navigate(question, beam):
frontier, calls = [(tree, ("Handbook",))], 0
while "children" in frontier[0][0]:
nxt = []
for node, path in frontier: # one LLM call per opened node
calls += 1
ranked = sorted(node["children"], key=lambda c: -score(question, c))
nxt += [(c, path + (c["title"],)) for c in ranked[:beam]]
frontier = nxt
leaf, path = max(frontier, key=lambda f: score(question, f[0]))
return " > ".join(path), score(question, leaf), calls
question = "How many paid training days do I get?"
for beam in [1, 2]:
path, s, calls = navigate(question, beam)
print(f"beam={beam}: {path} (leaf score {s}, {calls} LLM calls)")Output:
beam=1: Handbook > 1 Leave > 1.1 Vacation (leaf score 1, 2 LLM calls) beam=2: Handbook > 2 Expenses > 2.2 Training (leaf score 3, 3 LLM calls)
Now change it:
- Add the words
paid trainingto the summary of "2 Expenses". Predict the path that beam=1 takes now, and how many calls it needs. - Change the question to
"How many sick days do I get?". Predict whether beam=1 and beam=2 end on the same leaf, and what the extra call of beam=2 bought us. - Add
3to the list of beams. Every node has only two children. Predict the number of LLM calls for beam=3.
Pause and think: With beam=1 the walk went into "1 Leave" because the question shares the word "days" with its summary. A real LLM does not count words. Could it still take this wrong turn, and what in the index would cause it?
Yes. At the top level the LLM only sees the two chapter summaries, and the Expenses summary says "claims receipts travel meals" with no hint of training. Nothing tells it that training days are filed there, while "Leave" sounds like the natural home for a question about days off. The cause is an incomplete summary, and the fix is to make each summary cover what its children contain.
Pause and think: Here beam=2 cost 3 calls instead of 2. In a tree with 3 levels and 10 children per node, would a beam of 2 also cost just 1.5 times as much as greedy?
No. Greedy opens 3 nodes. A beam of 2 opens 1, then 2, then 4 nodes: 7 calls. The number of open nodes doubles at each level, so the extra cost grows with depth. Real systems keep the beam small, or keep only the best 2 nodes overall at each level (1 + 2 + 2 = 5 calls).
How this lesson connects
How this lesson connects Vectorless tree navigation is a form of agentic retrieval: an LLM decides step by step where to look. Compare it with the agentic RAG lesson (agents choosing tools) and the GraphRAG lesson (retrieval over explicit structure).
Key takeaways
- RAG only needs a way to retrieve relevant context; it does not have to use embeddings.
- Vector RAG can confuse similarity with relevance, break context with chunking, and is hard to explain.
- Tree-based vectorless RAG lets an LLM navigate a document's structure, like a table of contents, to the right section.
- Other vectorless options: BM25, agentic file search, SQL, graph queries and long-context stuffing.
- Vectorless suits long, structured documents with relaxed latency; vector search suits huge, unstructured, high-traffic collections.
Key terms
- Vectorless RAG: RAG that retrieves context without embeddings or a vector database.
- Tree index: A hierarchy of document sections, each with a title and summary, used for navigation.
- Reasoning-based retrieval: Retrieval where an LLM decides step by step which part of the content to read.
- Long-context stuffing: Putting a whole document directly into the LLM prompt instead of retrieving parts.
- Text-to-SQL: An LLM writes a database query from a natural-language question and reads the result.
← 10.12 GraphRAG: Combining Knowledge Graphs with Retrieval · 11.1 AI Agents: Autonomous Decision-Making Systems →