Lesson 10.11 · 27 min
Agentic RAG: Dynamic Retrieval with Multi-Step Reasoning
Standard RAG searches once and hopes for the best. What if the system could notice the first search was not enough, and decide what to look up next?
In short: Agentic RAG puts an LLM agent in charge of retrieval. Instead of one fixed "retrieve then answer" step, the agent reasons about the question, chooses which tool or source to query, reads the results, judges whether they are enough, and loops (rewriting queries, searching again, calling other tools) until it can answer. It handles multi-step and multi-source questions far better than standard RAG, at the cost of more latency, more tokens and less predictability.
The big picture
Our running example: a customer-support assistant for a laptop shop. A customer, Priya, asks: "Is the laptop I ordered still under warranty?" To answer, the assistant must (1) find out which laptop Priya ordered and when (that is in the orders database, not the documents), (2) find the warranty length for that product line (in the policy documents), and (3) compare dates. No single search returns all of this.
Standard RAG would embed the question, fetch the top 5 policy chunks about "warranty", and ask the LLM to answer. It might find a general warranty page but it never learns which laptop Priya bought. Agentic RAG lets the model work like a researcher: plan, look something up, read, decide what is still missing, look that up, and only then answer.
Think of it like a librarian vs a book chute Standard RAG is a book chute: you post a question and it drops out the five books whose titles look closest. Agentic RAG is a librarian: they ask themselves what you really need, check the catalogue, notice the first book only half answers it, walk to another section, maybe phone the records office, and come back with a complete answer.
A quick recap of RAG and of AI agents
RAG (Retrieval-Augmented Generation) combines search with generation. Documents are split into chunks, embedded and stored in a vector database. At question time, the question is embedded, the most similar chunks are retrieved, and an LLM writes an answer grounded in those chunks. It is a fixed, one-pass pipeline: retrieve once, generate once.
An AI agent is an LLM that runs in a loop and can take actions. In each turn it reads the conversation so far, reasons about what to do next, calls a tool (a function the developer exposes, such as a search API, a database query or a calculator), observes the tool's result, and repeats until it decides it is done. The pattern of interleaving reasoning and actions was popularised as ReAct (Yao et al., 2022).
Why standard RAG falls short
- One shot: if the first retrieval misses, there is no second chance. The LLM answers from whatever came back, or hallucinates.
- Multi-hop questions: "Which of our suppliers is based in the same city as our biggest customer?" needs one lookup to feed the next. A single search cannot chain.
- One source: the pipeline searches one index. Real answers often need documents and a database and a web search or an API.
- No judgement: it cannot tell that retrieved chunks are irrelevant or outdated, so it never decides to search again or to say "I could not find this".
- Same effort for every question: "hi" and a hard comparison question both trigger the same retrieval.
Pause and think: For "Is the laptop I ordered still under warranty?", which piece of information can standard RAG over policy documents never find, no matter how good the retriever is?
Which laptop Priya ordered and when. That lives in the orders database, not in the documents. Only a system that can call another tool, here an order lookup, can get it.
What is agentic RAG, and the agentic RAG loop
Agentic RAG is RAG where an LLM agent controls the retrieval process. Retrieval becomes a set of tools the agent can call as many times as it needs, with queries it writes itself. The agent decides whether to retrieve, where to retrieve from, what to search for, and when it has enough to answer.
One pass through the loop, in detail
- Understand the goal: The agent restates what a complete answer needs, sometimes splitting the question into sub-questions.
- Choose a tool and a query: Pick the source most likely to hold the next missing fact, and write a focused query for it.
- Read the result: Extract the useful facts; ignore irrelevant chunks.
- Check sufficiency: If facts are missing or contradictory, loop again; a step limit stops runaway loops.
- Answer with sources: When the notes cover the question, write the final answer and cite where each fact came from.
The three building blocks
Most agentic RAG systems are built from three parts:
- The reasoning model (the agent's "brain"): an LLM that plans, chooses tools, writes search queries, judges results and decides when to stop. It is given a system prompt describing its job and the tools it may use.
- Tools (retrievers and actions): functions with a name, a description and typed parameters, such as
search_policies(query),lookup_order(customer_id),web_search(query),run_sql(query). The LLM reads the descriptions to decide which to call, and the application actually executes them. - Memory / state: the running record of the task: the conversation, tool calls and their results, and extracted notes. Short-term memory lives in the context window; longer-term memory can be stored in a database and retrieved too.
Around these sit guardrails: a maximum number of steps, timeouts, permission checks on tools (read-only vs write), and cost limits. Many teams use frameworks such as LangGraph, LlamaIndex agents or provider tool-calling APIs to wire this together, but the loop itself is simple enough to write by hand.
A walkthrough with a real example, in code
Here is the warranty question as a runnable loop. Two tools: a keyword document search and an order lookup. The planner function is a hard-coded stand-in for the LLM's decisions so the script runs without an API; in a real system each planner call would be one LLM turn that returns a tool call.
agentic_rag_loop.py
from datetime import date
DOCS = ["Warranty length depends on the product line; see each line's page.",
"ProBook laptops come with a 2-year warranty from the purchase date.",
"AirLite laptops come with a 1-year warranty from the purchase date."]
ORDERS = {"priya": {"item": "ProBook 14", "bought": date(2025, 3, 10)}}
def search_docs(query): # tool 1: keyword retriever
words = set(query.lower().split())
return max(DOCS, key=lambda d: len(words & set(d.lower().split())))
def lookup_order(customer): # tool 2: orders database
return ORDERS[customer]
def planner(question, notes):
# Stand-in for the LLM's decision; a real agent asks the model each turn.
if not notes:
return ("search_docs", "warranty length")
if "order" not in notes:
return ("lookup_order", "priya")
if "years" not in notes:
line = notes["order"]["item"].split()[0]
return ("search_docs", f"{line} laptops warranty")
return ("answer", None)
question = "Is the laptop Priya ordered still under warranty?"
notes, today = {}, date(2026, 10, 1)
for step in range(1, 6): # hard cap on loop iterations
action, arg = planner(question, notes)
if action == "answer":
bought = notes["order"]["bought"]
end = bought.replace(year=bought.year + notes["years"])
print(f"step {step}: ANSWER -> covered until {end}: {'yes' if today <= end else 'no'}")
break
result = search_docs(arg) if action == "search_docs" else lookup_order(arg)
print(f"step {step}: {action}({arg!r}) -> {result}")
if action == "lookup_order":
notes["order"] = result
elif "2-year" in result or "1-year" in result:
notes["years"] = 2 if "2-year" in result else 1
else:
notes["hint"] = result # not enough yet: keep lookingOutput:
step 1: search_docs('warranty length') -> Warranty length depends on the product line; see each line's page.
step 2: lookup_order('priya') -> {'item': 'ProBook 14', 'bought': datetime.date(2025, 3, 10)}
step 3: search_docs('ProBook laptops warranty') -> ProBook laptops come with a 2-year warranty from the purchase date.
step 4: ANSWER -> covered until 2027-03-10: yesLook at the trace. Step 1 retrieves a general page that says "it depends on the product line". Standard RAG would stop here with a vague answer. The agent recognises the gap, looks up the order in step 2 (ProBook 14, bought 2025-03-10), then writes a new, more specific query in step 3 ("ProBook laptops warranty") built from what it just learned. Step 4 combines facts from two sources: covered until 2027-03-10, so yes.
Common patterns of agentic RAG
| Pattern | What the agent does | Example |
|---|---|---|
| Routing | Picks the right source or decides no retrieval is needed | Billing question → invoices DB; "hi" → answer directly |
| Query rewriting | Rewrites a vague or conversational query into a good search query | "and the cheaper one?" → "AirLite 13 price" |
| Decomposition (multi-hop) | Splits a question into sub-questions answered in order | Find the order, then the warranty for that product |
| Self-check / corrective retrieval | Grades retrieved chunks; if poor, searches again or elsewhere | Corrective RAG (CRAG) falls back to web search |
| Reflection on the answer | Checks the draft answer against sources before replying | Self-RAG-style critique of support and relevance |
| Multi-agent | Specialist agents (researcher, verifier, writer) coordinated by a lead agent | One agent per data source, one to merge |
Research names you may meet: Self-RAG (Asai et al., 2023) trains a model to decide when to retrieve and to critique its own output; Corrective RAG (Yan et al., 2024) evaluates retrieved documents and triggers corrective actions such as web search; Adaptive-RAG (Jeong et al., 2024) routes questions to no-retrieval, single-step or multi-step strategies based on estimated complexity.
Standard RAG vs agentic RAG, and when to use it
Use agentic RAG when questions need several lookups that depend on each other, answers span several sources (documents plus databases or APIs), queries are vague and benefit from rewriting, or wrong answers are costly enough that a self-check step pays for itself. Prefer standard RAG when most questions are answered by one passage, latency must stay low, traffic is high and cost-sensitive, or you need strictly predictable behaviour.
Common mistakes and how to spot them
Agentic systems fail in a way that one-pass RAG does not: small errors multiply. Suppose each step (pick a tool, write a query, read the result) goes right 90% of the time, and suppose the steps are independent. That is a simplification, but it shows the shape of the problem.
An 8-step chain gets everything right less than half the time unless something catches mistakes. There are two levers. Use fewer steps: route simple questions to a short path and give the agent tools that return what it needs in one call. And add checks that turn a silent error into a retry, such as grading each retrieved result before using it.
To find out which lever we need, we save every run as a trace and read the failing ones. Most bad traces show one of a few patterns.
| Pattern in the trace | What it means | Fix |
|---|---|---|
| Same tool, same query, several times | The agent does not notice that nothing changed | Keep past attempts in the notes; block identical repeats; cap the steps |
| Wrong tool first, right tool later | Tool descriptions are vague or overlap | Say what each tool is for and what it is not for |
| Right fact retrieved, missing from the answer | The fact is buried in long raw tool output | Condense each observation into a short note |
| Answer states a fact found in no tool result | The model answered from memory | Require a source for every fact; check the answer against the notes |
| Stops after the first weak result | No sufficiency check | Add a grading step before answering |
When we evaluate, we score two things separately: was the final answer right, and was the path reasonable (number of steps, cost, tools used)? A right answer reached in 14 steps is still a problem to fix.
Practice: try it yourself
The earlier code showed decomposition. Now we build the self-check pattern: retrieve, grade the result, and if it fails, rewrite the query or switch to another source. As before, simple functions stand in for the LLM: the plan of attempts is written out by hand, and the grader just checks that the chunk mentions the terms a useful answer must contain.
practice_corrective_retrieval.py
# Two sources. The knowledge base has no AirLite battery page; the spec sheets do.
KB = ["ProBook batteries are covered for 12 months.",
"Laptops can be returned within 30 days if unopened.",
"Battery care: avoid heat and keep the charge between 20 and 80 percent."]
SPECS = ["AirLite 13 weighs 1.1 kg and has a 13 inch screen.",
"AirLite batteries are covered for 24 months."]
def search(source, query): # retrieval tool: best word overlap
q = set(query.lower().replace("?", "").split())
return max(source, key=lambda d: len(q & set(d.lower().replace(".", "").split())))
def grade(chunk, must_have): # stand-in for an LLM relevance check
return all(term in chunk.lower() for term in must_have)
question = "How long is the AirLite battery covered?"
must_have = ["airlite", "covered"] # what a useful chunk has to mention
plan = [("kb", KB, question), # 1: first try
("kb", KB, "AirLite batteries warranty covered"), # 2: rewritten query
("specs", SPECS, "AirLite batteries covered")] # 3: another source
MAX_STEPS = 4
for step, (name, source, query) in enumerate(plan[:MAX_STEPS], start=1):
chunk = search(source, query)
ok = grade(chunk, must_have)
print(f"step {step}: search {name} with {query!r}")
print(f" got {chunk!r} -> {'PASS' if ok else 'FAIL, try again'}")
if ok:
print("answer from:", chunk)
break
else:
print("no good evidence found: say so instead of guessing")Output:
step 1: search kb with 'How long is the AirLite battery covered?' got 'Battery care: avoid heat and keep the charge between 20 and 80 percent.' -> FAIL, try again step 2: search kb with 'AirLite batteries warranty covered' got 'ProBook batteries are covered for 12 months.' -> FAIL, try again step 3: search specs with 'AirLite batteries covered' got 'AirLite batteries are covered for 24 months.' -> PASS answer from: AirLite batteries are covered for 24 months.
Now change it:
- Set
MAX_STEPS = 2. Predict the last line of the output, and decide whether that is a good or a bad outcome for the user. - Remove
"airlite"frommust_have, so the grader only asks for "covered". Predict at which step the loop now stops, and what answer the customer would get. - Add the sentence
"AirLite batteries are covered for 24 months."to the end ofKB. Predict at which step the search now passes: step 1 or step 2?
Pause and think: Step 2 reworded the query but searched the same source again, and failed again. What should a good agent conclude from two failures in the same place?
That the source probably does not hold the fact, so rewording a third time is wasted effort. The useful move is to change the source or tool. An agent that can only rewrite queries will loop on the same index until the step limit ends it.
Pause and think: The chunk found in step 2, "ProBook batteries are covered for 12 months", is about batteries and about coverage. Why is it still right for the grader to fail it?
Being on topic is not the same as answering the question. The chunk is about a different product. If the agent accepted it, the customer would be told "12 months" with full confidence, which is wrong for an AirLite. The grader checks for the specific thing the question is about, and that is what stops wrong-product evidence.
Limitations of agentic RAG, and quick summary
Limitations and common mistakes Latency and cost grow with every loop. Unpredictability: the same question can take different paths, which complicates testing. Error compounding: a wrong early step (a bad query, a misread result) can send later steps astray. Infinite or wasteful loops without step limits. Tool misuse: vague tool descriptions lead to wrong tool choices. Security: tools that write or send data need permission checks, and retrieved text can contain prompt-injection instructions the agent must treat as data, not commands. Evaluation is harder: you must judge the path as well as the final answer.
Pause and think: Our agentic RAG bot sometimes makes 15 tool calls for simple FAQ questions and users wait 30 seconds. Name two fixes.
Add a router so simple questions take a single-retrieval path (or no retrieval), and set a hard step limit with a "answer with what you have" fallback. Clearer tool descriptions and a system prompt that tells the agent to stop once it has enough evidence also help.
- Standard RAG retrieves once from one source; it cannot recover from a miss or chain lookups.
- Agentic RAG lets an LLM agent plan, call retrieval tools, observe, reflect and repeat.
- Building blocks: a reasoning LLM, well-described tools, and memory/state, plus guardrails.
- Patterns: routing, query rewriting, decomposition, corrective retrieval, answer reflection, multi-agent.
- It costs more latency and tokens and is less predictable, so use it where questions truly need it.
Key takeaways
- Standard RAG is one fixed pass: retrieve once from one source, then generate.
- Agentic RAG puts an LLM agent in a loop of plan, act (call a tool), observe and reflect.
- Its building blocks are a reasoning model, well-described tools (retrievers, databases, APIs) and memory/state.
- Common patterns: routing, query rewriting, decomposition, corrective retrieval, answer reflection, multi-agent.
- It shines on multi-step, multi-source questions but costs more latency, tokens and predictability; add step limits and routing.
Key terms
- Agentic RAG: RAG in which an LLM agent decides when, where and what to retrieve, in a loop.
- AI agent: An LLM that repeatedly reasons, calls tools and observes results until a goal is met.
- Tool: A function with a name, description and parameters that the agent can ask the application to run.
- Multi-hop question: A question that needs several lookups, where each depends on the previous answer.
- Query decomposition: Splitting a complex question into simpler sub-questions answered in order.
- ReAct: A prompting pattern that interleaves reasoning steps with tool actions and observations.
← 10.10 Semantic Caching: Skipping the LLM for Similar Queries · 10.12 GraphRAG: Combining Knowledge Graphs with Retrieval →