Modern AI Engineering

What Is RAG? Retrieval-Augmented Generation and How It Works

By Modern AI Engineering · · 8 min read

RAG stands for retrieval-augmented generation. It is a way to make a language model answer from documents you choose, instead of only from what it learned in training. The system first searches your documents for the passages related to the question. It then places those passages in the prompt, and the model writes its answer from them.

A simple picture is an open-book exam. The model does not have to remember everything. It looks up the right pages first and then answers. This article explains each stage and where it tends to go wrong.

What problem does RAG solve?

A language model has three gaps. It does not know anything after its training data ends. It has never seen your private data, such as contracts, support tickets or internal manuals. And when it lacks a fact, it may still write a confident answer that is wrong.

Retraining the model each time a document changes is slow and costly. RAG avoids that. The knowledge stays outside the model, in a store you control. To update what the system knows, you update the documents.

RAG also lets the system show its sources. A user can open the passage an answer came from and check it. For many business uses that matters as much as the answer.

How does RAG work, step by step?

A RAG system has two phases. Indexing happens ahead of time. Retrieval and generation happen each time a question arrives.

  1. Load the documents and turn them into plain text.
  2. Split the text into chunks, pieces small enough to search and to fit in a prompt.
  3. Turn each chunk into an embedding, a vector that captures its meaning, and store it in a vector database.
  4. When a question arrives, turn the question into an embedding with the same model.
  5. Find the chunks whose vectors are closest to the question vector.
  6. Optionally rerank those chunks, so the most relevant ones come first.
  7. Build a prompt that holds the instructions, the chosen chunks and the question.
  8. The model writes an answer from that text, with references to the chunks it used.

The parts of a RAG pipeline

Each part has one job, and each can be improved separately.

PartWhat it does
ChunkerSplits documents into pieces that keep their meaning.
Embedding modelTurns text into vectors, so similar meanings sit close together.
Vector databaseStores the vectors and finds the nearest ones quickly.
RetrieverRuns the search. It may combine vector search with keyword search.
RerankerScores the retrieved chunks again and puts the best first.
Prompt builderPlaces instructions, chunks and the question into one prompt.
GeneratorThe language model that writes the answer.

Why does chunking matter so much?

Chunking decides what the search can find. If a chunk is too large, it mixes several topics, its embedding becomes vague, and it wastes space in the prompt. If a chunk is too small, it loses the context that gives it meaning. A sentence that says the limit is thirty days is useless without knowing which limit.

Simple chunking cuts every fixed number of characters, with a small overlap between neighbours. Better chunking follows the structure of the document: headings, paragraphs, list items. Tables and code need special care, because a table cut in half means nothing.

There is no single correct chunk size. It depends on the documents and the questions. The honest way to choose is to try a few settings and measure which one retrieves the right passages most often.

Semantic search, keyword search and hybrid search

Vector search is also called semantic search. It finds text with a similar meaning even when the words differ. A question about refunds can match a passage about getting money back.

It has a weakness. Exact terms such as product codes, error numbers and names may not be matched well, because the embedding blurs them. Keyword search is strong exactly there.

Hybrid search runs both and merges the results. Many production systems use it for that reason. A reranker is often added afterwards. It reads the question together with each candidate chunk and gives a more careful relevance score than the first, fast search can.

Why do RAG systems fail?

When a RAG answer is wrong, find out which stage failed before you change anything.

  • The answer is not in the documents. No search can find what is not there.
  • The right chunk exists but was not retrieved. Look at chunking, the embedding model and the search method.
  • The right chunk was retrieved but ranked too low to be included. A reranker helps.
  • Too many chunks were included. The useful one is buried, and models can miss details in the middle of a long context.
  • The model had the right text and ignored it or misread it. Tighten the instructions and ask it to answer only from the sources.
  • The question was vague. Rewriting the question before searching can help.

Beyond basic RAG

Once the basic pipeline works, several extensions deal with harder questions.

HyDE asks the model to write a made-up answer first and then searches with that text, because an answer often looks more like the target passage than the question does. Agentic RAG lets the model decide when to search, what to search for and whether to search again. GraphRAG builds a graph of the people, things and relations in the documents, which helps with questions that span many of them. Vectorless RAG retrieves without embeddings at all, for example by letting the model walk through the outline of a document.

Add these only when your tests show a need. Each one brings extra cost and extra parts that can break.

How to learn RAG properly

Build a small system on documents you know well, so you can tell when an answer is wrong. Print the retrieved chunks for every question and read them. Most of what you learn about RAG comes from that habit.

Module 10 of the AI Engineering Bootcamp, Building RAG Systems, has thirteen lessons. They cover vector databases, approximate nearest neighbour search, semantic and hybrid search, rerankers, chunking, HyDE, caching, agentic RAG, GraphRAG and vectorless RAG.

Frequently asked questions

What does RAG stand for?

Retrieval-augmented generation. The system retrieves relevant text, adds it to the prompt, and the model generates an answer from it.

Does RAG train the model on my data?

No. The weights of the model do not change. Your documents are searched at question time, and the relevant passages are placed in the prompt.

Does RAG remove hallucinations?

It reduces them when the right passages are retrieved, because the model answers from text in front of it. It does not remove them. Poor retrieval or unclear sources still lead to wrong answers.

What is the difference between RAG and fine-tuning?

RAG adds knowledge at question time by changing the prompt. Fine-tuning changes the behaviour of the model by training its weights. RAG suits facts that change. Fine-tuning suits style and format.

Is RAG still needed with large context windows?

Often, yes. Putting every document into every prompt costs more, takes longer and can hide the important passage. Retrieval keeps the context short and relevant.

Learn it properly: the AI Engineering Bootcamp

More articles