Modern AI Engineering

Module 10 · Build

Building RAG Systems

In this module, we will learn how to give an LLM knowledge that it was never trained on. We will start with how vectors are stored and searched, then move to retrieval techniques, and finally to the advanced forms of RAG.

By the end of this module, we will be able to build a production-grade RAG pipeline and pick the right retrieval technique for our data.

Lessons

  1. 10.1 Vector Databases: Storing and Searching Embeddings at Scale: What is a Vector Database? · A quick recap of embeddings · Why normal databases fall short · What a Vector Database actually stores · How do we measure similarity? · Cosine similarity · Dot product · Euclidean distance · The nearest neighbour problem · Why brute force is too slow · Approximate Nearest Neighbour (ANN) and indexing · HNSW explained simply · IVF explained simply · PQ explained simply · A small code example · Real-world applications of Vector Databases
  2. 10.2 ANN Search: Finding Similar Vectors Without Brute Force: What is Nearest Neighbor Search? · How do we turn things into numbers (vectors)? · How do we measure "closeness"? · The naive approach and why it fails · What is Approximate Nearest Neighbor (ANN) Search? · The trade-off: speed vs accuracy · Approach 1: Trees (KD-Tree) · Approach 2: Hashing (LSH) · Approach 3: Clustering (IVF) · Approach 4: Graphs (HNSW) · A simple code example · Where ANN Search is used · Picking the right method
  3. 10.3 Semantic Search: Finding Meaning, Not Just Keywords: What is keyword search · Where keyword search fails · What is Semantic Search · What is an embedding · How similar meanings sit close together · How we measure closeness with cosine similarity · What is a vector database · The full Semantic Search flow · Approximate nearest neighbor for speed at scale · Semantic Search in the real world
  4. 10.4 Hybrid Search: Combining Sparse and Dense Retrieval: What is keyword search · Why keyword search alone is not enough · What is semantic search · Why semantic search alone is not enough · What is Hybrid Search · How Hybrid Search runs both searches · How the two result lists are combined · Reciprocal Rank Fusion (RRF) · Weighted score combination and normalization · Hybrid Search in the real world
  5. 10.5 Rerankers: Re-Scoring Retrieved Results by Relevance: What is a Reranker · Where a Reranker sits in a search / RAG pipeline · The two-stage retrieval idea · Why first-stage retrieval is fast but not precise · Bi-encoder vs Cross-encoder · How a Reranker scores documents step by step · The accuracy vs latency and cost trade-off · Late-interaction models like ColBERT · Real examples of Rerankers · Why Rerankers matter for RAG
  6. 10.6 ColBERT: Token-Level Late Interaction for Retrieval: What is the ColBERT paper? · The building blocks we must know first · The big picture: what ColBERT does · The two old extremes · Late interaction: the key idea · Encoding the query and document · The MaxSim operation · Why max, not average · Ranking documents · Training: positives and negatives · The loss with small numbers · Fast retrieval at scale · The cost of a bigger index · The Results · Where ColBERT led · Quick Summary
  7. 10.7 Document Chunking Strategies for RAG: What is RAG? · What is a chunk? · Why do we need chunking? · How retrieval actually works · What happens when we chunk badly · Fixed-size chunking · Chunking by sentence · Recursive chunking · Document structure based chunking · Semantic chunking · Contextual chunking · Small-to-big chunking · Agentic chunking · Chunk overlap · How to choose the chunk size · Comparison of all the strategies · Common mistakes · Conclusion
  8. 10.8 HyDE: Generating Hypothetical Documents to Improve RAG: What is RAG in simple words · What is the search problem in RAG · Why searching with the question is weak · What is HyDE · Why searching with a fake answer works better · How HyDE works step by step · A worked example of HyDE · A simple code example of HyDE · Advantages of HyDE · Disadvantages of HyDE · When to use HyDE · Summary
  9. 10.9 Embedding Caches: Avoiding Redundant Embedding Calls: What is an embedding · A quick recap of how we get an embedding · What is an Embedding Cache · Why we need an Embedding Cache · The core idea behind an Embedding Cache · The cache key, a hash of the text plus the model · The request flow, a hit and a miss · Eviction, LRU and TTL · Where the cache lives, memory or disk · The benefits of an Embedding Cache · An Embedding Cache in the real world
  10. 10.10 Semantic Caching: Skipping the LLM for Similar Queries: What is a cache? · The problem with traditional caching for AI apps · What is Semantic Caching? · What are embeddings? · What is similarity between embeddings? · How does Semantic Caching work step by step? · A numeric walkthrough · Setting the similarity threshold · Advantages of Semantic Caching · Things to keep in mind
  11. 10.11 Agentic RAG: Dynamic Retrieval with Multi-Step Reasoning: The Big Picture · A Quick Recap of RAG · A Quick Recap of AI Agent · Why Standard RAG Falls Short · What is Agentic RAG · The Agentic RAG Loop · The Three Building Blocks · A Walkthrough with a Real Example · Common Patterns of Agentic RAG · Standard RAG vs Agentic RAG · When to Use Agentic RAG · Limitations of Agentic RAG · Quick Summary
  12. 10.12 GraphRAG: Combining Knowledge Graphs with Retrieval: What is GraphRAG? · Why normal RAG is not enough · The big picture of GraphRAG · How GraphRAG builds the knowledge graph · How GraphRAG answers a question · Local search vs Global search · When to use GraphRAG · Trade-offs of GraphRAG · Quick Summary
  13. 10.13 Vectorless RAG: Retrieval Without Embeddings or a Vector Store: What is an LLM · What is RAG · How the normal Vector RAG works · Problems with Vector RAG · What is Vectorless RAG · How Vectorless RAG works · An example of Vectorless RAG · Other Vectorless approaches · Advantages of Vectorless RAG · Disadvantages of Vectorless RAG · Vector RAG vs Vectorless RAG · When to use which one · [AI Engineering Explained: LLM, RAG, MCP, Agent, Fine-Tuning, Quantization](https://www.youtube.com/watch?v=lnfWvX66FUk) (Video) · [Agentic RAG Explained](https://www.youtube.com/watch?v=6nSegpuWJVw) (Video)

← Module 9: The Art of Prompting · Module 11: Autonomous AI Agents →