Modern AI Engineering

Module 4 · Core

Transformers and How They Think

In this module, we will learn what Generative AI is and how the Transformer, the architecture behind every modern LLM, works from the inside. We will go from tokens to embeddings to attention, one piece at a time.

By the end of this module, we will be able to draw the Transformer from memory and explain every block inside it, including the math behind Q, K, and V.

Lessons

  1. 4.1 Generative AI: Creating Instead of Classifying: What is Generative AI? · Generative AI = Generative + AI · What does "generate" mean here? · How is Generative AI different from the old AI? · How does Generative AI learn? · How does Generative AI actually create something new? · What can Generative AI create? · What is a model in Generative AI? · The complete flow of Generative AI · Where do we use Generative AI every day? · The limitations we must know · Summary
  2. 4.2 Autoregressive Models: Predicting One Token at a Time: What is an Autoregressive Model? · The Chain Rule of Probability · The Generation Loop · Step-by-Step Numeric Example · Why GPT-style Models are Autoregressive · Why Autoregressive Models Need Causal Masking · The Connection with KV Cache · Autoregressive vs Non-Autoregressive Generation · Popular Autoregressive Models we should know · Pros and Cons of Autoregressive Models · Quick Summary
  3. 4.3 BPE Tokenization: How LLMs Split Text into Tokens: What is Tokenization? · The Problem: How to Break Text into Tokens? · What is BPE (Byte Pair Encoding)? · How BPE Works: Step by Step · How BPE Tokenizes New Text · Why BPE is Used in Modern LLMs
  4. 4.4 Embeddings: Encoding Meaning as Vectors: The problem: computers cannot compare meaning · What are Embeddings? · A simple example with two numbers · Why we need many more than two numbers · How do we measure closeness? · Where do embeddings come from? · The famous word math example · Embeddings are not only for words · What we use embeddings for · Things we must be careful about · Summary
  5. 4.5 RNNs vs Transformers: A Fundamental Architecture Shift: What both of them are for · What is an RNN? · The problem with an RNN · What is a Transformer? · The key difference in one line · RNN vs Transformer side by side · Let's tabulate the difference · When to use which one? · Summary
  6. 4.6 The Transformer Architecture: Built on Attention: Why the Transformer was needed · The two halves of the architecture · Tokenization, Embedding, and Positional Encoding · The Attention Mechanism and Multi-Head Attention · Feed-Forward Networks, Residual Connections, and Layer Normalization · How the Encoder and Decoder work · How data flows through the entire architecture · The three variants of the Transformer · Why the Transformer is so powerful
  7. 4.7 Encoder vs Decoder: Two Sides of the Transformer: What is a Transformer? · A word about tokens · What is an Encoder? · What is a Decoder? · The one big difference · The three types of Transformers · Let's tabulate the difference · When to use which one? · Summary
  8. 4.8 Self-Attention: How Tokens See One Another: What is Self Attention? · Why do we need Self Attention? · Query, Key, and Value vectors · Step-by-step working of Self Attention · A simple example walk-through · Why Self Attention works so well · Multi-Head Self Attention · Where Self Attention is used
  9. 4.9 Attention Math: Queries, Keys, and Values Unpacked: The Attention Formula · Setting Up: From Words to Vectors · Creating Q, K, and V Matrices · Computing Attention Scores (Q x K^T) · Scaling the Scores · Applying Softmax · Computing the Final Output (Attention Weights x V) · Putting It All Together
  10. 4.10 Scaled Dot-Product Attention: Why We Divide by √dₖ: The Attention Formula (Quick Recap) · What Happens Without Scaling? · Why Do Dot Products Grow with dₖ? · Understanding Variance of the Dot Product · Proving It Step by Step: Variance of the Dot Product is dₖ · What Large Dot Products Do to Softmax · Why √dₖ is the Right Scaling Factor · Seeing It with Real Numbers · Putting It All Together
  11. 4.11 Causal Masking: Preventing the Model from Seeing the Future: Without Causal Masking · With Causal Masking · Implementation of Causal Masking · The Causal Mask Matrix
  12. 4.12 Multi-Head Attention: Many Perspectives at Once: What is Multi-Head Attention? · A quick recap of Self Attention · Why do we need Multi-Head Attention? · Step-by-step working of Multi-Head Attention · A simple example walk-through · Where Multi-Head Attention is used · Advantages of Multi-Head Attention
  13. 4.13 Cross-Attention: Connecting Encoder Output to the Decoder: What is Cross Attention? · Why do we need Cross Attention? · Query, Key, and Value in Cross Attention · Self Attention vs Cross Attention · Step-by-step working of Cross Attention · A simple example walk-through · Where Cross Attention is used · Importance of Cross Attention
  14. 4.14 Rotary Position Encoding: Position Without Fixed Lookup Tables: The Big Picture · Why a Transformer Needs Position Information · Older Approaches and Their Problems · The Core Idea Behind RoPE · The 2D Rotation Math · How RoPE Is Applied to Q and K · Why the Dot Product Captures Relative Position · A Small Numeric Example · Real-World Use Cases · Quick Summary
  15. 4.15 Feed-Forward Networks: The Transformer's Memory Layer: What is a Feed-Forward Network? · Understanding Feed-Forward Networks with a Real-World Analogy · Where Does the Feed-Forward Network Sit in a Transformer? · How Does a Feed-Forward Network Work - Step by Step · The Expand-then-Contract Pattern · Why Does the FFN Expand and Then Contract? · ReLU and Activation Functions · What Does the Feed-Forward Network Actually Learn? · How Much of the Model is the Feed-Forward Network? · Feed-Forward Networks in Mixture of Experts · Why Feed-Forward Networks Are So Important · [Softmax Activation Function in Machine Learning](https://www.youtube.com/watch?v=2Zx6x01WwWM) (Video) · Inside ChatGPT: What happens token by token · [Tokenization in Large Language Models (LLMs)](https://www.youtube.com/watch?v=sK2s9I84EVI) (Video) · [Embeddings in Machine Learning](https://www.youtube.com/watch?v=LedXW6xl21s) (Video) · Positional embeddings: how transformers track order

← Module 3: Neural Architectures Deep Dive · Module 5: Inside the LLM Output Pipeline →