Modern AI Engineering

Module 8 · Adapt

Teaching and Shaping Models

In this module, we will learn how a pre-trained model is adapted to our own task, how it is made smaller, and how it is taught to follow instructions and human preferences.

By the end of this module, we will know when to fine-tune, how LoRA makes it cheap, and how RLHF, PPO, DPO, and GRPO align a model.

Lessons

  1. 8.1 Fine-Tuning: Adapting a Pre-Trained Model to Your Task: What is Fine-tuning? · Why do we need Fine-tuning? · How does Fine-tuning work step by step? · A simple worked example with numbers · Full Fine-tuning vs LoRA · When to use Fine-tuning · Tips before we Fine-tune · Summary
  2. 8.2 LoRA: Parameter-Efficient Fine-Tuning via Low-Rank Matrices: The Big Picture · Why Full Fine-Tuning Is Expensive · The Core Idea Behind LoRA · How LoRA Works Step by Step · A Small Numeric Example · Where LoRA Is Applied in a Transformer · Merging LoRA Back Into the Model · Real-World Use Cases · Quick Summary
  3. 8.3 Prefix Tuning: Learnable Context Prepended to the Input: What is a large language model? · The problem: why full fine-tuning is expensive · What is Prefix Tuning? · Prefix Tuning = Prefix + Tuning · How does Prefix Tuning work? · The prefix is not real words · Where the prefix is added · How the prefix is trained · How small is the prefix really? · A simple code example · Prefix Tuning vs Full Fine-Tuning · Prefix Tuning vs Prompt Tuning · Advantages of Prefix Tuning · Limitations of Prefix Tuning · Where Prefix Tuning is used
  4. 8.4 Knowledge Distillation: Compressing Large Models into Small Ones: What is Knowledge Distillation? · Why we need Knowledge Distillation · Hard labels vs soft labels · Dark knowledge · Temperature in the softmax · The distillation loss · A step-by-step training walkthrough · Types of Knowledge Distillation · Real examples of Knowledge Distillation · Wrapping up Knowledge Distillation
  5. 8.5 Continual Learning: Training Without Forgetting the Past: What is Continual Learning? · Why do we need Continual Learning in LLMs? · The big problem: Catastrophic Forgetting · Approaches to Continual Learning in LLMs · Challenges in Continual Learning · Real-world use cases
  6. 8.6 Deep RL from Human Preferences: The Foundational Paper: The building blocks we must know first · The big picture: what the paper does · Why it was needed: the reward problem · Trajectory segments, or clips · The human comparison · The reward model and the preference math · Training the reward model · Training the agent with reinforcement learning · The loop and the smart bits · The results: the backflip and beyond · The legacy: this is RLHF · What this looks like today · Quick Summary
  7. 8.7 InstructGPT: Teaching GPT-3 to Follow Instructions: What is the InstructGPT paper? · The building blocks we must know first · The big picture: what InstructGPT does · Why GPT-3 was not enough · Helpful, Honest, and Harmless · The three-step method · Step 1: Supervised Fine-Tuning · Step 2: The Reward Model · Step 3: Reinforcement Learning with PPO · The alignment tax · The results · What alignment looks like today · Quick Summary
  8. 8.8 RLHF: Aligning LLMs with Human Preferences: What is RLHF · Why we need RLHF · The Big Picture · Stage 1: Supervised Fine-Tuning (SFT) · Stage 2: Training the Reward Model · Stage 3: RL Fine-Tuning with PPO · The KL Penalty · Putting It All Together · Reward Hacking · Common Mistakes · Best Practices · Quick Summary
  9. 8.9 PPO: The Reinforcement Algorithm Behind Instruction Tuning: What is Reinforcement Learning? · What is a Policy? · The problem with simple policy updates. · What is Proximal Policy Optimization (PPO)? · The key idea behind PPO: Clipping. · The PPO objective function in simple words. · How PPO works step-by-step. · PPO in Large Language Models (RLHF). · Advantages of PPO. · Disadvantages of PPO.
  10. 8.10 DPO: Alignment Without the Separate Reward Model: What is RLHF and why do we need it? · The problem with RLHF. · What is Direct Preference Optimization (DPO)? · What is preference data? · The key idea behind DPO. · The DPO loss function in simple words. · How DPO works step-by-step. · DPO vs RLHF (PPO). · Advantages of DPO. · Disadvantages of DPO.
  11. 8.11 GRPO: Group-Based Preference Optimization Explained: What is GRPO? · Why do we need GRPO? · The problem with PPO. · How does GRPO work? · Step-by-step example. · The GRPO objective in simple words. · Advantages of GRPO. · Practical things to keep in mind. · When to use GRPO. · Conclusion.

← Module 7: The Language Model Zoo · Module 9: The Art of Prompting →