Modern AI Engineering

Module 3 · Foundations

Neural Architectures Deep Dive

In this module, we will learn how a neural network actually learns. We will understand the math behind gradient descent and backpropagation step by step, and the techniques that make training stable.

By the end of this module, we will be able to explain how a neural network trains, from the forward pass to the weight update, and why normalization and dropout matter.

Lessons

  1. 3.1 Neural Network Bias: What It Is and Why It Matters
  2. 3.2 Gradient Descent: Rolling Downhill to the Optimum: The Big Picture · What is a Loss Function · What is Gradient Descent · The Intuition Behind Gradient Descent · The Math Behind Gradient Descent · Step-by-Step Numeric Example · Gradient Descent with Multiple Parameters · The Role of Learning Rate · Types of Gradient Descent · Gradient Descent in Python · Putting It All Together
  3. 3.3 Backpropagation: How Neural Networks Learn from Mistakes: What is Backpropagation? · The Chain Rule of Calculus · Forward Pass · Loss Calculation · Backward Pass (Backpropagation) · Step-by-Step Numeric Example · Weight Update Using Gradient Descent · Backpropagation in Python
  4. 3.4 Cross-Entropy Loss: Scoring Probability Predictions: The Big Picture · What is Cross-Entropy · The Cross-Entropy Loss Formula · Why We Take the Negative Log · Binary Cross-Entropy Loss · Categorical Cross-Entropy Loss · Step-by-Step Numeric Example · Cross-Entropy Loss for Language Models · The Gradient of Cross-Entropy Loss · Quick Summary
  5. 3.5 Dropout: Controlled Forgetting as Regularization: What is Dropout? · The problem of Overfitting · Why do we need Dropout? · How does Dropout work? · A step-by-step example · Dropout during training vs testing · Dropout in code · Variants of Dropout · Advantages of Dropout · Where Dropout is used
  6. 3.6 Batch Norm vs Layer Norm: When to Use Each: What is Normalization? · Why do we need Normalization? · What is Batch Normalization? · What is Layer Normalization? · Batch Normalization vs Layer Normalization · When to use which one?
  7. 3.7 RMSNorm: Simpler Normalization for Transformers: Why normalization is needed in deep networks · A quick recap of Layer Normalization (LayerNorm) · What RMSNorm is and how it works · The math behind RMSNorm with a concrete numeric example · LayerNorm vs RMSNorm - the key differences · Why modern LLMs prefer RMSNorm · A code example · Where RMSNorm fits in a Transformer · Quick Summary
  8. 3.8 Recurrent Neural Networks: Processing Sequences in Order
  9. 3.9 PyTorch Internals: Dynamic Graphs and Autograd: What is PyTorch? · What is a Tensor? · The problem PyTorch solves · What is a Computation Graph? · What is Autograd? · A complete training example · What is the GPU and why does PyTorch use it? · Why is PyTorch so popular?
  10. 3.10 TensorFlow Explained: Static Graphs and Production ML

← Module 2: Learning from Data · Module 4: Transformers and How They Think →