Modern AI Engineering

Module 6 · Core

Next-Gen LLM Architectures

In this module, we will learn the improvements that modern LLMs add on top of the basic Transformer to become bigger, faster, and able to handle longer inputs. At the end, we will see all of these ideas together inside a real model.

By the end of this module, we will be able to read the architecture section of any new open-weight model and understand every design choice in it.

Lessons

  1. 6.1 A Timeline of LLM Architecture Improvements: What is an LLM Architecture? · Stage 1: Reading one word at a time (RNN) · Stage 2: Attention · Stage 3: The Transformer · Stage 4: Scaling · Stage 5: Mixture of Experts (MoE) · Stage 6: New Directions · Summary of the evolution
  2. 6.2 Mixture of Experts: Routing Tokens to Specialists: Why Mixture of Experts was needed · What an "expert" really means · The router and how it picks experts · Where MoE sits inside a Transformer · Sparse activation and why it saves compute · Load balancing across experts · Advantages and challenges of MoE · Why MoE powers many modern LLMs
  3. 6.3 Grouped Query Attention: Fewer KV Heads, Same Quality: The Big Picture · Quick Recap: Multi-Head Attention (MHA) · The Problem with Multi-Head Attention · What is Multi-Query Attention (MQA)? · What is Grouped-Query Attention (GQA)? · How Grouped-Query Attention Works · GQA is a Generalization of MHA and MQA · GQA vs MHA vs MQA · Real-World Use Cases · A Note on Terminology · Uptraining: Converting MHA to GQA · Quick Summary
  4. 6.4 Sliding Window Attention: Taming Very Long Contexts: What is attention? · The problem with normal attention · What is sliding window attention? · A simple step-by-step walkthrough · How information still travels far away · Comparing normal attention and sliding window attention · Where sliding window attention is used · Advantages and trade-offs
  5. 6.5 Attention Sinks: The Hidden Cost of Extended Context: What is a Large Language Model · What is attention · The problem of streaming with long conversations · The naive fix and why it fails · What is an attention sink · Why the first tokens become a sink · A step-by-step numeric walkthrough · The fix in code · StreamingLLM and modern attention sinks · Importance of attention sinks
  6. 6.6 Flash Attention: Memory-Efficient Attention at Scale: A quick recap of standard attention · Why standard attention is slow · How GPU memory actually works (HBM vs SRAM) · The core idea behind Flash Attention · Tiling: breaking the work into small blocks · Online softmax: computing softmax without the full matrix · Recomputation in the backward pass · Flash Attention 2 · Flash Attention 3 · Advantages and impact of Flash Attention
  7. 6.7 DeepSeek-V4: Anatomy of an Open-Source Frontier Model: The Big Picture · Two Models: DeepSeek-V4-Pro and DeepSeek-V4-Flash · Hybrid Attention with CSA and HCA · Manifold-Constrained Hyper-Connections (mHC) · Muon Optimizer · FP4 Quantization-Aware Training · Pre-Training · Post-Training: Specialist Training and On-Policy Distillation · Reasoning Modes · Putting It All Together · Quick Summary

← Module 5: Inside the LLM Output Pipeline · Module 7: The Language Model Zoo →