Modern AI Engineering

Module 13 · Production

Serving LLMs at Scale

In this module, we will learn how to make LLMs faster and cheaper to run. We will start with what happens during inference, then learn the caching, batching, and speculation techniques, then quantization, and finally the serving engines that put it all together.

By the end of this module, we will understand TTFT, TPOT, and throughput, and know which optimization fixes which bottleneck.

Lessons

  1. 13.1 LLM Inference Optimization: The Full Landscape: What is an LLM and how it writes text · What is Attention · What is the KV Cache · Why the KV Cache becomes huge · What is KV Cache Compression · Approach 1: Quantization · Approach 2: Token Eviction · Approach 3: Sharing Keys and Values across Heads · Approach 4: Low-Rank Compression · Comparison of the approaches · When to use which one
  2. 13.2 Prefill vs Decode: Two Distinct Phases of LLM Inference: What is LLM inference · The two phases: Prefill and Decode · Prefill explained in simple words · Decode explained in simple words · A diagram of the two phases and the KV cache flow · The KV cache as the bridge between the two phases · A step-by-step walkthrough of a few decode steps · Prefill vs Decode comparison table · Why this split matters: compute-bound vs memory-bound · The key metrics: TTFT, TPOT, throughput, and end-to-end latency · Optimization techniques mapped to each phase · Conclusion
  3. 13.3 Prefill-Decode Disaggregation: Splitting the Two Phases: How an LLM answers a request · What is the KV Cache? · Prefill is compute-heavy, Decode is memory-heavy · The problem when both run on the same GPU · TTFT vs TPOT · The naive approaches and their issues · What is Prefill-Decode Disaggregation? · How Prefill-Decode Disaggregation works · Walkthrough of one request · Advantages of Prefill-Decode Disaggregation · Disadvantages of Prefill-Decode Disaggregation · Where it works well and where it is overkill · Co-located vs Disaggregated serving
  4. 13.4 The KV Cache: Avoiding Redundant Attention Computation: How LLMs Generate Text · What Happens Inside the Model · The Problem: Repeated Computation · The Solution: KV Cache · Why Only Key and Value Are Cached, Not Query · How Much Faster Does It Get · The Trade-Off: Speed vs Memory
  5. 13.5 KV Cache Compression: Trading Some Accuracy for Speed: What is an LLM and how it writes text · What is Attention · What is the KV Cache · Why the KV Cache becomes huge · What is KV Cache Compression · Approach 1: Quantization · Approach 2: Token Eviction · Approach 3: Sharing Keys and Values across Heads · Approach 4: Low-Rank Compression · Comparison of the approaches · When to use which one
  6. 13.6 Paged Attention: OS-Inspired Memory Management for KV Caches: Quick Recap: KV Cache · The Problem: Memory Waste in KV Cache · What is Paged Attention? · How Paged Attention Works · Why Paged Attention Is So Effective · Memory Sharing Across Requests
  7. 13.7 Continuous Batching: Keeping GPUs Busy Between Requests: The Big Picture · Quick Recap: How an LLM Generates Tokens · Why Batching Matters for LLMs · The Old Way: Static Batching · The Problem with Static Batching · What is Continuous Batching? · The Ride-Share Analogy · How Continuous Batching Works Step by Step · A Numeric Example · Real Numbers and Speedup · Benefits of Continuous Batching · A Few Important Notes · Quick Summary
  8. 13.8 Speculative Decoding: Draft Fast, Verify in Parallel: What problem does Speculative Decoding solve? · The Big Picture · Why is LLM generation slow? · The core idea behind Speculative Decoding · Step-by-step walkthrough · The verification step · Real numbers and speedup · Where it is used · Trade-offs · Quick Summary
  9. 13.9 N-gram Speculation: Draft Tokens Without a Draft Model: How an LLM generates text · Why generating text is slow · What is Speculative Decoding · The cost of a draft model · What is an N-gram · What is N-gram Speculation · N-gram Speculation step by step · Why the output stays exactly the same · Where it works well and where it fails · N-gram Speculation vs Draft Model Speculative Decoding
  10. 13.10 Medusa: Parallel Decoding via Multiple Prediction Heads: What is Medusa · Why text generation is slow · A quick recap of speculative decoding · The problem with needing a draft model · The big idea: many heads on one model · How tree attention checks many guesses at once · The math behind the speedup with small numbers · The results · How Medusa lives on today · Quick Summary
  11. 13.11 EAGLE: Feature-Level Drafting for Faster Inference: What is EAGLE · A quick recap of speculative decoding · The problem with token-level drafting · The big idea: draft at the feature level · Resolving the uncertainty by feeding back the token · The math behind the speedup with small numbers · EAGLE-2 and dynamic draft trees · How EAGLE lives on today · Quick Summary
  12. 13.12 Model Quantization: Shrinking Weights Without Breaking Outputs: What is Model Quantization? · How numbers use bits (FP32, INT8, INT4) · Why fewer bits means less memory and faster speed · The core mechanic: scale and zero-point · Symmetric vs Asymmetric Quantization · Per-tensor vs Per-channel Quantization · Post-Training Quantization (PTQ) vs Quantization-Aware Training (QAT) · Weight-only vs Weight-and-activation Quantization · The outlier problem in LLMs · Popular methods: GPTQ, AWQ, bitsandbytes, GGUF / llama.cpp · The accuracy trade-off and running LLMs locally · Wrapping up Model Quantization
  13. 13.13 GGUF: The File Format Powering Local LLM Inference: What is a model and what are weights · What is local inference · The problem before GGUF · What is GGUF · What is stored inside a GGUF file · What is quantization · Understanding quantization names like Q4_K_M · How GGUF loads fast with memory mapping · Why GGUF is cross-platform and extensible · GGUF in the real world
  14. 13.14 llama.cpp: Running Large Models on Consumer Hardware: What is llama.cpp · Why we needed llama.cpp · A quick refresher on what an LLM is · The real problem: models are too big to fit · The first big idea: quantization · Understanding names like Q4_K_M · The GGUF file: everything packed in one box · Memory mapping: loading the model the smart way · Squeezing speed out of the CPU · Sharing the work with the GPU · The full journey of running a prompt · Where llama.cpp is used
  15. 13.15 vLLM: High-Throughput Serving with PagedAttention: What is serving an LLM · A quick recap of prefill, decode, and the KV cache · The problem: the KV cache eats GPU memory · Why naive serving wastes memory · What is vLLM · PagedAttention, the core idea · How PagedAttention shares memory · Continuous batching · The OpenAI-compatible API server · The benefits of vLLM · vLLM in the real world
  16. 13.16 SGLang: Structured LLM Programs for Efficient Inference: What is SGLang · A quick recap of how an LLM generates text · The problem SGLang solves · RadixAttention: the heart of SGLang · How RadixAttention reuses past work · The frontend language of SGLang · How the runtime and the frontend work together · Continuous batching in SGLang · Structured output and faster decoding · A simple end-to-end picture · More powerful features of SGLang · How SGLang compares to vLLM
  17. 13.17 TensorRT-LLM: NVIDIA's Optimized Inference Engine: What is inference · What is a GPU and what is a kernel · The problem: the GPU spends its time on the wrong things · What is TensorRT-LLM · The big idea: prepare the model ahead of time · The build step: from a model to an engine · Kernel fusion · Quantization · Custom attention kernels · The paged KV cache · In-flight batching · CUDA graphs · Speculative decoding · Running one model across many GPUs · How we actually serve the model · The PyTorch backend, the newer and easier path · The full journey of one request · TensorRT-LLM vs vLLM · Where it works well and where it fails · [LLM Inference Optimization](https://www.youtube.com/watch?v=jV2sCj4lHYk) (Video) · [The First-Token Latency Problem in LLMs](https://www.youtube.com/watch?v=XD8DD4cEHu0) (Video) · The full LLM inference optimization pipeline, end to end

← Module 12: Agent Patterns and Frameworks · Module 14: Measuring What Matters →