Modern AI Engineering

Module 17 · Production

Production AI Infrastructure

In this module, we will learn the hardware that runs AI models, where to deploy a model, how to send each request to the right model, and how to design a complete AI system end to end.

By the end of this module, we will be able to design an AI system end to end, from the hardware to the user.

Lessons

  1. 17.1 GPUs for Deep Learning: Parallelism at the Core: What is a GPU? · Why is the GPU perfect for deep learning? · CPU vs GPU · The math professor and the thousands of students · Why deep learning is mostly matrix multiplication · Serial work vs parallel work · GPU memory (VRAM) and memory bandwidth · Why the model must fit in VRAM · Tensor Cores and lower precision (FP16, BF16, INT8) · CUDA and the software stack (cuDNN) · Training vs inference on GPUs · Multiple GPUs working together · Why NVIDIA GPUs power modern AI
  2. 17.2 CUDA Kernels: Writing Parallel Code for NVIDIA GPUs: Why do we need a GPU? · What is CUDA? · What is a CUDA Kernel? · Threads, Blocks, and Grids · Host and Device · Writing our first CUDA Kernel · How a thread finds its own work · What happens inside the GPU when a kernel runs · Memory in CUDA · Why CUDA Kernels matter for AI · Where CUDA Kernels work well and where they fail
  3. 17.3 Google TPUs: Purpose-Built Hardware for Neural Networks: What is a TPU · Why Google built the TPU · A quick refresher: CPU and GPU · The one operation that matters most · The big idea: Systolic Array · How data flows through a TPU · The full journey of a TPU computation · Why a TPU is so fast and power efficient · Where TPUs are used · Limitations of a TPU
  4. 17.4 Language Processing Units: A New Approach to LLM Inference: What is an LPU? · How an LLM writes text, one token at a time · The real bottleneck is memory, not math · Why a GPU struggles here · Idea 1: Keep the model on the chip · The problem with on-chip memory · Idea 2: Remove all the guesswork · Idea 3: A network that never waits · The assembly line · What happens when we send a prompt · Why an LPU is fast, all in one place · Where an LPU works well · Where an LPU does not work well · LPU vs GPU · When to use which one
  5. 17.5 Cloud vs Edge: Where Should Your Model Run?: What is deployment? · Training and inference · What is Cloud Deployment? · What is On-device Deployment? · The one big difference · The round trip problem · Where does our data go? · How big can the model be? · Who pays the bill? · The shipping problem · What happens when the network is gone? · The hybrid approach · Some real examples · Let's tabulate the difference · When to use which one? · Summary
  6. 17.6 On-Device ML: A TensorFlow Lite Android Walkthrough
  7. 17.7 LLM Routing: Directing Each Query to the Best Model: The Big Picture · What is LLM Routing · Why we need LLM Routing · Anatomy of an LLM Router · Routing Strategies · A Full Trace Example · LLM Routing vs Mixture of Experts · When LLM Routing is Worth It · Common Mistakes and How to Fix Them · Quick Summary
  8. 17.8 Building a Real-Time Voice AI Agent from Scratch: What is a Voice AI Agent? · Why is Real-Time Voice hard? · Requirements · Back-of-the-envelope estimation · High-Level Architecture · Component 1: Audio Transport · Component 2: Voice Activity Detection and Turn Detection · Component 3: Speech-to-Text (STT) · Component 4: The Brain - LLM with Tools · Component 5: Text-to-Speech (TTS) · Approach 1: Cascaded Pipeline (STT -> LLM -> TTS) · Approach 2: Speech-to-Speech Model · Approach 3: Hybrid Approach · Cascaded vs Speech-to-Speech: Comparison · Latency Budget: Where every millisecond goes · Handling Interruptions (Barge-in) · Tool Calling in a Voice Agent · Memory and Context · Telephony: Connecting to real phone calls · Scaling the system · Edge Cases and how to handle them · Observability and Evaluation · Safety, Security, and Privacy · Cost · How to present this design in an interview
  9. 17.9 System Design Fundamentals for AI Engineers: What is System Design? · Why do we need it? · What are the required concepts?
  10. 17.10 Transport Protocols: HTTP, WebSockets, and SSE Compared: HTTP request · HTTP Polling · HTTP Long Polling · WebSocket · Server-Send Events(SSE)
  11. 17.11 How do Voice And Video Call Work?: Signaling · Peer-to-Peer Connection · STUN Server · TURN Server

← Module 16: Beyond Text: Multimodal AI · Module 18: The Edge of AI Research →