AI Engineering Bootcamp · Roadmap
AI Engineer Roadmap: what to learn, and in what order
A step-by-step path from your first machine-learning idea to designing production AI systems: 10 steps, 19 modules and about 63 hours of lessons. Every step links to the lesson that teaches it.
What does an AI engineer do?
AI engineering is the discipline of building real applications and systems on top of AI models, especially large language models (LLMs).
An AI engineer rarely trains a giant model from scratch. Instead, they understand how models work inside, choose the right model, give it the right context, connect it to tools and data, make it fast and affordable to run, measure whether it is doing a good job, keep it safe, and ship it to real users.
A machine-learning engineer mostly trains, tunes and deploys models. An AI engineer mostly builds products and systems on top of existing models, especially LLMs, using prompting, context engineering, RAG, agents, fine-tuning, inference optimisation and evaluation. The two overlap, and this course covers the foundations both need.
The skills an AI engineer needs
How LLMs work inside (Transformers, attention, tokenization); how to adapt them (prompting, context engineering, fine-tuning, LoRA); how to give them knowledge (RAG, vector search); how to make them act (agents, function calling, MCP); how to run them efficiently (inference, quantization, serving); how to measure and secure them (evaluation, observability, guardrails); and how to design complete systems.
The roadmap, step by step
Step 1: Start Here
Module 1: AI Engineering Starter Kit (1 lessons, about 1 hours). Before going deep, we meet the six words that come up in every AI engineering conversation: LLM, RAG, MCP, Agent, Fine-tuning and Quantization.
Step 2: Foundations
Module 2: Learning from Data (9 lessons, about 3 hours). In this module, we will learn what Machine Learning is, the different ways a machine can learn, and the basic terms we will keep using in every later module of this AI Engineering Course.
- 2.1 Machine Learning from First Principles
- 2.2 Labeled vs Unlabeled: Two Ways Machines Learn
- 2.3 Predicting Numbers vs Categories: Regression Compared
- 2.4 Feature Engineering: Turning Raw Data into Signal
- 2.5 Precision and Recall: Picking the Right Metric
- 2.6 L1 vs L2 Loss: Choosing Your Error Penalty
- 2.7 Regularization: Stopping Overfitting with L1 and L2
- 2.8 Reinforcement Learning: Teaching Agents Through Reward
- 2.9 Contrastive Learning: Training by Comparison
Module 3: Neural Architectures Deep Dive (10 lessons, about 4 hours). In this module, we will learn how a neural network actually learns. We will understand the math behind gradient descent and backpropagation step by step, and the techniques that make training stable.
- 3.1 Neural Network Bias: What It Is and Why It Matters
- 3.2 Gradient Descent: Rolling Downhill to the Optimum
- 3.3 Backpropagation: How Neural Networks Learn from Mistakes
- 3.4 Cross-Entropy Loss: Scoring Probability Predictions
- 3.5 Dropout: Controlled Forgetting as Regularization
- 3.6 Batch Norm vs Layer Norm: When to Use Each
- 3.7 RMSNorm: Simpler Normalization for Transformers
- 3.8 Recurrent Neural Networks: Processing Sequences in Order
- 3.9 PyTorch Internals: Dynamic Graphs and Autograd
- 3.10 TensorFlow Explained: Static Graphs and Production ML
Step 3: Core
Module 4: Transformers and How They Think (15 lessons, about 7 hours). In this module, we will learn what Generative AI is and how the Transformer, the architecture behind every modern LLM, works from the inside. We will go from tokens to embeddings to attention, one piece at a time.
- 4.1 Generative AI: Creating Instead of Classifying
- 4.2 Autoregressive Models: Predicting One Token at a Time
- 4.3 BPE Tokenization: How LLMs Split Text into Tokens
- 4.4 Embeddings: Encoding Meaning as Vectors
- 4.5 RNNs vs Transformers: A Fundamental Architecture Shift
- 4.6 The Transformer Architecture: Built on Attention
- 4.7 Encoder vs Decoder: Two Sides of the Transformer
- 4.8 Self-Attention: How Tokens See One Another
- 4.9 Attention Math: Queries, Keys, and Values Unpacked
- 4.10 Scaled Dot-Product Attention: Why We Divide by √dₖ
- 4.11 Causal Masking: Preventing the Model from Seeing the Future
- 4.12 Multi-Head Attention: Many Perspectives at Once
- 4.13 Cross-Attention: Connecting Encoder Output to the Decoder
- 4.14 Rotary Position Encoding: Position Without Fixed Lookup Tables
- 4.15 Feed-Forward Networks: The Transformer's Memory Layer
Module 5: Inside the LLM Output Pipeline (4 lessons, about 2 hours). In this module, we will learn how an LLM picks the next token, how we control its creativity, how the output reaches the user token by token, and where the context window fails.
- 5.1 Temperature Sampling: Dialing Up or Down Creativity
- 5.2 Nucleus Sampling: Top-k and Top-p Demystified
- 5.3 Token Streaming: Rendering Outputs as They Arrive
- 5.4 Lost in the Middle: Why LLMs Miss Central Context
Module 6: Next-Gen LLM Architectures (7 lessons, about 3 hours). In this module, we will learn the improvements that modern LLMs add on top of the basic Transformer to become bigger, faster, and able to handle longer inputs. At the end, we will see all of these ideas together inside a real model.
- 6.1 A Timeline of LLM Architecture Improvements
- 6.2 Mixture of Experts: Routing Tokens to Specialists
- 6.3 Grouped Query Attention: Fewer KV Heads, Same Quality
- 6.4 Sliding Window Attention: Taming Very Long Contexts
- 6.5 Attention Sinks: The Hidden Cost of Extended Context
- 6.6 Flash Attention: Memory-Efficient Attention at Scale
- 6.7 DeepSeek-V4: Anatomy of an Open-Source Frontier Model
Module 7: The Language Model Zoo (5 lessons, about 2 hours). In this module, we will learn that not every language model is a large, text-generating LLM. We will see the smaller, reasoning, recursive, diffusion-based, and decision-only models and when to use which one.
- 7.1 Small Language Models: Big Capability in Compact Form
- 7.2 Large Reasoning Models: Chain-of-Thought at Inference Time
- 7.3 Recursive Language Models: Self-Referential Generation
- 7.4 Diffusion Language Models: Text Generation Beyond Autoregression
- 7.5 Jev and System One: Fast vs Deliberate AI Thinking
Step 4: Adapt
Module 8: Teaching and Shaping Models (11 lessons, about 5 hours). In this module, we will learn how a pre-trained model is adapted to our own task, how it is made smaller, and how it is taught to follow instructions and human preferences.
- 8.1 Fine-Tuning: Adapting a Pre-Trained Model to Your Task
- 8.2 LoRA: Parameter-Efficient Fine-Tuning via Low-Rank Matrices
- 8.3 Prefix Tuning: Learnable Context Prepended to the Input
- 8.4 Knowledge Distillation: Compressing Large Models into Small Ones
- 8.5 Continual Learning: Training Without Forgetting the Past
- 8.6 Deep RL from Human Preferences: The Foundational Paper
- 8.7 InstructGPT: Teaching GPT-3 to Follow Instructions
- 8.8 RLHF: Aligning LLMs with Human Preferences
- 8.9 PPO: The Reinforcement Algorithm Behind Instruction Tuning
- 8.10 DPO: Alignment Without the Separate Reward Model
- 8.11 GRPO: Group-Based Preference Optimization Explained
Step 5: Build
Module 9: The Art of Prompting (5 lessons, about 2 hours). In this module, we will learn how to talk to an LLM so that it gives better answers, and how to manage everything that goes into its context window.
- 9.1 Chain-of-Thought Prompting: Making Models Reason Step by Step
- 9.2 Prompt Chaining: Decomposing Complex Tasks into Steps
- 9.3 Prompt Caching: Reusing Computation Across API Calls
- 9.4 Context Engineering: Curating the Model's Working Memory
- 9.5 Context Compaction: Fitting More Into a Finite Window
Module 10: Building RAG Systems (13 lessons, about 6 hours). In this module, we will learn how to give an LLM knowledge that it was never trained on. We will start with how vectors are stored and searched, then move to retrieval techniques, and finally to the advanced forms of RAG.
- 10.1 Vector Databases: Storing and Searching Embeddings at Scale
- 10.2 ANN Search: Finding Similar Vectors Without Brute Force
- 10.3 Semantic Search: Finding Meaning, Not Just Keywords
- 10.4 Hybrid Search: Combining Sparse and Dense Retrieval
- 10.5 Rerankers: Re-Scoring Retrieved Results by Relevance
- 10.6 ColBERT: Token-Level Late Interaction for Retrieval
- 10.7 Document Chunking Strategies for RAG
- 10.8 HyDE: Generating Hypothetical Documents to Improve RAG
- 10.9 Embedding Caches: Avoiding Redundant Embedding Calls
- 10.10 Semantic Caching: Skipping the LLM for Similar Queries
- 10.11 Agentic RAG: Dynamic Retrieval with Multi-Step Reasoning
- 10.12 GraphRAG: Combining Knowledge Graphs with Retrieval
- 10.13 Vectorless RAG: Retrieval Without Embeddings or a Vector Store
Module 11: Autonomous AI Agents (16 lessons, about 6 hours). In this module, we will learn how an LLM goes from answering questions to actually doing work. We will start with a single agent, see how it uses tools and memory, and then move to systems where many agents work together.
- 11.1 AI Agents: Autonomous Decision-Making Systems
- 11.2 Function Calling: Giving LLMs Tools to Act on the World
- 11.3 The Agent Loop: Observe, Think, Act, Repeat
- 11.4 ReAct Agents: Interleaving Reasoning and Acting
- 11.5 Plan-and-Execute: Tackling Complex Tasks in Two Phases
- 11.6 Reflection Agents: Self-Critique for Higher-Quality Outputs
- 11.7 Agent Memory: Short-Term, Long-Term, and Episodic
- 11.8 Model Context Protocol: A Standard Interface for Agent Tools
- 11.9 Agent Skills: Reusable Capabilities in Agentic Systems
- 11.10 Open Knowledge Format: Structured Agent-to-Agent Communication
- 11.11 Multi-Agent Systems: Dividing Work Among Specialist Agents
- 11.12 Subagents: Delegating Tasks Within an Agent Network
- 11.13 Agent Communication: Protocols and Message Formats
- 11.14 AI Orchestration: Coordinating Agents, Tools, and Flows
- 11.15 Sakana Fugu: Lessons from an Open-Source Agent Study
- 11.16 Computer-Use Agents: Controlling Interfaces with AI
Module 12: Agent Patterns and Frameworks (8 lessons, about 4 hours). In this module, we will learn the engineering practices for building reliable agents, and then see how the popular frameworks and coding agents are built.
- 12.1 Harness Engineering: The Scaffolding Around AI Agents
- 12.2 Loop Engineering: Designing Reliable Agentic Loops
- 12.3 Graph Engineering: Stateful Workflows for Agents
- 12.4 Defining Done: Why Exit Criteria Shape Agent Quality
- 12.5 LangChain: Composable Components for LLM Applications
- 12.6 LangGraph: Graph-Based Agent Orchestration Explained
- 12.7 Claude Code: AI-Powered Software Engineering at the CLI
- 12.8 Cursor: Inside an AI-Native Code Editor
Step 6: Production
Module 13: Serving LLMs at Scale (17 lessons, about 7 hours). In this module, we will learn how to make LLMs faster and cheaper to run. We will start with what happens during inference, then learn the caching, batching, and speculation techniques, then quantization, and finally the serving engines that put it all together.
- 13.1 LLM Inference Optimization: The Full Landscape
- 13.2 Prefill vs Decode: Two Distinct Phases of LLM Inference
- 13.3 Prefill-Decode Disaggregation: Splitting the Two Phases
- 13.4 The KV Cache: Avoiding Redundant Attention Computation
- 13.5 KV Cache Compression: Trading Some Accuracy for Speed
- 13.6 Paged Attention: OS-Inspired Memory Management for KV Caches
- 13.7 Continuous Batching: Keeping GPUs Busy Between Requests
- 13.8 Speculative Decoding: Draft Fast, Verify in Parallel
- 13.9 N-gram Speculation: Draft Tokens Without a Draft Model
- 13.10 Medusa: Parallel Decoding via Multiple Prediction Heads
- 13.11 EAGLE: Feature-Level Drafting for Faster Inference
- 13.12 Model Quantization: Shrinking Weights Without Breaking Outputs
- 13.13 GGUF: The File Format Powering Local LLM Inference
- 13.14 llama.cpp: Running Large Models on Consumer Hardware
- 13.15 vLLM: High-Throughput Serving with PagedAttention
- 13.16 SGLang: Structured LLM Programs for Efficient Inference
- 13.17 TensorRT-LLM: NVIDIA's Optimized Inference Engine
Module 14: Measuring What Matters (4 lessons, about 2 hours). In this module, we will learn how to measure whether our LLM and our agent are actually doing a good job, and how to see what they are doing in production.
- 14.1 Evaluating LLMs: Metrics, Benchmarks, and Methods
- 14.2 LLM-as-Judge: Automating Evaluation with Another Model
- 14.3 Evaluating AI Agents: Metrics and Methods That Work
- 14.4 Agent Observability: Traces, Spans, and Debug Signals
Module 15: Securing AI Systems (3 lessons, about 1 hours). In this module, we will learn how to keep an LLM application safe, how attackers try to break it, and how AI-generated text can be identified.
- 15.1 LLM Guardrails: Filtering Inputs and Outputs for Safety
- 15.2 Prompt Injection: Attacks Against LLM-Powered Systems
- 15.3 LLM Watermarking: Embedding Invisible Signatures in AI Text
Step 7: Frontier
Module 16: Beyond Text: Multimodal AI (6 lessons, about 3 hours). In this module, we will learn how AI works with images and other types of data, and the generative models that create images from noise.
- 16.1 Multimodal AI: Perceiving Text, Images, and Audio Together
- 16.2 Vision Transformers: Applying Self-Attention to Image Patches
- 16.3 Image Embeddings: Encoding Visual Content as Vectors
- 16.4 Diffusion Models: Iterative Denoising to Generate Images
- 16.5 GANs: A Generator and Discriminator in Constant Competition
- 16.6 Variational Autoencoders: Learning a Compressed Latent Space
Step 8: Production
Module 17: Production AI Infrastructure (11 lessons, about 5 hours). In this module, we will learn the hardware that runs AI models, where to deploy a model, how to send each request to the right model, and how to design a complete AI system end to end.
- 17.1 GPUs for Deep Learning: Parallelism at the Core
- 17.2 CUDA Kernels: Writing Parallel Code for NVIDIA GPUs
- 17.3 Google TPUs: Purpose-Built Hardware for Neural Networks
- 17.4 Language Processing Units: A New Approach to LLM Inference
- 17.5 Cloud vs Edge: Where Should Your Model Run?
- 17.6 On-Device ML: A TensorFlow Lite Android Walkthrough
- 17.7 LLM Routing: Directing Each Query to the Best Model
- 17.8 Building a Real-Time Voice AI Agent from Scratch
- 17.9 System Design Fundamentals for AI Engineers
- 17.10 Transport Protocols: HTTP, WebSockets, and SSE Compared
- 17.11 How do Voice And Video Call Work?
Step 9: Frontier
Module 18: The Edge of AI Research (3 lessons, about 1 hours). In this module, we will learn the ideas that are shaping the future of AI, from models that learn an internal picture of the world to systems that improve themselves.
- 18.1 JEPA: LeCun's Vision for World Model AI
- 18.2 World Models: Teaching AI to Simulate Its Environment
- 18.3 Recursive Self-Improvement: Can AI Improve Itself Indefinitely?
Step 10: Career
Module 19: AI Engineering Career Prep (1 lessons, about 1 hours). We have learned everything from machine-learning foundations to AI agents in production. Now we turn that knowledge into clear interview answers and confident system designs.
How long does it take to become an AI engineer?
About 63 hours of lessons, labs and quizzes.
- At hour a day: about 9 weeks.
- At hours a day: about 5 weeks.
- At hours a week: about 7 weeks.