Modern AI Engineering

Course guide

Everything you need before lesson one

A structured, step-by-step path to learn AI engineering from scratch: 19 modules and 149 interactive lessons, from the first idea of machine learning to designing complete AI systems.

About this course

We start with the basics of machine learning, then go inside the Transformer, see how an LLM generates text, learn how models are fine-tuned and aligned, build RAG systems and AI agents, make models fast and cheap to serve, evaluate and secure them, and finally design whole AI systems end to end.

Every lesson here is self-contained. It explains one concept in simple words, then shows it moving: animated diagrams, charts that draw themselves, widgets you can drag, real code with real output, and side-by-side comparisons. You should never need to open another tab to understand a lesson.

  • Open lessons: Every lesson and quiz is open to read. Your progress is saved in your browser, or in your account if you sign in.
  • Beginner first: Every term is defined the first time it appears. Math is shown with small numbers you can follow.
  • Deep, not shallow: Each lesson goes from “why do we need it” to “how it works step by step” to “where it is used”.
  • In order, with checks: Each lesson ends with a 5-question quiz. Score 4 or more to pass the lesson.

What is AI engineering?

AI engineering is the discipline of building real applications and systems on top of AI models, especially large language models (LLMs).

An AI engineer rarely trains a giant model from scratch. Instead, they understand how models work inside, choose the right model, give it the right context, connect it to tools and data, make it fast and affordable to run, measure whether it is doing a good job, keep it safe, and ship it to real users.

AI engineering = Understand the model + Build on top of the model + Run the model in production

Who is this course for?

  • Software engineers moving into AI engineering.
  • Backend, mobile and frontend developers who want to build AI-powered products.
  • ML engineers and data scientists who want to go deep into LLMs, RAG and agents.
  • Students and freshers starting a career in AI.
  • Engineering managers and tech leads who need to understand how modern AI systems are built.
  • Interview candidates preparing for AI engineer, GenAI engineer and LLM engineer roles.

What will we learn?

  • Module 1: AI Engineering Starter Kit: Six Concepts Every AI Engineer Must Know
  • Module 2: Learning from Data: Machine Learning from First Principles · Labeled vs Unlabeled: Two Ways Machines Learn · Predicting Numbers vs Categories: Regression Compared · Feature Engineering: Turning Raw Data into Signal · Precision and Recall: Picking the Right Metric · L1 vs L2 Loss: Choosing Your Error Penalty
  • Module 3: Neural Architectures Deep Dive: Neural Network Bias: What It Is and Why It Matters · Gradient Descent: Rolling Downhill to the Optimum · Backpropagation: How Neural Networks Learn from Mistakes · Cross-Entropy Loss: Scoring Probability Predictions · Dropout: Controlled Forgetting as Regularization · Batch Norm vs Layer Norm: When to Use Each
  • Module 4: Transformers and How They Think: Generative AI: Creating Instead of Classifying · Autoregressive Models: Predicting One Token at a Time · BPE Tokenization: How LLMs Split Text into Tokens · Embeddings: Encoding Meaning as Vectors · RNNs vs Transformers: A Fundamental Architecture Shift · The Transformer Architecture: Built on Attention
  • Module 5: Inside the LLM Output Pipeline: Temperature Sampling: Dialing Up or Down Creativity · Nucleus Sampling: Top-k and Top-p Demystified · Token Streaming: Rendering Outputs as They Arrive · Lost in the Middle: Why LLMs Miss Central Context
  • Module 6: Next-Gen LLM Architectures: A Timeline of LLM Architecture Improvements · Mixture of Experts: Routing Tokens to Specialists · Grouped Query Attention: Fewer KV Heads, Same Quality · Sliding Window Attention: Taming Very Long Contexts · Attention Sinks: The Hidden Cost of Extended Context · Flash Attention: Memory-Efficient Attention at Scale
  • Module 7: The Language Model Zoo: Small Language Models: Big Capability in Compact Form · Large Reasoning Models: Chain-of-Thought at Inference Time · Recursive Language Models: Self-Referential Generation · Diffusion Language Models: Text Generation Beyond Autoregression · Jev and System One: Fast vs Deliberate AI Thinking
  • Module 8: Teaching and Shaping Models: Fine-Tuning: Adapting a Pre-Trained Model to Your Task · LoRA: Parameter-Efficient Fine-Tuning via Low-Rank Matrices · Prefix Tuning: Learnable Context Prepended to the Input · Knowledge Distillation: Compressing Large Models into Small Ones · Continual Learning: Training Without Forgetting the Past · Deep RL from Human Preferences: The Foundational Paper
  • Module 9: The Art of Prompting: Chain-of-Thought Prompting: Making Models Reason Step by Step · Prompt Chaining: Decomposing Complex Tasks into Steps · Prompt Caching: Reusing Computation Across API Calls · Context Engineering: Curating the Model's Working Memory · Context Compaction: Fitting More Into a Finite Window
  • Module 10: Building RAG Systems: Vector Databases: Storing and Searching Embeddings at Scale · ANN Search: Finding Similar Vectors Without Brute Force · Semantic Search: Finding Meaning, Not Just Keywords · Hybrid Search: Combining Sparse and Dense Retrieval · Rerankers: Re-Scoring Retrieved Results by Relevance · ColBERT: Token-Level Late Interaction for Retrieval
  • Module 11: Autonomous AI Agents: AI Agents: Autonomous Decision-Making Systems · Function Calling: Giving LLMs Tools to Act on the World · The Agent Loop: Observe, Think, Act, Repeat · ReAct Agents: Interleaving Reasoning and Acting · Plan-and-Execute: Tackling Complex Tasks in Two Phases · Reflection Agents: Self-Critique for Higher-Quality Outputs
  • Module 12: Agent Patterns and Frameworks: Harness Engineering: The Scaffolding Around AI Agents · Loop Engineering: Designing Reliable Agentic Loops · Graph Engineering: Stateful Workflows for Agents · Defining Done: Why Exit Criteria Shape Agent Quality · LangChain: Composable Components for LLM Applications · LangGraph: Graph-Based Agent Orchestration Explained
  • Module 13: Serving LLMs at Scale: LLM Inference Optimization: The Full Landscape · Prefill vs Decode: Two Distinct Phases of LLM Inference · Prefill-Decode Disaggregation: Splitting the Two Phases · The KV Cache: Avoiding Redundant Attention Computation · KV Cache Compression: Trading Some Accuracy for Speed · Paged Attention: OS-Inspired Memory Management for KV Caches
  • Module 14: Measuring What Matters: Evaluating LLMs: Metrics, Benchmarks, and Methods · LLM-as-Judge: Automating Evaluation with Another Model · Evaluating AI Agents: Metrics and Methods That Work · Agent Observability: Traces, Spans, and Debug Signals
  • Module 15: Securing AI Systems: LLM Guardrails: Filtering Inputs and Outputs for Safety · Prompt Injection: Attacks Against LLM-Powered Systems · LLM Watermarking: Embedding Invisible Signatures in AI Text
  • Module 16: Beyond Text: Multimodal AI: Multimodal AI: Perceiving Text, Images, and Audio Together · Vision Transformers: Applying Self-Attention to Image Patches · Image Embeddings: Encoding Visual Content as Vectors · Diffusion Models: Iterative Denoising to Generate Images · GANs: A Generator and Discriminator in Constant Competition · Variational Autoencoders: Learning a Compressed Latent Space
  • Module 17: Production AI Infrastructure: GPUs for Deep Learning: Parallelism at the Core · CUDA Kernels: Writing Parallel Code for NVIDIA GPUs · Google TPUs: Purpose-Built Hardware for Neural Networks · Language Processing Units: A New Approach to LLM Inference · Cloud vs Edge: Where Should Your Model Run? · On-Device ML: A TensorFlow Lite Android Walkthrough
  • Module 18: The Edge of AI Research: JEPA: LeCun's Vision for World Model AI · World Models: Teaching AI to Simulate Its Environment · Recursive Self-Improvement: Can AI Improve Itself Indefinitely?
  • Module 19: AI Engineering Career Prep: Cracking the AI Engineering Interview

Prerequisites

  • Basic programming: Preferably Python. Most code examples are short Python programs, and every one shows its real output.
  • High-school math: The linear algebra, calculus and probability we need is explained inside the lessons, step by step.
  • Curiosity: That is all. No prior AI or machine-learning background is needed.

How to use this course

  1. Follow the modules in order. Each module builds on the previous one.
  2. Inside a lesson, play with every interactive: move the sliders, step through the animations, press Run on the code.
  3. Use the “Pause and think” checks. Predict the answer before you reveal it.
  4. Take the 5-question quiz at the end. You need 4 correct answers to pass the lesson. Wrong answers come with explanations, and you can retry as often as you like.
  5. Do not skip Module 2 and Module 3. Everything later is built on them.
  6. After each module, explain its ideas to a friend in your own words. If you can explain it, you have learned it.