Modern AI Engineering

Interactive lab

Play with every idea

42 hands-on simulations from across the course. Each one links to the lesson that explains it.

  • temperature: Slider for temperature; bars show softmax probabilities of next-token candidates reshaping live.
  • top k top p: Sliders for k and p over a next-token distribution; kept vs removed tokens and renormalised probabilities.
  • gradient descent: A ball steps down a 1-D loss curve; learning-rate slider, play/step/reset; shows converge, slow, overshoot, diverge.
  • vector similarity: Two 2-D vectors; drag angle and length; live cosine similarity, dot product and Euclidean distance.
  • precision recall: Threshold slider over scored examples; live confusion matrix, precision, recall, F1.
  • linear vs logistic: Toggle between a linear regression line on continuous data and a logistic sigmoid on 0/1 data.
  • loss l1 l2: Error slider and an outlier toggle; plots |e| vs e² and shows how each loss reacts to outliers.
  • regularization: Lambda slider; bar chart of model weights shrinking under L1 (many hit zero) vs L2 (all shrink smoothly).
  • rl gridworld: A Q-learning agent learns a path through a grid world live; play/step, reward-per-episode chart, epsilon slider.
  • contrastive: Anchor, positive and negative embeddings; step training to pull positives closer and push negatives away.
  • neuron: Inputs × weights + bias → activation (ReLU, sigmoid, tanh); sliders for weights and bias.
  • dropout: A small network; drop-rate slider and resample button; dropped neurons fade out, survivors are rescaled.
  • normalization: A batch × features grid; switch BatchNorm / LayerNorm / RMSNorm to see which axis is normalised and the numbers.
  • rnn vs transformer: Animated race: an RNN reads tokens one by one while a Transformer processes them in parallel.
  • tokenizer bpe: Step through real Byte Pair Encoding merges on a small corpus; vocabulary and token count update.
  • embedding space: 2-D map of word embeddings; click a word to highlight its nearest neighbours by cosine similarity.
  • attention heatmap: Click a token to see its attention weights over the sentence; toggle the causal mask.
  • qkv math: Step-by-step numeric walk-through of Q·Kᵀ, scaling by √dₖ, softmax and the weighted sum of V.
  • causal mask: An attention score matrix fills row by row as the upper triangle is masked to −∞ and softmax runs.
  • multi head: Switch between attention heads that learned different patterns (previous-token, subject, punctuation).
  • rope: Position slider rotates query/key vectors; shows that the dot product depends only on relative distance.
  • streaming: Side-by-side timer: a blocking response vs tokens streamed as they are generated (TTFT vs total time).
  • lost in middle: Move the relevant document through the context; accuracy follows a U-shaped curve.
  • moe router: Tokens flow through a router to their top-2 of 8 experts; shows active vs total parameters.
  • gqa: Toggle MHA / GQA / MQA; query heads share key-value heads and the KV-cache size bar updates.
  • sliding window: Window-size slider over an attention mask; shows how far information can travel across layers.
  • lora: Sliders for matrix size d and rank r; frozen W plus B·A, trainable parameter count and percentage.
  • kv cache: Step through decoding; keys and values append per token; memory calculator for layers, heads, context.
  • paged attention: Requests grow and finish; contiguous allocation fragments memory while paged blocks pack it.
  • continuous batching: Timeline of GPU slots: static batching waits for the longest request; continuous batching refills slots.
  • speculative decoding: A draft model proposes k tokens; the target verifies; accepted and rejected tokens animate; speedup estimate.
  • quantization: Bit-width slider; weights snap to the grid of representable values; shows error and memory size.
  • chunking: Chunk-size and overlap sliders over a document; coloured chunks and chunk count update.
  • ann search: 2-D points with a query; brute force vs IVF clusters with an nprobe slider; distance computations and recall.
  • hybrid search: Keyword and vector result lists merged with Reciprocal Rank Fusion; weight slider reorders the final list.
  • semantic cache: Similarity-threshold slider; incoming queries become cache hits, misses or false hits.
  • agent loop: Animated agent trace: think → call tool → observe → repeat until done, with the message log.
  • prompt caching: Slider for the shared prefix share; latency and cost bars for cached vs uncached requests.
  • diffusion: Step slider takes a small image from pure noise to clean (reverse diffusion) and back (forward noising).
  • guardrails: Pick a sample prompt; it passes through input checks, the model and output checks; see what is blocked.
  • gpu parallel: A CPU with a few fast cores vs a GPU with many simple cores working through the same matrix.
  • llm routing: Queries of different difficulty are routed to a small or large model; threshold slider trades cost vs quality.