Interactive lab
Play with every idea
42 hands-on simulations from across the course. Each one links to the lesson that explains it.
- temperature: Slider for temperature; bars show softmax probabilities of next-token candidates reshaping live.
- top k top p: Sliders for k and p over a next-token distribution; kept vs removed tokens and renormalised probabilities.
- gradient descent: A ball steps down a 1-D loss curve; learning-rate slider, play/step/reset; shows converge, slow, overshoot, diverge.
- vector similarity: Two 2-D vectors; drag angle and length; live cosine similarity, dot product and Euclidean distance.
- precision recall: Threshold slider over scored examples; live confusion matrix, precision, recall, F1.
- linear vs logistic: Toggle between a linear regression line on continuous data and a logistic sigmoid on 0/1 data.
- loss l1 l2: Error slider and an outlier toggle; plots |e| vs e² and shows how each loss reacts to outliers.
- regularization: Lambda slider; bar chart of model weights shrinking under L1 (many hit zero) vs L2 (all shrink smoothly).
- rl gridworld: A Q-learning agent learns a path through a grid world live; play/step, reward-per-episode chart, epsilon slider.
- contrastive: Anchor, positive and negative embeddings; step training to pull positives closer and push negatives away.
- neuron: Inputs × weights + bias → activation (ReLU, sigmoid, tanh); sliders for weights and bias.
- dropout: A small network; drop-rate slider and resample button; dropped neurons fade out, survivors are rescaled.
- normalization: A batch × features grid; switch BatchNorm / LayerNorm / RMSNorm to see which axis is normalised and the numbers.
- rnn vs transformer: Animated race: an RNN reads tokens one by one while a Transformer processes them in parallel.
- tokenizer bpe: Step through real Byte Pair Encoding merges on a small corpus; vocabulary and token count update.
- embedding space: 2-D map of word embeddings; click a word to highlight its nearest neighbours by cosine similarity.
- attention heatmap: Click a token to see its attention weights over the sentence; toggle the causal mask.
- qkv math: Step-by-step numeric walk-through of Q·Kᵀ, scaling by √dₖ, softmax and the weighted sum of V.
- causal mask: An attention score matrix fills row by row as the upper triangle is masked to −∞ and softmax runs.
- multi head: Switch between attention heads that learned different patterns (previous-token, subject, punctuation).
- rope: Position slider rotates query/key vectors; shows that the dot product depends only on relative distance.
- streaming: Side-by-side timer: a blocking response vs tokens streamed as they are generated (TTFT vs total time).
- lost in middle: Move the relevant document through the context; accuracy follows a U-shaped curve.
- moe router: Tokens flow through a router to their top-2 of 8 experts; shows active vs total parameters.
- gqa: Toggle MHA / GQA / MQA; query heads share key-value heads and the KV-cache size bar updates.
- sliding window: Window-size slider over an attention mask; shows how far information can travel across layers.
- lora: Sliders for matrix size d and rank r; frozen W plus B·A, trainable parameter count and percentage.
- kv cache: Step through decoding; keys and values append per token; memory calculator for layers, heads, context.
- paged attention: Requests grow and finish; contiguous allocation fragments memory while paged blocks pack it.
- continuous batching: Timeline of GPU slots: static batching waits for the longest request; continuous batching refills slots.
- speculative decoding: A draft model proposes k tokens; the target verifies; accepted and rejected tokens animate; speedup estimate.
- quantization: Bit-width slider; weights snap to the grid of representable values; shows error and memory size.
- chunking: Chunk-size and overlap sliders over a document; coloured chunks and chunk count update.
- ann search: 2-D points with a query; brute force vs IVF clusters with an nprobe slider; distance computations and recall.
- hybrid search: Keyword and vector result lists merged with Reciprocal Rank Fusion; weight slider reorders the final list.
- semantic cache: Similarity-threshold slider; incoming queries become cache hits, misses or false hits.
- agent loop: Animated agent trace: think → call tool → observe → repeat until done, with the message log.
- prompt caching: Slider for the shared prefix share; latency and cost bars for cached vs uncached requests.
- diffusion: Step slider takes a small image from pure noise to clean (reverse diffusion) and back (forward noising).
- guardrails: Pick a sample prompt; it passes through input checks, the model and output checks; see what is blocked.
- gpu parallel: A CPU with a few fast cores vs a GPU with many simple cores working through the same matrix.
- llm routing: Queries of different difficulty are routed to a small or large model; threshold slider trades cost vs quality.