Modern AI Engineering

Module 5 · Core

Inside the LLM Output Pipeline

In this module, we will learn how an LLM picks the next token, how we control its creativity, how the output reaches the user token by token, and where the context window fails.

By the end of this module, we will know exactly what happens between the prompt and the final answer, and which knobs change the output.

Lessons

  1. 5.1 Temperature Sampling: Dialing Up or Down Creativity: What is Temperature in LLMs? · How does an LLM pick the next token? · From scores to probabilities · Where does Temperature come into the picture? · Step-by-step example with numbers · Low Temperature · High Temperature · Temperature = 1 and Temperature = 0 · Why is it called Temperature? · When to use which Temperature? · Common mistakes while using Temperature
  2. 5.2 Nucleus Sampling: Top-k and Top-p Demystified: How an LLM picks the next token · The problem with always picking the best token · The problem with picking from every token · What is Top-k Sampling? · Step-by-step example of Top-k Sampling · The problem with Top-k Sampling · What is Top-p Sampling? · Step-by-step example of Top-p Sampling · Top-k vs Top-p Sampling · How Top-k and Top-p work with Temperature · When to use which one
  3. 5.3 Token Streaming: Rendering Outputs as They Arrive: What is token streaming · A quick recap of how an LLM generates text · Why we need streaming at all · What is SSE · How the HTTP connection stays open · The format of a streamed message · A full walkthrough from server to screen · The [DONE] marker that ends the stream · SSE vs WebSockets · Token streaming in the real world
  4. 5.4 Lost in the Middle: Why LLMs Miss Central Context: What is a context window · What is the Lost in the Middle problem · Let's understand it with an example · The U-shaped curve · Why does this happen · Where this hurts us in real life · How to test for the Lost in the Middle problem · How to solve the Lost in the Middle problem · Key points to remember · [Why is the context window limited in LLMs?](https://www.youtube.com/watch?v=CGIhxIaOg3M) (Video)

← Module 4: Transformers and How They Think · Module 6: Next-Gen LLM Architectures →