Reference
AI engineering glossary
Quick definitions of the most important terms. Each links to the lesson that teaches it in depth.
- Agent Skill
- An Agent Skill is a folder of instructions, and optionally scripts and reference files, that an AI agent loads by itself only when the task actually needs it. Lesson 11.9
- Agentic RAG
- Agentic RAG is a system where an AI Agent drives the retrieval process. Lesson 10.11
- AI Agent
- AI Agent = An LLM + Instructions + Tools + Memory + A loop that runs until the goal is achieved. Lesson 11.1
- AI Agent Observability
- AI Agent Observability is the practice of recording and understanding everything an AI Agent does internally, step by step, so that we can see why it behaved the way it did. Lesson 14.4
- AI Orchestration
- AI Orchestration is the process of coordinating multiple AI components, such as LLMs, tools, data sources, and agents, to work together to finish a complex task. Lesson 11.14
- AI SubAgent
- An AI SubAgent is a smaller, specialized agent that works under a main agent to handle a specific part of a larger task. Lesson 11.12
- Backpropagation
- Backpropagation is a method used to calculate how much each weight in a neural network contributed to the error, so that we can adjust those weights to reduce the error. Lesson 3.3
- BPE (Byte Pair Encoding)
- BPE (Byte Pair Encoding) is a tokenization algorithm that breaks text into pieces that are somewhere between characters and words. Lesson 4.3
- Chain-of-Thought (CoT) Prompting
- Chain-of-Thought (CoT) Prompting is a technique where we ask the model to write out its reasoning steps before giving the final answer. Lesson 9.1
- Chunk
- A chunk is a small piece of text that we cut out from a bigger document. Lesson 10.7
- Claude Code
- Claude Code is a coding agent from Anthropic that runs in the terminal. We give it a task in plain English, and it completes the task by reading our code, editing files, running commands, and checking its own work. Lesson 12.7
- Context Compaction
- Context compaction is the technique of shrinking the old conversation into a short summary, so the important facts stay while the box gets free space again. Lesson 9.5
- Context Engineering
- Context Engineering is the practice of designing, organizing, and managing everything that goes into an LLM's context window so that the model can do its task reliably. Lesson 9.4
- Continual Learning
- Continual Learning is the ability of a model to keep learning new information over time, without forgetting what it has already learned. Lesson 8.5
- Continuous Batching
- Continuous Batching is a way of running batches where the server does not wait for the whole batch to finish. The moment any request in the batch finishes, the server immediately replaces it with a new request that is waiting in the queue. Lesson 13.7
- Contrastive Learning
- Contrastive Learning is a way of teaching a model to learn good representations of data by comparing things. The model learns to pull similar things close to each other and push dissimilar things far apart in a representation space. Lesson 2.9
- Diffusion Language Model
- A Diffusion Language Model is a type of AI model that writes text by starting from a piece of pure gibberish and slowly cleaning it up into a clear, meaningful sentence. Lesson 7.4
- Diffusion Model
- A Diffusion Model is a type of AI model that learns to create new data, such as images, by starting from pure random noise and slowly cleaning it up step by step until a clear image appears. Lesson 16.4
- Embedding
- An embedding is a list of numbers that represents the meaning of something, arranged so that things with similar meaning get similar numbers. Lesson 4.4
- Fine-tuning
- Fine-tuning is the process of taking a model that is already trained and training it a little more on our own specific data so that it becomes good at our specific task. Lesson 8.1
- Function Calling
- Function Calling is a way to let an LLM use external tools, APIs, and functions to get things done. It is also called tool calling, and both names mean the same thing. Lesson 11.2
- Generative AI
- Generative AI is a type of artificial intelligence that can create new things, like text, images, audio, video, and code. Lesson 4.1
- GGUF
- GGUF is a single file format that stores everything needed to run a large language model for local inference, all in one self-contained file. Lesson 13.13
- Gradient Descent
- Gradient Descent means going downward in the direction of the steepest slope. Lesson 3.2
- Graph Engineering
- Graph Engineering is the practice of designing an AI system as a graph, where every step of the work is a node and every path from one step to another step is an edge. Lesson 12.3
- Grouped-Query Attention (GQA)
- Grouped-Query Attention (GQA) is a strategy where heads are divided into groups, and all heads within a group share the same Key and Value, while each head still has its own Query. Lesson 6.3
- Hybrid Search
- Hybrid Search is a technique that combines keyword search and semantic search, and merges their results into one final ranked list. Lesson 10.4
- Knowledge Distillation
- Knowledge Distillation is a technique where we train a small model to copy the behavior of a large model. Lesson 8.4
- KV Cache Compression
- KV Cache Compression is the set of techniques that make the KV Cache smaller while keeping the quality of the model output almost the same. Lesson 13.5
- LangChain
- LangChain is a framework that helps us build applications powered by Large Language Models. Lesson 12.5
- LangGraph
- LangGraph is a framework that helps us build applications powered by an LLM, where the work is organized as a graph of steps. Lesson 12.6
- Language Model
- A Language Model is a neural network trained to predict the next token (i.e. the next small chunk of text) based on the previous tokens. Lesson 7.1
- Large Reasoning Model (LRM)
- Large Reasoning Model = A Large Language Model that is trained to think first, and answer later. Lesson 7.2
- LLM Architecture
- An LLM Architecture is the blueprint of a large language model. It describes how the model reads text, how it remembers what it has read, and how it produces the next word. Lesson 6.1
- LLM as a Judge
- LLM as a Judge is a technique where we use a large language model to evaluate the output of another large language model. Lesson 14.2
- LLM Evaluation
- LLM Evaluation is the process of measuring how well a Large Language Model performs on the tasks we expect it to do. Lesson 14.1
- LLM Guardrails
- LLM guardrails are safety checks that sit around an LLM to control what goes in and what comes out. Lesson 15.1
- LLM Routing
- LLM Routing is the practice of choosing the right LLM for each user query, instead of sending every query to the same LLM. Lesson 17.7
- Loop Engineering
- Loop Engineering is the practice of designing the repeating cycle that an AI agent runs, so that the agent keeps making real progress on a task and stops at the right moment with the right result. Lesson 12.2
- LoRA
- LoRA is a way to fine-tune a large model without updating all of its weights. Instead of changing the original weight matrix, we keep it frozen and learn a tiny pair of extra matrices on the side. Lesson 8.2
- Lost in the Middle
- The Lost in the Middle problem is the behaviour where an LLM pays strong attention to the information placed at the beginning and at the end of a long input, and pays very less attention to the information placed in the middle. Lesson 5.4
- LPU
- An LPU is a chip that is built for one single job, running a large language model that is already trained, and producing text as fast as possible. Lesson 17.4
- MCP (Model Context Protocol)
- MCP, which stands for Model Context Protocol, is an open standard that defines one common way for AI applications to connect to outside tools and data. Lesson 11.8
- Model Quantization
- Model Quantization is the process of storing and computing a model's numbers at lower precision, so the model takes less memory and runs faster. Lesson 13.12
- Multi-Head Attention
- Multi-Head Attention is a mechanism that runs many Self Attention operations in parallel, each with its own set of Q, K, and V projections, and then combines their outputs into a single richer representation. Lesson 4.12
- OKF (Open Knowledge Format)
- OKF, which stands for Open Knowledge Format, is an open standard for writing down what an organization knows about its data and systems, as a folder of plain markdown files, so that any AI agent or any tool can read that knowledge without custom work. Lesson 11.10
- Paged Attention
- Paged Attention is a technique that manages KV Cache memory more efficiently by breaking it into small, fixed-size blocks called pages. Lesson 13.6
- Prefill
- Prefill is the phase where the model reads and processes your entire input prompt in one single pass and produces the very first output token. Lesson 13.2
- Prompt Caching
- Prompt Caching is a technique where the model saves the work it already did for a repeated part of a prompt, so that next time it can reuse that saved work instead of doing it all over again. Lesson 9.3
- Prompt Chaining
- Prompt Chaining is a way of breaking one big task into smaller prompts, where the output of one prompt becomes the input of the next prompt. Lesson 9.2
- Prompt Injection
- Prompt Injection is an attack where someone slips their own instructions into the text that an AI application sends to the model, so that the model follows the attacker's instructions instead of the developer's instructions. Lesson 15.2
- RAG (Retrieval-Augmented Generation)
- RAG stands for Retrieval-Augmented Generation. It is a way to make an AI model answer using our own documents instead of only using what it already knows. Lesson 10.8
- ReAct Agent
- A ReAct Agent is an AI Agent built using the ReAct (Reasoning + Acting) pattern - the most common pattern for building AI Agents. Lesson 11.4
- Recursive Self-Improvement
- Recursive Self-Improvement is a process in which an AI system improves its own abilities, and then the improved version improves itself further, and this cycle keeps repeating. Lesson 18.3
- Reinforcement Learning
- Reinforcement Learning, often called RL, is a type of machine learning where an Agent learns to make a sequence of decisions by interacting with an Environment, with the goal of maximizing a Reward over time. Lesson 2.8
- Reranker
- A Reranker is a model that takes a list of documents and reorders them, putting the most relevant ones at the top for a given question. Lesson 10.5
- RLHF
- RLHF (Reinforcement Learning from Human Feedback) is a training technique where we teach a Large Language Model (LLM) to produce responses that humans prefer, by collecting human preferences and converting them into a reward signal that guides further training. Lesson 8.8
- Self Attention
- Self Attention is a mechanism that allows every token in a sequence to look at every other token in the same sequence, including itself, to understand the context. Lesson 4.8
- Semantic Caching
- Semantic Caching is a cache that matches questions by their meaning instead of their exact words. Lesson 10.10
- Speculative Decoding
- Speculative Decoding is a technique where we first guess the next few tokens quickly, and then ask the big model to verify all those guesses in one single run. Lesson 13.9
- Token Streaming
- Token streaming is a technique where the server sends the model's reply to us piece by piece, as each piece is produced, instead of waiting for the whole reply to be ready. Lesson 5.3
- Tokenization
- The first step is to break the text into small pieces called tokens. Each token is then converted into a number. This process of breaking text into tokens is called tokenization. Lesson 4.3
- Top-p Sampling
- Top-p Sampling is a decoding strategy in which we keep the smallest group of top tokens whose probabilities add up to at least p, throw away all the others, and then pick randomly from that group. Lesson 5.2
- Transformer
- A Transformer is the architecture behind most modern AI models that work with language. Lesson 4.7
- Variational Autoencoder (VAE)
- A Variational Autoencoder, also called a VAE, is a special type of Autoencoder that learns a smooth and organized latent space, so that we can pick any random point from it and generate brand new, meaningful data. Lesson 16.6
- vLLM
- vLLM is a high-throughput engine for serving LLMs. It is built to serve as many requests as possible on a GPU by managing the KV cache memory very efficiently. Lesson 13.15
- Voice AI Agent
- A Voice AI Agent is a software program that we can talk to using our voice, and it talks back to us, just like a phone call with a human, but the one on the other side is an AI. Lesson 17.8
- World Model
- A World Model is an AI that learns an approximate internal copy of how an environment behaves, so that it can predict what happens next, given the current state and an action. Lesson 18.2