Modern AI Engineering

Module 14 · Production

Measuring What Matters

In this module, we will learn how to measure whether our LLM and our agent are actually doing a good job, and how to see what they are doing in production.

By the end of this module, we will be able to build an evaluation suite and trace every step of an agent in production.

Lessons

  1. 14.1 Evaluating LLMs: Metrics, Benchmarks, and Methods: What is LLM Evaluation? · Why do we need LLM Evaluation? · Types of LLM Evaluation · Automatic Metrics · Benchmarks · Human Evaluation · LLM as a Judge · Task-Specific Evaluation · Safety and Red-Teaming Evaluation · Challenges in LLM Evaluation · Best Practices · When to use which method
  2. 14.2 LLM-as-Judge: Automating Evaluation with Another Model: What is LLM as a Judge? · Why do we need LLM as a Judge? · How does LLM as a Judge work? · Types of LLM as a Judge. · Steps to build an LLM Judge. · A prompt template for LLM as a Judge. · Chain-of-thought judging (G-Eval). · Biases in LLM as a Judge. · Best practices for LLM as a Judge. · Real-world use cases of LLM as a Judge.
  3. 14.3 Evaluating AI Agents: Metrics and Methods That Work: What is an AI Agent? · What is AI Agent Evaluation? · Why do we need AI Agent Evaluation? · How is AI Agent Evaluation different from LLM Evaluation? · Types of AI Agent Evaluation · Outcome Evaluation · Trajectory Evaluation · Tool Use Evaluation · Planning Evaluation · Key Metrics for AI Agents · Agent Benchmarks · Methods to Evaluate AI Agents · Frameworks and Tools for AI Agent Evaluation · Challenges in AI Agent Evaluation · Best Practices
  4. 14.4 Agent Observability: Traces, Spans, and Debug Signals: What is an AI Agent? · What is Observability? · What is AI Agent Observability? · Why do we need AI Agent Observability? · How is AI Agent Observability different from traditional Observability? · The Three Pillars of Observability · Traces and Spans · What should we observe inside an AI Agent? · Key Metrics for AI Agent Observability · How AI Agent Observability works · Tools and Frameworks for AI Agent Observability · Observability vs Evaluation · Challenges in AI Agent Observability · Best Practices

← Module 13: Serving LLMs at Scale · Module 15: Securing AI Systems →