Modern AI Engineering

LLM Evaluation: How to Test an AI Application Properly

By Modern AI Engineering · · 8 min read

LLM evaluation is how you find out whether an AI application is good enough, and whether a change made it better or worse. The core of it is simple. Collect a fixed set of realistic inputs. Decide what a good output looks like for each. Run the system on all of them after every change, and score the results in a consistent way.

Without this, every decision about a prompt, a model or a retrieval setting is a guess. This article shows how to set evaluation up in small steps.

Why is testing an LLM application different?

Ordinary software is tested with exact checks. A function that adds two numbers either returns the right sum or does not. An LLM application breaks that pattern in three ways.

First, the output varies. The same input can give different text on different runs. Second, there are many correct answers. Two summaries can use different words and both be good. Third, small changes have wide effects. A prompt edit that fixes one case can quietly break five others.

So you need tests that judge qualities such as correctness and relevance, not exact strings. And you need to run them across many examples at once, not one at a time by hand.

Model evaluation vs application evaluation

Two different things are both called LLM evaluation.

Model evaluation compares models on public benchmarks. It answers a general question: how capable is this model on standard tasks? It is useful for making a shortlist.

Application evaluation tests your whole system on your own task: the prompt, the retrieval, the tools and the model together. A model that leads a benchmark may not be the best for your documents and your users. Only your own test set can tell you. The rest of this article is about this second kind.

How to build an evaluation set

The evaluation set, sometimes called a golden dataset, is the most valuable thing you will build. Start small and grow it.

  1. Collect thirty to fifty realistic inputs. Use real user questions if you have them.
  2. Include the hard cases: vague questions, questions with no answer in your data, and inputs that try to misuse the system.
  3. For each input, write the expected answer or the points a good answer must contain.
  4. For cases where the system should decline, write that down as the expected behaviour.
  5. Have someone who knows the subject review the expected answers.
  6. Add a new case each time you find a failure in real use, so the same mistake cannot come back unnoticed.

What should you measure?

Pick a few measures that match what your users care about. The table lists the common ones.

MeasureThe question it answersHow it is checked
CorrectnessDoes the answer match the expected answer?Comparison with a reference, by a person or a judge model.
FaithfulnessIs every claim supported by the supplied sources?Each claim is checked against the retrieved text.
RelevanceDoes the answer address the question asked?A judge model or human review.
CompletenessIs anything important missing?A checklist of required points.
FormatIs the output valid and well structured?A check in code, such as parsing the JSON.
SafetyDid it refuse what it should refuse?Rules and test cases written for that purpose.
Cost and speedWhat does a request cost, and how long does it take?Logged token counts and timings.

Three ways to score an answer

There are three kinds of grader, and a healthy setup uses all of them.

Checks in code are the cheapest and most reliable. Is the output valid JSON? Does it contain the order number? Is it under the length limit? Use them wherever a rule can be written.

Human review is the most trusted and the slowest. Use it to define what good means, to label a starting set and to audit the automatic graders.

LLM-as-a-judge sits between the two. A second model reads the input, the output and a rubric, and gives a score. This scales to thousands of cases. But a judge is a model too. It can be inconsistent, it can favour longer answers, and it can be too generous. Give it a narrow question with a clear rubric. Ask for a pass or a fail before you try fine scales. And compare its verdicts with human labels on a sample before you rely on it.

How do you evaluate a RAG system?

Test retrieval and generation separately. If you only score the final answer, you cannot tell which stage failed.

For retrieval, each test question needs the passages that should be found. Then you can measure recall, the share of the right passages that were retrieved, and precision, the share of retrieved passages that were right. Rank matters as well. A correct passage in tenth place may never reach the prompt.

For generation, hand the model the right passages and check faithfulness and completeness. A high faithfulness score means the answer stayed within the sources. It does not mean the answer was complete or useful, so measure those separately.

How do you evaluate an AI agent?

An agent takes many steps, so there is more to judge. Look at three levels.

  • The outcome. Was the task completed? Check the final state, such as the file written or the record updated, not only what the agent said.
  • The path. Did it choose sensible tools with correct arguments? How many steps did it take?
  • The cost. Tokens, time and the number of tool calls for each task.
  • Consistency. Run each task several times. An agent that succeeds on some runs and fails on others is not ready.

Evaluation after launch

Testing does not end at release. Real users send inputs you did not predict. Keep traces of requests: the prompt, the retrieved text, the tool calls and the output. Review a sample on a regular schedule. Watch signals such as user ratings, retries and questions that got no answer. Feed every new failure back into the evaluation set.

Run the full set again whenever you change the prompt, the model or the retrieval settings. Model providers update their models too, and an update can shift behaviour.

Module 14 of the AI Engineering Bootcamp, Measuring What Matters, has four lessons: evaluating LLMs, LLM-as-a-judge, evaluating agents and agent observability. The module on securing AI systems adds guardrails and prompt injection.

Frequently asked questions

What is LLM evaluation?

It is the process of measuring how well a language model or an application built on one performs. In practice it means running a fixed set of test inputs and scoring the outputs in a consistent way.

What is a golden dataset?

It is a set of test inputs paired with trusted expected answers or labels. You run your system against it after each change to see whether quality went up or down.

What is LLM-as-a-judge?

It is using a language model to grade the output of another model against a rubric. It scales well, but its verdicts should be checked against human labels on a sample.

How many test cases do I need to evaluate an LLM application?

Start with thirty to fifty realistic cases that cover normal use and hard cases. Add more over time, especially from failures you see in real use.

What is the difference between evals and benchmarks?

A benchmark is a public test used to compare models in general. Evals for an application are your own tests, built from your own task and data.

Learn it properly: the AI Engineering Bootcamp

More articles