Lesson 14.2 · 23 min
LLM-as-Judge: Automating Evaluation with Another Model
If we need to grade 10,000 chatbot answers by tomorrow, could another LLM do the grading, and could we trust it?
In short: LLM as a judge means using a capable language model, given a clear rubric, to evaluate the outputs of another model or system. It scales open-ended evaluation that string metrics cannot handle and that humans cannot afford to do at volume. Judges come in pointwise, pairwise and reference-guided forms; they work best with explicit criteria and step-by-step reasoning, and they must be checked for biases such as position and length preference and validated against human labels.
What is LLM as a judge, and why do we need it?
LLM as a judge is an evaluation technique where we prompt a large language model to assess the quality of a text, usually the output of another LLM application. The judge receives the input, the output to grade, a rubric (the criteria and scale) and sometimes a reference answer, and it returns a score, a label or a preference, ideally with a short justification.
Why not just use metrics or people? Our support bot writes free-form answers. String metrics such as exact match or ROUGE only count overlapping words, so they punish correct paraphrases and reward fluent wrong answers. Human reviewers understand meaning, but grading thousands of answers after every prompt change is slow and expensive. An LLM judge sits in between: it reads for meaning like a person and runs at the speed and price of software.
Think of it like a teaching assistant with a marking scheme A professor cannot grade 800 exams alone, so teaching assistants grade using a detailed marking scheme, and the professor spot-checks a sample to make sure the assistants grade fairly. The LLM judge is the teaching assistant, the rubric is the marking scheme, and our human-labelled sample is the professor's spot-check.
The approach was popularized in 2023 by work such as the MT-Bench and Chatbot Arena paper (Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena), which reported that a strong judge model agreed with human preferences about as often as humans agreed with each other on their test, while also documenting the judge's biases.
How does LLM as a judge work?
The judge is just another LLM call. Everything that makes it reliable lives in the prompt design, the choice of judge model and the validation we do around it.
Types of LLM as a judge
Judges can also be reference-free (judge faithfulness to retrieved documents without any gold answer) and multi-criteria (separate scores for correctness, completeness and tone, each with its own rubric line).
Steps to build an LLM judge
From idea to a trusted judge
- Pick one criterion at a time: “Good answer” is too vague. Choose concrete criteria: correctness, groundedness, follows refund policy, polite tone. One judge prompt per criterion is often more reliable than one prompt for everything.
- Write a rubric with anchors: Define every score level with an example of what it looks like. Prefer binary pass/fail or a small 1–3 or 1–5 scale over 1–10.
- Collect human labels: Have people grade a sample (for example 100–200 outputs) with the same rubric. This is the ground truth for the judge.
- Run the judge and measure agreement: Compare judge and human labels: accuracy, Cohen's kappa, a confusion matrix. Read every disagreement.
- Iterate: Fix the rubric or examples where the judge disagrees, try a stronger judge model, then re-measure.
- Deploy and re-check: Use the judge at scale, and re-validate on fresh human labels from time to time or whenever the judge model changes.
A prompt template for LLM as a judge
A practical pointwise template for our support bot's correctness criterion. Placeholders in curly braces are filled by code.
judge_prompt.txt
You are an impartial evaluator of customer-support answers.
Task: grade the ANSWER for CORRECTNESS against the POLICY DOCUMENTS.
Scale:
1 = states something that contradicts the policy, or invents a policy
2 = partly correct but misses or misstates an important condition
3 = fully consistent with the policy and answers the question
Rules:
- Judge only correctness. Ignore length, style and politeness.
- If the policy does not cover the question, an answer that says so scores 3.
<question>{question}</question>
<policy>{retrieved_docs}</policy>
<answer>{answer}</answer>
First write your reasoning in 2-4 sentences, then output JSON:
{"reasoning": "...", "score": 1|2|3}- Role and task tell the judge what it is grading and against what.
- Anchored scale: each score has a concrete description.
- Explicit exclusions (“ignore length”) fight known biases.
- Delimiters separate the data from the instructions, which also makes it harder for an answer to manipulate the judge.
- Reasoning before the score and JSON output make verdicts more accurate and easy to parse.
Chain-of-thought judging (G-Eval)
Asking the judge to reason before scoring usually improves agreement with humans, for the same reason chain-of-thought helps other tasks: the model works through the criteria instead of jumping to a number. G-Eval (Liu et al., 2023) is a well-known recipe built on this idea.
- Give the judge the task and the criterion (for example “coherence of a summary”).
- Let the model generate detailed evaluation steps for that criterion (chain of thought), once.
- Use those steps in a form-filling prompt to grade each output.
- Instead of taking the single score token the model writes, read the probabilities it assigns to each possible score and compute a weighted average. This gives finer-grained, less tie-prone scores.
The probability trick needs an API that exposes token log-probabilities; many judge setups simply use “reason first, then score” without it.
Biases in LLM as a judge
| Bias | What happens | Mitigation |
|---|---|---|
| Position bias | In pairwise mode, the judge favours the answer shown first (or second) | Judge both orders; count it as a win only if both agree, else a tie |
| Verbosity (length) bias | Longer, more detailed-looking answers get higher scores even when not better | Rubric says “ignore length”; compare length-matched answers; check correlation of score with length |
| Self-enhancement bias | A model may rate outputs from itself or its own family higher | Use a judge from a different model family, or several judges |
| Leniency and scale compression | Most scores land on 4 out of 5; small differences vanish | Binary or 3-point scales with anchors |
| Style over substance | Confident tone, formatting or citations impress the judge even when facts are wrong | Reference-guided judging; separate criteria for correctness and style |
The code below simulates a pairwise judge with a built-in preference for whichever answer is shown first, and shows how running both orders exposes it.
position_bias.py
import numpy as np
rng = np.random.default_rng(42)
N = 400
gap = rng.normal(0, 1, N) # > 0 means answer A is truly better
human = (gap > 0).astype(int) # human label: 1 = A wins
def judge(g, first_bonus=0.6, noise=0.7):
# Simulated judge: sees true quality, plus a bonus for the answer shown first
return int(g + first_bonus + rng.normal(0, noise) > 0)
a_first = np.array([judge(g) for g in gap]) # order: A, B
b_first = np.array([1 - judge(-g) for g in gap]) # order: B, A (flip verdict back)
consistent = a_first == b_first
print(f"A wins | A shown first: {a_first.mean():.2f}")
print(f"A wins | B shown first: {b_first.mean():.2f}")
print(f"A wins | humans: {human.mean():.2f}")
print(f"order-consistent verdicts: {consistent.mean():.2f}")
print(f"agree with humans, one order only: {(a_first == human).mean():.2f}")
keep = consistent # inconsistent pairs -> 'tie'
print(f"agree with humans, consistent pairs: {(a_first[keep] == human[keep]).mean():.2f}")
# Cohen's kappa: agreement corrected for chance
po = (a_first == human).mean()
pe = a_first.mean() * human.mean() + (1 - a_first.mean()) * (1 - human.mean())
print(f"Cohen's kappa (one order): {(po - pe) / (1 - pe):.2f}")Output:
A wins | A shown first: 0.67 A wins | B shown first: 0.33 A wins | humans: 0.51 order-consistent verdicts: 0.59 agree with humans, one order only: 0.74 agree with humans, consistent pairs: 0.91 Cohen's kappa (one order): 0.48
Pause and think: From the output: single-order agreement with humans is 0.74, but on order-consistent pairs it is 0.91. Why do we not simply report 0.91?
Because 0.91 is measured only on the 59% of pairs where both orders agreed. The other 41% become ties or need human review. Swapping order gives us a trustworthy verdict on fewer pairs, plus a flag on the rest, which is honest; reporting 0.91 as overall accuracy would hide the ties.
Best practices and real-world use cases
- Validate before trusting: measure agreement with human labels and read disagreements.
- Small, anchored scales, ideally binary pass/fail per criterion.
- One criterion per judge call; combine scores in code.
- Reason, then score, in structured JSON, at temperature 0.
- Swap order in pairwise mode; treat inconsistent verdicts as ties.
- Use a strong judge, often stronger than the model being judged, and ideally from a different family.
- Version the judge (model + prompt). Changing either changes the scores, so old and new numbers are not comparable.
- Keep humans in the loop for high-stakes decisions and periodic recalibration.
Real-world use cases Regression testing of prompts and models in CI (“did groundedness drop?”); scoring RAG answers for faithfulness to retrieved documents (evaluation libraries such as Ragas and DeepEval provide judge-based metrics for this); A/B comparisons between two model versions; online monitoring that samples live conversations and flags low-scoring ones for human review; filtering or ranking synthetic training data; and producing preference labels for training reward models (often called RL from AI feedback).
The common mistake Writing “Rate this answer from 1 to 10” with no rubric, using the numbers immediately, and never checking them against people. Such scores look precise but often measure length and confidence, not correctness.
Worked example, step by step
“Our judge agrees with humans 91% of the time” sounds like a pass. Let us check that claim by hand. We have 100 support answers. A human reviewer marked 90 as pass and 10 as fail. Our judge graded the same 100. The table of counts is below; the numbers are illustrative.
Is 91% agreement good?
- Raw agreement:
(88 + 3) / 100 = 0.91. This is the number that sounded good. - How often each says “pass”: The human passes 90 of 100 answers (0.90). The judge passes
88 + 7 = 95(0.95). The judge is more generous. - Agreement by chance: Two graders who pass that often would agree a lot even if they ignored the answers:
0.90 × 0.95 + 0.10 × 0.05 = 0.855 + 0.005 = 0.86. - Cohen’s kappa:
κ = (observed − chance) / (1 − chance) = (0.91 − 0.86) / (1 − 0.86) ≈ 0.36. On a scale where 0 is chance and 1 is perfect, this judge is only a third of the way. - Look at the row that matters: Of the 10 answers the human failed, the judge failed only 3. If the judge is meant to catch bad answers, it misses 7 of every 10.
The lesson of the arithmetic: when most answers are good, raw agreement is dominated by the easy passes. A judge that says “pass” to everything would score 90% here.
The off-diagonal cells also tell us what to fix. Seven answers in “human fail, judge pass” and only two the other way means the judge is too lenient. We read those seven, find what the human saw that the rubric does not mention (perhaps an invented policy), and add it to the rubric as an explicit fail condition. Then we measure again, on answers the judge prompt was not tuned on.
Practice: try it yourself
We will write a small function that takes the four counts from a human-vs-judge table and returns raw agreement, chance agreement and kappa. Then we compare three judges on the same 100 answers, including one that passes everything.
practice_judge_agreement.py
def kappa(both_pass, human_only, judge_only, both_fail):
# Confusion counts between human labels and judge verdicts (pass / fail)
n = both_pass + human_only + judge_only + both_fail
observed = (both_pass + both_fail) / n # raw agreement
human_pass = (both_pass + human_only) / n # how often the human passes
judge_pass = (both_pass + judge_only) / n # how often the judge passes
# Agreement we would expect if the two graded independently, by chance
chance = human_pass * judge_pass + (1 - human_pass) * (1 - judge_pass)
return observed, chance, (observed - chance) / (1 - chance)
judges = {
# name: (both pass, human pass / judge fail, human fail / judge pass, both fail)
"always-pass judge": (90, 0, 10, 0),
"lenient judge": (88, 2, 7, 3),
"careful judge": (85, 5, 2, 8),
}
print("judge agreement by chance kappa bad answers caught")
for name, counts in judges.items():
observed, chance, k = kappa(*counts)
caught = counts[3] / (counts[2] + counts[3]) # of the 10 human fails
print(f"{name:<18} {observed:9.2f} {chance:9.2f} {k:5.2f} {caught:18.0%}")Output:
judge agreement by chance kappa bad answers caught always-pass judge 0.90 0.90 0.00 0% lenient judge 0.91 0.86 0.36 30% careful judge 0.93 0.80 0.66 80%
Raw agreement barely separates the three judges: 0.90, 0.91, 0.93. Kappa and the last column separate them clearly: 0.00, 0.36, 0.66, and 0%, 30%, 80% of bad answers caught.
Now change it:
- Add a strict judge:
"strict judge": (70, 20, 0, 10). It fails every bad answer but also 20 good ones. Predict: is its raw agreement above or below the lenient judge’s? And its kappa? - Rebalance the data: change the careful judge to
(45, 5, 2, 48), a test set that is half fails. Predict how the gap between agreement and kappa changes. - Change the always-pass judge to an always-fail judge,
(0, 90, 0, 10). Predict its agreement and its kappa before running.
Pause and think: We want to use the judge as a release gate that blocks bad answers. The lenient judge has 91% agreement. Why is that number the wrong one to look at, and what is the right one?
Agreement is dominated by the 90 easy passes. A gate is only useful if it catches failures, and the lenient judge catches 3 of 10. The number to watch is the share of human-failed answers the judge also fails (its recall on failures), with kappa as a summary. We should also check the opposite error: how many good answers it blocks.
Pause and think: A teammate validates a new judge on 20 hand-picked answers that are all obvious passes or obvious fails, and reports perfect agreement. What is wrong with this check?
It tests the judge only where judging is easy. Real traffic contains borderline answers: mostly right with one invented detail, or correct but incomplete. That is where judges and humans disagree. The sample should be drawn from real outputs with a realistic mix, include hard cases on purpose, and be large enough that a few disagreements do not swing the result.
Key takeaways
- An LLM judge grades outputs with a rubric, scaling meaning-aware evaluation far beyond what humans can label.
- Main setups: pointwise scoring, pairwise comparison and reference-guided grading.
- Clear criteria, anchored small scales, reasoning before the score and JSON output make judges more reliable.
- Judges have biases (position, verbosity, self-preference); swap orders and use rubrics to counter them.
- Always validate judge verdicts against human labels before trusting the numbers.
Key terms
- LLM as a judge: Using a language model with a rubric to evaluate the outputs of an LLM system.
- Rubric: The written criteria and score definitions the judge applies.
- Pairwise comparison: A judging mode where the judge picks the better of two outputs for the same input.
- G-Eval: A judging recipe using generated evaluation steps and probability-weighted scores.
- Position bias: A judge's tendency to prefer an answer because of where it appears, not its quality.
- Verbosity bias: A judge's tendency to rate longer answers higher regardless of quality.
- Cohen's kappa: An agreement score between two raters that corrects for agreement expected by chance.
← 14.1 Evaluating LLMs: Metrics, Benchmarks, and Methods · 14.3 Evaluating AI Agents: Metrics and Methods That Work →