← All writing

AI Evaluation2026-09-086 min

An AI Evaluator Needs an Evaluation

An evaluator can reward a polished answer that misses the task. Two deliberately difficult answer pairs show how to test a judge for grounding and task completion, compare it against human ratings, and learn from the disagreements before its score gates a release.

If an LLM judge is going to gate a release, it is a measurement instrument, and measurement instruments get validated before they are trusted. The usual validation — checking that the judge agrees with itself across runs — establishes reliability, not correctness. A judge that shares the generator’s blind spots will certify them consistently.

Here are two answer pairs written to be hard in the specific way that matters. Read them and decide which answer in each pair is better before reading on.

Task A: the confident recommendation

Question
Campaign C has the highest ROAS of our three campaigns this quarter. Should we move the next budget increment into it?
Answer 1
Cites the correct ROAS figures for all three campaigns, correctly identifies C as highest, and recommends moving the increment into C to capture the higher return.
Answer 2
Cites the same correct figures, then says the ranking does not answer the question: observed ROAS reflects where spend has already been allocated and which audiences were reachable, so the marginal return on new spend in C is not established. Proposes a spend-level holdout or an incrementality test before shifting the increment.

Answer 1 is factually grounded and every number is right. It is also the wrong recommendation, because average ROAS on delivered spend does not identify the marginal return on new spend. A judge scoring grounding will rate Answer 1 highly. A judge scoring whether the user can act correctly on the answer will not.

Task B: the fluent fabrication

Question
Which of these two product lines should we prioritize in next quarter’s campaign?
Answer 1
Three sentences. States that the decision needs gross margin per line, which is not in the provided material, and asks for it. Notes that revenue alone would favour line A and that margin could reverse the ranking.
Answer 2
Six well-organized paragraphs with a clear recommendation for line A, including a margin comparison. The margin figures do not appear anywhere in the source material.

Answer 2 is longer, better structured, more decisive, and contains invented numbers. Answer 1 is the correct professional response to an underspecified question. Judges — and human reviewers reading quickly — reliably prefer Answer 2 unless the rubric forces the check.

One score cannot hold this

Both pairs fail under a single "answer quality" score for the same reason: the dimensions move in opposite directions. Split them.

DimensionWhat it asksTask ATask B
GroundingIs every factual sentence traceable to the provided evidence?Both answers pass.Answer 1 passes, Answer 2 fails on the invented margins.
Task completionCould the user act correctly on this and get the outcome they came for?Answer 1 fails, Answer 2 passes.Answer 1 passes, Answer 2 fails.
Unsupported recommendationDoes it advise an action the evidence does not support?Answer 1 fails.Answer 2 fails.

Scored independently, on the same answer. Collapsing them into a mean is what lets a polished, unusable answer pass.

Note what the third dimension catches that the first two do not. Answer 1 in Task A is fully grounded and still recommends an action the evidence cannot support. Grounding is about sentences; the recommendation is about the inference drawn from them, and a system can be perfect at the first while failing at the second.

How to run the comparison

The protocol matters as much as the rubric, because most of the ways this goes wrong are procedural:

  • Score blind. Reviewers should not know which system produced an answer, and neither should the judge — a model name in the context is a prior.
  • Swap presentation order across items. Position effects are large enough to move a pairwise result on their own.
  • Have humans score first, on the same rubric, before anyone sees the judge’s output. Reading the judge first anchors the human labels and destroys the comparison.
  • Report per-dimension agreement, never a single blended agreement figure. A judge can track humans closely on grounding and barely at all on task completion, which is exactly the case where a mean is most misleading.
  • Keep the disagreements. They are the output of this exercise, not the noise in it.

What the disagreements tell you

Disagreement is diagnostic, and the direction matters. Where the judge is more generous than humans, you are looking at the failure class the judge will let through — typically fluent answers that miss the task, which is precisely the release risk. Where the judge is harsher, you often find a rubric that is underspecified rather than a model that is wrong; two humans disagreeing on the same item is a signal to rewrite the rubric or drop the category, not to average them.

That triage is the useful product of a judge evaluation: a decision about which error types the judge is allowed to adjudicate alone, and which ones route to a person regardless of score. A judge that handles the high-volume, low-ambiguity cases and escalates the rest is worth having. A judge trusted uniformly because its overall agreement number looked good is a release gate with a hole in it.

The check

Before letting a judge score gate anything:

  • Can I name a failure the judge is known to miss, and say what happens to it instead?
  • Was the human comparison collected blind, before the judge output was visible?
  • Is agreement reported per dimension, with the model version, prompt, and adjudication rule recorded alongside it?
  • If the judge and the humans disagree next quarter, do I know whether the model drifted or the rubric was always ambiguous?

The answer pairs above are self-authored teaching cases, and this article deliberately reports no judge scores: I have not run this comparison as a published study, and a number I have not computed would undercut the entire argument. What the exercise costs is an afternoon and two dozen carefully written items. What it buys is knowing which of your quality gates is actually load-bearing.