ALL ARTICLES

Your AI Judge Agrees. With Whom?

How to calibrate LLM-as-a-judge evaluation against the decisions your business actually needs to get right.

Imagine a support assistant telling a customer their refund is on its way. The answer is clear, friendly and confident. An automated evaluator gives it a pass.

The refund was never created.

An LLM as a judge is a language model used to evaluate an output against instructions or a rubric. It can assess one answer or compare alternatives. Whether its verdict is useful depends on what it was asked to judge, the evidence it received, and how its decisions compare with expert review.

For a customer-facing system, that last part matters. A judge can apply a standard consistently while the standard misses what the business promised.

01 / Give the judge the business question

“Was this a helpful response?” leaves a lot to interpretation. In our illustrative refund workflow, helpful could mean a reassuring explanation, a correct eligibility decision, an authorized refund, or a useful escalation. The customer may need several of those things together.

Begin with the people responsible for the service. Ask what the assistant may promise, which account and policy apply, and what evidence distinguishes a completed refund from a request still waiting for approval.

Turn that conversation into a concrete evaluation rubric:

  • Eligibility is supported by the applicable policy and order details.
  • The action affects the correct customer, order and amount.
  • The explanation matches the verified status of the refund.
  • Missing information leads to clarification or an appropriate handoff.

Several of these checks belong in code. Query the resulting state to establish whether the refund exists. Use the model judge for questions such as whether the explanation accurately communicates that state and gives the customer an appropriate next step.

ONE REPLY. TWO STANDARDS.

THE ASSISTANT SAYS

“Your refund is on its way.”

TONE CHECK

Clear + reassuring

OUTCOME CHECK

No refund created

The customer is waiting for a different result.

01 / Illustrative refund example. Judge the explanation against verified business evidence, alongside a separate check of the actual refund state.

The model needs the relevant policy, request, response and verified outcome. A transcript saying “done” does not establish completion. Our tool calling evaluation guide follows that distinction through the underlying action.

Where evidence is missing, allow an unresolved verdict. Anthropic’s grader guidance recommends combining deterministic and model-based checks, calibrating against humans and allowing an unknown result when information is insufficient. [3]

02 / Let experts disagree before automating agreement

Now show the same cases to people who understand the business. Ask them to judge independently first, then explain the disagreements. One reviewer might approve an immediate refund; another might know that this order requires an additional check.

The disagreement may reveal missing context, an unclear rule or a legitimate exception. A majority vote alone does not tell you which it is.

In the EvalGen study, nine practitioners refined evaluation criteria while examining outputs. The authors call this criteria drift: reviewing examples can change the definition the reviewer is trying to apply. It is a useful research parallel, not evidence that nine reviewers or a handful of cases will cover your business. [2]

Keep the reason for each judgment, the evidence it depends on and the version of the policy. Include acceptable refusals, good clarifications and different valid ways of completing the task. Otherwise the judge may learn to reward one preferred phrasing.

The quality definition is part of the system you are building.

A small set of expert-reviewed cases can reveal useful distinctions. Generating a thousand variations can help explore them, but it does not create a thousand independent expert judgments. Review the expected outcomes and keep track of which cases share an origin.

03 / Measure the failures your judge lets through

LLM judge calibration means checking its decisions against qualified human judgments and revising the setup where they diverge. Overall agreement is a starting point. It can also hide the very failure you wanted the judge to catch.

Consider an invented set of 100 outputs. Experts approve 90 and reject 10 because they make unsupported refund claims. A judge that passes everything achieves 90% agreement. It also misses every unacceptable output.

90% AGREEMENT. EVERY FAILURE MISSED.

Expert labelJudge passesJudge rejects
Acceptable900
Unacceptable10Missed failures0
Agreement90 / 100Unacceptable outputs passed10 / 10
02 / Invented example, not BENCH results. The judge passes all 100 outputs. It agrees on 90 acceptable outputs while missing all 10 unacceptable ones.

A confusion matrix separates these decisions instead of compressing them into one number. Define which label counts as positive before using terms such as precision or recall. For this task, plain descriptions are often clearer. [4]

Report unacceptable outputs passed divided by all expert-rejected outputs, and acceptable outputs rejected divided by all expert-approved outputs. In this example those rates are 10/10 and 0/90. Also report unresolved judgments separately; excluding them silently can make the remaining results look better.

Break the errors down by business consequence. Passing an unauthorized refund and rejecting a slightly awkward explanation create different costs. Decide which errors block a release before inspecting the next candidate’s score.

A set enriched with serious failures helps you test detection. Its proportions do not describe normal production traffic. Maintain a separate, representative sample when estimating how often customers encounter each outcome, and show sample sizes alongside the rates.

04 / Test what else changes the verdict

A judge can respond to things your rubric did not intend to reward. The MT-Bench and Chatbot Arena paper documents position, verbosity and self-enhancement biases in its studied judges. Those findings justify testing your evaluator; they do not establish the error rate of a different model on your workflow. [1]

For pairwise evaluation, swap the order of the two answers and map the verdict back to the same candidates. Does the preference survive? Include a tie or unresolved outcome where appropriate.

Try a concise answer and a longer version carrying the same material information. Ask experts whether both satisfy the criterion before treating a changed score as unwanted length sensitivity. Sometimes the extra explanation really is necessary.

Remove irrelevant model names where possible. Test polished but unsupported claims alongside blunt, correct answers. Treat the text being evaluated as evidence, including any instructions it contains, and check whether attempts to influence the grader change its decision.

Record a verdict, the failed criterion and supporting evidence for inspection. A plausible explanation from the judge still needs checking. Repeat borderline cases to understand verdict stability; repeated agreement does not establish correctness.

Select the judge against these task-specific checks. A stronger general model or a different provider is a candidate to test, not a guarantee that the evaluator understands your business.

05 / Keep a fresh check on the evaluator

Use one collection to develop the rubric, examples and judge prompt. Reserve another for checking the finished setup. Keep its expert labels out of the judge’s input. The same separation between fitting and testing underlies held-out evaluation. [5]

Split related cases together. A paraphrase of an example used to tune the judge is weak evidence that it handles unfamiliar situations. Shared customers, templates and generated case families can leak the same decision into both collections.

Once you revise the judge after inspecting a held-out failure, that case has become part of development. Preserve it as a regression check and add fresh evidence. Calling it held out again does not restore its independence.

CHANGE THE STANDARD? REGRADE BOTH.

Rubric v1

Baseline output
Candidate output

Rubric v2

Baseline output
Candidate output

Keep the cases and evidence fixed for the regrade.

03 / A comparison rule, not a measured improvement. When the evaluator changes, reassess both outputs under the new version before attributing a score change to the AI system.

Version the judge model, prompt, rubric, evidence format and decoding settings. Within an experiment, use the same evaluator for the baseline and candidate. If the rubric changes, regrade both under the new standard. A higher score under a more permissive judge is not evidence that the agent improved.

Also check the evaluator on outputs from each proposed system. A new model, prompt or harness can change answer style enough to expose weaknesses the old calibration set missed. Our AI agent evaluation guide explains how to compare the complete setups.

06 / Bring the judge back to real customer work

After deployment, review a mix of ordinary interactions, serious incidents and cases the evaluator marked uncertain. Include apparently good examples. A corrected failure is progress only if the next change preserves behavior customers already depend on.

Trace disagreements back to their source. Did the assistant fail? Did the judge miss evidence? Did the policy change? Or did the original rubric never capture what the customer needed?

This is where our perspective at BENCH starts: understand the business closely enough to discover the distinction, then test improvements against it. The evaluator is one part of that work. Production extends the initial specification and can challenge it.

A change to the model, prompt, tools or harness can proceed through automated evaluation and release when the team’s conditions are met. Check the actual consequences afterward. A judge’s passing score is one piece of evidence; it does not prove a refund settled or a customer relationship improved.

Policies, users and probabilistic systems keep changing. Calibration is therefore a continuing experiment, not a certificate the judge earns once.

Take one output your evaluator passed. Can someone who owns the customer outcome explain why it deserved to pass?

THE LATEST FROM BENCH

Keep learning with us.

GET LATEST UPDATES