LLM evaluation

You cannot score a paragraph with accuracy, so the field reaches for two things: multiple-choice benchmarks, and asking a stronger model to judge. Both work, and both fail in ways that are measurable rather than philosophical. Below, a judge with no position or length bias still contradicts itself on 22.2% of comparisons, a mild length preference makes it pick the wordier answer 60.8% of the time when length is unrelated to quality, and memorising 15% of a test set adds 6.8 points to a score that has not improved.

LLM-as-judgeposition bias contaminationbenchmarks pairwise comparison

The judge is a model, so it has failure modes

The standard setup is pairwise: show the judge a prompt and two answers, ask which is better. It is cheap, it correlates with human preference well enough to be useful, and it is now how most post-training is evaluated. The simulation below gives the judge a genuine signal about quality plus two biases it is not aware of — a preference for whichever answer it sees first, and a preference for whichever is longer — and then scores it against the truth we generated.

4000 pairwise comparisons, scored against known ground truth

Look at the self-consistency number before anything else. With both biases set to zero the judge still gives a different verdict on 22.2% of pairs depending on which answer came first — and that is pure sampling noise, not bias, because a judge is a stochastic model being asked a close question. Any evaluation that runs each comparison once is carrying that variance into its conclusion, and a 2-point difference between two models is well inside it.

The fix that actually works: ask twice, both ways

Present the pair in both orders and keep only the comparisons where the judge says the same thing each time. You lose the pairs it cannot decide, and on what remains you get a much better answer.

Single pass against swap-and-agree

That is a real trade, not a free win. You are throwing away roughly a quarter of your comparisons and paying double for inference. What you buy is a verdict that no longer depends on presentation order, and the discarded pairs are not noise — they are the genuinely close ones, which is exactly the population where a single-pass judge was least trustworthy. Report the coverage alongside the accuracy or the number is incomplete.

Benchmarks have the opposite problem

A multiple-choice benchmark is reproducible, cheap and unambiguous — and published, which means it ends up in the next model's training data. A model that has memorised part of the test set does not need to be better to score better:

share of the test set memorised true abilityreported score gain
0%55%55.0%—
5%55%57.3%+2.2
15%55%61.8%+6.8
30%55%68.5%+13.5

Arithmetic, not simulation: a memorised item is answered correctly, the rest are answered at the true rate. Leaderboard gaps between frontier models are routinely smaller than the 6.8 points that a 15% leak buys — which is why held-out and freshly-written evaluations (LMSYS Arena's live votes, private test splits, benchmarks released after a model's cutoff) carry more information than a number on a public set.

What to actually do

Take away: both evaluation routes measure something slightly different from what you want. A judge measures what a model finds persuasive, which correlates with quality and also with length and position; a benchmark measures a fixed set of questions, which the next model may have read. Neither is a reason to skip evaluation — they are reasons to swap the order, control for length, keep a private set, and report the uncertainty rather than the point estimate.
Next: Fast fine-tuning — how to change a model cheaply enough that you can afford to evaluate every attempt.