You cannot score a paragraph with accuracy, so the field reaches for two things:
multiple-choice benchmarks, and asking a stronger model to judge. Both work, and both fail in
ways that are measurable rather than philosophical. Below, a judge with no position or
length bias still contradicts itself on 22.2% of comparisons, a mild length preference
makes it pick the wordier answer 60.8% of the time when length is unrelated to quality,
and memorising 15% of a test set adds 6.8 points to a score that has not improved.
The standard setup is pairwise: show the judge a prompt and two answers, ask which is better.
It is cheap, it correlates with human preference well enough to be useful, and it is now how most
post-training is evaluated. The simulation below gives the judge a genuine signal about quality plus
two biases it is not aware of — a preference for whichever answer it sees first, and a
preference for whichever is longer — and then scores it against the truth we generated.
4000 pairwise comparisons, scored against known ground truth
Look at the self-consistency number before anything else. With both biases set to zero the
judge still gives a different verdict on 22.2% of pairs depending on which answer came
first — and that is pure sampling noise, not bias, because a judge is a stochastic model being
asked a close question. Any evaluation that runs each comparison once is carrying that variance
into its conclusion, and a 2-point difference between two models is well inside it.
The fix that actually works: ask twice, both ways
Present the pair in both orders and keep only the comparisons where the judge says the same thing
each time. You lose the pairs it cannot decide, and on what remains you get a much better answer.
Single pass against swap-and-agree
That is a real trade, not a free win. You are throwing away roughly a quarter of your
comparisons and paying double for inference. What you buy is a verdict that no longer depends on
presentation order, and the discarded pairs are not noise — they are the genuinely close ones,
which is exactly the population where a single-pass judge was least trustworthy. Report the
coverage alongside the accuracy or the number is incomplete.
Benchmarks have the opposite problem
A multiple-choice benchmark is reproducible, cheap and unambiguous — and published, which means
it ends up in the next model's training data. A model that has memorised part of the test set does
not need to be better to score better:
share of the test set memorised
true ability
reported score
gain
0%
55%
55.0%
—
5%
55%
57.3%
+2.2
15%
55%
61.8%
+6.8
30%
55%
68.5%
+13.5
Arithmetic, not simulation: a memorised item is answered correctly, the rest
are answered at the true rate. Leaderboard gaps between frontier models are routinely smaller than
the 6.8 points that a 15% leak buys — which is why held-out and freshly-written evaluations
(LMSYS Arena's live votes, private test splits, benchmarks released after a model's cutoff) carry
more information than a number on a public set.
What to actually do
Build a small evaluation set from your own traffic. Fifty real prompts with
hand-written expectations beat any public benchmark for deciding whether a change helped
your product. Public benchmarks tell you about general capability; they cannot tell you
about your distribution.
Always randomise order, and swap. It is the cheapest bias fix there is, and the panel
above measures exactly what it buys.
Control for length — either instruct the judge explicitly, or truncate both answers to
a similar budget. A judge that likes long answers will happily reward a model for padding.
Prefer graded rubrics to raw preference when you can. "Score factual accuracy 1–5 given
this reference" is far more stable across runs than "which is better", because it does not
require the judge to hold two candidates in mind at once.
Never let the judge be the model under test. Self-preference is well documented, and it
is the one bias you cannot correct by swapping.
Report the interval. With a judge that self-contradicts a fifth of the time, a
single-run difference of two points is not a result. Bootstrap it.
Take away: both evaluation routes measure something slightly different from what you want.
A judge measures what a model finds persuasive, which correlates with quality and also with
length and position; a benchmark measures a fixed set of questions, which the next model may have
read. Neither is a reason to skip evaluation — they are reasons to swap the order, control for
length, keep a private set, and report the uncertainty rather than the point estimate.
Next:Fast fine-tuning — how to change a model
cheaply enough that you can afford to evaluate every attempt.