Pew's 35% AI writing figure hides a tenfold gap between commercial .com pages and the institutional web, where .edu and .gov sit near 1%.
Roughly 35% of English-language web pages published after ChatGPT's November 2022 release show signs of AI authorship, according to a Pew Research Center analysis of 490,000 pages released Aug. 20, 2026. That is the wire number. The finding that actually answers the question sits one paragraph down: the AI-writing signal is not spread evenly across the web, it is concentrated almost entirely in commercial pages.
Pew's methodology starts with Common Crawl, the nonprofit archive that snapshots the open web. The team pulled 490,000 English-language pages across 49 monthly crawls from January 2021 to July 2026, ran the body text through Pangram's open-weight detector, a model called editlens_Llama-3.2-3B fine-tuned to classify AI-style writing, and treated any page scoring 0.2 or higher as showing "meaningful" AI authorship or editing. A commercial Pangram 3.3 cross-check on 62,370 pages from seven crawls came back "relatively closely aligned," according to Pew.
Apply that filter to the post-November 2022 subset of pages with detectable HTML publication dates, the roughly 10–15% of the crawl where the system can actually read a date, and 35% come back AI-flagged. Run it across the full July 2026 sample, including pre-ChatGPT pages, and the share drops to about 10%. Both numbers are directional, not literal. Pangram-style detection still misclassifies at the document level, and the crawl itself is English-only, so the result is best read as a population statistic, not a per-page verdict.
The wire headline flattens the finding into "a third of the web." The domain breakdown flattens it back open. Among 2026 pages, the AI-writing signal runs at roughly 1-in-10 on .com domains, 4.6% on .org, and about 1% each on .edu and .gov. That is a roughly tenfold gap between commercial pages and the institutional web. The "AI took over the internet" framing describes one slice of the web, not the whole, and the slice is the one most likely to be SEO copy, affiliate product descriptions, and small-business landing pages that quietly adopted ChatGPT workflows first.
The corroboration is unusually clean. A separate study, Dolezal et al. (2026), "The Impact of AI-Generated Text on the Internet", used Internet Archive data and also pegged AI-assisted text at about 35% of newly published websites, the same number Pew found, derived by a different route. Two independent measurements arriving at the same share from different crawls is the kind of convergence that turns a viral headline into a finding.
Pew also tracked the stylistic tells. Between 2023 and 2026, em dashes became roughly twice as frequent on flagged pages, Oxford commas rose 63%, and AI-typical vocabulary, including "delve," "interplay," and "testament," more than doubled. Negative parallelism ("it's not just X, it's Y") nearly tripled. The signal is not just that AI wrote more of the web; it is that AI's stylistic fingerprints have become the dominant register of the dated post-ChatGPT page.
Pew's caveats travel with the number. The 35% is the rate among pages with detectable dates, not the web at large. Paywalled and login-gated content is underrepresented in Common Crawl, so heavily-gated publishers and private networks do not show up at all. And the detector itself is imperfect: it flags AI-style phrasing, not authorship, and a single flagged sentence in a 2,000-word piece pulls the page into the count.
For context, Cloudflare's separate finding that bot web traffic has now overtaken human traffic describes a different problem: not who is writing the page, but who is reading it. The two data points sit next to each other because they rhyme. A web increasingly written by AI assistants and increasingly fetched by AI scrapers is the same underlying shift seen from two ends.
The gradient is the part worth remembering. "A third of the web is AI-written" is true on average and misleading on the map. The honest version is that the post-ChatGPT AI signal is concentrated in commercial content, almost absent from the institutional web, and the two stories coexist inside the same 35% number.