One-third of new web pages show signs of AI authorship, Pew finds

One-third of new web pages show signs of AI authorship, Pew finds
News

A new Pew Research Center analysis finds that more than one-third of web pages published after ChatGPT’s public launch show significant signs of AI authorship or heavy AI editing. The center published the analysis on August 20 after examining almost half a million English-language pages collected from the Common Crawl web archive. The sample covers the period from before ChatGPT’s November 2022 release through July 2026.

Pew used Open Pangram, an open-weight AI detection model, to look for statistical patterns associated with machine-written text. In a random sample of 10,000 pages collected in July 2026, about 10% showed significant signals. That overall figure includes older pages that predate modern AI writing tools. After Pew filtered those pages out and considered only pages published after ChatGPT’s release, the share rose to 35%.

The results vary considerably by domain. Pages on .com sites showed signs of AI authorship at roughly ten times the rate measured on .edu and .gov domains, both of which were around 1%. The rate for .org pages was 4.6%. Pew also found that language patterns often associated with AI, including certain punctuation and recurring phrasing, became more common over the same period.

The study is important, but its number is not a direct census of all AI-written material online. Detection systems can misclassify human writing, and the analysis covers English-language pages available through the archive. Pew’s result is therefore best read as a large-scale estimate of a trend, not a definitive authorship label for every page.

For readers, creators and businesses, the finding makes provenance more important. A page can look polished while offering little information about who produced it, which sources were checked or how much editorial judgment was involved. Publishers may need clearer disclosure and stronger human review, while search and recommendation systems face a harder task in distinguishing useful automation from low-value repetition. The wider web is not simply becoming more automated; it is becoming more dependent on trust signals that users can understand.