A Third of Post-ChatGPT Webpages Show Significant Signs of AI Authorship, Pew Finds
Key Takeaways
- •Pew analyzed roughly 490,000 pages from Common Crawl spanning January 2021 through July 2026 using the Open Pangram AI detector.
- •About 10% of a July 2026 snapshot showed significant signs of AI authorship, while the share rose to 35% for pages published since ChatGPT’s launch.
- •Pages on .com domains showed AI-authorship signals at about 10 times the rate of .edu and .gov sites and roughly double the rate of .org sites.
- •Pew said the detector can misclassify individual pages and that the flagged material may have been AI-assisted rather than entirely machine-generated.
- •The study noted that search engines and AI companies are increasingly moving toward content provenance and anti-spam measures as AI-written text becomes more prevalent.

One in 10 English-language webpages now shows significant signs of being written by AI, according to a Pew Research Center study published Thursday. Narrow the sample to pages published only since ChatGPT—launched by OpenAI in November 2022—and the share jumps to 35%.
The findings come after researchers pulled roughly 490,000 pages from the Common Crawl web archive, spanning January 2021 through July 2026, and ran the text through Open Pangram, an AI detection model. Ten percent of a July 2026 snapshot of the corpus shows significant signs of AI authorship. The choice of corpus carries extra weight: Common Crawl has long served as a source of training data for large language models, so the archive used to measure AI's spread across the web is also raw material for the models writing that web.
An uneven fingerprint
The AI fingerprint isn't spread evenly across the web. Pages on .com domains show signs of AI authorship at roughly 10 times the rate of .edu or .gov domains, which sit near 1% each, and about double the 4.6% rate on .org sites. In 2021, before generative AI tools reached mainstream use, all four domain types looked nearly identical.
How Pew spotted the machines
The detector doesn't flag any single word or phrase but instead weighs statistical patterns across large batches of text. Still, some individual tells have become sharply more common since 2023: em dashes now appear roughly twice as often, Oxford commas are up 63%, and AI-favored words like "delve," "interplay," and "testament" have more than doubled in frequency.
"Negative parallelism"—constructions like "it's not just X, it's Y"—has nearly tripled since 2023, though Pew notes it remains rare overall. Decrypt has tracked these same tells before, and Merriam-Webster made it official in December, naming "slop" its word of the year.
The limits of detection
Pew is careful to note the study's limits: detection models can misclassify individual pages in both directions, and "significant signs of AI authorship" doesn't mean a page was written entirely by a machine—plenty of the flagged text was likely AI-assisted rather than AI-generated outright.
Open Pangram comes from Pangram Labs, whose detector has also turned up in separate research finding AI-generated text in roughly 9% of U.S. newspaper articles this year, including opinion pages at outlets like The New York Times.
AI detection may become an easier task with time, however—not because of linguistic patterns but because of the underlying technology. Major AI companies like Anthropic are already working on model-level text fingerprinting, which would make AI text easy to identify while minimizing errors. That points toward provenance—proving where content came from—over after-the-fact guessing, the same principle behind the C2PA standard that media and technology companies have adopted for attaching credentials to content.
A steep trend line
The domain gap tracks who's publishing and why. Academic and government sites run through editorial review, institutional sign-off, and slower publishing cycles; .com domains include everything from newsrooms to affiliate-marketing content farms churning out pages at a pace no human editor could sustain. Search engines have already begun responding to that flood: Google started enforcing a "scaled content abuse" policy in 2024 that penalizes mass-produced pages whether they were written by humans or machines.
At scale, the trend line is steep regardless: the .com AI-authorship rate climbed from roughly 1% in January 2021 to 9.35% five years later, in January 2026 alone.