Pew: 10% of English webpages show AI authorship signs

Abstract glass surfaces reflecting digital text create a mysterious tech ambiance.

In brief

  • Pew Research scanned 490,000 webpages (Jan 2021–Jul 2026) using Open Pangram AI detection model
  • 10% of English-language pages show AI authorship signs; 35% of post-ChatGPT pages (Nov 2022+)
  • Commercial .com domains show AI rates ~10x higher than .edu or .gov sites (both ~1%)
  • AI text patterns: doubled em dashes, 63% more Oxford commas, words like 'delve' doubled since 2023
  • Detection models can misclassify in both directions; 'significant signs' doesn't mean entirely AI-written

The .com advantage

The distribution isn't random. Pages on .com domains show signs of AI authorship at roughly 10 times the rate of .edu or .gov domains, which sit near 1% each. The researchers attributed this gap to institutional friction—academic and government sites undergo editorial review and institutional sign-off with slower publishing cycles compared to .com domains, which face fewer gatekeeping barriers.

The .com AI-authorship rate climbed from roughly 1% in January 2021 to 9.35% in January 2026, a near-tenfold surge in five years. This trajectory matters. The velocity suggests that AI-assisted or AI-generated publishing has become routine on commercial sites, not fringe.

Linguistic fingerprints

The researchers used Open Pangram, an AI detection model, to analyze the text. What they found was distinctive: AI-written text leaves measurable traces. Em dashes now appear roughly twice as often in AI-written text since 2023. Oxford commas are up 63%. AI-favored words like 'delve,' 'interplay,' and 'testament' have more than doubled in frequency since 2023.

Even grammar patterns betray the machine. Negative parallelism constructions have nearly tripled since 2023—the kind of rhetorical flourish ("it's not just X, it's Y") that LLMs lean on heavily.

Caveats and detection futures

Pew's researchers flagged two important limits. The Pew study notes that detection models can misclassify individual pages in both directions. And "significant signs of AI authorship" doesn't mean a page was written entirely by a machine—plenty of the text was likely AI-assisted rather than AI-generated outright.

Detection may improve. Model-level text fingerprinting being developed by companies like Anthropic could make AI text easier to identify. Meanwhile, Open Pangram's parent firm, Pangram Labs, has found AI-generated text in roughly 9% of U.S. newspaper articles this year in separate research—suggesting the contamination isn't confined to .com noise.

Frequently asked questions

What is Open Pangram and how does it detect AI text?

Open Pangram is an AI detection model developed by Pangram Labs that analyzes linguistic patterns in text. It identifies AI authorship by measuring the frequency of distinctive markers like em dashes, Oxford commas, and specific vocabulary choices (such as 'delve' and 'interplay') that LLMs favor. The model has also found AI-generated text in roughly 9% of U.S. newspaper articles this year.

Why do .com domains have so much more AI-written content than .edu or .gov sites?

Academic and government sites undergo editorial review, institutional sign-off, and slower publishing cycles compared to .com domains, which face fewer gatekeeping barriers. This institutional friction slows down AI-generated publishing on .edu and .gov, while commercial .com sites can publish at scale with minimal friction.

Does 'significant signs of AI authorship' mean the entire page was written by AI?

No. 'Significant signs of AI authorship' doesn't mean a page was written entirely by a machine—plenty of the text was likely AI-assisted rather than AI-generated outright. Detection models can also misclassify individual pages in both directions, so the measure captures a spectrum of AI involvement, not binary machine authorship.