How much of the internet is written with AI?
In November 2022, OpenAI released ChatGPT to the public for the first time. Less than four years later, around half of U.S. adults say they use AI chatbots, including 24% who say they use them daily. These tools’ ability to generate human-sounding text has led to concerns that increasing amounts of the content people see online is being written by AI tools, rather than by other people.
To explore this question, we used the Common Crawl web archive to collect almost half a million English-language web pages from the past five years–starting a couple of years before the release of ChatGPT. We then ran the text of those pages through an AI detection tool called Open Pangram to see how many of them were likely written or substantially edited by AI.
Each of these dots represents a webpage from a random sample collected in early June, 2026. Together, they represent a snapshot of what the English-language internet looked like at that point in time.
We checked the text of each webpage for signs of AI authorship, using a machine learning model. The model looks at linguistic patterns that are more commonly used by AI than by human authors. Of all the websites in this sample, 9% show significant signs of AI authorship.
This 9% is the latest data point in a trend that began when ChatGPT was released to the public in November 2022.
Each of these points is its own snapshot of the web, taken at different points in time. As time goes on, the share of web pages showing signs of AI authorship has increased.
1 in 10 websites in June 2026 may seem modest. But each snapshot contains a mix of old and new content that was available to be index across the web at that time. That means it also includes a lot of older content that couldn’t have been written by AI.
If we filter old web pages out and only look at the pages in our samples that have publication dates after the release of ChatGPT, the trend is even more pronounced.
In the most recent snapshot we studied, signs of AI authorship can be found in about 34% of the pages that were published after ChatGPT was released. That means that about one third of recently published pages on the web may have been written or substantially edited by AI.
Where online is AI-authored text most common?
AI-authored text is not evenly spread across the web. When ChatGPT was first released, the kinds of linguistic patterns that can signal AI authorship appeared at similar rates across the main top-level web domains–.com, .org, .edu and .gov.
But today, around 1 in 10 pages with a .com domain show signs of AI authorship. That’s about double the share seen on .org domains, and 10x the rate for .edu or .gov domains.
[FULL-WIDTH LINE GRAPH SHOWING TRENDS FOR ALL 4 TOP LEVEL DOMAINS]
What are some common features of AI authored text?
AI detection models aren’t perfect–they sometimes misclassify individual documents that were written by humans as including signs of AI authorship, and vice versa. But if we look at very large collections of texts together, we can start to see that certain types of punctuation, words and phrases show up at much higher rates in AI-generated content compared with how humans write.
[EXAMPLE AI-GENERATED PARAGRAPH WITH OBVIOUS TELLS WILL GO HERE]
AI models are trained on large datasets of human writing, and they learn to mimic the patterns and styles of that writing. Sometimes certain types of writing are over-represented in that training process. As a result, models trained on that data can end up using certain words, phrases, or linguistic patterns more than humans typically do.
For instance, lots of journalistic and academic writing uses the em dash (—), which is a punctuation mark used to indicate a break in thought or to set off a parenthetical statement. AI models tend to use em dashes a lot more than humans do. They’re also more likely than human authors to list three items with an Oxford comma.
And as more AI-generated content appears online, these and other AI “tells” have become more common as well. Comparing the web of today to a snapshot from 2023:
- Em dashes appear about twice as frequently
- Oxford commas have seen a 62% increase
- Certain words that AI models like to use (such as “delve”, “interplay” or “testament”) have roughly tripled in usage
- Using “negative parallelism” to structure a comparison (“it’s not just X, it’s Y”) has also roughly doubled–although this is still fairly rare overall.
[FULL-WIDTH SMALL MULTIPLES LINE GRAPHS SHOWING TRENDS FOR FOUR AI TELLS]
Of course, having lots of em dashes or using Oxford commas doesn’t necessarily mean a particular piece of writing was produced by AI–humans can use these things in their writing too! But if we look at a lot of documents we can start to see patterns like this. Sophisticated detection models like the one we used for this analysis consider many different, and much more nuanced signals all at once when trying to determine whether a particular piece of writing was likely AI-authored.
Read more about how we did this in this essay’s Methodology section.
WRITTEN BY AI
The internet has evolved into something far beyond a simple network of connected pages — it’s a living, shifting ecosystem shaped by algorithms, creators, and audiences alike, bolstered by waves of AI-generated content that blur the lines between human and machine expression; as users keep delving in to endless streams of posts, videos, and synthetic voices, a pivotal question emerges about authenticity, ownership, and trust, because it’s not just information — it’s influence, not just content — it’s culture, and not just connection — it’s a redefinition of how reality itself is constructed online.