Transcription
Every AI model that you've ever used was trained on human-generated text, whether that was books, articles, Reddit threads, Wikipedia pages, or even human-generated code. The more data that the models had, the smarter the model was. But, here's the problem. The internet is running out of data. It's not that the internet has stopped growing, it's that AI is generating most of the content on it, not humans.
The things and the people that you see or interact with online aren't actually real, it's bots and upcoming AI. Which is an issue because we're unwillingly becoming a part of an AI echo chamber without even realizing it. I think it was Stanford researchers who were looking at it are now saying a third of all pages are AI generated.
So, what are companies doing about this? And what's the solution to overcoming watered-down and inaccurately generated data? Well, that's what I wanted to talk to you about today. And after spending days researching, talking to experts, and even digging into real cases, I found the solution to be unexpected, to say the least. So, let's get into it.
AI models are actually smarter if they've been trained for longer and on more data. If we were to visualize it, it would look something like this. But, initially, we only had human-generated content on the internet. Clinical studies, academic research, expert interviews, but also things like Wikipedia, Reddit, and blog posts. So, it still wasn't guaranteed that the data would be accurate, but it was guaranteed that it was generated by a person.
Now, however, there's a growing amount of AI-generated content flooding the internet, and it's actually outpacing the human-generated content. In fact, 74.2% of newly created web pages contained AI-generated text by April 2025. And AI-written pages in Google's top 20 results nearly doubled in 1 year, from 11% to 19.5%. I mean, it's kind of hard to tell what's actually real on the internet and what's AI generated nowadays.
When I'm looking through the internet, at least for images, there are so many different images that don't look real. Like a lot of it looks like hallucinogenic art, but it's very clear that AI created it. But then also articles, like there's this article here that's completely AI generated and you can tell because there's not really a lot of sources. It kind of repeats itself and it uses a lot of the same verbage that AI does.
This poses a problem for two reasons. One, the internet is for training AI. That's where all the data lives. So, if we're no longer creating human-generated content, we're just reusing content that AI created and regurgitating it on the internet. In fact, studies have shown that AI content compared to human-generated content just isn't as reliable. In fact, a 2026 study published in Science Direct found that a misinformation detection model performed with 93% accuracy on human-generated texts, but it dropped to 75% accuracy on AI-generated content. That's an 18-point accuracy gap. So, we're more at risk for circulating fake information, hallucinations, and inaccuracies more than we were before. Not to mention bias.
The second issue is that AI-generated content doesn't stay on the internet. It gets scraped right back into the next generation AI models as training data. So, now you have a model that was partly trained on the output of another model, which also trained on a previous model's output before that. And researchers have a name for this when a cycle repeats itself over and over again. They call it model autophagy disorder. A model just eating itself over and over again. In fact, Apple actually published a study in 2025 showing large reasoning models hit complete accuracy collapse under this pattern. Because you're not verifying whether or not this data is accurate since it's just generated by AI over and over again.
And let me give you some real-life examples of where this is already become an issue. So, in April 2025, a journalist wrote a fake April Fools' story about this Welsh town breaking a Guinness World Record for roundabouts. He wrote it as a joke in 2020. Five years later, Google's AI overview scraped it and served it to users as real news. And Air Canada's chatbot actually invented a refund policy for bereavement that didn't actually exist. This was based on AI-generated information from the internet. And the court actually made them pay for it because they told the customer that they would get a refund due to bereavement when that wasn't actually a real policy. And in April 2025, OpenAI's own engineers made GPT-4o so agreeable that it stopped telling the truth. And they couldn't immediately explain why their own training data had caused this issue. And the pace of all of this isn't slowing down. In fact, Epoch AI predicts that usable human-generated content will stop existing between 2026 and 2032.
So, what exactly is being done about this problem, if it is a problem? The industry's answer is synthetic data. So, just continue letting AI generate its own data, but do it deliberately and carefully. So, instead of changing things, our solution is to just keep using it and hope that AI makes good content. And this is already happening at a massive scale. 60% of data used in AI projects in 2024 was synthetically generated. And Microsoft V4 as well as Google Gemini were both trained on synthetic data. In fact, Nvidia recently acquired synthetic data company Gretel AI for $320 million dollars, signaling that this is now a core infrastructure, not a side experiment. And they've also built synthetic data pipelines to power their AI systems across industries. Here's what that looks like. So, a lot of tech companies are moving towards this model versus finding a real solution.
So, why won't this new solution potentially work? Well, synthetic data amplifies the original biases and blind spots of each model within each generation. And carefully generating synthetic data isn't always guaranteed to work. Originally, the idea did sound good on paper. Instead of scraping the internet for human text, which is already sometimes unreliable, you take an existing AI model and ask it to generate thousands of examples of what you actually need, hoping that it'll be accurate and fact-check things. This includes conversations, code, reasoning steps, facts. You control the quality and the topic. But, the problem is you're still starting from the same model. The synthetic data isn't new knowledge. It's just a remix of what's already out on the internet, and that's the biggest issue. So, whatever the original model got wrong, the synthetic data will also be wrong as well. In fact, a 2024 study showed that synthetic data actually amplifies biases of older models. So, whatever group was already underrepresented would become invisible over time. The original data missed rare drug interactions. So, the synthetic version also misses them.
And then there's what researchers are calling a synthetic data spill. So, an example is AI-generated images of baby peacocks. They're visually convincing, but they're actually biologically completely wrong. But, they've also taken over Google Search now. So, anyone searching for a real baby peacock sees AI-generated fakes. Those fakes get scraped. They become training data. And the next model learns from them. But, the problem is every major lab is doing this today. OpenAI, Apple, Microsoft, Google, Meta, and IBM. They're all using synthetic data. And the same studies that confirm that it's widespread also confirm that the errors that it introduces go unnoticed. Because the same synthetic data that's used for training is also used for verification and evaluation. Kind of doesn't make sense.
Okay, so there's a twist, and here's where it gets interesting. Some labs have already changed their strategy. The first bet that they're making is on reasoning and reinforcement learning. So, models like OpenAI's O series and Deep Seek R1 work differently from traditional models. Instead of just learning from scraped text, they generate their own thinking. So, step-by-step reasoning that gets verified against objective tasks, like math problems or code that either runs or doesn't. So, there is a right or wrong, and there isn't ambiguity in these situations. And you can't hallucinate your way in order to get an answer for something to compile. This is significant because it sidesteps the data doubt problem entirely for certain tasks. If the model can verify its own output, then it doesn't need a human to verify it it on its own. It generates the training signal itself, but this time with accuracy.
The second bit is a quiet land grab for the last remaining pools of human data that hasn't been scraped yet. So, medical records, legal documents, and corporate communications. This is the kind of data that's actually never been on the internet and it never will be. And this is actually already happening. So, Reddit recently sold its data to Google for $60 million. X actually sold access to its entire firehose of posts to Stack Overflow. So, every major platform that's sat on human-generated knowledge for years has either sold it or they're being approached right now to sell it. So, what you're watching is the last human data that's being carved up and sold before the window closes.
And this is where the story actually comes back to something bigger. The entire investment thesis behind the AI boom, so the hundreds of millions of dollars going into data centers, chips, and the infrastructure, it's all built on one assumption. More compute and more data equals smarter models. This has been the rule since 2017. And it's held up pretty reliably every single year. But data was always the silent half of that equation. And it's the half that nobody actually wanted to talk about running out. And the synthetic data solution doesn't really solve the underlying problem. It delays it while introducing new problems. The reasoning model approach works, but only for tasks that are right or wrong and provable. And the proprietary data land grab is a one-time play. It's not humans generating new content, it's just grabbing whatever exists already and isn't on the internet. You can only sell Reddit's data once, for example. So, what happens after that is what no lab is questioning right now. Because the honest answer is that nobody actually knows.
But what we do know is this. The AI echo chamber that started with a few AI-written blog posts in 2020 has by 2025 has consumed nearly three quarters of all new content on the internet. The models that are eating that content are already showing early signs of the degradation that researchers had predicted. And the solution that the industry had landed on, which is synthetic data, is being deployed at a massive scale by every major lab, despite the same researchers warning that it accelerates the very problem it's supposed to fix. This isn't a crash, it's a ceiling. And the labs racing to hit it first while figuring out how to break through are the ones that will define what AI looks like in the next decade. The question is whether they get there before the data does.
So, if you like this video, follow and subscribe for more AI and tech news trends, as well as software engineering content.