Transcription
In 2023, Stack Overflow was the largest programming knowledge base on Earth. Built over 15 years by millions of developers voluntarily sharing information, it was often the fastest way to solve coding problems.
By 2025, Question Volume had collapsed by 78%. The platform announced mass layoffs and scrambled to pivot its business model. The cause wasn't a rival platform, but ChatGBT, which had trained on Stack Overflow's own content, and was now giving developers those same answers while sending very little back to the source.
Stack Overflow isn't alone. Check, the homework help platform used by millions of students, saw its stock price fall 99% since Chat GBT launched. Press Gazette reported that publishers globally lost a third of their search traffic during 2025 as AI summaries began intercepting queries before users ever reached a website.
And in late 2024, Samman himself acknowledged what had been a fringe conspiracy theory for years, the dead internet theory. The idea that genuine human content is vanishing from the web, replaced by an increasingly synthetic flood of AI generated material accessed by bots. Alman called it basically right.
But the real problem goes deeper than dying platforms or disappearing traffic. AI researchers have now documented something with far more serious implications. When AI models are trained on content generated by other AI models, their outputs degrade rapidly, and the internet that these systems depend on for training data is filling up with exactly that kind of content. AI may be devouring the very ecosystem it needs to survive.
To understand why this matters, we'll need to explore how large language models are built. Training an AI like GPT or Claude requires staggering quantities of text. The models are fed essentially the entire indexable internet plus digitized books, academic papers, code repositories, and news archives. From this ocean of human generated language, the model learns to predict what word comes next in the sequence billions of times over, gradually developing the ability to write essays, solve problems, and hold conversations.
The quality of that training data is everything. If the data is rich, diverse, and genuinely produced by humans working through real problems, real arguments, and real creative impulses, the resulting model is more capable. If the data is thin, repetitive, or derivative, the model reflects that, too.
For decades, the internet provided an extraordinary substrate for this. Humans created content because there were real incentives to do so. Developers answered questions on Stack Overflow to build professional reputations. Journalists wrote articles because publishers monetized traffic. Academics published papers to advance careers. Bloggers, hobbyists, and forum users shared knowledge because they cared about their communities. All of this taken together represented the largest collaborative knowledge project in human history.
But AI companies harvested all of it. The entire economic and social architecture of online knowledge creation became in effect an unpaid training data pipeline. The problem is that this pipeline was never designed to survive what happened next.
When AI began answering the same questions that those platforms existed to answer, it severed the economic feedback loop. Traffic declined, revenue fell, fewer people had reason to contribute. The platforms that generated the training data began to hollow out. The speed of this hollowing has stunned even pessimistic observers. CHG's CEO told investors in May 2023 that Chat GBT was having a significant impact on customer growth. Within 18 months, the company had gone from a market capitalization of over 12 billion to under $und00 million.
Stack Overflow once so central to software development that the joke was every programmer's real skill is knowing how to search Stack Overflow, laid off 28% of its workforce in October 2023 and began a desperate pivot towards enterprise AI products. AI companies built their products by harvesting the open web and are now undermining the ecosystem that made that web worth harvesting. The result is a slow motion collapse of the internet's knowledge generating infrastructure and it has implications that extend well beyond any individual platform.
In July 2024, a team led by Ilia Schumalov at the University of Oxford published a paper in Nature that gave this problem a name: model collapse. The researchers demonstrated what happens when AI models are trained on data generated by other AI models. Across multiple architectures, including large language models and image generators, the same pattern emerged. Each successive generation of model produced outputs that were slightly more generic, slightly more repetitive, and slightly less diverse than the last. The statistical distributions narrowed rare but valuable outputs. The unusual ideas, creative phrasings, and edge case solutions that characterize genuine human thoughts disappeared fast. The tales of the distribution collapsed, leaving only a bland statistical center.
A separate term from Rice University published the same year termed this phenomenon model autotoagi disorder or MAD. MAD, a deliberate medical suggesting that AI systems training on AI generated content are in a clinical sense consuming themselves. The mechanism explains why simple filtering won't fix the problem. When a language model generates text, it tends to favor high probability outputs, the statistically average response. Unusual phrasing, minority viewpoints, and creative leaps are less likely to appear. When the next model trains on this output, it learns an even narrower distribution. Each generation amplifies this effect, progressively stripping out the diversity that made the original human generated data valuable. The researchers compared it to repeatedly photocopying a photocopy. Each generation loses detail until the image is unrecognizable.
The implications become clearer when you consider the trajectory of the internet itself. A 2024 projection by Epoch AI estimated that high-quality text training data could be effectively exhausted by 2028. Meanwhile, research from the University of Waterloo found that synthetic content on the web is growing exponentially. In some categories, AI generated text already outnumbers human written content. These two trends are converging. The supply of genuine human content is shrinking while the flood of AI generated material is growing. Future models will inevitably be trained on data sets increasingly contaminated by the outputs of their predecessors and the contamination is difficult to filter out.
Researchers at the University of Maryland found that the current detection tools for AI generated text are unreliable, particularly as models improve. The synthetic content that degrades training data is becoming harder to identify and remove. There is a grim irony here. The better the AI models become at mimicking human writing, the harder it becomes to protect future models from training on their output.
The situation has been likened to Ourorus, a mythical snake depicted eating its own tail. AI systems consumed human knowledge to become capable. That capability is now destroying the sources of human knowledge. And as those sources disappear, the models are left consuming their own outputs. A recursive loop that the research suggests leads to progressive degradation rather than improvement.
The immediate casualties, Stack Overflow, CHEG, and traditional publishers are visible and measurable. But the deeper concern among researchers is something harder to quantify, the possibility that this cycle could slow the pace of genuine knowledge creation itself.
Consider what Stack Overflow actually represented. Every day, developers encountered novel bugs, undocumented edge cases, and integration problems that no existing resource addressed. They posted these problems and other developers proposed solutions. The platform was a living record of new knowledge being created in real time. When a developer now asked chatbt instead, the AI draws on its training data to produce an answer. Often that answer is adequate. But the question itself and whatever novel solution it might have generated never enters the public record. The knowledge loop that sustained the platform and contributed to future training data is broken.
Emily Bender, a computational linguist at the University of Washington and a persistent critic of large language model hype, has argued that these systems create an illusion of understanding that mass a fundamental limitation. They can recombine existing knowledge with remarkable fluency. They cannot generate genuinely new knowledge. If the ecosystems that produce new human knowledge atrify, AI is left recombining an increasingly stale corpus. This points toward what might be called epistemic stagnation, a future in which AI systems are ubiquitous and superficially impressive, but the underlying knowledge they draw from has stopped growing. The answers feel authoritative. The pros is fluent, but the substance is a hall of mirrors, each reflection slightly more distorted than the last.
There is a historical parallel. In 2011, Google launched its Panda algorithm update, which penalized low-quality content farms that had flooded search results with cheaply produced SEO optimized articles. Companies like Demand Media, which had built billion-dollar businesses generating thousands of formulaic articles per day, saw their traffic collapse overnight. The update was celebrated as a victory for quality, but the underlying dynamic is familiar. An information ecosystem was flooded with low-quality content produced at scale, degrading the experience for everyone until a drastic correction became necessary. The difference now is that AI generated content is orders of magnitude harder to identify than the crude output of content farms and there is no Google Panda equivalent on the horizon for synthetic text contaminating training data.
There are also changing habits in how people access information. Reddit has grown significantly since 2023 in part because users append Reddit to searches specifically to find human perspectives rather than AI generated summaries. Substack has thrived as readers seek verified human voices. These platforms are surviving precisely because they offer something AI summaries cannot: the texture of genuine human experience and opinion.
But this creates a split. Those who know how to seek out human generated sources and can afford subscription-based platforms will have access to higher quality information. Those who rely on default AI answers and free web content will increasingly encounter a degraded information environment. Access to authentic human knowledge could become a kind of luxury good stratified by digital literacy and a willingness to pay.
The economic spiral compounds the problem. Less traffic to original sources means less revenue, which means fewer professional writers, journalists, and experts creating content, which means less high-quality training data, which means worse AI outputs, which drives even more users away from original sources. The feedback loop is self-reinforcing, and no single actor has either the incentive or the ability to break it unilaterally.
Against all of this, there are reasons for cautious optimism, and they deserve serious engagement. The most significant is the synthetic data revolution. Deep Seek's R1 model released in early 2025 demonstrated that AI systems can develop novel reasoning capabilities through self-play and reinforcement learning with minimal reliance on human generated examples. If models can improve by generating and evaluating their own training data, the dependence on the open web's content pipeline may matter less than the model collapse research implies.
Some researchers, including Yan Lun, formerly at Meta, argue that the data wall is a temporary constraint that synthetic data and new training paradigms will overcome. There is also a historical argument. The internet disrupted previous knowledge ecosystems too. Newspaper sales fell. Encyclopedia publishers went bankrupt and entire categories of expertise were devalued. And yet, new forms of knowledge creation emerged. Wikipedia replaced Britannica with something more comprehensive and more current. Although Britannica does still exist, YouTube created an entirely new category of educational content. It is possible, perhaps likely, that AI's disruption will similarly give rise to new platforms and formats we cannot yet anticipate.
And some existing platforms are adapting rather than dying. Stack Overflow has integrated AI tools into its own platform. Publishers are experimenting with licensing deals and paywalled content. The market in its messy and uneven way is responding. The question is whether these adaptations will be sufficient and whether they will arrive quickly enough.
The model collapse research describes a degradation that compounds over time. If the contamination of new training data reaches a critical threshold before new paradigms mature, the damage may already be locked in. And unlike previous technological disruptions which played out over decades, this one is measured in months. The stack overflow that took 15 years to build was gutted in under three. The check that served millions of students went from market leader to near bankruptcy in 18 months.
Whatever new knowledge ecosystems might eventually emerge, the speed of this transition is unprecedented. And speed matters when you're dealing with a recursive process that degrades with each cycle. The internet was built by humans sharing what they knew. What happens when the systems trained on that sharing make the sharing itself unsustainable is a question that will shape the information landscape for decades. And it is a question that for now only humans can answer.
You've been watching Absolutely Agentic. We're a media consultancy and startup focused on AI and we create weekly YouTube videos and a newsletter we send a few times a week. In a world of AI hype, we're trying to strike a balance looking at a range of opinions to cut through an increasingly complex world. The best thing you can do to support us is to sign up to the newsletter in the description or subscribe to the channel. But also don't forget to like and leave us a friendly comment. See you next time. It'll be absolutely agentic.