Transcription
You were promised “infinite intelligence” for just $20 a month. That promise is already breaking. AI was supposed to be as cheap and limitless as electricity, but now Silicon Valley is pulling it back. Across models and platforms, users are now facing tighter message caps and shrinking access. It’s like an all-you-can-eat buffet… where you’re told one bite every few hours.
“The era of infinite AI is fading. The reason is a frantic new game inside Big Tech called “Tokenmaxxing.” So, if AI is getting more powerful… why is access getting smaller? For a while, anyone paying for ChatGPT Plus could open OpenAI's newest reasoning model, o1-preview, and use it as and when it was required. Then, in September 2024, a new limit was set. Just 50 messages a week. For the people paying the most, that came out to a handful of real conversations spread across 7 days. Compared to the previous generations of models, it was a cut of roughly 98%. Once the quota ran dry, the screen stopped. There was a way to bypass this. But it came with a price. In December 2024, OpenAI introduced a tier called ChatGPT Pro at $200 a month, 10 times the cost of Plus. It came with near-unlimited use of the very model everyone else was now being rationed on. Unlimited intelligence still existed. It had simply moved behind a much bigger paywall.
OpenAI framed the move as housekeeping, a way to protect the quality of the service. But for people who’d built it into their daily work, it felt like a tool had suddenly been put out of reach without much warning. The model wasn’t gone, but it was harder to get to. People started rationing themselves. They hoarded their most difficult questions or work for the hour the quota reset. If they ran out, they’d just stop and wait. The companies that had spent years talking about intelligence getting cheaper and more widely available were restricting access to the very thing they’d sold as the future. The gap between the big promises and what people were actually using was hard to ignore.
The reason was obvious. Running these models at full capacity, the way the industry had scaled them, wasn’t adding up financially anymore. So Big tech came up with the perfect work around. Tokenmaxxing. A token is just a scrap of language: a word, a small step in the machine's train of thought. So, when you can no longer make a model smarter the old way, you make it work harder. You force it to chew through far more words, or tokens, for the very same answer. Instead of giving an instant reply, the system now talks to itself first, privately, at length. This can be tens of thousands of words, before a single sentence ever reaches your screen. And instead of learning only from what humans wrote, the newest models are fed enormous piles of text that older models generated. The whole point is to keep the curve from going flat by shoving more and more computing through the same machine.
The scale is mindblowing. Meta's Llama 3 ate through more than 15 trillion tokens of text, about 7 times the data used on the version before it. Stack that human equivalent on a bookcase shelf and it would be close to 1,800 miles (2,897 km) long. Tokenmaxxing takes that number and piles more on top of it. Models loop through their old output for every hard question they’re asked. It sounds like a good idea. More intelligence is squeezed out of the same chips, with no new supercomputer required. The longer the machine talks to itself, the more it can check its own work, try an idea and throw it away. That extra thinking is bought one word at a time. All you see is a brief pause, then an answer. But this was never a success to be heralded. It was born out of sheer panic.
Something in the AI world had broken, and the people closest to it knew exactly what it was. For 10 years, the whole industry ran on one core belief. Multiply the computing power by 10, and the model gets dramatically smarter. More compute means more capability, year after year. Entire business models were gambled on this assumption. The results said different. Inside multiple labs, the data said the same thing. A 10 time jump in compute was only producing marginal gains. Around 10% to 15%, sometimes less. What used to buy a leap in intelligence was now buying a sliver of improvement.
OpenAI saw it firsthand. The model supposed to be its next leap forward, code-named Orion, reached the level of the previous flagship after only a fraction of its training. Then it stalled. The people who tested it said the improvements were smaller than expected. When it finally shipped, it wasn’t the long-promised GPT-5, but GPT-4.5, a downgrade in name that said everything. Not everyone agreed. When news of the slow down leaked in 2024, OpenAI chief Sam Altman publicly dismissed them. There is no wall, he said. The researcher Gary Marcus shot back that the limitations had been obvious for months. Even Andreessen Horowitz, with billions riding on the outcome, admitted that more computing power was no longer delivering the same leaps in intelligence. To some, it was a temporary plateau. To others, the end of an era. Either way, the money kept flowing… just in a different direction.
In public, the message remained the same, progress was just around the corner. In private, more and more people were coming to the same conclusion. The strategy that had fueled a decade of breakthroughs was running out of room, and the next gains would have to come after training. The early jumps were hard to miss. Each new model seemed smarter than the last, it was obvious to anyone who used them. The newer releases felt different. Better, yes. Smoother and more reliable. But not the kind of leap that justified the billions being spent. Admitting that would have meant questioning the foundations the entire industry was built on. So the industry grabbed onto the one thing that still seemed to work: giving models more tokens.
Back when models were smaller and answering a question cost next to nothing, ChatGPT launched at $20 a month. It was set when AI was cheaper to run, and years later it’s still the same price. The economics, though, had changed completely. Some of the most active users were costing OpenAI more than $120 a month in computing power while paying just $20 for the privilege. Every one of those users deepened the gap between what the service cost and what it charged. It’s like a grocery story treating a loyal customer well, then slipping them a $100 bill on the way out the door. But it doesn’t stop there. It looks for a thousand more customers exactly like them. The more people who fall in love with the product, the faster the cash burns away. This is not a rough patch that fixes itself. It’s baked into the system.
Every time the model stops to think a little longer, the bill goes up. What used to be a quick response can turn into a long chain of reasoning running behind the scenes before the answer ever reaches the user. That's great for accuracy. It's less great for costs. The hardest questions demand the most compute, and those are exactly the questions people come to the best models to solve. It affects the whole industry. The chips were bought and the data centers went up, but the revenue to justify them has yet to arrive. In 2024, the venture firm Sequoia framed it as AI's $600 billion question, the distance between the tens of billions being poured into AI hardware and the money the industry could ever earn back. Those data centers need hundreds of billions a year just to break even, and subscriptions cover only a small portion. Adding more $20 subscribers can’t close it, because each new heavy user only adds to the problem.
So the cap makes sense. Letting everyone run the most powerful model all day was never going to work. The numbers don't allow it. So the most expensive workloads are increasingly being pushed toward higher-priced tiers and enterprise customers, where the economics make sense. Corporate contracts can hide the cost of a top model inside the hours it saves and the staff it replaces. A single $20 subscriber can’t. From the perspective of the labs, that subscriber was never the real customer. They were the proof the product worked, the hype that made the enterprise deals possible. A deliberate loss-making base, scaled to a level the industry has never tried before.
Money was only part of the issue. The bigger problem is that the labs are starting to run low on the very thing needed to build the next generation at all. But what happens when scaling stops delivering, and a deadline is beginning to loom? There’s only a finite amount of writing in the world. It sounds strange to say, but it’s true. Strip out the spam and low quality slop and the amount of genuinely useful human-written text left on the internet shrinks fast. Researchers at Epoch AI estimate it at roughly 300 trillion tokens. At the fastest training schedules, that supply could be mostly gone by 2026, with most projections landing around 2028. A few years after that, it runs dry entirely. But not all of it is equally valuable. The best material got used first: the carefully edited books, quality journalism, peer-reviewed research. A lot of it has already been seen multiple times in training runs. What’s left is everyday web pages that add less and less each time you go back to them. This isn’t something you solve by scrapping more pages. The human text that drove the last big jumps is running out. And it’s not being replaced fast enough to keep up.
So the industry turned to the only source that still looked bottomless. The machines themselves. If humans have run out of words, let the machines write their own and feed those to the next model. Close the loop, and let the machine teach the machine. It has been tried. What happens next is called model collapse, and in 2024 the journal Nature published the autopsy. Train each new generation mostly on machine-made text, and it rots in a very specific way. The range of what it can say shrinks inward. The identity and personality that exists in real human writing fades out of every new version. The model drifts toward a flatter, more repetitive copy of itself. It only gets worse. The Nature team fed a model a passage about medieval church towers, trained the next version on its answers, then the next on those, over and over. By the 9th generation the model had forgotten the question entirely and was discussing jackrabbits in a passage that had started out about architecture. Each generation trained on the last drifts a little further from the real thing, and pouring in more machine text does not slop the decline. It speeds it up, dragging the system toward a flat, hollowed-out echo of what it once knew. Even slipping a little human writing back into the mix can’t save it. The machine-made portion sneaks in errors that are almost impossible to find. Eventually, the model starts trusting its own guesses as fact, doubling down on its blind spots, repeating its own mistakes and poisoning the well it drinks from.
There’s no escape from it. If more data only makes things worse, what happens when AI runs into the physical limits of the world itself? Just outside Loudoun County in Northern Virginia, miles and miles of windowless buildings dominate the landscape. It might sound like an unremarkable place, but it’s estimated that 70% of global internet traffic flows through here. Data centers brought in roughly $875 million in tax revenue for the county in 2024 alone, enough to fund schools and build roads. It has become the densest concentration of AI hardware on the planet. And the amount of power flowing into it has reached a scale that’s hard to grasp. A few years ago, a large data center might have asked the grid for around 30 megawatts of power. Today, a single campus can ask for hundreds. The biggest ones now push toward gigawatts. It’s so much electricity that one site can draw the output of two full nuclear power plants. That power has already been requested and approved. What’s missing is the infrastructure to move it. Dominion Energy, the utility behind most of the region, has publicly stated it can’t deliver new capacity on the expected timelines. Getting a large site connected can now take 4 to 7 years, and even getting into the queue for review takes longer than that. The bottleneck is no longer chips or models. It’s the physical grid.
Tokenmaxxing just makes it worse. Every advanced query runs longer than the simple chatbots that came before it. That means a longer drain on the energy grid. So even if the data never ran out and the models never decayed, the power to supply them would be locked behind permits and an imaginary infrastructure. It puts a hard ceiling on how many AI projects can come online at one time. If the system can’t expand fast enough for everyone, what decides who gets in and who doesn’t? The best AI won’t end up as a cheap utility for everyone. It will go to whoever can afford what it actually costs to run. The simpler models will stay widely available because they’re cheap enough for companies to absorb. But the systems that think longer and spend more compute per answer will be pushed into higher tiers, reserved for customers who can cover the costs attached to them. Over time, the gap between those two worlds will widen, as the cost of the most capable systems keeps rising. The top tiers will carry prices that match real computing. It wouldn’t be a marketing friendly price. Corporations will keep signing bigger and bigger contracts, because they can spread the cost. Anyone who relies on the best model for serious work will need to ask if it’s really worth the price tag.
The promise sold to the world was simple. AI would get cheaper every year until it became universally available. What’s emerging instead is something else: a system where the best models are deliberately limited. For companies, it’s economics. For users capped without warning, it feels like exclusion. There is no universal AI anymore. There are tiers of it. AI models aren’t just changing subscriptions, they’re reshaping careers. As machine learning scales up, the next generation of workers is feeling it first. Watch How AI is Causing a White Collar Purge. Or watch this video.