Transcription
Many people expect that the AI bubble will pop in spectacular fashion because of a recession. But what if it simply deflates because China gives the world an alternative? A cheaper version of the AI that America is spending a trillion dollars to build. So stay with me because by the end you'll know how close China really is, what it could mean for your money, and the surprising part that America itself is playing in it.
Now picture that whole US AI bet as one big balloon. America's hyperscalers such as Amazon, Microsoft, and Alphabet are pouring vast sums into centralized AI. The data centers that you rent your intelligence from, and of course, the hardware where your favorite AI is running. And the air being pumped in is capital and demand.
But I'd suggest that the AI the world is after is a bit different. It's sovereign, which means that every government wants its own version, and it's personal and private, running on your own hardware, with no one else able to turn it off on a whim. Chinese open models helped along by Nvidia itself are now cheap enough and good enough to inflate those other balloons. And that means less air goes into the American balloon.
Now, you already know that your tracker leans heavily on a handful of US giants. I've made that concentration point in a couple of recent videos, so I won't go over it all again, but just to put a number on it, in a global tracker, the top 10 holdings make up a quarter of the index, usually led by Nvidia. So, this isn't a niche bet. It's probably people's biggest one. And the real question is, what could undermine that bet that those giants are making? And that's where China comes in.
Now, let's meet the challenges. For years, the comfortable assumption was that the best AI is American and everyone else can just rent it. I think that assumption is quietly breaking. Chinese open models, things like Deep Seek, Quen, Kimmy, GLM, and Miniax have basically caught up and they're a lot cheaper.
Let me give you the number that surprised me. AI is charged based on tokens. Now, these are the small chunks of text or data that a model consumes and produces. Deepseek's flagship model costs about 87 cents per million output tokens. Now, that's roughly 60 times cheaper than Anthropic's Fable 5. 60 times. Now, for a lot of real work, that's the difference between an idea being worth doing and not.
Now, cheap wouldn't matter if the quality wasn't there, but it is. As of March, the top US model led the top Chinese model by just 2.7%. In other words, the frontier gap has shrunk to almost nothing. If we look at this graph from Epoch, you can see the same story from another angle. It tracks the US and Chinese frontiers over time. And what you can see is that since 2023, Chinese models have trailed the US by about seven months on average, sometimes as little as four. So, it's not that China has pulled ahead, it's that the lead has become a short lag.
And what's really remarkable is that nearly all the leading Chinese models are open weight, while the top US ones are still closed. So you can download the Chinese ones and run them yourself. And this isn't just benchmarks. The switch is already showing up in real usage. In the first two weeks of June, four of the five most popular models on Open Router, which is where developers go to pick an off-the-shelf model, were Chinese. Among the global top 20, Chinese models processed twice as many tokens as their American rivals.
Now that open router data is debated. It's one platform and it leans towards developers. So I wouldn't hang the whole argument on it. But it points the same way as a lot of other evidence. The quality, the cost, the open licenses. So developers with real money on the line increasingly seem to be voting open and Chinese.
Now you might be wondering, how do you actually use one of these? It is pretty straightforward. So if you just want to use a model, you go to an aggregator like open router and route your request to whichever one you like. We actually use open router for the AI back end of the pension craft website. And if you want absolute control over your privacy with no dependence on anyone's servers, you could go to hugging face, download a model, and run it yourself. the tools to run them locally, things like Alama and LM Studio are mature now. In other words, this used to be a research lab thing and now it's just basically a download.
And remember that cost gap of 60 times for Anthropics Fable 5 versus Deepseek. For a business running millions of these requests every day, that price gap could be transformative.
Now, today's video is sponsored by Lightyear, and I open my stocks and shares ISA with Lightyear because I've always believed that in long-term investing, the one thing you can actually control is costs. And Lightyear, who I've been working closely with for a few years now, is one of the lowest cost brokers in the UK right now. Their stocks and shares ISA and general investment account are commissionf free, have no account fees, and offer a decent range of sterling denominated ETFs with thousands of US and European stocks. And though fund manager fees and a.1% conversion fee may apply if you're primarily investing in sterling denominated stocks or ETFs, that's largely a non-issue. If you'd like to give Lightyear a try, use the code pensioncraft to receive up to 100 pounds in a fractional share or ETF in your general investment account. The link with full terms and conditions is in the description below. And as always, investing involves risk. The value of your investments can go down as well as up.
Now, let's turn to the part I think matters most in terms of its impact. You might not need the cloud at all. The hardware to run a serious model privately now fits on your desk. Take Nvidia's RTX Spark. Now that's got about a paflop of AI compute. To show you how mad that number is, back in 2008, the first machine on Earth to break one petlop was an IBM supercomput called Roadrunner. And that was built to watch over America's nuclear stockpile. Now that filled a warehouse. It ran nearly 20,000 processors. It used 57 miles of fiber optic cable inside it. And it took 21 articulated lorries just to deliver. And it drew over two megaww. Now that's enough to power thousands of homes. The RTX Spark hits that same one petlop mark from a box around 14 mm thin under 2 kilos sipping about 80 watt.
Now I should be straight with you that's not quite apples to apples the comparison. Roadrunner's petlop was full 64bit high precision maths and that's what you need to do physics simulations. The spark in contrast is low precision. So it's a stripped down 4-bit format called FP4. So far rougher arithmetic. But the clever part is that AI doesn't need the precision. Running a neural network is closer to rough pattern matching than exact sums. So you can use cruder, faster maths and pack far more of it onto the same chip. So it's the same headline number, a petlop, but a different kind of calculation. Even so, a National Lab supercomputers worth of throughput now sitting on your desk. And with 128 GB of memory, it can run 120 billion parameter model locally. Now, that's a properly capable AI with your data never leaving the room. And that privacy angle is really the point. NVIDIA's own software can route your sensitive queries to the local model and disguise your personal details on anything that does get sent to the cloud. So, private by default stops being a slogan and becomes what we consider to be everyday and normal.
Now, you might think this is just a desktop curiosity, but the same shift is already happening at scale on our phones. In 2025, about one in three smartphones shipped were Gen AI capable, over 400 million of them. Now, that's up from about one in five in 2024. And the obvious objection, which is that companies won't let staff run AI on their own laptops, actually cuts the other way. And that's because if the model runs on the device, the data can stay on the device. So for a privacyconscious business, local AI isn't the risk. It's AI running on data center that's the risky option.
So let's look at the stakes because they're enormous. If demand drifts to open, local, and private AI, there's a very big question. Who fills the trillion dollar data centers that have already been built? And the spending here is genuinely staggering. This year alone, Meta, Alphabet, Amazon, and Microsoft have together set aside about $725 billion for AI infrastructure. Amazon on its own is around 200 billion. If we look at JP Morgan's capex graph, you can see those curves climbing steeply. The forward estimates push towards $900 billion a year. So, this is one of the biggest capital bets in corporate history.
And riding on top of it is a power bet. Data centers used about 415 terowatt hours of electricity. That was in 2024. And that was roughly 1.5% of the world's power. The IEA reckons that could more than double to around 145 terowatt hours by 2030. But even that's far from certain. And in the IEA's high efficiency case, demand in 2035 comes in about 20% lower. So the electricity story rests on assumptions that just might not hold. Add politics, and that makes things even more uncertain.
Now, when I say the bubble deflates, let me be really clear about what that does and doesn't mean. I don't think these corporate giants are going to go bust. They're not pure AI plays. Microsoft, Amazon, and Alphabet. All of them sit on huge profitable businesses underneath in cloud, ads, software, and retail. So, this isn't a 2000 style wipeout. The risk is narrower, and it's really about returns. All that AI spending has to earn its keep. And right now, it mostly isn't. The FT reports a Bane estimate that the industry needs to generate roughly $2 trillion a year in AI revenue just to justify what's being spent already. Even back in 2024, Sequoia flagged a similar version, its famous $600 billion question. So there's a big gap between what AI earns and what it needs to earn to pay for all of this. And that gap is not closing. An MIT study found that about 95% of company AI pilots show no measurable return at all.
And so we have richly valued companies pouring a fortune into AI that isn't paying off yet with no signs of slowing down. If huge profits don't justify the expenditure, that doesn't bankrupt the companies, but I think it does mean they'll underperform. The market gets tired of paying a premium for spending without returns. And falling share prices are the punishment. And for a diversified investor, that is of course a feature, not a bug, because this is how concentrated markets always resolve. The expensive leaders cool off. The market gets cheaper and it gets less topheavy and leadership passes to more profitable companies.
And here's the accelerant I mentioned. A few weeks ago, Washington reached in and froze America's two most powerful AI models. The Trump administration put export controls on Anthropic's two newest models, Fable 5 and Mythos, and that was to stop anyone outside the US from using them. But Anthropic, a $900 billion company, was given just 90 minutes to comply. And it couldn't carve out the foreign users that fast. So, it pulled both models for everyone. 90 minutes. And its two best models went dark worldwide. Now, to be fair, a business already running on Anthropic wasn't switched off. It could still use the older models. It just lost the powerful new ones.
But here's the precedent that really matters. Washington showed that it can reach in and cut off a frontier US model for an entire country overnight. And the freeze didn't even touch its rival, Open AI. So if you're a government or a business anywhere outside America, that lesson lands very hard. As one FT journalist put it, companies may decide the best AI isn't the smartest one. It's the one they can actually log into every Monday morning. So, a kill switch like that doesn't pop the US balloon. It just makes every other balloon look a lot more attractive.
So, what does all of this mean for your money? I don't think the message is sell everything. AI isn't becoming worthless. The risk is specifically to this centralized bet, the rented US-hosted model. And I think much of that value doesn't vanish. It simply moves. And you can already see it happening. More than 200 companies tied to the data center buildout have beaten the world index over the past year. Corning, which makes the cabling, is up more than 270%. an air conditioning maker comfort systems is up about 260%. Because these chips run really hot and someone has to cool them. So the air leaving the big balloon is going somewhere. At the moment it's being pumped into suppliers, into devices, and into the grid.
But I'd look for the weight of evidence pointing the same way. the quality gap between open-source and proprietary AI models narrowing faster than the cost differential AI moving on to devices and the revenue to justify the spend not showing up when several of these independent measures line up like that I think that's the sign that there will be a market adjustment so what's the disciplined response I'd argue that it's patience not trading a broad capweighted global index already does the hard part for you. If the centralized winners deflate and the value moves to suppliers and devices and the grid, well, the index is going to trim the losers and top up on those winners automatically. So, you don't have to guess which balloon wins. Markets are really good at sorting out this capital misallocation eventually, and the job mostly is to be patient and let them.
So in conclusion then I think the AI story is transformative but the centralized bet on how we'll consume AI is the fragile part and the Chinese AI models are the most likely thing that lets the air out of the US AI bubble. Now personally I'd rather own the whole balloon factory through a global index than gamble on which particular balloon stays inflated. So, if you want to understand exactly how to do that, the companion video on the best index funds for global stocks is the natural next watch, and I'll link to it here. And as always, thank you for listening.