📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Elon and xAI BREAK the Coherence Bottleneck--with HUGE Consequences!

Dr. Know-it-all Knows it all25:42

Transcription

Well, Elon Musk has apparently done it again, creating something that industry experts said was impossible to do. The All-In Podcast takes a deep dive into it, so let's take a look.

Hey y'all, it's Dr. Know-It-All. I have sliced and diced about an 8 and a half minute clip out of the hour and a half All-In Podcast, specifically the 20-minute segment starring guest Gavin Baker. I will leave a link to the original in the description. I highly recommend listening to the entire podcast, of course, but in this video, I want to focus specifically on coherence in a large supercluster and how the experts in the field thought this was impossible. Apparently, Elon Musk was able to come up with a solution that uses Ethernet of all things.

Before I start in on this, I do want to say that the All-In Podcast gives 100% credit for this to Elon Musk. They suggest he was sitting in a room, came up with the idea, and solved the problem instantly. That is not the case. I'm sure he may have been the one who came up with the spark of the idea, but we don't know at this point. Hopefully, we will find out in the future. Elon, if you want to talk to me on my channel about all of this, I would be happy to do that, of course, so you can give me the inside scoop.

Anyway, even if he came up with the original spark of the idea, we're talking about 100+ engineers who had to sit and think about how to actually implement this, write the code, install the hardware, and all of that kind of stuff. So, X.AI in this particular instance, but also Tesla and its Cortex supercomputer, which I'm sure is using the same basic architecture for how to do all of this stuff, were the ones who came up with the actual solution and implemented it. I just want to be clear that we can't give Elon Musk all the credit for this because this is way too complicated for any single human being to do.

With that caveat in mind, let us proceed. A great next place to go would be to talk about the supercomputer being built by friend of the pod Elon. He's now got the world's largest supercomputer, and he's going to 10x it. I would just say this is a very important moment for this entire AI trade in public and private markets. Everybody, I'm sure, who watches your podcast is very aware of scaling laws. We have not had scaling laws for training, or if you 10x the amount of compute used to train a model, you significantly improve the intelligence and capability of that model. Often, there are these emergent properties that arise alongside that higher IQ.

No one thought it was possible to make more than 25,000, maybe 30,000, or 32,000 Nvidia Hoppers coherent. Coherent means that in a training cluster, each GPU, to simplify it, knows what every other GPU is thinking. So, every GPU in that 30,000 cluster knows what the other 29,999 are thinking, and you need a lot of networking to make that happen.

First of all, Peewee's word of the day is coherence. Anyway, that is the ability of a large cluster of compute nodes to be able to talk to each other fast enough to maintain the coherence of a complex computation. Oddly enough, in this case, quantum mechanics might help us understand this. They, of course, use the term coherence as well. You could think of two entangled particles. You have a couple of particles; this happens in quantum computing. This is actually how quantum computing works. You have two electrons, two protons, whatever, and you entangle those two particles, which means that they coexist. If one of them has spin up, the other one's going to have spin down. All of that kind of stuff comes along with that. That is coherence on the submicroscopic level, and those two entities have coherence with each other; they're connected to each other.

The problem is, as you add a third electron or proton, and a fourth, and a fifth, and a thousandth, and a ten millionth, they start to decohere. They can't maintain coherence anymore, and they become spread out. Essentially, they become macroscopic; they don't have coherence with each other anymore, so they're unable to talk to each other instantaneously. You could consider that an analogy to the way these compute clusters work. If you have a single computer, like what I've got, even this is complicated. Let's dial it way back to the 1980s, where you had one CPU. That CPU can easily talk to itself because there's only one of them.

In my Mac Studio, with 10 CPUs, 16 GPUs, and a bunch of neural network processing, we've already got a significant problem with all of these things talking to each other. Now, of course, multiply that out by a thousand, and then a thousand again, and you see the problem of building a gigantic supercluster. Basically, all of the industry experts thought that it was impossible to maintain coherence past about 25,000 or 30,000 of these GPU nodes. This is where Elon's genius apparently came in. We don't know exactly what he did, but somehow he managed to figure out how to maintain coherence not just at 50,000 or 100,000, but apparently a planned 1 million of these nodes.

Now, probably by then, we'll be talking about Blackwell, so it will be H100 equivalent or something rather than an actual million H100s. Still, we're talking about a massive number of these GPU nodes connected together in real-time enough through slow bandwidth memory and things like that to be able to talk to each other fast enough to maintain coherence as they're doing training and inference. Of course, if that works, we can get potentially incredible emergence. The way our brains work, apparently, is we have relatively simple connections between neurons, but we have so many of them, and they maintain coherence across long periods of time. That allows the emergence of complex behavior, including intelligence.

So, what we might see is something this big and this complicated maintaining coherence. If it actually works, and Gavin talks about this, this could be a complete game changer to the nature of AI itself and might actually enable something to become semi-conscious made out of silicon. That's the kind of thing we could be looking at if this actually works.

With that in mind, let's take a look at this Tom's Hardware article, which I will also leave a link to in the description. The first in-depth look at Elon Musk's 100,000 GPU AI cluster, X.AI Colossus, reveals its secrets. I'm just going to touch on this one paragraph because this is the really important part for our discussion here. Because of the high bandwidth requirements of an AI supercluster constantly training models—and I will say also doing inference, because inference compute time is going to be important going forward—that means the amount of time that you can think about a problem, as well as the amount of training that went in.

Anyway, X.AI went beyond overkill for its networking interconnectivity. It sounds like this was the plan that Elon had long-term to be able to 10x this, you know, and then 10x it again. Each graphics card has a dedicated NIC, or network interface controller, at 400 GB, with an extra 400 GB NIC per server. That means that each HGX H100 server has 3.6 terabits per second Ethernet. So, that's the speed at which it's able to communicate with the outside world, with the rest of the cluster. In other words, if you think of this analogously as a neuron in the brain, that's the speed it can communicate with the rest of the neurons in the brain. That is a game changer; it is radically faster than other supercomputers can talk to each other.

I think this is the way that these things are able to communicate and maintain coherence across such a large number of GPUs. Interestingly enough, the last sentence here states that yes, the entire cluster runs on Ethernet—the thing that I'm connected to right now, that you're connected to—all of that kind of stuff, rather than InfiniBand or other exotic connections, which are standard in the supercomputing space. I will throw in here that probably Elon Musk's and Tesla's experience with doing Ethernet for the Cybertruck and, of course, the Cybercab coming up likely has a huge impact on their knowledge of Ethernet and their ability to use it for such a high-performance use case as building Colossus and, of course, Cortex in Austin at Tesla's Gigafactory.

This reminds us of the synergies between Elon Musk's companies. We're also going to talk about the synergy between X.AI and X in just a couple of minutes. Getting back to the podcast, just to down for the audience here, Gavin maybe explain why transporting information between the GPUs is important. That's what these H100s do particularly well; they'll move a couple of terabytes a second from one processor to the next processor. Picture a server; in the case of a GPU, it looks like maybe three pizza boxes stacked on top of each other, and it has eight GPUs together. You can think of the speed of communication on-chip as the fastest, chip-to-memory next fastest, you know, chip-to-chip within a server next fastest.

So, you take those units of servers, which are connected. The GPUs are connected on the server with a technology called NVSwitch, and you stitch them together into a giant cluster. Each GPU has to be connected to every other GPU and know what they're thinking. They need to be coherent; they need to kind of share memory for the compute to work. The GPUs need to work together. No one thought it was possible to connect more than 30,000 of these with today's technology. From public reports, Elon has, as he so often does, focused deeply on this, thought about it from first principles, and he came up with a very, very different way of designing a data center. He was able to make over 100,000 GPUs coherent.

No one thought it was possible, but I would have said there were all these articles that were being published in the summer saying that no one believed he was going to be able to do it. It was hype; it was ridiculousness. The reason the reporters felt comfortable writing those silly stories is because engineers at Meta, Google, and other firms were saying we can't do it; there's no way he can do it. I think the world really only believed it when Jensen did that podcast and said what Elon did was superhuman; no one else could have done it.

Now we will see if someone else is able to do it. It was really, really hard, and as a result of that, Grok 3 is in training now on this giant Colossus supercomputer, the biggest in the world—100,000 GPUs with a lot of MegaPacks around it—and the City of Memphis is all in on supporting this. But you have not had a real test of scaling laws for training, arguably since GPT-4, and this will be the first test.

The interesting upshot of this, of course, is we don't know, including Elon Musk, if scaling this thing up to the kind of size that we're talking about—a 50,000 cluster right now, building to 100,000, then eventually to a million—we don't know if this is actually going to work. There has been a lot of talk back and forth about how scaling laws are actually breaking down, and throwing more and more compute at these problems is getting tinier and tinier incremental gains. But maybe this is going to break the whole thing apart, and it's going to be able to work much better.

I did a video recently on version 13 of Tesla's software; you can check it out up here. The full self-driving software—I haven't got it myself yet; I can't wait to do a first test drive, of course, so stay tuned for that. The upshot of this is that version 13, which is the first one that's really been trained on the Cortex cluster in Texas, the big new cluster, is a step change by all accounts. It is a step change from the quality of FSD 12, which is already very, very good.

So, we're talking about emergent properties—the ability to get much better with the appropriate kind of scale. This is what Elon, X.AI, and also Tesla are banking on: if they can scale this compute, if they can scale the training and also inference compute enough, they're going to leapfrog everybody else. The proof is going to be in the pudding in January to February when Grok 3 comes out. If that is significantly better than what anybody else can do, then everybody else is going to be playing catch-up.

At the end of this episode, Gavin talks about the prisoner's dilemma, so we'll get back to that later on in this episode. If scaling laws for training hold, Grok 3 should be a significant advance in the state of the art. From a Bayesian way to look at the world, that is an immensely important data point. But if that car doesn't work—and I think it is going to work; I think Grok 3 is going to be really good—they've raised a tremendous amount of capital, a lot of it from the Middle East, and they're supposedly going to build Colossus to a million GPUs, ten times bigger than it is currently.

Do we start to see a shift in how the architecture of the systems are run? Meaning, do we start to build models of models, and that starts to resolve a higher-level architecture that unlocks new performative capabilities? I would just say we're already building models of models. Lots of very clever things are being done. Every AI application company has what's called a router, so they can swap out the underlying model if another one is better for the task at hand.

There's been a big debate that we were hitting a wall on these scaling laws and that the scaling laws were breaking down. I just thought that was deeply silly because no one had built a cluster bigger than 32,000 H100s, and nobody knew. It was a ridiculous debate. Grok 3 is the first new data point to support whether or not scaling laws are breaking or holding because no one else thought you could make 100,000 Hoppers coherent. Based on public reports, they're going to 200,000 Hoppers, and then the next check is a million. It was reported they're going to be first in line for Blackwell, but Grok 3 is a big card and will resolve this question of whether or not we're hitting a wall.

I want to jump in here quick and say that I actually think version 13 of Tesla's full self-driving is the first data point, and that Grok 3 will be the second data point. Even though it's a different company we're talking about, it's got to be more or less the same kind of structure that they've built out, the same architecture. Elon is famous for sharing, you know, for cross-pollinating between different companies. For something this expensive and this cutting edge, of course, they're going to be talking to each other.

If version 13 of Tesla's full self-driving lives up to what it looks like it's going to live up to, then that clearly shows that scalability, that next step in AI training compute, is going to have a massive effect. We should see a huge step forward. Grok 3 should be significantly better than what we're seeing out of the competition if this is true. If it's not, then we have a data point that shows that scaling doesn't work quite so well and that maybe people need to turn things back a little bit and not spend a hundred billion dollars on these data centers and everything like people are talking about.

Of course, Gavin also talks about—I've cut most of this stuff out—but he also talks about creating better models of models, creating better architecture, thinking about things in different ways. That is all going on at the same time. So, even if these scaling laws don't work, we still have a lot of room to develop all of the technology that has been built in the last few years and kind of scale it out and make it work even better. But if this scaling law actually works, and a hundred billion dollars is table stakes now, that just makes things all the more interesting.

Who can afford to keep doing this stuff? Well, only the richest companies in the world and maybe the richest individual in the world. By the way, we should note there is now a new axis of scaling. Some people call it test time compute; some people call it inference scaling. Basically, the way this works, you just think of these models as human. The more you speak to one of these models, the way you would speak to your 17-year-old going off to take the SAT, the better it will do for you.

We have been giving these models the same amount of time to think, no matter how complicated the question was. What we now learned is if you let them think for longer about more complex questions—test time compute—you can dramatically improve their IQ. So, we're just at the beginning of this new scaling law.

With that in mind, let's remember that Colossus might not only be used for training. Right now, it's being used for training, but it might be used for inference as well. I mean, if you throw 100,000, 200,000, or a million of these GPUs at a fundamental physics problem, like what is the nature of time or how can you integrate general relativity and quantum field theory, if you can throw that kind of inference or test time compute at a problem like this, can you actually solve the problem? Can you actually come up with a theory of everything?

I always think about that because it's just a problem that I want solved before I die. I really want to know what the answer to this question is, or at least our best guess at it. I think that is probably the best option we have at this point. But just imagine if you threw a million of these GPUs at a problem like that and gave it as long as it needed to come up with the answer. Would it be able to come up with the answer? That would be a fascinating question, and it would be answered by test time compute, not training compute.

There's a context window shift underway as well, which also creates a new kind of scaling axis, arguably in terms of the potential set of applications. So, networks of models, think time, context window—there are multiple dimensions upon which these tools ultimately resolve to better performance.

Can you kind of theorize on what the buildout that's being done with Colossus does to the advantage that OpenAI has today? How long until we kind of catch up there with X.AI, and how much is going to be disrupted, and how quickly here?

If scaling laws hold, the best information I have is the largest cluster Microsoft has, after panicking, is still smaller than the X.AI cluster in Memphis. If you didn't believe it was possible, you weren't even working on it. Grok 3 should take the lead if scaling laws hold in January or February. I think there are a lot of reasons, if scaling laws hold, to be optimistic about Grok 3.

Right now, you have like a friend in your pocket who has an IQ of 115, 110 maybe, but has all of the world's knowledge accessible to it, and that's what makes it amazing. You will have a friend in your pocket with an IQ of maybe 130 that knows everything, has more up-to-date knowledge of the world, and is more grounded in factual accuracy. Grok, because of the Twitter dataset, exactly knows what is happening at the moment in the world today.

So, let's pause on that for a second. This is kind of the genius, the Trojan Horse nature of Elon Musk's companies—at least Tesla and X.AI. You think they're one thing, but they're actually another. Tesla is not a car company, at least not anymore; they are an AI company, and the cars are the way that they gather the data to train the AI. The AI is the real product; that is the future of Tesla. If you don't believe that, you should probably not own Tesla stock. Again, not financial advice—ever.

As far as I can tell, Tesla is not a car company; it is an AI company. That's what they are. X.AI, and I know officially it's two separate companies, but they're basically rolled into one thing because they're both privately owned by Elon Musk. X and X.AI—it's not a social media company; it is an AI company. I don't know if Elon knew that when he purchased Twitter and turned it into X, but he knows it now.

X is just the way that you ingest the information that you need to train a large language model, a large multimodal model. I would assume Grok 3 is going to be able to do vision and sound and things like that as well; otherwise, why would you have it? The next generation frontier models can do that, and Grok needs to be able to do that as well.

Effectively, as Jason said, we're talking about X—Twitter being the mouth that feeds this AI. You've got automobiles from Tesla that feed Tesla's AI, and you've got all of us humans that interact with X, just throwing tons and tons of information at it, training it up—not just in past historical information and how to communicate, but also in current events and everything like that. That is going to give Grok a massive leg up over other AI because it will have access to real-time data.

That's something that these other things—they train, and then, you know, they're like, "Our cutoff was April of 2024," or something. You're like, "Well, that's great, but what about the election? What about things that have happened in the last six months?" It doesn't know. It might have a tool to go out and search the web, but it's not going to know as efficiently as Grok 3 will be able to know.

So anyway, at least for X/X.AI and Tesla, think of them as AI-first companies. The other stuff is just a way of feeding the AI. I'm sure both of you have come across these—there are lots of companies that are just these thin wrappers over a foundation model, and they go from zero to 40 million instantaneously. Yeah, and they're profitable, and for their customers, they are replacing labor budgets.

I'm sure you guys are noticing this too, but startups today, at a given size, are employing fewer people than they would have three years ago. It's funny; people were very skeptical. I would say 50% less. Yeah, and that's the ROI on AI. You're seeing real ROI on AI from startups the same way they saw real ROI from cloud computing before anyone else.

It's crazy. These companies are in a classic prisoner's dilemma. They all believe, to varying degrees, that whoever gets there first to artificial superintelligence is going to create tens or hundreds of trillions of dollars of value, and I think they may be right. They think that if they lose the race, their company is at mortal risk. So, as long as one person is spending, I think they will all spend, even if the ROI decelerates. It is a classic prisoner's dilemma.

All right, so a lot to unpack there at the end. The first thing is that thin wrapper—he kind of throws that off like it's nothing. I have been working on, you might call it, a thin wrapper around something for about four or five months now. It is not easy.

On the other hand, our little startup company has a team of six people working on something that easily would have taken double or triple. I call it 30 people; you probably would have had to work on something like this five years ago, in 2019 or 2020. It's incredible how much AI has reduced the need for labor in a circumstance like this. Our team is a startup team; we're very, very small, we're very, very lean, but we can be much, much more lean than we used to be—not just in human capital, but also in terms of compute and stuff.

There would have been a time before AWS and other entities like that that we would have had to build out a server farm just to do the work, to do the training, to be able to do the compute that we need. We don't have to do that right now; we can just throw that off onto a solution like AWS, and that takes care of it for us.

Now, of course, they make a lot of money on this, and it costs us money ultimately to do that, but it's a much more efficient way to build out your ideas and build out your products than it would be to have to invest all of that money. You'd need hundreds of thousands of dollars to build out your compute cluster, not to mention all of the extra people that you would have to hire. Everything would be much more expensive.

From experience, I will say what Gavin is talking about here right now in terms of ROI, especially for startups, is huge. Now, the prisoner's dilemma actually gets flipped. The larger companies, of course, we are not going to be investing billions of dollars in creating these gigantic frontier server clusters. That's for somebody else to do, and we just make use of it after it comes out, which is great for us.

It's very complicated for these larger companies that have to raise and then invest tens to potentially a hundred billion dollars to build these things over time. Very, very complicated. So, circling back to these large clusters, the question, the crux of all of this was how big can you build these and still have them act as a single entity? How long can they maintain coherence?

The wisdom on the street about six months ago was somewhere around 30,000 of these GPUs. Elon Musk apparently figured out a way to break this. It sounds like it was very difficult to do, but they're now at 50 to 100,000. The goal is to go to 200,000, then to a million, perhaps more. If he really has figured out how to break through this bottleneck and scaling laws hold—in other words, you can still get more performance out of larger scale—both of those things we don't know the answer to yet.

But Elon Musk, he's the bet-the-farm kind of guy. That's what he's always done. He's like, "I'll get this stuff; I will go all in on this; I will try to make this work, and if I'm right, I will get the reward. If I'm wrong, of course, I reap the reward of failure, which is not having your company be in business anymore."

But if he's right—and he tends to be right—if those two things hold, then X, X.AI, and Tesla are going to be worth tens of trillions to a hundred trillion dollars. I don't know; it's difficult to know how big the market can get if these kinds of things hold. At that point, a ten or hundred billion dollar investment—who cares? It's just table stakes at that point; it's not really a big deal.

So, with that, I'll close this video. I would love to know what you all think about this in the comments. Please do let me know, and while you're at it, if you don't mind liking and subscribing, I would love to get to 100,000 subscribers before my birthday in late January. It's my 60th birthday; it would be a fantastic birthday present. So, thank you all in advance for helping me out with that, and in the meantime, I will see you in the next video. Bye-bye!

[Music]