Transcription
We did not raise 1.5 billion. That's revenue, actually about 30% of the revenue of OpenAI. Your job is not to follow the wave; your job is to get positioned for the wave. You can almost say we're one of the best things that ever happened to Nvidia because they can make every single GPU they were going to make and sell it for training—high margin gets advertised across deployment. We'll take the low-margin, high-volume inference business off their hands, and they won't have to sell either margin. We are growing faster than exponential, and when you are growing faster than exponential, there is no amount of profit that you can make that matters. What matters is getting a toehold in the market and becoming relevant. Ready to [Music] go.
Jonathan, thank you so much for agreeing to do this. Imparis, you look fantastic, by the way. I feel so underdressed, but you look great.
Thank you. I could take the tie off if you want. I'll never be able to tie it again; I don't know how to tie a tie. No, literally, my chief of staff has to tie it for me. It's a struggle because he's putting it on himself, he's tying it. I literally only bought this suit recently.
Well, I mean, you look fantastic. I don't think I have a suit, so you one-up me. I want to split the show into two parts. There, I want to talk about the landscape where we're at and then I want to dive specifically into Grock, where you're at. You've announced a massive new deal that I think everyone's slightly misunderstanding.
What we're just talking about—um, I just want to start on where we're at in terms of like scaling laws. Everyone says we are at the limits of scaling laws, and then there seems to be exponential innovation happening with the likes of DeepSeek and others. Where are we at in terms of the limits of scaling laws?
So, scaling laws is a paper that was published by OpenAI. What it does is it effectively says the more parameters your model has, basically the better it can absorb information. So you'll see these curves that they draw, and they're amazing. You should show it if you can. But effectively, you have these sort of asymptotic drop-offs where you keep getting better and better, but you get a logarithmic improvement when you put a linear number of tokens in. This is why you see people doing 15M trillion tokens of training and whatnot. But they're misunderstood because the assumption is that all of the data is the same quality. So you have a kid now, right? So eventually, you're going to be training your kid, and you're going to say—and play along with me here—what's 1 + 1? Two. What's 2 * 3? Six. What's the second derivative of the square of the hyperbolic tangent? Yeah, yeah, good question. But that's how we train these models. We give them really simple problems to solve, and then we give them these really hard ones. We don't really train them up; we don't do it smart. So what some people do is they will train on the dregs of the internet, and then they'll save some high-quality data for the end to make them better. But what you can do—and this is where I think everyone's getting confused—is it's sort of like with AlphaGo Zero, where it generated its own data and trained. You could have an LLM generate synthetic data, and when it generates the synthetic data, the data is better. You then train on that synthetic data. So what you do is you train. Why is synthetic data better than real data? Because the model is smarter. So you know Reddit is great, but not necessarily high-quality as talking to someone with a PhD in a topic. Sure. And so just like with more expert people who are more knowledgeable and more capable, if you have a better model, it generates better data. So you train the model, it gets better, you produce better data, and you produce a range of data here, and you get rid of all the parts that are wrong. So now it's the best part. So it's a little better than the model is because you're pruning it because you get to do this offline, right? And then you train the model, and the model comes up here, and then you do this again, and then you keep the better data, you train it again, you just keep moving up. So when you do that, the actual scaling laws don't look like these asymptotics; they actually—but they have to be a ceiling on efficiency. No, does there? So there's a mathematical limit. So if you study computer science, you probably heard of something called Big O complexity. Big O complexity is—you know, if I'm solving a problem and I look at how I solve it, I might need to take more steps if I solve it with one algorithm versus another. So, for example, quicksort versus bubblesort. Quicksort, I need n log n steps; bubblesort, I need n squared. What's the difference if I'm sorting 1,000 numbers? N log n, that's 10,000 steps, but with n squared, that's a million steps because it's either 10,000 or 1,000,000. One of the reasons that these LLMs struggle to multiply large numbers is because multiply is not linear. These LLMs can do anything linear without, you know, needing to think, but just like on a piece of paper how you need to write out all those intermediate steps, these LLMs need that intermediate space in, in those steps in order to compute these things. It's a mathematical requirement. There's nothing you cannot train a model enough so that it'll see any arbitrarily large number just be able to multiply it, but you can choose bigger and bigger groupings of numbers for to memorize, in which case it can do it in fewer steps. And effectively, as you are training the model on more and more data, it's seeing more and more examples, so now it just has the answer for more specific situations, so it doesn't need to do as much reasoning, but it still needs to do reasoning for some of these problems.
So what does that mean for the next step in terms of what happens now if we have no efficien—like if we have no efficiency ceiling? What does that actually mean? You need both. So the training of the model makes it more intuitive; it means that it can sort of just come up with the answer like that, more stream of consciousness. The reasoning part is different; the reasoning is the algorithm on top, right? The Big O complexity portion. So it's system one, system two thinking, or thinking fast, thinking slow, like Daniel Kahneman's book. And so when you pair them together, when you make it more intuitive, you know, you get better this way, right? But when you start adding in the system two portion, you start to get this, right? You hear the volume is very little, but when you do this, and so you get this polylinear—is the term—but you could think of it as geometrically increasing improvement in the model when you combine it with that improved training, but also the improved what they call test time compute, or runtime compute. Totally get that.
So just so I understand, so when we think about bottlenecks, if we have synthetic data that powers training, it gets more intuitive; it gets to the answer more quickly, sort of like a grandmaster in chess just seeing the right moves. Sure. But synthetic data is not constrained in terms of its supply side. If we think—if we think about the other bottlenecks, there is hardware, there's energy efficiency, there's algorithmic limits. What is the—but but if I'm telling—if your job is to get better at multiplying numbers and I tell you that I want you to be able to do it with fewer steps, more intuitively, for you to be able to multiply three-digit numbers versus two-digit, you need 10x the data and you need 10x the examples, right? And so as you get better on the intuitive part, you need more examples to train on. Makes sense. Totally. And so what is the bottleneck then? Is it the hardware quality? Is it compute? Is it algorithms? Because it's not data. It is the compute; it is the data; it is the algorithms. It's all three of them. But so people misunderstand the concept of a bottleneck. Compute has been more of a—a less of a bottleneck and more of a, you know, soft neck or something, right? Where when you provide even more compute, you can sort of overpower the lack of data, the lack of improvement in algorithms. So it's not a hard bottleneck; it's a soft bottleneck. But ideally, you would improve all three; you would be getting better data, you would be getting better algorithms, and the algorithm improvements are going to be there; the data improvements are going to be there, but compute has always been the easiest lever because it's so fungible. If I just give you more compute, it works better. Has DeepSeek not showing you that actually we don't need the compute and you can do more with less? Not exactly. There was an algorithmic improvement on that, and the algorithmic improvement, as I explained, you know, is this seemingly silly thing where they just wrote the answer in a box and then they knew what to look for rather than having to have a human being check it or something like that, right? It was very simple, um, but that was an algorithmic improvement, and it made it easier to generate the data that was then trained on.
Can I ask—I think there's misconceptions around compute, data, uh, especially kind of synthetic data, as you said, there are algorithms. When you think about the biggest misconceptions that people have around AI and specifically kind of inference, what do you think they are?
When we started, the first misconception, which people don't hold anymore, is that training was more expensive than inference. At Google, anytime we would train a new model, we would end up using 10 to 20 times as much compute on the inference as the training. So we always ended up—inference was always the critical infrastructure piece that we needed. Um, but then after getting, you know, past that, now everyone understands inference is important. Um, but I think one of the—do you think they fully do? Because when you look at Nvidia's stock price post DeepSeek, it was down 15%. If you understood the value of inference, it shouldn't be down 15%. And Jean's Paradox and all that. And yeah, I don't agree that Nvidia stock should have gone down for that. I think that was a misunderstanding on most people's part, but it also shows—I think that shows more like everyone keeps saying Nvidia stock can't possibly go higher, right? And they were looking for an excuse for, "Oh, now that's it; that's why we were wrong, and we need to sell now." But that has nothing to do with the—that that's just a sort of popularity contest side of the market that had nothing to do with the weighing machine of the market.
So, Grock is building stage. Should they build with the assumption that scaling laws will continue? Should they build with what we have today? How do you advise them on that?
I would advise you to build based on things getting better, but I would also focus a little more on the sort of big quantum steps. So the analogy that I like is if you look at the information age we went through, we had the printing press, we had the telephone, we had the telegram, we had the internet, and we had smartphones, right? And if you had built Uber back when we had internet, it wouldn't have worked because you'd book a ride, you'd go somewhere—how do you get home? Exactly, right? And we're in the same sort of space now. So we don't—the models hallucinate, so it would be hard to build a medical diagnosis company; it would be hard to build a legal company, right? However, if you were doing that and the algorithmic enhancements happen that get the hallucination rate down, you were perfectly positioned, just like Grock. We were around for seven years before we had product-market fit, right? We were around—our bet was scaled inference, that inference was going to be the bottleneck that we were going to need to run really big, heavy models. Like everyone was assuming you would have a single PCIe card running inference because training was the complicated part, right? The reality was we made the right bet ahead of time, and then we were perfectly positioned. Your job is not to follow the wave; your job is to get positioned for the wave. And that's the hardest thing to do because everyone is trying to talk you into coming onshore again, right? Almost everyone was telling us, "Don't do LLMs; they're going to be terrible for you." We're like, "This is literally what we built for."
Did you ever doubt yourself? Seven years is an incredibly long wait time.
Well, it just—doubt—there was doubt, but there was never a pause. And the reason was so even back before—before starting the TPU, I was concerned that AI was going to be a technology that would allow some people to have outsized control, outsized influence. If you allow that to just happen in potentially not the best hands, it doesn't really matter how rich you are; it doesn't matter—nothing matters. It's the most important technology, so it didn't matter how hard it got; there was no choice but to be successful. And our goal is to preserve human agency in the age of AI, right? If we don't do that, we have failed. And so it wouldn't matter whether there was doubt or not, and yes, there was plenty of doubt. There was a point where we were so close to running out of money; we did this thing that we called Grock bonds, so you know—um—war bonds from World War II, of course, but for anyone that doesn't—what as a war bond? So, war bond—World War II was funded with bonds—the US government. They had these posters; it was like, "Fund your troops," and whatever, and you'd buy them, and they would pay you a return, and that funded the war effort. We were very close to running out of money at one point. Rather than trying to pretend to be strong, you know, we were vulnerable with our employees, and we said, "We're going to run out of money; we need you to trade equity salary for equity." We literally took pictures of the war bonds, and we put Grock bonds on it instead, and we had an all-hands where we said this, and we were worried everyone was going to leave. Instead of leaving, about 80% of the employees participated; 50%, I think, went to the statutory minimum salary by law. When we finally raised the first bit of our $300 million round, we had so little money in the bank left that it was less money than we saved doing Grock bonds. So had we not done that, we would have literally run out of money. So there were some really hard times, and I know every founder has these, and from the outside it's so hard to understand; it's like watching a TV show—you're not in it—but you mult—you know, it's like when you are there, everything is 10 to 100 times more intense because people left their jobs, they left their careers, their families are banking on this, and you have to make decisions like, "Go out there." What hap—what would have happened if we went out there and asked everyone to do Grock bonds and everyone quit? Then the shareholders would have been like, "You have all of these people depending on you." But if you lean towards that vulnerability, people are often going to go with you on it.
So what is a world where inference is so crucial and 20 times more important than training? What does that world look like?
I think the simplest way to understand it is equate an LPU or a GPU to an employee, right? If you—you have enough of them—the LPUs or GPUs—you can do work, just like with an employee, but it's a little different in the sense that they can't quit and take another job; you don't have to retrain. Once you get a model to a certain capability, it'll always be at least that capability, right? It's not going to regress. You—you know—so you get the consistency out of it, but now imagine that you're a startup, and rather than having to go out and hire 100 people, you hire 10, and you buy the amount of compute equivalent to 90 employees' worth. That's a very different way of thinking about the world because now CAPEX—or in some cases, different types of OPEX—can be used instead of just employees. And so that—and in terms of inference, just to give you a sense of our scaling, we started 2024 with about 640 chips in production; we ended with over 40,000. This year, we want to be at over 2 million, and next year the number is much, much, much larger.
Are we seeing constraints on chip supply? I mean, that is an unbelievable scaling story.
Yeah. So for us to hit our numbers next year, which I'm not sharing publicly, we're going to need almost all of the capacity of the fab that we're—we're using. The biggest issue—so, seven partners—we love seven partners, right? Hamilton Helmer, okay. Um, you don't normally think tech companies as having a cornered resource, but Nvidia has a cornered resource; they're a monopsony—the opposite of a monopoly—a single buyer for HBM and the interposer—the COOs.
So what is HBM?
So HBM is high-bandwidth memory. Okay. And what GPUs—and who produces HBM? I'm sorry for the dumb questions.
There's three companies in the world that do this: SK Hynix, Samsung, and Micron. Okay. And it's a specialty memory; it's only used in high-end servers. So there's a limited quantity that's built; it's very expensive to ramp up; it's a very technically challenging type of memory to build—more so than others. So there's a very limited supply, and GPUs are so fast computationally that if you were using regular memory, it'd be like drinking out of a martini straw; it would just take forever. This is why you see people preferring to do even inference, but especially training on GPUs rather than CPUs because the memory bandwidth is too limited, and CPUs rarely use HBM; they're mostly regular memory.
Our architecture—so the observation that we had when we started Grock—everyone knows Moore's Law—every 18 to 24 months, like clockwork, double the transistors means double the compute. But we noticed that AI was getting better faster, and it clearly wasn't the algorithms because algorithms have sort of discontinuous jumps; it also didn't seem to be the data because there wasn't that much more data, and the transistors were only doubling every 18 to 24 months. So where was all of this capability coming from? Turns out the number of chips was also doubling every 18 to 24 months. So rather than 2x, it was 4x. So the question we asked was, "If you're effectively going to have an unlimited number of chips, do you do something architecturally different?" The answer is absolutely. So rather than using external memory, we just use a large number of chips and keep all of the parameters of the model in the chips live, and then we just have this pipeline where the computation flows through it, sort of like an assembly line, right? So imagine if you were trying to build a factory, and the factory was only 1/100th of the size needed for the assembly line, so you'd run a bunch of cars through 1/100th, tear it down, set up the next 1/100th assembly line; you just do this over and over again. That's the way a GPU works. LPUs are very different; we actually just have the computation flow through a whole bunch of chips. So rather than using eight chips, we'll use 600 or 3,000 for a model.
How does that change energy efficiency?
It improves at about 3x. And the reason is—how does it improve it? When you use more—so because you use less for more per token. So the footprint is higher; think of it as the difference between a factory or a backyard sort of garage. The backyard garage is not going to be as efficient; however, it has a lower energy footprint. Or another example would be if you were trying to transport a ton of coal from one side of the city to the other, and you did it on mopeds or you did it with freight trains—which one would be more efficient? The moped would use less energy per trip, but it would need more trips and therefore would use more energy overall. In fact, this is one of the things most people misunderstand; they think that edge computing is lower energy; actually, edge computing is less energy efficient than computing in the data center. Why is that? When you're computing in the data center, it's a little bit like that freight train; you're actually getting to do a whole bunch of jobs simultaneously. So the fact that we don't have to read from that external memory means that we don't have to spend the energy doing that. Even with GPUs, you get to batch, but going back to why it's so energy efficient—the amount of energy used in a chip—there are these physical wires, and the physical wires have a width, and when you look at the width and you look at the length, you charge that wire up to set it to a one and then you discharge it to set it to a zero, which means it's sort of like charging a capacitor and discharging a capacitor; you're using energy. The longer that wire, the more charge. When you have HPM here and another chip here, you're actually having to charge a wire between the chips and then discharge it every time you send a bit, and so that's a long distance to travel, but also the wires are wider than the wires that are inside the chips, so you just use a lot more energy. When we keep that memory in the chip, it's only traveling a little distance, using much thinner wires, and therefore it uses a lot less energy.
So do we see a world of LPU and GPU—GPU usage in combin—like how—how does that distribution look between LPU usage and GPU usage?
There's a couple of things. The first is training should be done on GPUs, and actually, I think Nvidia will sell every single GPU they make for training right now. About 40% of their, you know, market is inference. Um, I think if we were to deploy a lot of much lower cost inference chips, um, what you would see is that same number of GPUs would be sold, but the demand for training would increase because the more inference you have, the more training you need, and vice versa. Um, the other use case is we're actually so crazy fast compared to GPUs that we've actually experimented a little bit with taking some portions of the model and running it on our LPUs and letting the rest run on GPUs, and it actually speeds up and makes the GPU more economical. So since people already have a bunch of GPUs they've deployed, one use case we've contemplated is selling some of our LPUs to sort of nitro boost those GPUs.
This is my question, which is that people have bought GPUs so far ahead of time that by the time you get them, they're deployed and installed, they're almost out of date.
Actually, we've—we've spoken with some customers that put orders in over a year in advance; they paid a year in advance and still haven't gotten them. The recent deployment we did in Saudi Arabia—51 days from contract to the first tokens being served in production in country.
How are you able to do it so quickly? 51 days is astonishing.
Yeah. Um, part of it is architecturally; things are much simpler for us. We don't have a bunch of other hardware components; we actually don't use switches to communicate between our chips; we just plug our chips into our chips; our chips are the switch, and we don't have all of this network tuning. Think about it this way: when you're going across town in France, how long does it take to get from one side to the other?
A long time.
A long time. But a variable amount of long time, for sure. If you do it in the middle of the night, it might be fast; if you do it in, you know, middle of the day during an event like we've got going on with AI Summit—slow. Terribly. Exactly. But it's unpredictable. However, certain modes of transportation, like trains, can be predictable. With what we're doing, it is 100% predictable. Given the energy efficiency, given the predictability—why is Nvidia not being more proactive on LPUs?
Um, what makes you think that they don't want to be more proactive on it? It's—they don't talk about it.
Well, why would they talk about it? That would be like talking about something you don't have when you're trying to project strength rather than vulnerability. Well, I think if you wanted to protect shareholder value and wanted to protect a Wall Street image of dominance and being ahead of the game, you'd at least say, "Oh, we are, of course, working on LPUs as well." But then, until they had that ability, until they had LPUs, they would effectively be exposing that there's something missing. Like if you look at the last GTC, there was an announcement that the latest GPUs were 30x faster than the previous generation, and when you look at how it was done, there was this curve that looked kind of like this, and then it basically ended here, and then there was another curve that was kind of like this. Now, that 30x was from the end of this curve to this curve. If you moved it here, it would have been less than 30x; if you moved it here, it would have been infinite. So their chip is infinitely faster than the previous one, but that wouldn't have sounded reasonable, right? There's a history in this market of specsmanship because it's so hard to, like, get access to chips, and this is a lesson on enterprise sales. I think in enterprise sales, people rely on specsmanship better than your specs—"My chip is faster than your chip; I get more teraflops per second than you do," right? But who cares? Like, just tell me what the tokens per dollar is and tell me what the tokens per watt is; nothing else really matters. But people will find all of these other weird things to measure that they might be better on, sort of like, "I'll sell you a car with better RPMs." RPMs don't matter, right? What matters is miles per gallon and maybe the speed that you can drive that, although speed limits kind of render that, you know, moot, right? But in the case of enterprise sales, people often—well, there was a time when the way that you would buy soap or you would market soap—the billboards would say, "Our soap has more bubbles than this other brand soap." Who cares? And what they figured out was, "Let's put really happy people up on a billboard after they use the soap, and then maybe people associate that happiness," right? Lifestyle marketing. Sure. For some reason, enterprise still hasn't learned this lesson; it's still—"We have more bubbles; we have more teraflops; we have more"—whatever things that people just literally don't care about.
So you think Nvidia's "30 times faster" is not good marketing?
I think it worked because it's what people are used to, but our counter was we did a press release to that that said, "Grock—still faster." That was it, and people went gaga over it, right? Because it was just—"We are—we're still faster," so who cares? I totally get that.
Do you think Wall Street ups and downs that way?
I think they're starting to, yeah. But again, I—I don't think there's real competition here. I think if you are competing, you have done something seriously wrong. If you're competing, it means that you haven't found an unsolved customer problem because if you're competing, someone else has already solved the problem, so why are you spending time on it?
So you don't view Nvidia as a competitor?
No, they—they don't offer fast tokens, and they don't offer low-cost tokens. It's a very different product, but what they do very, very well is training; they do it better than anyone else and by such a wide degree—it's a solved problem. Why would we bother trying to solve a problem that's already been solved?
So you like to seed the training market to them; you'll own the inference market?
Yeah. And they're saying that we also want the inference market, of course. It's the way it always works.
So what do we do now? So now we are competing in the influence market, but are we—yeah. So we don't really have people saying, "You know, we're going to buy GPUs instead of you." We do have people saying, "We're going to buy both," that happens, but we don't care because we—all—like I showed a demo to someone, and he's like, "Should we just not buy any more GPUs?" I'm like, "No, you should buy every single GPU you can get your hands on." And he's looking at me very perplexed, and I'm like, "Well, how are you going to do training?" We don't do training; buy the GPUs; get every single one you can because I want your models running on us to be really good. But for inference, they don't need to buy Nvidia anymore; they don't need to buy GPUs for inference. But if you can get them, I mean, they're a little expensive, but if you're used to it, why not? Plenty of people still sell mainframes, but if you want lower cost and faster, then you want an LPU.
How much lower cost is it?
More than 5x lower.
More than 5x lower. Just the memory alone in the latest GPUs costs more than our fully loaded CAPEX per chip deployed. And—and on top of that—so we talked about the energy efficiency—so we use about a third of the energy per token—about over three—period—one-third of our cost is the OPEX, which is mostly energy and data center rent, and 2/3 is the CAPEX, which means that since we're one-third of the energy, the cost to run that GPU to produce the same number of tokens for inference is the same as our total cost. Just the OPEX for the GPU is the same as our CAPEX plus our OPEX.
Why is 40% of their revenue inference then, and why have you not taken so much more of that?
At the beginning of 2024, we only had 640 chips; at the end, we had 40,000. We're not at that scale yet. So you have to—you have to provide quality; you have to provide low cost; you have to provide speed, but you also have to provide capacity. And so this is where that most important part of not using HBM came in; it means that we effectively have no scale limits. So the GPU itself is actually manufactured using the same process that you use for your mobile phone, right? So the same silicon that's in your mobile phone is the same silicon for the GPU. In fact, they build the mobile phone chips first because they're smaller, so they're better. So Nvidia actually gets it after Apple. The difference is that memory—that's the only difference—but that memory is the hard part to manufacture; that's that's limited in scale. So by us avoiding that, we effectively have almost no limit on how much we can scale up, and that's important for inference.
What is Nvidia's margin?
70 to 80%.
70 to 80%. So they can take 70 to 80% off and be radically more—yeah—comparatively cost-effective compared to you. Like you could destroy their margin, but why would—so I—in that same vein, you can almost say we're one of the best things that ever happened to Nvidia because they can make every single GPU that they were going to make and they can sell it for training—high margin, right?—gets amortized across the deployment, and you know, we'll take the low-margin, high-volume inference business off their hands, and they won't have to sell either margin.
What's low margin?
Anywhere from—depending on the deal—we do get some on the back side, but up front it's about 20%.
About 20%. Yeah. Okay. So there's their 80; yours is 20, but then you're looking at a 20x—but then we get more later off of it. So we take some of the risk.
What do you mean you get more later?
Sorry. So the deals that we do—the partner will off—because we don't deploy—we don't spend money for our own CAPEX—the partner will put up the money for us to deploy; we pay back with a, you know, decent IRR, and but we split, and most of it goes to the partner, and then once we hit the IRR, it flips the other way. So others are putting the CAPEX up for us.
What does it look like at the end then?
It's a little—it's not like other business models. So we—we didn't just innovate on the chip; we also innovated on the business model, and we're limited in how much money we can make based on how much we can deploy, not how much money we have because the partners are putting that money up. So when I'm looking at what we can do, it's all about how much we can scale.
What are the limits to your deployment? Is it purely chip constraints?
Mostly.
You're asking about misconceptions in AI. I think one of them is about power. So it is true that there is a mismatch in the market between people with chips and people with power, but that's partially because you need a data center in the middle, and there aren't enough data centers. Those aren't the hardest thing in the world to build; they're not easy, but they're not the hardest thing. It's harder to build up the power. However, because of that mismatch, you have big hyperscalers going around and saying, "I need a gigawatt of power," and they'll say this to 60 different potential data center builders, and then all of a sudden you hear this echo—"Well, I heard that, you know, there's a gigawatt here and a gigawatt here and a gigawatt here"—and all of a sudden there's like 60 gigawatts of demand, and it's this echo from that first gigawatt. The thing is, I am aware of about 20 gigawatts of power that people want to make available for data centers right now. Right now, there's about 15 gigawatts of data centers worldwide. So more than double the current capacity. Concern that I have is that people are now building up more power, and what's going to happen in the next 3 to 4 years is people are going to be like, "I built up all this power, and no one's using it, and this was like a complete waste, and we're never going to do this again." Then what's going to happen? Remember that doubling of chips every 18 to 24 months? Well, 3 to 4 years, you double that 15 gigawatts twice, and now you're talking about what—120 gigawatts? There isn't that much power available, and then another one after that—now you're at 240. And so what's going to happen is we're going to overbuild slightly right now just because of that mismatch and the miscommunication that's going on right now, and then we're going to dampen our building, and we're going to, you know, close down on that, and then
Right, it's at seven years now on the data centers, and then the people building the data centers then need a long, seven-year commitment. Yeah, that's the kind of thing they're asking for. So you've got this complete mismatch throughout the ecosystem.
But the funny part about it is, while they all want to take zero risk and have a committed, um, you know, sovereign wealth level, um, sort of credit rating on the other side of it with long commits, the longer the payoff time, the more generic the infrastructure is. A model has a pretty specific use, but accelerators like LPUs and GPUs can be used for other things besides generative AI or LLMs. The data center can be used for other things besides the accelerators; the power can be used for anything. So while they're looking for the least risk over here, it's the place where there is the least risk because if we don't use it for AI, we'll use it to power all of the electric cars.
Is this a case where incumbents SN because they're one of the only ones who are able to match the durations required by data center providers? Well, and this is why we've partnered with Aramco and this new entity in Saudi Arabia because they have an enormous ability to fund this over the long term. They have a very long-term perspective, and they have an amazing credit rating. I mean, listen…
So when you say they have an ability to fund it, and this is why the misconception was: people think it's a funding round of a billion and a half. It's not a funding round of a billion; no, we did not raise 1.5 billion. That's revenue; that's actually about 30% of the revenue of OpenAI.
Can you just walk me through how that deal is structured? Yeah, so we started off last year, right, and we got to 19,000 of our chips deployed. We did that in about 51 days, and the question was what could we do this year? So they've gone off, they've collected up a bunch of power in the country, and the deal is structured so that they will put up the CAPEX for us to deploy our chips in that data center or those data centers, and we pay back based on the money that we make. So it's sort of… it's a little bit different than debt in that they participate in the upside, but it is similar in nature. But it is revenue because we actually make profit upfront.
How does that change what you can do? Well, we are not limited by capital anymore. And one of the unique reasons we can do this: there is one misconception around Grok. There was a paper that was written that said that we couldn't be profitable while being lowest price. We could charge more, but actually we have a very positive contribution margin right now, and so as far as we know, we're the only ones that are actually making money running these open-source models. Because with the open-source models, everyone's sort of competing with VC dollars trying to take market share, Uber style, right? Meanwhile, we're sitting here going, "We could do this all day long because we're making money," and we're able to even pay off an IRR and make our partners money. So the difference here is there's another part of the model. So we're also working with some proprietary model providers, so we actually showed off the first one at Leap on Sunday where we did a voice model with Play.AI. That one is also a rev share, but the thing is they get to make money off of that whereas most others in the industry are losing money because of the commoditization of the models.
Do you have cheaper pricing over time as you bluntly have less monopoly power, or do you have higher prices as your monopoly increases? Well, we want the margin to stay about the same, but we want the prices to go down because then we get into Jeevan's Paradox, and life gets great because we're going to scale, and our focus is on getting to scale, right? To preserve human agents in the age of AI, we need to be one of the most important compute providers in the world, and our goal by the end of 2027 is to be providing at least half of the world's AI inference compute. We think we could be further than 2x, given that we don't have all the constraints, but in order to get there, we do need to be very aggressively building out, and we need to give people no excuse for not running their models on us and using the models that are on us by charging extra. And what I keep telling the team over and over again, because you have to remind them sometimes, is we are growing faster than exponential, and when you are growing faster than exponential, there is no amount of profit that you can make that matters. What matters is getting a toehold in the market and becoming relevant.
What would prevent that? We used to be worried that someone would try and price below us, and then we realized that wasn't a concern because there's so much money going into this that people are going to want to lose less money by running on us. So that isn't a concern, but that was the big one early on until we realized that when we see Zuck investing $65 billion in data centers, what does that actually mean? That means he's internalizing all of the margins that he would have had to have spent on data centers with the providers that we mentioned earlier. Facebook is doing full stack. What does that mean? So Meta is doing $65 billion a year. I think Google said 70 or 75, and Sacha said Microsoft's doing 80, and then you've also got Stargate, right? Yeah, these are crazy sums of money, and this is all for data center builders. No, it also includes the stuff that goes in; includes the chips as well, the systems, everything. Okay, we've never seen money like this. No, no, there's never been anything like this, but there's never been a case where it was so clear that there was going to be value at the end, right? If you knew how successful search was going to be, right? Remember Google stayed private as long as they did because they were afraid that Microsoft would figure out how much money search was making and then would try and replicate it. In the moment that they went public, Bing, right? They called that perfectly. Everyone knows how much money there is in AI, so everyone's going after it.
Do you think that value is distributed amongst many players or concentrated towards one or two? I completely agree with you in terms of the clear value when assigned, but is it distributed to some levels evenly or concentrated? It's a power law, and the more value there is in the economy, the more risk there is of a single entity being so far on one end that they just dominate. And you see this with the MAGS, right? And it's predicted: just the bigger the economy gets, the more you will have big swings in economic outcomes. Right now, the hyperscalers are all sort of even in their market caps. It's strange; you would expect one of them to just be killing it and taking it much further, and so I don't understand why they're so closely grouped. So when we think about that distribution, like how do we think about changing that? Then, like obviously with Grok, you know, you want to be one of the MAG 7; you want to be one of the most important companies in the world. How do you see that? There's so… the way that you get there and the way that you stay there are two very different things, and there's a sort of a circle of life that happens in startups. The first circle, the first stage, is solve an unsolved problem; that's how you go viral, that's how you do well. The second stage is the marketing stage, which is now other people are trying to copy what you've done, right, because they can't think of something themselves, and now you have to fight it out in advertising and marketing and whatnot. And you see CPG companies often get stuck there, right? And it becomes more about where on the shelf they are than anything else. And then the final stage is the seven powers; it's once you've found some of those and you've really started improving it and you have sort of systemic advantages. And then what happens is someone solves an unsolved customer problem, and the whole cycle of life continues. Right now, Google has to redo this because LLMs are better than search, right? So the way that you start off to become a MAG 7 is you solve that unsolved problem. The way that you stay there is first you find one of those seven powers or multiple, but then you have to be ready for when you get disrupted to continue fighting back and solving customer problems.
Can we mentioned the different huge amounts of money that's being spent here? Is this a good bubble that bluntly lays the foundations for an incredible next 10 to 20 years where, bluntly, the capital actually turns out to be productive but not seemingly so on paper, or is it where actually just a huge amount of money is incinerated on depreciating assets? I can guarantee you that a huge amount of money will be incinerated, but I also bet that in total more money will be made than will be put in. And so this is the problem: you have to look at it either in aggregate or individual bets, right? When everyone is making investments in the market, some people are going to lose money because not every company's going to be successful. So what you always see is when there is some real tech improvements or things coming, you've got the things that were early that people are investing in heavily that are super successful, and then everyone else wants to get in on it, and you know it goes from you have AI chips and AI models to now you've got AI, you know, t-shirts, and next thing you know you've got AI thermal grease, right? It just… like people just start applying AI to everything. Next thing you know you'll have an AI condo.
Sure, yeah. And so the trick is discerning what is real and what isn't. You're always going to have all of these really obnoxious charlatans coming in whenever there's something real, and that's unfortunate, but eventually they get cleared away once people start to understand the technology and what's real and what isn't. And so the job is to start educating, and the more educated people are, the less they'll invest in AI thermal grease.
What is the largest individual bet that will lead to the largest incineration of cash? I'm not going to call anyone out in particular, but I actually think it will happen across every single discipline.
Are you aware of the Keynesian beauty contest? No? Okay, so John Maynard Keynes, the economist, he has this great… this will explain everything you need to know about VC.
I'm nervous, but keep going. So, um, take a magazine full of models, human models, like, you know, good-looking models, and have a whole bunch of VCs in the room, and they're allowed to make bets on who the most beautiful model is. And in the end, whoever has the most money on them is the winner, and based on the proportion that you put on that particular model's face, you get the share of all of the money. So if you put money on one that isn't the most beautiful by dollars, then you lose your money to the people who bet on that one. And that was sort of the bet that SoftBank was making, which was they could win the Keynesian beauty contest. I'm just going to put more money in, and I'm going to win. That is problematic when you have true technological advantages as opposed to marketing. When you're solving customer problems, it's a weighing machine. Once the customer problem has been solved, you then get into this sort of popularity contest of marketing. Now something unusual has happened this time around, which I don't think has ever happened in VC before, which is you see people raising billions of dollars who have competitors who've raised billions of dollars. It usually there is a clear winner in the Keynesian beauty contest. You don't have this like fight where, you know, it's sort of like, well, I got to put a little more money in, I got to put a little more… I got to, you know, put 10 billion in, I'm going to put 20 billion, I'm going to put 500 billion in, right? Because the Keynesian beauty contest has gone completely amok, and this has never happened before. And so now people don't even understand how to react because it used to be if someone had raised a billion dollars, you're like, "Oh, they're the winner." Now it's like there's three or four competitors who have a billion dollars, so who wins and who loses? Like, is Masa going to incinerate the largest amount of cash ever? I think the Keynesian beauty contest no longer applies here because there's so much money available being spread out, and I think you're going to see that the people who have the best products are actually going to be the winners because everyone can be capitalized, but there will be problems for the winners because of this. The problems are going to be of the sort: you had this employee that you were going to hire, and someone offered them a ridiculous amount of money. Yeah, you see this all the time now, and they could have gone and contributed to the winner, but now they're contributing to a competitor that shouldn't exist or is equally likely to win, and now you're splitting the talent.
What do you also do when you have such high salaries? We've seen a million, two million for kind of junior to mid-level in some of these companies, and they are living an amazing life, actually, in great places. You think they're living that amazing life in Guangdong when they're working for DeepSeek or any other Chinese alternative? I don't think so. I think they're actually getting paid much less, working their ass off 20 hours a day and not getting kombucha and being paid two million a year. Fair? Not only fair, we have a policy that we never offer the highest because we want people to choose us, not choose the salary. If we win in a bidding war, then that means the next time someone comes along with a higher salary, that's it; they're just going to go take that other job. There's no loyalty; they don't believe in the mission. Instead, we focus on: look, we're going to build this; this is your opportunity; you're going to get to work with amazing people; spend some time with the team; are these the people you want to be working with? Because frankly, you're going to make so much cash it doesn't matter. But bet on the equity, the outcome, right? Help us make this thing valuable. And people who buy into that, they're so much easier to manage because they're mission-oriented; they all want to do the same thing; they're not there because they want the kombucha, and they're not going to complain because the cappuccino machine is broken; they'll just go and buy their coffee next door.
Will you and Nvidia move into the model area? Everyone talks about model LRs becoming application providers; will infrastructure providers become model providers? We have decided that we're not going to train our own models. We'll do a little fine-tuning for specific cases or whatnot, but we don't want to compete, and that's really important because people are putting their models with their weights on us, right? And they don't want us to learn from and take that stuff for our own benefit. This is the problem you have when you work with a hyperscaler because you know they're also doing everything that you are doing. So we've decided: model providers, you make the model; we don't do that. I think there's also the data side of the users in the queries. So the other thing that we could do that we do not do is log the queries, and then we've got data if we want to train. We don't train; we have no reason to hold the data, so we only temporarily store things in the DAM. So there's no persistent storage; if the power went out, everything's gone, and DAM is limited, so we can't hold things for a long time. So you know that we don't have your data. Now people who are building businesses on top of us, you can obviously keep the data from your customers if you want; we have no control over that; that's fine, but we don't take any data.
Do you think Nvidia will move into model providing? It's possible, but I think… I mean, if I was them, I would avoid it because I wouldn't want to give the customers of mine… I mean, Nvidia is great at training, right? It's crazy. It would be like, you know, being an automotive, a car company, and then creating your own taxi service; you're now competing directly with your customer, right? And I think tech companies love to do this. We have a management philosophy, and it's based on Big O complexity, and we only do things that require a sublinear number of employees. So what I mean by that is if someone comes to me and says, "I need 10 people to go do this thing," a lot of people would say, "Well, why can't you do it with five?" I would say, "Okay, you're supporting customers; if we double the number of customers, do you need 20 or do you need 11?" Because I want to know what's that growth rate; are they automating everything, right? We completely automated our compiler; we completely automated everything that, you know, large portions of our cloud, and that means that we can scale with a small team. We have 300 people; we have 300 people. We built our own chip; we built our own networking hardware and software; we built our own runtime; we built our own orchestration layer; we built our own compiler; we built our own cloud. We built all this with 300 people. Now we would only be able to do this with a small number of people because you don't have the communication overhead. But if we… if you have to decide what your constants and variables are, what are the things that you want to preserve? And one of the constants is talent density. We want to stay small; we want to stay nimble. And the other side of this is growth is a problem, so we measure our growth in what I call problem units. A problem unit: every time you triple something, you have about the same number of problems as the last time you tripled. Going from 100 employees to 300, 300 to 1,000, 1,000 to 3,000; each one of those has the same number of problems. We scaled from 620 or 640 LPUs last year at the beginning to 40,000; that's four problem units; that's four triplings of the number of chips. If we were also tripling the number of employees, that would be another problem unit. Management bandwidth is limited; you can only solve so many problems, so you have to decide where you're going to allocate them. If you build things really well from the beginning and you can scale up with the number of employees you have, then you can scale over here if you want to triple the number of customers; there's another problem unit that you have to solve.
What's the biggest challenge when you are scaling at that rate, but then the team is not scaling in conjunction with it? There's this common belief that the people that you have early on are right for the job, and the people that you get later, maybe they're better in a sort of more corporate environment. I don't think that's the case. I think you should always try and get generalists, otherwise you get stuck in a specific way of doing things because that's what that one person knew how to do. But there are people who burn out; being in a startup is hard. Like, there are people who just literally burn out. There's also people who were the best that you get at the time, and then there are people who are just unmanageable wild children, and they should go off and start another startup, and they shouldn't be scaling with you. That happens, but it's the rarer of them. I think saying that you're going to hire B players because you've gotten large enough is laziness and an excuse, and it's a lack of creativity in your business model and how you're going… the algorithm of how you're going to scale. Think of it this way: Walmart versus Amazon. Walmart: Walmart wants to double the number of customers; they have to double the number of stores and employees. Amazon does not need to double the number of websites; that's a fundamental advantage. But Amazon still has to double and improve the logistics, right? They don't have as many problems where they have to scale linearly, but they have some. You wanted to disrupt Amazon; what you would do is you'd build a completely robotic logistic system and bring the comp… the overhead and complexity of that down, and then you can outmaneuver them, right? That's how you need to improve; don't just say, "I need more people"; focus on the algorithm of your business.
The last time we spoke, we discussed DeepSeek. I think more has come out over the last few weeks about their innovations, some of the distillation that they used. Where is China better than us today? Well, as we discussed, they're more willing to use things that maybe they shouldn't be using, you know, they distilled the OpenAI model. A lot of people have the opinion, well, OpenAI was scraping the internet, so you know, good for DeepSeek, but whether that's right or wrong, most of the model providers had considered that a red line they didn't want to cross. I don't know if that's going to change, but it might. But the other… the open-source nature of DeepSeek, OpenAI now benefit from the innovations that they did. Also have… well, and they also probably have all the data that DeepSeek paid them to generate, so yeah, I… but but they also were clever; they innovated. I think the biggest thing is this is a shot in the arm for morale in China, and it gives them a sense… but but again this, you know, as I said, Sputnik 2.0; it's also woken up the US totally has…
How do you compare Stargate to the $128 billion that China has now committed? China has a more complicated situation and a simpler one at the same time. The problem is they don't have the technology that we have in terms of the chip efficiency. On the other hand, they have scale. If they wanted to deploy 150 nuclear reactors, is I think the plan is, no big deal; they just do it. So if the chips aren't as efficient, they can just deploy more of them. On the other hand, if they want to go out into the world and deploy chips like they did with Huawei and networking gear, that's going to be complicated because people aren't going to have the power around the world to run more expensive accelerators. The home… I don't think anything is a problem; I only think as they're trying to expand, it's going to be an issue. China is quite opaque in everything.
What do we not know about China that we would like to know? I think the most important thing to understand is where they're going to end up on the censorship and privacy of these models. We come from democratic countries; we have an expectation that companies can build something that says anything. Are they going to be permissive and allow models to make mistakes and hallucinate, or are they going to shut it down? Because I think if you know that… you know, whether or not China has a shot… one of the biggest nightmares that they have is free speech; it's the exact opposite of that vulnerability we talked about earlier. Can you imagine Xi Jinping going out and saying, "Country, we've lost our advantage in AI; I need your help"? Never, ever. It's always going to be, "We're the greatest; we're the best." Everyone's going to know differently, but they're all going to have to tow the party line, right? Because of that, I think it's really hard for them to just allow these models to say anything… say, you know, "The US is great and better at that." That's a bad thing for them, and so that's going to really tell you a lot about the AI story in China. And so if they aren't permissive of more open, truthful models, then they're inherently disadvantaged.
You're saying, well, look what… I forget if it was… was it Jack Ma who got in trouble with the CCP? Yeah, yeah. Um, if they aren't more permissive, then if you are running a Chinese tech company, your fear is that you become Jack Ma. That's really going to stifle innovation. If I was in China right now, I'd be looking for the exit; for how I could do… like, if your craft is AI, I would want to do that someplace that's supportive.
Do you really buy that they don't have access to Blackwell? This is China. I think Xi Jinping's like, "Well, sorry, no Blackwell." Well, I don't… I don't think it matters whether or not they physically have it because right now most of the cloud providers are happy if you swipe a credit card to rent it to you, but there is limits to renting. No, I… I think if you… so one of the concerns right now is about Malaysia or Singapore, that region over there, being a place where people are deploying GPUs with the wink-wink, like, "We're not going to rent it to China," right? But that's… that's a belief that a lot of people are doing that. Otherwise, that's a lot of GPUs for that region. That feels like it's even more of a safety net just in case the tap ever gets turned off at the hyperscalers because right now you could just write a check to any of the hyperscalers and say, "I need these chips," they'll deploy them, and you can run on them; doesn't really matter where you're coming from. I mean, if you're a sanctioned country… no, China's not sanctioned.
Okay, so we have China, where China is obviously in terms of innovation and actually proving that they are in the race. We have the US, and then we have Europe, which feels like it's languishing. Yeah, is this the ultimate nail in Europe's coffin? We talked about how Grok almost died, but we had the right technology all along; we're just waiting for the thing… for the LLMs to arrive. And I think Europe's very similar. I think Europe has amazing talent, amazing talent, but that talent leaves and goes to the US or other places. So the question is how do you have Europe's LLM moment? How do you position yourselves? And it's not that complicated; the problem is when you surround yourself… you become the average of your five closest friends, right? If your five closest friends are like, "That'll never succeed; ah, you should just keep your job; ah, startups, they're terrible," then you're going to be risk-averse. But if your five closest friends say, "You should do it; that's great; I support you," then you're going to be more likely to do a startup. And even in Silicon Valley, people make that transition from the big tech company to the startup, and it's hard, right? They're, you know, comfortable, right? They're making those crazy salaries; the big companies take care of them, and they have a fiduciary obligation to their family. How do they make that leap? And it's because you've got tons of entrepreneurs trying to hire them, and they hear the pitch all the time, and they get used to it, right? They also see the success around them, and then VCs come in and try and close some of the candidates in the early stage, too, right? Europe needs the same thing; you need a place where people are surrounded… just surrounded by entrepreneurial people who are risk-on and who aren't going to try and talk people out of joining a startup.
From a regulation perspective, Europe is, you know, unbelievably efficient in the masters of regulation. You know, I was speaking to someone the other day; the EU has supposedly hired 1,500 people for AI safety and policing. What would you do if I put you in charge of European AI regulation? Well, I wouldn't waste my time regulating something that doesn't exist. Instead of regulating, what are you going to promote? You want to promote risk-taking; you want to promote that enclave of people are risk-on. So I was just visiting Station F yesterday; amazing; Macron was there; it was like full of people, right? Vibrant; you feel it. I would… and I was talking to the person who runs Station F, Roxanne, and Xavier Niel, and with Roxanne we were talking about, what about a City F? What about a place where there's… you start off with like 10,000 people in the center, right, a little radius, and then once that's full, you expand it; once it's full, you expand it, and so on to get to like a million people in Europe who are all risk-on; the little Silicon Valley here. And I would give it special economic dispensations; I would allow everything that employers need; I would make it simple, and I would say, "You know what? If you don't want to buy into that, that's fine; go to other regions in France; go to other regions in Europe. But if you want to participate in what is going to be the biggest technological revolution in human history, this is the city for you."
You know, inherently punishing incumbents then, and what I mean by that is if we are talking about… I'm just using this as an example… AI insurance underwriters, startups… yeah, there's many companies that are going after insurance underwriting in AI, and you are giving them benefits like that; you are inherently punishing some of the biggest providers of insurance in your region; you're inherently punishing people who hire 200,000 people. That feels unfair. So there is no right to be an incumbent, especially a slothful incumbent that is not reacting to disruption, and you want to encourage disruption. And this is one of the things in Silicon Valley: you can move from one place to another; there are no… I mean, we had non-solicits when I started, but even that's gone, right? So that free movement of people is very important.
Are you allowed to start work straight away? Straight away, but not before. If you start before… that's… we have months… months… six months. There's no such thing. And so in that region, I would say you can immediately start, like literally the next day. That is so good. We have to wait six months; it's not good. If you are a company right now, it feels like, okay, well, it's harder to poach, but what does that do? It suppresses wages; it's harder to hire someone; they're less likely to move; there's less competition; it suppresses wages, and by the way, the company has to pay for the six months anyway; it makes no sense at all. I… so I totally get you and understand that.
Can I ask you, you know, you mentioned like what would you promote? A lot of people would promote… I love the way you said risk-on. Being a European, I actually thought first safety and regulation, but specifically safety. So sticking with that, all that Dario will talk about these days is safety. Is he losing a step by being so focused on safety when, bluntly, his competitors are talking about product? So safety matters in AI; it's a little bit like nuclear power; lots of pros, lots of cons. I'm worried about different things than I think… than Dario is worried about. I'm more worried about people voluntarily giving up their decision-making authority because it's so easy. And this is what I mean by preserving human agency in the age of AI. A good analogy is: you probably know plenty of wealthy people and the struggles they have bringing up children with wealth. I refer to it as financial diabetes, right? You have children who aren't incentivized to… they're not… they're not going to strive to succeed.
I was very fortunate when I was growing up, and so I actually just told this story for the first time today, so no one's heard it, but I was fortunate because my father lost all of his money multiple times, just… and I've heard you say the same thing. Yeah, and he would sell a billion-dollar life insurance policy, and he would get all the commissions from that, and you would have tons of money, and then you would spend it all. And so there was one time we were living in a $20 million mansion, and there was a couple of times where we ordered food; he would talk to the delivery guy, and he would convince him to give us the food, and he would pay him back later because we'd get money later. But this time he was like so despondent, he sort of locked himself in his office and wouldn't come out. My little brother came to me and said I had to go and talk to the Chinese food delivery guy and convince him to give us the food, and I was like mentally preparing how to convince him to do it, and I walk out and I walk up to him and I'm like getting ready to do my whole spiel, and he hands it to me, and I'm like, "I don't have the money right now." He's like, "Oh, yeah, pay me later." I didn't have to… fortunately didn't have to do anything; he just trusted because we're living in a $20 million mansion. But that happened multiple times, and when that happens multiple times… like I have a friend who was homeless once for a couple of weeks, and he'd almost been homeless a couple of times, and he said the best thing that ever happened to him was that he was homeless for a couple of weeks because he survived it, and he's like, "I've been through it; I always viewed this as the worst thing that could ever happen in the world, but now that I've been through it, I can survive it; I'm not worried anymore." I think we live incredibly comfortable lives, way too comfortable; most people don't have to go through that. And so we have the sort of financial diabetes as a society, and I think it's going to get worse with AI. I think we're really going into an age of abundance; very few people have to worry about food security now, but what happens if you don't need to worry about home security or ending? What happens if you can just live a life without working, and what is that going to do to your psychology? And so as we enter an age of abundance, how do we get people to still be making their own decisions and have a fulfilled life? Do we get better, or do we get accepting of "good enough"? And what I mean by that is, you know, now, bluntly, with the majority of schedules, we will start with OpenAI, and we will do deep research, and then we will kind of use different prompts depending on different guests, and then we supplement it with a huge amount of research from speaking to Chath and speaking to Scooter and speaking to everyone in between. We care about it being good enough first and then great later with all the references. Most people will actually just be happy with good enough and get away with it. Do we, as a human society, get happy with good enough?
When we hire, we hire for something that we call booking the win early. So one of the most important driving forces for people is loss bias. When you have something, you don't want to lose it. People are less likely to go after something that they've already had. And you did grow up in a family that was well off, and then you lost that; that might be part of the drive because you want to get back to it. When we have an engineer that we're hiring and there's a room full of people who are saying, "You know, if we do this thing, we could be twice as fast," I want that engineer to hear, "Wait, if we don't do that, we're going to be half the speed we could have been." The loss bias, right? Book the win early because it's possible; it must be done. I think that's a smaller segment of the population; those are the people who deliver amazing things that no one else is going to
I know gold, but I'm made decisions, exactly well, this was a very made decision because I had to consolidate everything we were doing into one very simple message: we're going to get to 25 million tokens per second. And then I engraved it on a coin, on this tiny amount of space right here, and gave it to everyone at Grock. Now, whenever we're in a meeting and something doesn't help with this, they can just tap their coin on the table and be like, "No, no, no, that's not the way this is going to go."
So is everyone wrong on Founder mode? Then I think that's what you do when you don't have the quality of people working for you; you need the right gearing ratio between you and your direct reports. It's a really unfair question, but I have to ask it: how do you analyze Elon's attempt to buy Twitter—uh, not buy Twitter, to buy OpenAI?
I was sitting at the Elise Palace, or however I pronounce it, um, at dinner with McCon and Sam Altman. So it was McCon, it was J.D. Vance, and it was Sam Altman. Frankly, I think Elon was a little jealous that Sam Altman was sitting next to J.D. Vance and it wasn't him. Because it was right around the time that Sam Altman was speaking that he announced it. And frankly, I thought Sam's tweet response, part of it, was pretty good. I would have probably said, instead of whatever he said about 9 billion, I would have said, "Yeah, I'm going to take Twitter public at $420 a share." It was just—it was attention-grabbing. Some people can't stand to not be getting attention, and so my revenge on this is to give as little attention as possible. So let's move on.
What would you do if you knew you couldn't fail? I would put in 100% of the orders for every single chip we could possibly manufacture, because right now the demand is unlimited. But every time you triple, you find the same number of problems, and so you got to keep—you got to do it a little judiciously. But if I knew that no matter what problem was going to come up, that we didn't need to be safe at all, I would just go, "Great, we're going to go build 20 million chips."
Done. In 10 years. Is NVIDIA 3x bigger, 10x bigger, or 50x bigger? I think they will be bigger. I couldn't tell you a number. Training will become more important. I wouldn't be surprised if they were 3x bigger. I also wouldn't be surprised if they stayed around the same. Wow, it's so hard to tell where things are going because remember a lot of assumptions in the investment in Nvidia were that they were going to run away with the entire market, including the inference market, including the inference market, and they just haven't built the right thing for inference. I do think that as a weighing machine they should increase in value, but so much popularity contest applied to it that I don't know if—if they're—they might need to grow to get to where they are. You know, they might need to grow their revenue to get to where they are, but it's a pretty fair multiple given everything going on, so I couldn't tell you. Like, the popularity contest skews everything.
What's a crazy AI prediction you have that everyone else thinks is science fiction? I would assume that in the next 10 years—and I know this is going to be crazy—but you—you saw that picture of me and my weight loss, right? Unbelievable, dude. 70 pounds. 70 pounds. Yeah, but I was on Mounjaro. So if you know anyone is overweight and it's hurting their health, get them on Mounjaro as soon as you can. It works. What is Mounjaro? It's one of those GLP-1 inhibitors, one of the weight loss drugs that have become popular recently. It works. But my crazy AI belief is that if it is possible—if it is possible—to significantly slow or stop aging, I think that you will have a Mounjaro moment in maybe the next 10 years. Because that came out of nowhere. All of a sudden, you know, you could just lose weight. Something finally worked, and it's worked for a bunch of people. You probably know a bunch of people who've lost weight. Yeah, exactly. And I don't know if it is possible to slow or stop aging, right? Some wear and tear is a real thing, and it might just be impossible. But if it is impossible to slow or stop aging, then I think in the next 10 years we will do it, and it will be sudden. It'll be like the Mounjaro—um—and the other one as well, the other—it'll be like that moment. I don't see how it is not possible. Like when you look at the advances that will come in medical research, I don't see how it's not possible that we will at least extend, you know, longevity by 60 years. I mean, diary have will live to 150. I don't see why that's impossible. I don't either, but I also don't know that it is possible, and until I know that, I'm going to—I'm going to stick that conditional in there and say, "If possible."
What have you changed your mind on in the last 12 months? And this is less of a mental one and more of an emotional one: we didn't have product-market fit for seven years. Yeah, like that is terrible. Like the morale—like when you find product-market fit, the world is brighter, the birds sing, like I feel like hugging people. You sleep—I sleep—life is better. And you know, I forget if it was you or someone else, someone was talking about type one and type two happiness. Yeah, and I think there's a—I think there's a third. So as a Founder, the only type of happiness you get is this third type, which is future happiness. The other two, the common ones, are—the present is happy, right? And the other one is you went through some real crappy stuff, but the memories are make you happy, right? So there's past, there's present, and there's future. As a Founder, you're living 100% in future happiness. When you get product-market fit, you start to get—you start to get that past happiness, and when you start to get that revenue and everything, then you end up getting the—the present happiness, and it changes everything. I love that.
If you had to bet on one company other than Grok to define the AI era, who would it be? I would probably focus more on the companies that you haven't heard about. Um, and I don't know what the companies are, but I can tell you what they're going to do, and I can tell you what each one of them will be. Co—the first one will be the one that solves the hallucination problem. The second one will be the one who is best able to break down sub-goals for agentic. I think agentic comes after you solve the hallucination problem, because otherwise you got these long chains where you can introduce hallucinations. It'll kind of work, but it'll work much better after. I think the next one is what I—what I call the invent stage. So right now the way LLMs work, they make the most probable prediction. It's actually kind of amazing. It's like, I'm going to take an entire novel, I'm going to delete you, and you've got—a detective, you know, murder mystery, and you get to the point where the detective says, "And the murderer is," and it can actually predict it. It had to understand everything, right? But it's going to give you the most probable answer, and that's not good for invention. It's not good for, you know, art, writing. The reason that the writing from LLMs is terrible is because it's predictable. So how do you actually say something that's non-obvious but is obvious when you see it? We don't even have the right word for it, right? Non-obvious but obvious. And that is going to unlock invention. And then the final one is what I call the proxy stage, when someone makes it so that models can just make decisions for you. You can proxy your decisions, like the decision to do this interview, right? Like other things had to be canceled, the flight had to be booked, we had to get a ride over, right? You would trust an EA or a chief of staff to make that decision; you wouldn't trust an LLM, and that's the final stage I think before you get to general AI. But each company that does that is going to be a defining company.
Can I ask you—you said that we're going to have to fix hallucination before we get like efficient agentes. Does that mean that money going into AGI say will be burned? No, I—let's take an example on hallucination. So the examples I gave you were medical diagnosis and law were two areas that will be unlocked once we get rid of the hallucinations. But there are—startups like Perplexity that are doing just fine right now, even though there's hallucinations, because it's not high-risk and it, you know, it's for entertainment only. But—but if you click those links, you can check them, and it works kind of okay. It depends on how risky the industry is that you're in on whether or not you can get started trying to position for the wave early and generating—but if you're in the right position—we were in the right position for seven years, and the wave came. So that money isn't incinerated. In fact, that recent deal we just announced is more revenue than the money we've raised. So does that—how does that cash hit? I know it's a ner—throughout the year—throughout the year. Yeah, but it's this year; there's potentially more next year, lot more. What is that contract in three years? If we sell everything that we possibly can this year, it is many billions, but just from the capacity alone, there's—there's tens of billions of capacity of hardware that we could build next year in these sort of deals. But we're also doing it at, you know, high volume, low margin. So if we were talking about GPU sales and GPU prices, I mean, we'd be talking about hundreds of billions; we're just not charging that much.
Final one for you: the thing that I'm singly most excited for is actually like—disease discovery in terms of drugs. You obviously—my mother has MS, and that's incred—it was always taught to me that it was incurable, and actually like, now it's like, actually maybe not. What are you singly most excited for? We went from a phase where people were hardware engineers to they were software engineers. To be a hardware engineer is ridiculously difficult. The training—you have to get things right; there's a real expense if you get it wrong. Becoming a software engineer so much easier. All you have to do is get a little bit of time on a machine, and you can teach yourself. Nowadays, you can just download—manuals from the internet—or tutorials from the internet—or whatever. I think prompt engineering is going to unlock a huge swap of human society. There's 1.3, 1.4 billion people in Africa who know how to—who know how to speak, and if you were to give them access to a tool that they could create applications live just by speaking to it, that would be another 1.3, 1.4 billion potential entrepreneurs. There's 8 billion people on the planet, and the difference is hardware was just ridiculously difficult—it was arcane knowledge that's hard to get. Software was plentiful; language you already know it; you don't have to learn a thing. What's that going to do for venture? What's that going to do for entrepreneurialism, Jonathan? I love talking to you; it's always such a broad and wide-ranging discussion. Thank you so much for putting up with me in person, and I've loved it. Awesome. So glad to be here.