📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Andrew Feldman, Cerebras Co-Founder and CEO: The AI Chip Wars & The Plan to Break Nvidia's Dominance

20VC with Harry Stebbings1:14:37

Transcription

Our AI algorithms today are not particularly efficient. In a GPU, most of the time it's doing inference, it's 5 or 7% utilized. That means it's 95 or 93% wasted. We won't be as dependent on Transformers in 3 years or 5 years as we are now—100%. The fundamental architecture of the GPU, with off-chip memory, is not great for inference now. They will continue to do well in inference, but it can be beaten, and I think they know it.

Ready to go. [Music]

Andrew, it is such a pleasure to meet you. Man, I've wanted to do this one for a while. I've heard so many good things from Eric for a long time, so thank you so much for joining me.

Well, Harry, thank you for having me. I appreciate it.

Not at all. This will be a fantastic conversation. I have my pen ready. I feel like this is going to be a learning experience for me. Um, I want to go back to 2015. What did you and the team see in the AI landscape in 2015 that led to the founding of Cerebrus?

We saw the rise of a new workload, and this is every computer architect's dream. We saw a new problem to solve, and what that means is, is maybe you can build a new machine better suited to that problem. And so in 2015—and the credit goes to Gary and Shan and JP and Michael, my co-founders—they saw on the horizon the rise of AI, and that what that meant was there'd be a new problem for computers. That what the AI software would ask from the underlying chip, the processor, would be different. And we came to believe that we could build a better machine for that problem. That's what we saw. You know, obviously we didn't see it exactly right. I underestimated it. You know, this is my fifth startup, and the first time I underestimated the size of the market by a lot. Um, you know, it uh, but what we did get right was that this was going to be big, and it would put a different type of pressure on a processor, and that it would put pressure on the memory bandwidth, that it would put pressure on the communication structure. So that's what we saw. We dove in. It's been an extraordinary 9 years.

It's been an extraordinary 9 years. Can you just help me understand how does the movement into an age of AI change the requirements from a chip perspective of what is needed for a provider, and how that then resulted in how you built Cerebrus?

The way to think about a chip is it does two things: it does calculations and it moves data. Right? This is what what what what a chip does, and uh, sometimes along the way it stores data. And so uh, what AI presented was a very unusual combination of challenges. First, the underlying calculation is trivial; it's a matrix multiplication, and an FMac can be developed by any second-year electrical engineering student. So you say to yourself, holy cow, this has a huge number of very, very simple calculations. The hard part with AI work is results and intermediate results have to be moved a lot, and therein is the most complicated part. They have to be moved to memory and from memory, and they have to be broken up and moved among GPUs. And what we saw was that this was going to be the hard problem, and that if we could solve for that problem, we would build an AI computer that was faster and used less power.

When we think about how we're going to build and what we're building for, there's to me kind of a couple of core elements, which is like where you're going to focus. Are you focusing on, you know, fine-tuning? Are you focusing on, you know, training? Are you focusing on inference?

The three. You chose all three. Yeah. Why? And I'm sorry for my base questions, but but I thought like GPUs were specialized towards training and they weren't specialized towards inference. Can you have a mono architecture that does three best?

The first step in computer architecture is deciding what you're not going to do. Right? That that's what are we not going to be good at? Is really the first important question I think to answer. To your question you say, is the computational work for training from scratch different from fine-tuning? And the answer is it's not different; it's approximately the same. Now, inference and training have some different requirements, and generative inference, in particular, has some very challenging requirements on exactly the communication dimension that I mentioned. In generative inference, you have to move all the weights from memory to compute to generate a single word, and you have to move them again to generate the next word, and again and again. So if you have a 70-billion-parameter model—not a giant model—and each weight is 16 bits, you're moving what, 140 gigabytes of data to generate one word? This is an enormous amount of data movement across memory, and that's called what what what's consume that needs is is memory bandwidth. And if you have an architecture like we saw in the GPU that is your fundamental limitation, it's a fundamental architectural limitation, and that was what we went to—way for scale to solve. They use memory, a memory called HBM, to type a DE, and it is phenomenal memory, but it's slow. It is slow and high capacity, and when they set the architecture for graphics, that's what you wanted. You didn't have to go back and forth to memory very often. SRAM, on the other hand, is unbelievably fast but has low capacity. And so we wanted to use SRAM, but if you build a normal Siiz chip, you can't hold a model. And so by going to wafer scale, we were able to put down a huge amount of SRAM and get the benefits of speed and enough capacity. If you build a normal Siiz chip with SRAM and you want to do a 400-billion-parameter model in inference, you might need 4,000 chips, or if you want to do a deep SE 671, you might need six or 8,000 chips.

What an administrative nightmare! I'm sorry, you can can keep it on on as much as you can on one wafer or two wafers or four or 10. You get all the benefit of the SRAM, and because you've been able to use the wafer, you get this tremendous capacity as well.

Can I ask you first, I totally get you on HBM and kind of the slowness of it. Why is it then that, bluntly, so much of the market just continues to use it? And 40% of Nvidia's revenue is using that—that ships for inference.

Look, uh, there wasn't really—unless you went to wafer scale—there wasn't really a credible other choice. Um, that this is the way GPUs had always been made. It's called a graphics processing unit, right? That that that's that's the way they were built, and they it was part of their advantage against a CPU was they were built this way. But now they're dedicated chips like ours, and what used to be their advantage is now weakness. And that's that's a fun market to be in when uh, over a very short period of time, what you're good at becomes your weakness. With a market cap like they do and with Jensen as good as he is, which I'm sure we both agree with, they must know—one of know this. They do know this. There's not a lot of choice. Well, I mean, a they don't make memory, so they're a consumer of other people's memory, right? And that's SK, right? The high guys or Samsung, I mean Micron, there're only so three or four or five companies that make huge amounts of of memory. Not many choices, but it's it's part of a a complex architectural tradeoff. You know, the flip side you could say it's worked really well for them, right? Look at where it's taking them. Um, but in comparison to those of us who wafer scale, it's a small set; it's a set of one us. Right? We have we have real advantage against them on on on inference.

How do GPUs fit into this? We've got HBM, we've got SRAM with you, and bluntly, having many more of them to make it work and scale, where do GPUs fit into this mix?

In our business, there there there are a lot of ways to skin a cat, and uh, you know, our way is is different than Nvidia's way. It's different than the TPUs, different from Trainium; they're different. Um, right now, and every day since August 26th when we launched inference, our way has been the fastest way uh, across a whole set of models tested by artificial analysis and others.

Can ask when we think about kind of that speed? I am—you said that kind of you're one of one with wafer and kind of the architecture associated. What does that mean in terms of cost? With such efficiency, is it inherently more expensive, and what does that look like from a cost profile?

This isn't our our our first dance. We've been building computers for for a long time, and when you make a choice like wafer scale, you you have to weigh the tradeoffs. We use less power. We use less power because one of the most power-hungry things on a on a chip are the IOs, right? Are moving data off-chip. And so if you are moving data off-chip frequently, you're using more power than if you can keep it in the silicon domain on-chip. So we knew we would use less power. We knew if you went to wafer scale that you had to solve some problems that people said were impossible to solve, like yield. So we had to invent techniques that allowed us to yield wafers. In fact, we invented techniques that allow us to yield as well or or better than others who are building much smaller chips.

Can I just interrupt and ask what is yield and why is it impossible to solve?

Okay, uh, a wafer becomes—it's a uh uh 12-inch uh diameter circle, a slice of of silicon, and your your chip is punched out of this, the way your mother might take a cookie cutter and cut out cookie dough. And uh, during the process, at some point, just like your mom might have done, she lifts up the edges, and all the little bits are removed, and what's left are just the cookies. Those are your chips. Um, now what happens is there are a set of naturally occurring flaws, and that's like your mother closing her eyes and throwing up a handful of of M&Ms. Now, the bigger the cookie, right, the higher probability you hit an M&M. The bigger the chip, the higher the possibility that you have a flaw. And traditionally what you did when you had a flaw was you threw away the chip or you sold it as a less valuable part. You shut down part of the chip and sold it as a less valuable part—something called binning. So every wafer is going to have flaws. The bigger your chip, the higher probability you hit a flaw, and the more wa—the more part of silicon is wasted when you throw it away. This is what everybody thought was known truth, and one of the the things our our team realized was that there are other ways to handle flaws, like what what if instead you built your computer, you built your processor out of hundreds of thousands of identical tiles, and say there was a flaw, say you just shut down that tile and worked around it, say you had a row or a column of redundant tiles that when you needed them you could just pull in. Now that had been traditionally the the the technique used in memory making, and the memory yields are extraordinary. And so it occurred to us that if we could build a computer, build a processor, built of hundreds of thousands of identical tiles, we could use redundancy such that when there was a flaw, we could just leave it there, shut it down, work around it, and pull in one of the redundant tiles. And that had never been done in a in a computer before, and that's at the heart of our architecture that allowed us to yield and deliver whole wafers. Nobody ever been able to do that in the 70-year history of our industry. Really, really smart people struggled. I mean, Gene Amdahl, one of the fathers of of our industry, had a company called Trilogy that that crashed and burned trying to do this. Um, and uh, we we figured it out.

When you speak about kind of being the fastest and across all benchmarks being the fastest, what matters the most? Is it being the fastest? Is it being the most efficient? Is it being the least costly? How do you think about the stack of prioritization for your customers?

I think it varies. I I think uh um, look, if if if you go to to to to get a cancer diagnosis on, God forbid, your mother or uh your wife, I think uh 93% accuracy is just plain not as good as 94% accuracy, and you pay a lot and wait another week to understand what the accuracy is, right? You pay a lot. Right now, on the other hand, uh, if you want Llama 405B to generate data to help you tune Llama 70B, uh maybe you can you can wait a few days, three days a week more. You don't—there's no urgency there. On the other hand, if you want an answer from Perplexity, right, you don't want to wait 45 seconds for a search answer, right? You don't want to wait in a chat; you don't want to wait 3 minutes for R1 on GPUs to give you a an answer. What we know is that in interactive mode, milliseconds matter. In interactive mode, what's—holes over Google years ago showed was that you can destroy your user's attention with milliseconds of delay. So being the fastest matters everything in that domain. So I I think what you have to do is sort of be thoughtful and say, in some cases, being the fastest doesn't matter; we'll call those batch uh lots of maybe cheapest matters there. In other domains, there is no search if you got to wait eight minutes to get an answer, right? That's not a product that—when you go fast, a whole set of of new opportunities open up. I mean, right, Netflix used to to mail DVDs, right? That's what happened when the internet was slow. They'd mail DVDs. The internet, right? I look, I look, I look young, Andrew. I'm not that young. I remember Blockbuster. If you remember Blockbuster, right? First, I mean, let's look at the history of that. You're exactly right. First, we used to drive to Blockbuster to get a DVD, right? Or uh a video. Uh, then in Netflix was mailing them to us, right? And then we got broadband, and there suddenly Amazon's a studio, right? It changed everything, and speed in inference does the same thing.

When we chatted before, you gave this great equation for inference. What was the equation that you gave for inference, because it was really helpful for me and understanding it?

It begins with the following: Um, training makes AI. That's how we make AI, and inference is how we use or consume AI. And so understanding how big the inference market is is understanding the number of people who are going to use it, how often they're going to use it, times how much compute each use takes. And right now we are in this rare time where the number of people using AI is growing, the frequency with which they use it is growing, and the amount of compute used in each instance of use is growing. That's why you're getting this extraordinary growth, and that's why it's off the charts right now.

When we think about the distribution of resources between training and inference, what will that look like in 5 years' time, because we've seen all focus go to training—well, not all, but a lot of focus go to training and not as much go to inference? What does that look like?

What we made until the middle of 2024, right? What we made in AI was a novelty; it wasn't very useful. Late in 2024, what we made began to be useful. What was that—chk—what was the turning point? I think the—if you look at the models, they became—I mean, ChatGPT was not really a technical innovation; it was a user uh user interface invention, and but it gave more people access, but we we didn't really right away know what to do with it. It was cool, right? That's what I mean by novelty—like, whoa, this is cool. But now if your marketing team isn't on an LLM each person several times a day, they're not doing their jobs. That difference between novelty—it's cool—and this is part of everyday workflow, and that's what changed starting sometime in Q4 last year and running into this year is AI became useful, not just to a select group in Silicon Valley, but to my dad, to my brothers, the doctors, to ordinary people uh who aren't buried in the the Silicon Valley discussion. And when you get them, then the market is is ripping.

Do you not still think we are so incredibly early, though? Going back to your point, how many—for sure, in five years' time, then where are we? Are we a hundred times bigger? Are we a thousand times bigger?

Think we're way over 100 times bigger.

Yeah. What does that mean in terms of what we need to equip ourselves to deliver? These are incredibly uh energy utilizing; it is incredibly difficult. Our our our industry consumes a lot of power. Yeah, does, and and a lot of water. Um, and some water, and we're seeing that come down, but are we equipped from an energy and a data center standpoint to deliver the inference requirements for a population that is as AI hungry as we are?

I I think a couple things. I think the the first thing is to admit that that this is a power-intensive problem. We we consume—our industry consumes an enormous amount of power. The second thing to say is therefore the burden is on us to deliver exceptional value as an industry. I mean, that that that's—you take both the good and the bad, right? In order to make it worthwhile from a societal perspective to to expand all this power, you better deliver the goods, right? We we we better use AI to find cures for uh for diseases; we better use AI to find a bunch of different uh solve a bunch of different societal problems. So that's the the macro view. Um, do I think that uh we are equipped? I think we are in a very unusual situation in the US where we have plenty of power, but it's in all the wrong places, right? We we we we have power in Niagara; we have power—what what we don't have is power where you want to build data centers, where we have good fiber. And what we don't have is a national way to relax the local regulations that make getting power difficult. And so when you go to Silicon Valley, if you want to build a data center, you're dealing with local government and installed interests, and that is not an efficient way to decide if you want to build a power plant or or put a new data center in, um, especially if it's large. And uh, I think those places that have ripped out some of that burden uh in taxes, for example, are getting a huge amount of data centers built. You know, when I spoke to, you know, Jonathan at Coreweave before, he said there were a huge amount of data centers being built that were not actually really equipped properly, and that we've seen this massive supply of data centers that are really kind of done by tourists, so to speak, and that is a massive problem, and that the provisioning of these data centers isn't there.

Do you agree?

I I think the following: I I think uh a data center is a is a construction project to begin; it's access to power, and it's a construction project, and it's got a design engineering component. I think there's been a a huge push for new construction data centers, and I I think uh we we will see—we we don't know if they're going to be good enough. I I think many of them will be fine. I think the guys who were there early were some of the Bitcoin mining companies like TeraWulf, the guys at Crusoe and and others, uh guys in Europe uh uh they were early in building buildings near low-cost power in order to run compute uh that used a lot of power, and they are some of the leaders now in some of the largest projects. Now, those are certainly not tourists; those are extremely sophisticated data center builders. Now, sure, there's some tourists, but there are a lot of of very, very knowledgeable data center builders building huge facilities right now—I mean, gigawatt-scale facilities both domestically and internationally.

How do you think about how the cost of inference goes down with the surge of demand that we mentioned—you over 100x? Does the price reduce 100x? Does it follow Moore's Law continuously? How do we think about the ever-reducing price of inference?

Look, there there are uh the cost of inference is built up of of several pieces, right? There's the power and space that is consumed to generate the response, right? That that that's a data center cost; that's an OpEx item number one. Number two, there's the uh cost of of the the computer. We can drive down the cost of the computers with each generation by driving up their performance, etc. But the other thing we can do is we can develop more efficient algorithms. Our AI algorithms today are not not particularly efficient. There's a tremendous amount of room uh in a GPU. Most of the time it's doing inference, it's 5 or 7% utilized. That means it's 95 or 93% wasted. So over time, I think as an industry we get better at things, we can drive the cost of compute down, we can build more efficient data centers with lower PUEs, and our algorithms will get more efficient, so that our utilizations on our now cheaper computers are higher, so you get a higher percentage of the maximum number of FLOPS, you get more tokens per unit time for the same power.

When you look at the inefficiency of the algorithms as you mentioned there and what that means for the utilization of the chips, why are people suggesting that we're at scaling laws already? That seems to suggest that there is so much room for improvement. How do you think about what you just said in conjunction with the idea that scaling laws—we're we're hitting this asymptote point? How do you reconcile the two?

I think that uh I don't think there's a lot of of debate among uh sort of senior ML thinkers that that that we have tremendous room for algorithmic improvement. I I I don't think there's uh a lot of debate there. There's even debate about whether the scaling laws are over—whether we run out of—ran out of mojo to keep making data or gathering data to fill these ever-bigger models. But OpenAI's work on 01 shows me that the scaling laws certainly for inference are are fully functional, right? The more compute you put on inference, the better answer you get. And so uh I I think that uh, you know, many of the leading models are now sparse, right? So they're not presenting all of uh all of the weights to to to to each token, and and that's one way to do it—present the important stuff, not the unimportant stuff. There are other ways to do it that we will invent and learn over time. But uh, I I think we know the—we have—we have sort of human models that aren't all-to-all connected. Many of our models today are all-to-all connected. That that's a lot of unnecessary connections—connections that don't produce anything that we still end up doing math over.

I'm sorry, what does all-to-all connected mean?

In many of the layers in an neural network, every element is connected to every other one, and that's not the way actually the the the learning happens—happens uh some are more valuable and some are not valuable at all, right? And imagine you you're going to read 50 books; you want to learn something. You can read all 50 books, or you could read three books that are really important, or you could read summaries of the three books that are the most important. The problem is we don't know which they are at the beginning, and there's a process that you could learn—there's things called dropout and all these other techniques to to use sparsity to help help solve these problems. We are early in the evolution of AI.

Plays right into this point that we'll get better at these algorithms. You know, Transformers aren't the end of the world, right? We we we'll get better—better will mean faster, more accurate, and more efficient. And I I think that's what's exciting about an ever-changing industry. I mean, that's not—that's why I'm not in all these other industries that don't change quickly—same nine years ago as they are today. But this show is kind of strange to me because I speak to a lot of people, and they think about the three pillars, and they're like, you know, compute, algorithms, and data, and you actually—a lot of the uh common refrain is that actually we're very far along in all of them. Um, and that has been the refrain, and when I hear you, it's like, actually, it's very exciting.

I think they're wrong. I think they're wrong. Uh, I I I don't think we're very far—far along, and it's very difficult to say that we are early in an industry but we're far along on all this underpinnings, then what, right? I think we are early in all of them.

If we just take them one by one, in 5 years' time, how much synthetic versus human data will be used to train models? If you were to put a percent on it?

Almost all synthetic, and the utility value of synthetic is the same as human. I think this—I think when you teach a pilot to fly in a simulator, right, there is a lot of potential data that isn't very useful in teaching a her a fly, right? They spend a lot of time going straight, doing nothing as a pilot. Now, takeoff and landings are where you want to spend your time, and that's why when we put them in simulators, that's what we have them doing. And in simulators, we can create data where engines blow, where there are a whole set of problems where learning can take place—that's simulated data. And in the same way as we think about creating data uh whether it's for self-driving, whether it's for other forms of AI, what we want is the data that's hard to gather, right? Otherwise, we just have a bunch of data of people driving straight on a freeway—not not difficult; we've been able to do that for a decade. What we want is an unprotected left turn in the snow—it's snowing, it's hard to see; you got an unprotected left turn—that's a difficult thing, and you want that thousands of different ways, millions of different ways. That's where the synthetic data comes along is to use it to fill in the empty parts where it's really expensive or painful to get that type of data. Think of the pilot, right? You want them spending a huge amount of time on things that are rare in their training. Same with a surgeon—a huge amount of time on things that are rare; most of the time it's carpentry, but their expertise is only when it's rare—something happens, the unexpected occurs, right? That's when their mettle is shown. And I think we we will get better synthetic data by by a great deal.

I love it. I get it. From a consumer perspective and from an expectations perspective, if we move the needle on compute, algorithms, and data, what does that mean for the experience of AI?

It gets faster, cheaper. Faster and cheaper, faster and cheaper is the first answer. The second is is when things become faster and cheaper, new applications emerge. So it it's used everywhere, right? When uh when computers became faster and cheaper, suddenly they were in cars, and then you were in your pocket, and then they were in your dishwasher and in your TV and right—that's what happens. I mean, 30 years ago you're like, I need a computer in my TV? You kidding me? I need one in my pocket. Now you've got powerful computers in your pocket; you've got them in your TV; you've got in your kids' toys; you've got in the car. That's what happens—diffusion of innovation accelerates when you make things faster and cheaper. This is Jevons' paradox and Satch's belief there.

No, yeah, yeah. That uh I know in the VC community you got to you got to cite 19th-century English economists. Uh, I I think J—I'm English. I'm English. Come on, if I'm not allowed to cite an English philosopher, what am I here for? Look, I mean, I I think uh are you just like, oh…

He VC being like, "Oh, Java's power off, that's right. It's like, look, make stuff cheaper and faster." There are no—there are very few examples in our industry, actually none in compute, in 50 years in which by making things cheaper, faster, the market got smaller. Market always gets bigger, um, always.

Can I ask from an architectural standpoint, you mentioned Transformers? Is there a world where we move past Transformers?

Transformers 100%. We won't be as dependent on Transformers in three years or five years as we are now. 100%. They're not the end-all be-all. Why is it what will replace it, and what does that look like?

I don't know. I don't know whether they're going to be States-based models. I don't know whether they're going to be other types of models. But what I know for sure is that Innovation doesn't stop and that uh that the Transformer uh has some weaknesses that that people are desperate to overcome. Um, there's a quadratic effect in the attention head. Um, the there's all sorts of things that that could be improved, but it's pretty darn good now, the best we have. And that's what you run with; you run with the best you have. And the minute it's not the best you have, you drop it in favor the best you have. And I I think that's that's what we're seeing. We're seeing I the number of of innovative companies designing models is large. And what Deep Seek showed us is you don't need 5,000 people and you know billions of dollars a gear; you can do it with 200 smart people and more gear than Deep Seek said they had, but less gear than others had.

Were you very impressed with Deep Seek, and what impressed you most?

I think it was a result of focused engineering, and that impressed me. It was designed to be better, and uh they they weren't confused about being sort of model intellectuals, or they weren't confused about uh whether it was important to break new ground, or they were interested in being better. And uh from an invention standpoint, that's a little boring, but from an engineering standpoint, that was sweet effort. They really built a model that was just plain better at many, many things, and that's cool. I I like good engineering projects. Now that they chose to announce it right around Trump's inauguration and the politics of it, that that's all a separate matter, and and we can talk about that later. But um, is distillation wrong?

People up. I don't think distillation is wrong. I mean, is summarization wrong? Um, I'm a VC; are you kidding me? That's what we do. If you didn't summarize, you wouldn't know anything, right?

That's right. Exactly. I don't think distillation is wrong, and if distillation is wrong, then certainly using people's copyrighted data is wrong, right? That's the problem. The problem is, right, is you got to be a little bit consistent. Uh, well, S's been guessed many times, and we hope he will be again. So I hope he to no. I'm I'm I I think neither are wrong, actually, but I think you have to be consistent.

Well, the thing it was with it bluntly, Deep Seek is open, so everything that they did innovate on, OpenAI can learn from and take too. That that's look, I think there are few examples of an open-source anything having the sort of immediate impact that model had, right? I mean, that that model had a giant impact in a technical community of really smart people, and there are very few examples of other open-source uh software projects that had that type of impact in that amount of time. Usually, and you know you're you're in the business of betting on these guys; they ramp up and they oh, look, 10,000, now it's 100,000 users, it's now a million users, and we better better start a company around that, right? Get those grad students. But this had a loud boom uh in the industry immediately. It's like who the the thing I have to think as venture investors, where is enduring and defensible value, simply, and how do I get in early and and build that over time in hardware here?

Well, this is my question, like, I mean, you have to be a very smart investor like Eric Vish to do hardware, to be clear, um, but on the model side, do you think there is value when you look at the sheer number of players or with relatively comparable models?

I think to to demonstrate enduring value, you need both immediate value and a trajectory for more, right? I I think the problem is in some industries you you are capable of demonstrating a leadership position for a short period of time, and then someone else, maybe the next Generation, they generate the next, and the next Generation the next. And I I think that ends up in the soft world being, you know, you're competing against other people's release cadences; you're four months ahead, they're six months. That if that's really where you are, there's not a lot of value. But if you can stay at the top uh over years, right, even if you're not the best, even if you're, you know, top decile over years, and the people above you are changing constantly, I think there's a lot of value. I think uh very large Silicon Valley companies have been built uh with sort of not the most compelling technology; it might have started the most compelling technology, and then it got to a point where it's good enough; it was easy enough to use; it was well—right, that's that's when you're at the mature market, but we're a long way from there. Right now, right now we are in the early phases. You I you know you you characterize my position exactly right: data, compute, algorithm. I think we have a ton of room for improvement on all of them.

You said that you computing hardware, that's where the value is. How does that value distribution shake out? You know, we've obviously got the 800B gorilla that is NVIDIA. How do you think about how the distribution of value shakes out in hardware and in compute over the next 5 years?

Historically, um, one of the one of the barriers to entry was sort of the capital intensivity of a product of a project, and in in the world of building chips, there's both scarce resources in expertise, and it's very expensive. Um, and historically, it it hasn't fit very comfortably in a software company, and the things that software, modern software companies value are not entirely conducive to chip making. And so when I look down the road, I mean, I I think uh who has endured in uh in much of infrastructure tech—people who build systems: Cisco, Juniper endure. Um, uh chip makers have endured. There's a reason that Apple and NVIDIA are among the most valuable companies on Earth. There is what they do is hard, and I I think it's that's why it's worth challenging, right? That that's if it weren't hard, if it wasn't enormous and difficult, you know, why spend time being the underdog and and challenging it?

A lot of people place defensibility around NVIDIA is kind of CUDA locking. To what extent is that real versus hype in inference?

It's not real at all. There's no CUDA locking in inference, none. Uh, you can move from OpenAI on an NVIDIA GPU to Cerebrus to Fireworks service on something else to together to Perplexity with 10 keystrokes. I mean, anybody who actually uses AI knows there's no CUDA locking in. In I think there is there was a fundamental effort to disintermediate CUDA first by Google with TensorFlow and first by some grad students with Caffe and some of these early efforts, but later but Google with TensorFlow and then Facebook or Meta with PyTorch. I I think today most most AI is written in PyTorch. You ought to be able to compile it and run it on on your hardware. Uh, I I think NVIDIA has many moats. I think when you are a dominant market share leader, that in itself is a moat; that you're the default solution is a moat; that everybody learns to to to think about AI in your structures; those are moats. Um, the software, you know, compilers are hard, but they're tractable. I completely agree with you in terms of kind of being the leader is a moat in itself. It is; it's never talked about that way. And um, I mean, look at put OpenAI in that same—it is the leader; everyone's mother knows ChatGPT.

That let's look at Intel, right? Intel has made until hiring Pat Gelsinger prior to that nearly a decade of catastrophic decisions, right? And they still own 80% of the x86 market, 75% of the market, right? AMD is worked up to like 25% or 30%, and after a decade of screwing up, and you ask yourself, right, that that's a moat, right? How big is my moat? I can make a bunch of bad decisions for a decade and only lose 20% share—that's extraordinary. The moat was just unbelievable. Um, we'll see. I mean, I'm a huge fan of Pat Gelsinger; he's an investor in our company. I I I wish him well, um, and I I think if anybody can change that company, he can. But uh I I think we rarely talk about what being the market share leader means in terms of a moat in in the right context because as a challenger, we have to think about it exactly because it's exactly that that we need to; we need a bridge for, right? It's exactly these characteristics of the moat that we need to get over.

In 5 years' time though, is it Uber or is it like AWS and cloud, and what I mean by that is like cloud is an interesting market where like yeah, a couple of players or several players have relative segments, 25, 30%, and it's shared relatively evenly between them, not exactly, but relatively, or is it one like Uber where Uber has 90%, Lyft has five, and then there's alternative providers with the other five, right?

I think it's going to be between those two. I I think uh you in five years from now, NVIDIA is going to have 60, somewhere between 50 and 60% of market, right? I think right now they have approximately all of it. Um, I think they will come down over time.

Of NVIDIA's usage, what percent will be training versus inference?

I think they will continue to have a meaningful business on both sides. I think they're exceptional at uh uh at training. I I think uh they will not roll over and play dead in inference. I I think they are uh really like they're a world-class company. I mean, they've had one of the great decades of any company in history. I mean, from 2014 they were worth what, 10 billion, to where they are right now; it's one of the great decades in corporate history, uh, and I I don't think that they're going to roll over and and and oh yeah, we're we're not going to be in the in the inference market; that's not going to happen. They're going to have a meaningful share, but the market's growing, and we'll have a piece. I think others will have a piece. I think there'll be some very big companies made in this 100x growth; some very big companies.

Do you think chip providers will be far larger than model providers in terms of Enterprise Value in the 5-year timeframe?

Yes.

How does that prediction change in a different timeline?

I think in a shorter timeline, I think you know when you price an option, variance and uncertainty increases the options value, right? If you if you look at the the way Black-Scholes works, or if you look at any option pricing model, uncertainty is a a a friend; variability is a friend of the value of the option. And when people are paying these extraordinarily high prices for our model companies right now, I think part of that is this extraordinary uncertainty; is this wild variance. Um, and so in the shorter run, it it it might not be the case, but in the longer run, as markets mature, as we begin to understand the value of these models, we understand what their businesses look like, what their long-term uh net profitability looks like—what did Warren Buffett say about markets? In the short-term, they're a voting mechanism, and in the long term, they're a weighing mechanism, right? At at some point, the weighing kicks in, and uh usually it's in the public markets, and then investors say which which is likely to give me better growth in the future.

I mean, listen, you mentioned the word public there. I do want to just hone in on your business; your cash flow positive in a world where everyone else literally bleeds cash. Help me understand what do you do to making cash flow positive when everyone else is bleeding or hemorrhaging cash?

Traditionally, your your gross margins were a measure of of your technical differentiation, right? And I I think if you're running a negative gross margin business, I I think you're it speaks for itself; you're you're selling commodity; you're you're not uh your your value creation isn't being recognized in the market. And so I I think our technology is creating an opportunity for us to maintain margins where some others can't.

A lot of your revenue is concentrated to the G42 deal. To what extent is that a strength or a weakness?

It's both. Like the way you the way you catch three large customers is to catch one first; the way you build three large strategic partners is learn to be a strategic partner; that that's a learned skill. I I think uh we we didn't arrive knowing how to be a strategic partner at G42, and now that we've worked at it and worked at it, it's a muscle we can replicate. We could be a better partner to any of a dozen different companies in the world.

What have you learned in the G42 relationship build process that that makes C a good partner in a way that you work?

We we've deployed uh tens of exaflops of compute, vastly more than than anybody else that that that isn't AMD or NVIDIA, right? I mean, and a huge amount of compute. Um, we uh our software has been hardened on some of the largest AI clusters in the world. Uh, we've gone through the Growing Pains of increasing manufacturing 2X and 5X and 2X gam through unbelievable growth in manufacturing. We've worked with our supply chain partners to uh uh to be sure that they're ready for this extraordinary growth. We've I mean, I I think when you uh work with a strategic partner um of this size uh your organization comes out different on the other side, and there are things you've learned and there mistakes you've made and and you know you I hadn't done a big relationship in the Middle East; there was a huge amount to learn, and um you know, I think you come out a a much better company and much better prepared to do uh business with a hyperscaler, to do business with another massive partner, to do business with uh another sovereign. But it it takes real work, and your team has to learn.

You said you come out better. Yeah. Why go public when you did? When this happened, I was like, it seemed preemptive, respectfully. And my question now to companies is why go public at all?

Know bluntly, there is so much private capital; the decisions have shown I think very clearly that you can stay for a lot longer than you plan to. DataBricks certainly shown that, right? I mean, there those were historically uh public market valuations, and uh you know, the valuations that Anthropic and OpenAI and some of the others are getting are historically public market only valuations. And like you said, you your S-1's live; anyone can read it. I wouldn't want people reading mine. We have nothing to hide. I mean, I I think no, but your your competitors have got asymmetric information. Yeah, we've got asymmetric technology, right? I mean, I I think you you have to be pretty to to to be public. I think they're uh you have to be ready organizationally; you have to be ready with your processes; you need to be ready to to forecast and predict, to be held accountable uh in a way that private companies historically haven't been. We think that there's tremendous value; we think that we will be among the first in the category; we think we uh that some of our uh our largest targets uh would have a stated preference for doing business with public companies. Large Enterprises in the US have done that historically. Um, those were some of the reasons that led us to—

How many G42 relationships shall G42 will you have in the next 24 months? How fast can you ramp them?

That's a good question. Several.

Sorry, remind me how big is the G42?

It's 87% of revenue.

I know that it was big. I mean, when we announced it, it was some estimated it was north of a billion.

Wow. Well done. That must be a bit of a high five.

Look, I I think uh come on, you million-dollar first F for first, there's tremendous excitement, and then there's sort of every entrepreneur's reality is I got to make a lot more gear, right? Right. I I I need to make, and I got you got, you know, you make a list of your top 10 vendors, and you fly it to the mall and say, "Big orders are coming; be ready," right? You work with all your partners to get ready because you need to make a great deal more stuff. And that that's one of the real differences between hardware and software is uh when when we get when we grow fast uh the number of people you need to work with in your supply chain uh and that the amount of collaboration that needs to happen is truly extraordinary.

I totally agree with you on that. Sorry, are NVIDIA going to have a cluster of unhappy customers who bluntly have waited so long for chips, by the time they get them, the chips are outdated, and they're going—

What? All of that's an opportunity for uh for us and others. I think that's that's opportunity. I think being uh a market share leader isn't easy either. Um, but when you're late, right, uh when the bully falls, everybody wants to give him a kick, right? I mean, that's a lot of that happened at Intel. Um, they'd been the the dominant player, and when they fell, everybody was happy to jump in and kick them when they were down. And so uh I think there is a a real opportunity in the potential for NVIDIA customer unhappiness for sure for those of us who who are competing with them. I mean, if you can't get your gear, you may as well test somebody else, and that that's a huge opening. Head over to Cerebrus and use the promo code Harry20 for your chips today.

Do that. There we go. I'm here for you, baby.

Influence appr—yeah, no, no worries, it's fine. I I hope if we could do a 20% take and on the billion-dollar deal, I'm happy. That's fine. I I I know this venture business isn't been so good to you, Harry, and you got to get get shoes for your kids and the like, and and yeah, we're happy to to donate to to to to to the 400 million fund and fees when when I have no kids as well. No, no, look, two and 20 is a rough way to make a living, Harry. Dude, you don't you you don't get it, okay? You in hardware, can you said about the complexity of hardware? Yeah. Are export controls being implemented properly? Do you think that is a good idea? You know, everyone was going with Deep Seek. Wow, how did this happen? They must have stolen chips. How could this be? What do you think about that?

It turns out that they probably did use chips in Singapore. Um, I think the following: I I think managing software and managing hardware compliance are are extremely different things because their their vector of diffusion is different; there's different weight, right? If you sell a a server that weighs five or 600 pounds, arrives on a pallet, you can go visit it, right? You you want to deploy it in Kazakhstan; you can put a data center, and you can have somebody from the embassy visit it, take photos of it once a month; it's not going anywhere, right? You you can keep track of who uses it and provide logs, and that's much, much harder with software, and open source is whole another level, right? And so um that's the first observation; the second is that we had I got to know the the the the leadership in Commerce and in the previous administration. I didn't always agree with their policies, but it is a world of unintended consequences. You uh sought to limit Chinese access to EDA tools to delay the growth of a Chinese chip market, and so US venture capitalist backed tons of Chinese companies in Shenzhen to build EDA tools, right? I mean, right, right. That that that this is an unbelievably slippery dynamic, challenging problem, and I I don't know if it's a tractable problem. To delay another nation's progress uh on a technical trajectory is an enormously challenging thing, and uh I I certainly came to appreciate just how difficult it was for well-meaning people to predict the impact of policy during the last two years. For sure.

Do you think this Administration is better for AI than the prior Administration?

I don't think there's any doubt that's the case.

What makes you say that?

I think the past administration lined itself up against Big Tech, and um that that that was a mistake. AI is also in a different place, so it's easier to be for it, right? It's less scary now than it was in I don't know, 21, right? It it we we sort of have a better picture of the trajectory, both the risks and the uh the benefits. I think the this Administration sort of had the foresight to to put in place a an AI czar or leader uh to to be a focal point for discussions. Yeah, I think it's probably net a fair bit better.

You said it's very challenging to kind of hinder a nation's um development, adoption, progression of a technology. Respectfully, you chose to not sell to China. Yeah, why was that, and does that not go against the difficulty in hindering progression?

No, I I think uh I I have a a very simple rule, and I encourage your team to to to use it. I mean, you don't need a a big sort of handbook for to help you make good decisions in a company; just ask yourself, "Would my mother be proud, and would she be proud if I did this? Would she be proud if I explained exactly the situation, and would she look at me and say, 'I'm proud you're doing this, son'?" And I asked myself that, and I I came to believe that that the deal on the table was wouldn't be used for good, and I wasn't comfortable with that, and I wouldn't have been able to explain to my mother. And that's a moral compass, and I think that's a uh it wouldn't have been used for good.

I'm sorry, I'm just naive. They they use it to what power do—and you know some of do facial recognition to identify minorities uh for persecution, to build military equipment, to to things that I either couldn't see or what I saw didn't didn't feel right. And so we it's more important than money.

Do you think we fundamentally underestimate the Chinese capabilities?

100%. I I think, and it is one of the most obvious and frequent uh errors in judgment, is that you underestimate the other side. I think uh you have to to look carefully at what they're doing, and their investment in infrastructure has been extraordinary. Uh, the rate at which they generate engineering talent is exceptional. The government's ability to have a policy and implement it—you know, they're that's not a democracy; they weren't designed to have checks and balances there, right? Um, the uh funding that flowed into the development of of AI technology, that their venture capitalists were backed up by their government, that uh they have national champion companies, that they've developed a a belt-and-suspenders strategy to sort of make much of the third world dependent uh on them and their technologies. I I I think they absolutely should not be underestimated. They have a lot of people, and we see a tiny fraction of it. I I think they have produced industrial policy that has moved their nation forward.

What was the most significant, do you think?

I think they they the creation of economic zones like Shenzhen was clearly a visionary move. Um, they had uh they they knew that that their own system was in the way, and so they created zones that that relaxed their own system.

Could the US learn from them in that way?

We did some of the same things that in Trump's won Administration, right? What what did we do? We relaxed our our uh our own rules in the the development of vaccines. We uh we knew that in in this time it would be very difficult to go through the the steps that we always go through, and we tried to implement some thoughtful shortcomings or workarounds rather. You know, why are they committed to trains as a mode of transportation, and and we can't build a decent train system in the US or in California, or why we have three different standards for train rails, and the rest of the world can can build extraordinary high-speed trains linking important cities. Um, why what are we doing wrong in the building of of our infrastructure that that our bridges and our freeways are in disarray? Um, I think those are are questions we got to ask ourselves when we see other people doing it differently. I mean, you know, if if if you watch a good football team and you say, "Whoa, that's interesting offense," and you're not thinking to yourself, "How could our team learn? What what could we do? Why did that work? What was it about the people they had or the talent or the structure or something that made that a successful series of plays?" Um, and what what what can I take away from that? How can that inspire me to do to do better? I mean, I'm I'm always looking for for inspiration in in others and competitors and uh partners. I think you know we we have some of our partners at G42; I mean, the work ethic is unbelievable; it inspires me. Um, and the the scope of the challenge that undertaken inspires me, and and I I think I'm always looking for that.

Andrew, I could talk to you all day. I do want to do a quick fire with you, so Isa short statement, you ready?

Yeah, sure.

What do you believe that most around you disbelieve?

I think we're closer to peace in the Middle East than people believe because I I think uh there is a rise of a of a moderate, business-focused Arab state that it wasn't there 25 or 30 years ago, and I I think uh if you visit the UAE or Qatar or or even KSA uh what you see is amazing transformation, and I I think uh there is a a desire for uh to to be included in the West in their own way, but also to to to enjoy the benefits of it. Um, I I think it's uh yeah, I I think there's um we are closer than than maybe people think.

What's the most underrated threat to NVIDIA's market share dominance?

The fundamental architecture of the GPU uh with off-chip memory is not great for inference. Now they will continue to do well in inference, but it's it can be beaten, and I think they know it.

What's a crazy AI prediction you have that most people would call science fiction?

You Dario Anthropic's that will live to 150. I don't think we're going to live to to 150. I I don't think—what? 90% of our code will be written by machines in this year. Um, but I do think that within a year or two, AI's penetration will be approximately the same as telephones, cell phones.

What have you changed your mind on in the last 12 months?

Oh, I there are lots of things. Many decisions I made turned out to be wrong. I mean, I I think what what was the most wrong decision? There are two ways you can be wrong: you can actively be wrong or you can fight against what was right. Um, in uh 2016, JP, one of our co-founders and chief system architect, laid out a plan that would have us doing water cooling and for our systems, and nobody else was doing it, and I fought so hard, and I was so wrong. Um, JP was right. Uh, a year or two later, Google announced that the TPUs were going to be water-cooled. You know, we we were first, and uh now NVIDIA's only selling water-cooled parts. I mean, I I was dead wrong, and JP was right. I mean, I I think many, many instances when you make a lot of decisions every day where where you're wrong. I've been wrong about people. Um, people I thought were pretty good turned out to be extraordinary. People I thought would be extraordinary uh were really smart but couldn't finish projects and get stuff done. Um

I I think if you're not prepared to be wrong a fair bit, you ought not to be making a lot of decisions, because it comes with the territory. As a venture capitalist, I'm never wrong, so I, as a venture capitalist, you're wrong nine times in 10, and everybody forgets as long as you're really right. And I get a picture of you signing the term sheet with me, good. And yours is a perfect industry in which nobody cares about the average; on average, you're wrong all the time, and what they care about is the occasional time you're really right, and that that's what moves a fund. And I I think that's different than than than being a CEO. I think we got to be mostly right most of the time, but if you're making a lot of decisions, you're still making a ton of mistakes.

This is your fifth startup. Yeah, I mean, you are a sucker for punishment, aren't you? I mean, really, like five times? Like, Christ, Andrew, did you not get beaten alive enough? Uh, my question to you though is like, I believe in the value of serial entrepreneurship. I've spoken to many who don't, respectfully. How do you think about the inherent benefits that you have having done it four times before?

Okay, I think if you are in a business in which running a business is a benefit, then experience matters a great deal. Uh, I think if you are in a business in which you look like your customer, there was a reason why social social networks were started by people right out of college or in college, because dating is top of their mind, right? I mean, that that that's, and they look like their customers, and that was more important than knowing anything about running a business. And so in that environment, it will certainly select for people who are of the demographic that that their customers are; they know that backwards and forwards. But if you want to have a business that has manufacturing in it, that has a supply chain, that uh has you managing hundreds or thousands of engineers to a timeline to a schedule, I I I don't think anybody would turn around your statement and with a straight face say, "You know what, I'm looking for is an engineering leader with no experience," right? No, and I don't want somebody who's led a team of four or 500 who has experienced the challenges of growth; what I'm looking for is somebody with no experience. Naivity is a bonus here.

That's right. I think the people who sell that sometimes are consultants, right? Oh, look, my guys have no experience in your industry; they're not biased, right? Maybe a little bit of experience in the industry would help, uh, right? Come on, where are people investing today in AI across the stat? You can choose any pot where you're like, "What, why is so much cash going to that part?" I'm not saying that company. I think part the the dynamic in your industry is is sometimes money needs needs to find a home, right? Some guys have raised really really big funds, and they got to find a home for that money, um, and some people don't like to be left out; they're willing to to make investments for maybe for some status purposes or other reasons that that don't seem to make sense. I don't know; I haven't thought about it. I mean, I I think there's some underappreciated places of investment. I'd say in the chip world, the the sub-mowatt, really tiny tiny little chips that live next to sensors that do uh inference, these are tiny little things that will uh only send back useful data, is an extremely interesting market, and they will sell enormous volume. Now, I it's not a part of the market I I love to play in; I like to build bigger things and sell them to the data center, but I think that part is extremely interesting. I think they'll be fundamental for robotics. I think um uh that's an area where uh I I think it's it's extremely underappreciated.

If we think about Cerebrus in 10 years' time, where do you envision the business in 10 years' time if everything goes well? Where are we in business having that conversation? So 10 years ago, Nvidia was worth 10 billion, so uh that's a that's a long run in in our world right now. I think in 3 to 5 years, I would like our technology to have been used to to solve two important societal problems. I would like it to be used to to have found a therapeutic for uh an an affliction that impacts more than a million people a year. I I would like uh I would like our inference to be powering a collection of of apps that don't exist today, and I would like uh that when uh that a meaningful portion of the population in the US and and in Europe inadvertently uses our technology, so use is something that we power and that they don't even know it.

Andre, I've wanted to make this show happen for a long time, as I said. I heard so many good things from Mar for for many years. Um, there's been so many requests to have you on the show; my team is just like, "Just get Andrew on the show, Harry." I'm like, "Okay, okay." I like tweeted it, obviously, which is how we got this. Thanks for joining. You tweeted it, and like 40 people sent me, not saying, "How come you're avoiding Harry? How come he has to go tweet it?" I was just like, "All right, just call me; it's good. Send me a note; happy to come on." Really thoughtful questions, Harry, really thoughtful and interesting um really conversation.