📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Inference Revolution: Groq, Nvidia and the Future of AI

Sohn Conference Foundation15:59

Transcription

Uh, thank you to the Irish Open conference for having me back again this year. This is really a a wonderful conference for a great cause and I'm honored to be invited back with Jonathan.

Uh, Jonathan, you guys might have seen has been in the news a little bit lately. Uh, he's the he's the chief software architect of Nvidia, founder and CEO of Groq, and inventor of Google's TPU. He's had three, uh, no, uh, four first-time right silicones. Um, he's one of my dearest friends and someone I'd consider to be one of the greatest engineers of our time.

Well, thank you. And and for those of you who don't know John, although probably most of you do, uh, uh, John's one of those rare, uh, uh, hedge fund managers who both gets great returns and provides immense entertainment value. Watching you hold those shorts like you've got diamond hands. I don't know how you do it. Everyone else is like sweating for you. Not everyone's supposed to know I'm like a crazy short seller, so >> [laughter] >> Oops. Yeah. >> [laughter] >> Crazy. Can't believe you just docs me like that. >> [laughter] >> Um, so, you know, Jonathan called everything that's happening today like over 10 years ago. Everyone thought he was crazy. I was one of the lucky few that believed him back then. And, you know, up until a few years ago nobody even knew what inference was and so, I think that's a good place to start given it's now one of the most probably the most important like thing about AI today.

Well, for once I don't have to explain what inference is, that's nice. >> [laughter] >> Um, yeah, and and so, um, you want me to get into the economics of what do you want >> Yeah, like, you know, one thing I think is very important is that, you know, portability across architectures is diminishing and porting is no longer really an engineering inconvenience but is involved into more of like an economic handicap. So, maybe we can talk through some of the tokenomics of you know, inference and Yeah, I I A good parallel here is we landed people on the moon about 60 years ago, roughly, and um that was 60-year-old technology that got us there. What we're doing in AI is brand new technology. It's cutting edge even for today. Uh we finally have enough compute to do these things with models. We've had the algorithm since the 1970s. Uh so, what I'm seeing a lot of is people will look at one small part of the stack and think that that is the most crucial part, that's the bottleneck. And I don't think people understand just how much you have to build to make AI work. It's the chips, the packaging, the systems, the networking, the data centers, the servers, the racks, the you know, power, all the stuff. But then also on top of that, it's both training and inference, which are two completely different problems. A little bit like, you know, getting to the moon and then landing on the moon were two very separate problems. And um you know, in terms of like bottlenecks and and all this, the the the focus I see a lot in in, you know, investors are asking questions is what is the bottleneck? What's the next thing that's going to be constrained? What should I I do next? And that's not how this industry works. Every time a bottleneck gets big enough, people solve it. So, as a person running a business, you're always going, "Hey, what is my biggest problem and can I solve it?" And when you look at some of the the components that are limited, when they're a problem but not a huge problem, people can charge a lot of money for it. But as they start to become a bigger problem, people start solving that problem. And so it it's this constant shifting um dynamic of like where in the supply chain is the biggest problem, but don't become too big of a problem because then it'll get solved.

>> Yeah. So like memory is like the big theme today, right? Uh it's the biggest bottleneck in AI. But so what you're saying is um if memory continues to become a larger and larger bottleneck and continues to be extremely supply constrained that you think it's going to be solved?

Yeah, exactly. So so memory used to be a commodity. Correct. Yeah, it was the most commoditized uh segment of the semiconductor supply chain, pretty >> Yeah, maybe maybe a good way to explain this. Um there there are two ti- kind of goods that you can charge a lot for. One is a um Veblen good and the other is a Giffen good. Veblen goods, the more you charge, the more desirable it becomes, like luxury goods. Giffen goods are a little bit different. A Giffen good is An example is uh rice. So if you start to charge more for rice, some people can no longer afford steak, so they actually spend more money on the rice. So by raising the price, not for a luxury good, but by raising the price for a staple, you can actually increase the value. And that happens to some extent until someone goes, "Let's stop eating rice and let's eat corn or something else, right?" And so there's this inverse um uh sort of headwind where if memory is too expensive and if people don't build enough of it, people are going to solve that problem technologically.

So you think algorithmic efficiencies like you know, for example, like Deep Seek's latest model release, the V4 release, um it compressed KB cache by 90%. But But the argument around that has been Jevons paradox, Jevons paradox, Jevons paradox. What >> those engineers, them working on that problem is an opportunity cost. They could have been working on something else. If they had enough memory, they would have worked on something else. If it wasn't so expensive, they would have worked on something else. You make it the big enough problem It's the tall poppy. As soon as it gets too tall, it gets chopped down. Got it. Yeah. Cool. Um so >> [laughter] >> Um So let's >> Start start building more memory fabs. That's the solution. Yeah.

So what what do you think um We talk We talk about diminishing returns on intelligence a lot, right? Um you and I I mean, the way we prepared for this presentation is we literally just went through our conversations and thought, "Oh, that was a good topic. That was a good topic." And then we picked >> long-standing disagreement between John and I. >> Yeah. I think intelligence has diminishing returns and at a certain point the models get so smart above PhD level and human beings, you know, can't really understand the differences between one or the other. And then, you know, couple that with the closed-source versus open-source uh frontier landscape where open-source models are roughly 6 months behind the closed-weight model labs. So, you know, you can kind of make the assumption in the you know, if you buy my argument that there's diminishing returns over time that um the open-weight landscape will catch up to the closed-weight landscape. Now, there's a lot of nuances around that. So, I want to hear Yeah. I know your thoughts, but like share your thoughts with us.

Well, I might surprise you a little. So, there are a lot of uh areas in the economy where if you produce more of something, it becomes less desirable. And then there's others where it becomes more desirable. And I would argue that intelligence there's no way to satiate the appetite for intelligence. Uh the more intelligence you get, the more intelligence you want. Now, I'll break it down for like Let me break it down to first of all, is intelligence going to plateau? Because that's important for decisions. And then second of all, um uh would not not only would intelligence plateau, if it didn't plateau, would we get to enough of it where it'd be like we don't need any more? So, in terms of having enough intelligence, the first thing is as long as cancer isn't cured, as long as people still die of old age, as long as uh we don't have enough compute to run some of these AI models, we don't have enough intelligence. So, there's an economic incentive to keep building smarter and smarter machines to help us solve bigger and bigger problems. The second is competition. Everyone in this room probably could retire. You probably don't need to make more money, but you keep investing and you keep competing with each other. Why? It's competition. We just do it. We're humans. And competition isn't going away. And so, if my AI is less intelligent than your AI, and I can't tell the difference directly, I'm still going to be able to tell the difference in my returns based on which one I'm using. I'm going to want the better AI. And so, whenever I I um program, I actually use >> it's a gift to the Earth that Jonathan's programming again and writing code again. Just So, everyone should say thank you to him. I But I I'm actually writing code which is being used for for real things now again, thanks to AI, right? And I'll use multiple models. Each one is better at different things. Um but, you know, even though one model might be better than the other, I'll I'll still have one use that model. And I think AI will know that this other AI is smarter, even if we can't tell the difference and will like This is the whole like agentic thing. So Why or what is agentic? Agentic is Do you get better productivity by using AI? Yeah. Well, so does AI. AI likes to use AI. So it calls to AI to do some task for it and return the result just like you do. It's just more of that. And the AI is going to recognize smarter AI and use that smarter AI. They There was a actually a conversation we had outside [snorts] with one of the hosts. I don't know if you remember um about resumes. Oh, yeah. Yeah. Yeah. So >> funny. >> Yeah, I So I didn't know this. This is something I just learned. >> someone did a study and showed that resumes generated from one LLM are preferred by that same LLM over the resumes from the other. Recruiters are now using LLMs to determine like who to interview. But you got to figure out which LLM the recruiter is using. So So you should you should build an one one resume with Claude code or Claude Opus 47 and one with chat GPT and you'll have the highest probability of being selected basically.

So then the other question is is intelligence going to saturate? Or is are we just going to need more and more intelligence and this build out is going to make total sense? My argument for why it's not going to saturate is as follows, which is um There's two two components to intelligence and there's a great easily digestible book called Thinking Fast Thinking Slow that many of you have read by Daniel Kahneman that explains exactly what AI does. Thinking fast is the intuitive part. It's the You're You're given a problem, do you have an answer immediately? Thinking slow is you iterate on it. So if you think about chess, speed chess is thinking fast, regular chess is thinking slow. You're evaluating multiple opportunities and recognizing better moves when you when you string moves together, right? And even though AI is coming from computers and and so this is sort of a little bit hard for us to see. AI is very intuitive. It's actually better at being intuitive than we are. And that's because it's been trained on so much data. When when Waymo is sending all their cars out the amount of data that they get in a day is about I don't know if it's now at the amount of experience that a human being gets in a lifetime of driving, but it's starting to approach that at least. When you're getting a lifetime of driving data in a day, you're seeing every possibility. You're You don't need to figure out how to deal with the fact that some I-beam is falling off the back of a freight truck and about to hit you. You've seen it happen two to three times and you know exactly what to do. The thing is the the more these models produce data, the more they're able to just intuitively deal with the situation cuz they've already seen it, right? And that's intuition when you just have the answer. So, when you train these models, what you do now, they used to be trained on just data pulled from the world and and that human beings were producing. And now what we do is we use the models to generate the data that they get trained on. So, you have a model at this level of capability and it produces data at these levels of capability and it's gotten good enough that it can tell what good is. It keeps this data, trains, moves up to here. Then it produces data of this quality, prunes it to here, trains, and goes up to here and just keeps moving up. And so, at this point, we're seeing these models improve at a pretty linear rate. And so, there's no reason to believe that they're not going to get any smarter. We may not recognize the difference between two really smart models, but one will be much smarter than the other. And that matters in uh the context of competition. Competition and solving big unsolved problems. Yeah. Right.

Do you want to talk about sentience? Oh, jeez. >> [laughter] >> Okay. I have a hobby. My hobby is to take words that people have used for centuries that don't have a good concrete meaning meaning and try and ascribe a definition to it that helps me understand the world. And so uh I'm going to I'm going to give you my definition of sentience. So, first of all, intelligence is your ability to make a prediction or or influence an outcome to what you want to have happen. But it's stationary. It's like you have a an amount of intelligence that is you know, fixed in these models. But sentience is your rate of improvement in your intelligence. That makes sense because when we talk about sentience, we're talking about the ability to sort of self-reflect and get better and that's an important part. So, I I say that intelligence is, you know, your your your capability and sentience is your rate of change. But rate of change doesn't have to be binary. You don't You're not sentient or not sentient. It's how sentient are you? Are you linearly sentient? Are you asymptotically sentient? Right? You look at um you know, the world's best Go player and he asymptoted. He He stopped getting better because he didn't have better players to play against. So, people often conflate LLMs with just being intelligent, but there's something else that language gives you beyond intelligence. It gives [clears throat] you the ability to transfer information. Going back to the example I gave with Waymo, [clears throat] if if you have an entire civilization producing information, each participant in that civilization gets to benefit from that distillation of knowledge. And so while intelligence is a property of an organism or an individual, sentience is a property of a civilization. AI is producing more intelligence. It's getting smarter. You interact with it, you get smarter. You ask better questions. You're making the AI smarter. And so there's this feedback loop of sentience that's accelerating in our society. And AI is contributing to that. And so as that happens, I I would I would expect our kids to get much smarter than we ever were, just like we're probably smarter than our parents were because we had the internet.