📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Daniel Guetta on the Guts of AI, Agentic AI & Why LLMs Hallucinate | The Real Eisman Playbook Ep 46

Steve Eisman1:02:40

Transcription

Hey, it's Steve Eisman, and we're going to talk about AI today with our guest. This is a topic that we have explored for many, many months. It is crucial to the future of the US economy. Recently, hyperscalers have announced that they are increasing their budgets enormously. The top four hyperscalers are going to literally spend $650 billion alone on tech stuff, all related to AI. The future of the US economy is really at stake.

And a few weeks ago, we had on Gary Marcus, who's a big critic of LLMs and argues that they are losing their efficaciousness. And so I wanted a second opinion because this is just too important a topic. And so today, we have Columbia professor Danny Guetta, who, as you'll see, agrees with Gary on certain points but disagrees on others, and we're going to explore all of that. And afterwards, I'll be back to talk about what we've learned.

Hi, this is Steve Eisman, and welcome to another episode of the Real Eisman Playbook. And today, we're going to explore the world of AI, but not per se from a business perspective, but from the within the guts of AI. You know, is AI a bubble? Is it going to change the world? You know, these are questions that we've explored over in our podcast over the last many, many months. And to help us go further is our guest, Professor Daniel Gua, who was a professor at the Columbia Business School. Daniel, welcome.

>> Thank you so much for having me, Steve. This is going to be great. Really excited for the conversation. Really excited. So before we get started, why don't you just give us a little bit about your background to explain why you're here?

>> Yeah. No, absolutely. I mean, my whole sort of career has been working on kind of AI even before it was cool. Uh, so I started off doing my PhD in data-driven supply chain management a while back, and I was really lucky. I got to spend some time at Amazon doing that, which obviously is a great place to study uh supply chains. After that PhD, I went to work at Palantir Technologies. Um, yeah, I was not a US citizen at the time, so they didn't let me close to the government stuff, but I did get to work a lot on their commercial business, and that just involved going around the world working with companies in just a whole range of industries and helping them figure out how to do what they do better, but using data, AI, analytics, that kind of stuff. And then around eight years ago, uh, the dean of the business school at Columbia, uh, Cassis McLaras, he kind of already back then had the foresight to realize that this kind of data analytics wave was, you know, not going anywhere, that it was here to stay. And so he asked me if I would come back to Columbia to kind of help build up the data and analytics curriculum. And that's kind of what I've been doing there. And, you know, I get to work with companies, help them figure out how to get value out of AI. I get to teach classes on AI, operations, coding, all that kind of stuff. And I'm very lucky to get to teach our MBA students, engineering students, executives, but then also um we even launched this open program with Wall Street Prep. So now anyone can take that program. The reason I mention it is because we were introduced by a student of mine from that program. So hat tip to Ari. Thank you for the introduction.

>> Yeah. So let's let's start with um the big topic, which is large language models. Before we talk about how efficacious they are, you know, when when you watch business news, they all they're really talking about is how many chips are being bought? Who's who's ahead? Is it Google? Is it OpenAI? But let's go to let's get to basics. What exactly is a large language model? And how does it work?

>> It's a great question, and I think if I may, I think it makes sense to even take one step back and to say, what even is AI? Like, I think large language models have almost become synonymous with AI in people's minds, but they're not. I mean, that is not the only thing AI is. And I think it's actually helpful to think of two different kinds of AI. You have kind of predictive AI that you also hear called machine learning or predictive analytics. And that's been around for a while and very successful for a while. And then you have Gen AI, so including those large language models that have been uh kind of more recent. And it's actually really helpful to dig in to what each of those two mean because, as we'll see, it's really like understanding what those two things are really makes a difference to understanding how large language models work and how they sort of fit into the broader landscape.

So if we start with this predictive AI, been around for a while since the 80s or '90s, and and the idea really is to just say, can we look at a data set? Can we identify patterns in the data set? And can we kind of do something useful with those patterns? And so my favorite example is uh the Zestimate. You know, when you go on Zillow and you look at a property, it gives you the the estimated price of what that property would be. I always joke with my students, right? There's like all kinds of lofty reasons you might use it, like real estate development, to buy a house, sell a house. Really, people are using it to find out how rich their neighbors are. So, that that number, right, that you just get when you go, >> How expensive is the house that's next door to me? >> How expensive is the house that's next door to me? That's exactly right. So, you know, that is a machine learning model. That is a predictive AI model. It uses stuff about the property to predict what its price is going to be. And if you were thinking, you know, how does it work? Obviously, it's quite complex. But if you were to build a slightly simple version, maybe just with houses on Upworthy, what could you do? Maybe you'd look at the square footage of the property plus, you know, the number of bedrooms, number of bathrooms, and then maybe you'd say every square foot is worth $2,000, every bathroom is worth $10,000, whatever. Add them all up. >> It'd be a fairly simple model. >> Fairly simple model. Now, what makes it a machine learning model rather than just like a sum you come up with yourself is the fact that the way Zillow creates it is they train it using historical data. And what do I mean by that? They're going to take a huge trove of data, right, from uh uh previous properties that have sold. They're going to look at data about those properties, and they're going to start with totally random numbers for the value of each element. So, they might begin by saying a square foot is worth $1. Totally incorrect. That's going to give really bad predictions. But then they tweak that number, okay, until it fits the data as closely as possible, right? And that's absolutely fundamental, as we'll see later, to even how large language models are trained. And just in terminology, this is called training. The kind of tweaking of those numbers, and you often hear something called parameters or weights. That is what those numbers are. They're kind of parameters or weights of those of those models. Uh, and so you're tweaking them as you sort of create uh create the model.

>> Okay. So that's been around for a long time.

>> That's been around for a long time. Absolutely. And just it's worth just saying it's been around for a long time, but I think today, if you tell me like, if you look at the value generated by AI, what percentage of it comes from that? It's huge, and it's still massive. Every time you swipe a credit card, fraud analytics, uh, figuring out if someone should be able to borrow money, sort of uh uh, you know, uh creditworthiness, um, when you order a Lyft, right, how much it should cost and what the length of the trip is going to be, all that is really sort of based on that AI. So, you know, how's that different from the new kind of AI, the Gen AI, the LLMs? The problem with that old kind of AI is that it only really works with numbers, with numerical data, stuff you can put in Excel. So, thinking back to this Zestimate, you can use the square footage, the number of bedrooms, number of bathrooms, but what if you want to use the image, right, the picture of the house, or you want to use the text that you see inside the description?

>> You can't really put that in a model like that because you can't multiply text. It's not a number, right? You can't do anything with it. And uh uh, you know, uh so that kind of was a limitation of those uh models, and really that's where the deep learning, the Gen AI revolution sort of came along and overcame that limitation, right? So this was maybe in the mid-2010s. I mean, it's started at various times, but let's say sort of the mid-2010s, this new field called deep learning came along that introduced these models called deep neural networks, and large language models are an example of those deep neural networks. And what they can do is just like as a human, you're, if we're reading a piece of text, we kind of understand what it says. We get kind of the concept behind the piece of text. Those models can also do that, right? And the way they do that is by using an enormous number of parameters, an enormous amount of data. And maybe we'll get into how that works, right? Sort of mysteriously, but they're effectively using just enormous amounts of data to understand what's going on behind the text. But I think it's important to point out, and you discussed this with Gary Marcus a few weeks ago, and we'll touch on it today again, maybe understanding is a misnomer, right? Because all they're doing is they're training themselves on historical data and trying to mimic those patterns. And so is that understanding? Is that not understanding? And maybe we'll we'll get into that, but that's how these models work. And sort of they overcame that limitation of machine learning because they're now able to use all this unstructured data.

>> So what can a large language model do that the old probabilistic AI can't do? Give us an example.

>> Yeah. So first of all, I don't know that it necessarily makes sense to say large language model versus probabilistic AI. These large language models are a kind of probabilistic AI, right? I mean, they they sort of, you know, are a much bigger version. They have many more parameters. They have many more sort of uh so that's the first thing to point out. But generally, it is that idea of looking at completely unstructured data. I mean, that is sort of, you know, their magic. That is what they're able to do. A historical sort of one of those historical probabilistic AI, machine learning, predictive AI models couldn't look at a piece of text and understand what it says. Now, people had ways to get around that. They would say like, oh, if I want to understand a piece of text, let me look at the number of happy words in the piece of text. Like, you know, if the text says wonderful and amazing and fantabulous and whatever, then that means it's probably positive. But those were very sort of uh simplistic ways of looking at text. These large language models can actually look at text and kind of almost understand it.

>> So why do large language models hallucinate? And let's talk about exactly what's a what is a hallucination?

>> That's a great question. If you'll allow me, this is going to be a very long answer, but I think to kind of get to kind of I I want because I've, you know, I gave an example about a recent hallucination which I just discussed with Gary Morris, which was on the day that uh Maduro was taken out of Venezuela, within like the first hour, um people went onto chat GPT and said, "What's going on with Maduro getting taken out of Venezuela?" And chat GPT responds, "Maduro is still in Venezuela." Right? Because because apparently large language models have a problem with novelty. But let's talk about hallucination.

>> Let's just tell you how they work. I mean, how those models work, how they end up hallucinating. Long story short, I'll tell you the answer already. What you should be surprised by is that they ever don't hallucinate, right? The fact they do hallucinate is kind of, you know, that's like not the surprising part.

>> You got the right answer.

>> The fact they ever get it right, right, is is crazy. So, okay, how do these models work? I'm going to explain it at various levels. Let's start at the top level. This is something your listeners may have heard already, but it's worth repeating. These are just autocomplete engines, right? So, I don't know if you've ever played a parlor game where you say a word, the next person says a word, the next person says a word, and you kind of build up a story together. Yes, that is basically what they're doing. So, for example, if you ask a model, you know, "What is the capital of Argentina?" It's going to look at the question, it's going to say, "Huh, I wonder what the next word would be. Maybe the capital of Argentina." And it would say, "Oh, Buenos" might be the next word. And then, and this is key, it's going to look back at the full conversation. "What is the capital of Argentina, Buenos?" And then it's going to say, "If that was a useful conversation between a chatbot and a human, what would the next word after that be?" And it would say, "Aires." And it would continue until it decides, "I'm done." Now, by the way, just a side note, unrelated to hallucinations, but this is why part of the reason they're so expensive is every time you want the next word, you have to reprocess the entire conversation from the beginning, right? Which is kind of wild. So if you've already spoken to it for like, you know, you know, 10,000 words to generate the 10,000 and 1 word, it has to reprocess all those 10,000 words through the model. So that's kind of crazy.

>> That eats up a lot of energy.

>> Eats up a lot of energy, which is why, you know, I mean, we could talk about energy. Certainly get certainly get energy.

>> But okay, so now I understand why it eats up so much energy.

>> That's why it eats up so much energy. Now, how does it do that? How does it know what the next word is? And the key concept you need to get, and I promise you this is as technical as this gets, but I think it's just it really helps to kind of know this is a concept called an embedding. And what an embedding is is you take a word and you basically turn it to numbers, right? You remember I said that the limitation of those machine learning models was they couldn't look at words? Well, those embeddings convert the words to numbers so the computers can understand them. Okay? And maybe the easiest way to kind of get what an embedding is, imagine you take uh every word in the English language and you give it two scores.

>> The first score is how alive it is. So maybe a human gets a 10, a rock gets a zero, uh uh I don't know, a spider gets a five, a carrot gets a two. You get the idea.

>> Yes.

>> And then the second score is maybe how loud it is. So maybe a loud. Yeah. Maybe a baby gets a 10. Maybe, you know, rock gets a zero. Maybe a carrot gets a one because it's like quiet, but if you crunch into it, it makes a sound.

>> So, every word has two scores.

>> Every word has two scores. Now, you could imagine taking those scores and putting them on this kind of XY axis, right? On a piece of graph paper,

>> Every word.

>> Every word based on those scores. So, the X axis is how alive it is, and the Y axis is how loud it is.

>> And if you think of that piece of paper, you're going to start

>> Who gives the score?

>> Love the question. I promise you we're getting there in like 30 seconds. This is like

>> There's is there a referee?

>> Oh, so it's the key key question, right? And that's the first question.

>> Who's making that decision?

>> Whenever I tell students, they always ask me that question. And the answer is is crazy. But we we'll get there. So, you know, you have that piece of paper, you got the scores, and you can imagine you're now going to get these kinds of clusters, right? Like the the vegetables are going to be, you know, bunched up in one side, the humans are on the other side, etc. And the profound thing about that is that now those scores kind of give you they allow the computer to look at numbers, and those numbers kind of tell you, ah, I know something about the word, right? If I tell you a word is like zero on the live scale and nine on the loud scale,

>> Guess what? It's going to be close to the cluster that contains like foghorn and bell and all these sort of loud things. So it really gives you a vibe. Now, to your question, right, who the hell decides these words? And of course, right, two scores is not enough. In reality, if you want to really know a word, you got to go for thousands. And the answer of course is uh no one. Uh, the answer is machine learning. So back, remember this Zestimate case? You remember I told you in this Zestimate case, you don't decide how much a square foot is worth or how much a bathroom is worth, you just give the model sort of the the the the the data, and it figures it out. It tweaks the numbers until you get to a good sort of result. And it turns out that's what happens with these embeddings. And this is truly a wild idea, but have you heard this fact? People say that these large language models are trained using all the text on the internet.

>> Yes.

>> You've heard that said. That is the training data they use. So, let me explain how it works. They'll take something like two words, let's say like king or queen, two words. Those two words, if you look on all the internet, often they're going to be close to each other when you see king. And so the algorithm is going to realize, wait a second, those two words or my X and Y on my piece of graph paper, they should be close to each other because they often occur together. And so they're going to tweak the scores by a few decimal points to get to get them closer to each other.

>> On the other hand, a word like, I don't know, Pringle and existentialism, right? Probably far apart. And so they're going to be tweaked to get further apart.

>> Okay.

>> And Steve, it's actually astonishing. But if you do that,

>> So the scores are constantly changing.

>> Constantly changing while the model is being trained. So while OpenAI is training that model, and amazingly, if you do that enough, you end up with something that actually captures the meaning of words. So you actually get clusters of words that mean the same thing. You get these crazy relationships where like, if you look at the difference numerically on the graph paper between foccacia and baguette, it's kind of the same difference as between France and Italy. So it kind of understands the concept of like country, right? And by the way, I mentioned this before our podcast, we'll put a link in the show notes, but I created a little online tool that your listeners, if they're interested, they can go actually check this out and try with a bunch of words. Um, and so at a high level, that is sort of how a large language model understands text. There's another complexity, which is that of course, one word is not enough. You want words in context, right? So if I say "I like figs and dates," "I have a date tonight," and "What is today's date?" The word "date" means very different things. And so there and just again, to maybe link this to something some words your listeners may have heard, I don't know if you've heard of something called the transformer? So it's this uh uh uh uh mathematical technique that Google published in 2017 that was kind of a big uh breakthrough. All the transformer is is Google realized how to make these embeddings pay attention to each other.

>> Okay.

>> Okay. So that was my big explanation. So how does an LM work and how does that link to hallucinations? I know you are.

>> So I promise we're getting back to it.

>> I'm waiting with bated breath.

>> On the edge of your seat. I the answer is going to be a little disappointing. I'm warning you. But

>> Okay.

>> Basically, when you ask a question of an LLM,

>> Yeah.

>> It first takes your question, it converts.

>> I got to ask you, why was Sandy Koufax such a great baseball player?

>> What was Sandy Koufax such a great baseball?

>> Perfect question to ask.

>> You're picking someone who was born in France and grew up in England. So I may or may not have heard of Sandy Koufax, but we're not going to tell the listeners that.

>> Great. He was the greatest baseball pitcher in the early '60s.

>> Great baseball pitcher. Great.

>> Okay.

>> So what the model is going to do is it's going to take all the machinery I just described. It's going to get this big list of numbers that captures the essence of your question. So if you think again of our graph paper, it's going to put it in a position that kind of implies, oh, this is the vibe of the question that Steve just asked. And then it's literally going to say, what are the words around? What are the words?

>> What are the words around Sandy Koufax's greatest pitcher?

>> Exactly. What are the sort of? And then it's going to pick one of those words and then it's going to respond with that one.

>> And then it'll add another one.

>> But it's going to go back to the beginning and search again and then add another one.

>> Create those numbers again. Look at a word nearby and another one, another one, another one. So to your answer, why does it hallucinate? The real crazy question is why does it ever not hallucinate, right? All it's doing is it's getting that essence of your question based on all the text it's seen on the internet and everything it knows about where text is on the internet, and then it just generates the next word, next word, and the next word.

>> This is not what we would think of as thinking.

>> Well, that's a absolutely amazing question. Um, there's actually a uh trick that I love doing uh and I think this really brings it home for people whether this is thinking or not because, you know, I think

>> Take a step back. A big question people are asking is, is this going to get more intelligent than humans? Right? Is this going to eventually scale? How much can it scale? Is that thing?

>> So, but what you're saying is it's it's almost like a miracle that you ever get the right answer.

>> Exactly. And I'll give you an example that really shows what how much of a miracle it is. So if you ask a question to this model and you say, "I have a bag with five balls and the balls are labeled A, B, C, D, E."

>> Okay.

>> "Pick one at random and tell me the letter on the ball." So, as a human, if I ask you this,

>> Yeah.

>> You could tell me, I could ask a 10-year-old and they would tell me, "Oh, well, you know, 20% of the time I'm going to get A. 20% of the time I'm going to get B. 20% of the time I'm going to get C."

>> Right?

>> It turns out what you can do with these large language models is, and you can't do this on chat.openai.com, but if you can code using something called API, you can actually do this. You can ask chat GPT, when I ask you that question in your internal LLM brain, what are the probabilities you're thinking of? Like, what do you think the next word should be? And it goes completely off the rails. So chat GPT will tell you 50% of the time the answer should be C, and 20% of the time it should be A, and 5% of the time it's basically wrong. Now, if you think about why that is, it makes sense because all chat GPT is doing, it's taking your big question, it's embedding it. So, it's creating that essence, those numbers, right, that that get the vibe of the question. And then it's saying, "What is the closest sort of word that maybe fits that answer?" But that is not how a human works. That's just not how we think. It's just a very, very different mode of thinking. Now, I need to be clear.

>> It's not obvious to me that that mode of thinking is in any way inferior to a human mode of thinking. I mean, who knows? But it's definitely different. It is not the same thing. Okay. Yeah.

Hi, Steve Eisman here. When my kids were little, I woke up one morning in a panic. I realized that if something happened to me, my family would be completely unprotected. I immediately went out and got life insurance so that my family would be financially secure. And I have never regretted that decision. So stop procrastinating. Getting life insurance is much easier than you think. Fabric by Gerber Life makes it easy to start the year knowing your family has more financial protection this year. Fabric by Gerber Life is term life insurance you can get done today. Made for busy parents like you. All online on your schedule, right from your couch. You can be covered in under 10 minutes with no health exam required. If you've got kids, and especially if you're young and healthy, the time to lock in low rates is now. Fabric has flexible, high-quality policies that fit your family and your budget. Like a million dollars in coverage for less than a dollar a day. Fabric has partnered with Gerber Life, trusted by millions of families like yours for over 50 years. There's no risk. There's a 30-day money-back guarantee, and you can cancel at any time. Join the thousands of parents who trust Fabric to help protect their family. Apply today in just minutes at meetfabric.com/isman. That's meetfabric.com/isman. meetfabric.com/isman. Policies issued by Western Southern Life Assurance Company, not available in certain states. Prices subject to underwriting and health questions.

Hi, Steve Eisman here. I run a small business, so I know that running a small business is tough. And when it's time to get a loan, it can feel impossible to find a lender you actually trust. Big banks say no, and the internet is full of sketchy offers with sky-high rates and fine print you can barely read. Whether you need help covering payroll, managing cash flow, or investing in growth, you deserve better. That's why I recommend the Small Business Marketplace, powered by NerdWallet. It's a free, easy-to-use platform that lets you compare real financing offers from trusted lenders all in one place. What I like is that you don't need perfect credit to get started. No spam, no bait and switch, just personalized options that fit your business needs. In the future, when my business needs a small business loan, NerdWallet is where I'm going to go. And here's the best part. For a limited time, when you visit nerdwallet.com/isman and fill out the no-obligation form, you'll get VIP treatment and talk with a real person who knows all the ins and outs of small business lending. Don't risk your business on unreliable lenders. Go to nerdwallet.com/isman to find the funding you deserve. Fender, Inc. NMLS ID number 1240038.

>> So let's just explore for a second um Gary Marcus for a little bit.

>> Yeah, totally.

>> Because you you watched the podcast.

>> I did listen to your fascinating.

>> It was wonderful. I learned a ton. But if I could maybe sum up his argument, I think it would be two parts. He would argue that the improvements that we are getting as large language models, the model that's out there, is we're going to keep scaling large language models and they're going to continue to get much, much, much, much better until we reach whatever we're supposed to reach. And his his argument is that we're already at a point where large language models, the improve, the marginal improvement from chat GPT 5 to 4 is less than 4 to 3, which was less than 3 to 2. And so you're getting diminishing returns. And so that if we're really going to take this thing to where it needs to be, we have to go back almost to the to the drawing board and and put in what he called world models to to train these things on much better. That that's one part of his argument. And his other part of his argument is that given the way LLMs work, they're always going to hallucinate, which is a problem because hallucinations are a problem because if you use it, how do you know it's hallucinating?

>> Right?

>> So, what do you think about all that?

>> I got so many thoughts. Uh, so let's first, I think uh, you know, just definitionally for your listeners, just so that they sort of get where we are when we say scaling a model,

>> Right? When I described those embeddings and those transformers and those parameters that you look at data and you tweak the parameters, scaling just means more data. You just create more, not so much, well, more data as well, that's one way of scaling it, that's one dimension. But the other dimension is actually more parameters, more complexity in the calculations, right? And that's why by the way, it takes so many more chips and so much more energy and so on.

>> So in other words, a a word doesn't have two variables, two scores, it might have a thousand variables?

>> That is exactly one way you could scale. And another way you could scale is those transformers, which we didn't spend too long talking about, but they also have their own set of parameters, and there's many more than in the variables and the words, and you can add to those. We don't need to get into the details, but basically adding parameters is a big thing, okay? So that's just the first thing I want to put out there, that's what we mean by scaling, and as you said, the sort of at least the guys for a while was that the more we scale, the better these models are going to get. Now, I do want to say one thing, and Gary almost touched on that in your conversation, and I want to, you know, and I think he probably would have talked about it further.

>> I want to point out what we mean by better, because we never ask like, what exactly does that mean, better? What does it mean, better? Uh, better typically, when a new model comes out, better means that it performs better on a very, very specific set of benchmarks. The computer science community has a whole bunch of these benchmarks that these models get tested on, and you know, there's a bunch of examples. One of them is called the Massive Multitask Language Understanding benchmark, MMLU. It's basically, you can think of it as like an SAT exam. You give it, you give the large language model this exam, can it answer the question? So I just want to put that out that, you know, doing well at an SAT exam is not necessarily, it is part of intelligence, but it's not the only thing about intelligence. And so there's actually a lot of research going on now on what it means to actually test those models better, like what other benchmarks could we use that maybe are uh sort of, you know, more revealing than the current ones we're using. I actually have a colleague at the business school, Hong Namkung, and I'm very fortunate to get to work with him a bit. He's trying to build benchmarks based on like work in Excel, right? Can we give it an Excel spreadsheet and can it actually build the DCF or can it figure out what the error is in the formula? So anyway, a lot of research there on figuring out how those models can sort of uh can sort of uh be made uh sort of uh be tested sort of better. So that's the first thing I would say. Second thing I would say is that, you know, Gary, in some ways, I he has incredible knowledge of what actual intelligence is. I mean, he's done so much research that's been through, you know, his life's work. One thing that mystifies me a little, someone who doesn't have Gary's knowledge, is I think we easily throw around phrases like, "Oh, are these models going to get better than humans? Are we going to get super intelligence?" I'm like, I don't know what that means exactly in the sense of like,

>> Well, I think they're all thinking, you know, we all watch Terminator growing up, and we're just trying to figure out when that's going to happen, we'll have to go into hiding. You know, I can't actually overstate this because all the people in this world are like science fiction geeks, and they they grew up on this. They take this very seriously. They're sincerely worried about this.

>> You know, Steve, it's funny. It reminds me of a, have you seen Good Will Hunting? So, you know, for those who don't know, it's a movie about there's this genius, genius kid, but he's very led a very troubled life, and his therapist is played by the late Robin Williams. And there's this scene in that movie that I just love where the therapist is sitting with with with Matt Damon, who's playing the who's playing the kid, and he's like, you know,

>> "You're so smart. If I ask you about Michelangelo, you could tell me everything about him, right?"

>> "But can you tell me what it smells like in the Sistine Chapel? Or if I asked you about love, you could like recite..."

>> "But do you know what it is to really love a woman?"

>> Exactly. So it's like it's an amazing scene, and sometimes I think about that when it comes to large language models, where it's so difficult to get your finger on what it means for a model to be human. And sure, the model sort of, you know, destroying all of us is maybe an extreme version of that, but there's plenty of steps in between. And so it's kind of it's kind of hard to uh to uh to define. Gary clearly thinks that they can't scale to sort of, you know, a level where they have the same intelligence as humans. And the truth is, I don't think I disagree. I certainly that most people I respect would probably agree that the current way large language models, just like this predict the next word, predict the next word, predict the next word, might not scale to sort of a level of uh of superhuman intelligence. But I do just want to point out, it's very difficult to really get your finger on what that means and to kind of know sort of, you know, when you've uh when you've reached uh uh that point. Um, a few more things to say here. I'll just give you the headlines, and you tell me if you want to hear about them. Uh, another thing I'd want to probably talk about is the fact that, you know, I think maybe we'll get there at the end, but one thing we sometimes miss, you know, in a conversation about, is this going to get super intelligent? Is even if it isn't,

>> Maybe it's worthwhile anyway.

>> Maybe there's so, and I'm going to take away the maybe from that. There is still so much value in the models. If you told me today, these models will never get better. They will hallucinate forever. They will, whatever sort of the shortcomings are today, they will remain forever. I have personally seen in my work with companies, there is still so much value that you can capture.

>> Let's explore that.

>> Yeah. I just want to say maybe we'll talk about that at the end. There's a lot of research going on now on just whole new paradigms. So you mentioned world models. That's one example. Maybe at the end, we'll talk about those. But I'm saying like the research, it's not like it's not like people at OpenAI and Anthropic are just sitting there being like, well, we give up. They're just, you know, scaling is done and, you know, that's it. We're done. There's plenty of other places people are going. But maybe let's get to value. Yeah. What do you think these large language, given the project that ex as as it exists today, what do you think these, and we'll get to the agentic stuff. We'll come to that. Let's just start with the large language. What are these things good for?

>> When I think of the value that I get out of them, that people I've worked with get out of them, that companies get out of them, I think about them in three buckets, okay?

>> Right?

>> One bucket is probably uh the least obvious. That's why I say it first, is taking those classical machine learning models, which as you correctly pointed out, have been around for like years, and making them better using those LLMs. And we can talk about some examples of that.

>> The second bucket is AAI, and we'll talk about that too. And the third bucket is, and this sounds very, it's the most obvious one, but just using them as chatbots. Like, I think they've been around now for a few years, and so it's easy to just like lose track of how magical they actually are.

>> So let's let's explore all three. Let's start with the first one.

>> Let's start with the first one. So probably easiest to explore with examples, but I've seen so many places where this Gen AI can take classical machine learning and supercharge it. So let's start with the first example, okay?

>> Suppose you run a website where there is um user content that is posted, and this could be anything. Could be a social media network. Could be an e-commerce platform where people write reviews, anything like that, a Reddit, whatever it might be. One problem that bedevils websites like that, companies like that, and that's been the case forever, is content moderation. Right? You've got to, if someone posts a comment to your website, you've got to make sure it's not illegal, it's not abusive, it's not, you're going to have a bunch of rules as to what what comments can go on there. And you know, it may sound like a kind of esoteric operational thing, but this is huge. I mean, people pay enormous teams of people, and these are expensive teams, to really look through those websites. And it turns out the quality of content, like if you look at Amazon, >> a lot of the value Amazon provides, not the only value, but a lot of the value is the reviews, right? I mean, there's a lot, for example, super, super, super valuable. And so forever, for a long time, these companies have been using machine learning models to try and identify, automatically flag what comments should go to a human, right? How do we look at something and figure out, ah, this might be suspicious. Let's send it to a human to review. So, for example, they might look for specific words, right? They might look for maybe the length of the message, the time it's posted, the location it's posted from, etc. The problem with that, and it goes back to the problem from your very first question, what can these LLMs do that that machine learning can't? They can't actually understand the meaning of the message, right? They can look for a word in it, but that doesn't tell them what the, you know, the full meaning of the message is. So, for example, right, this is a law. I don't know if it's true, but people used to think that if you write a message or post a video with the word "kill" on Instagram, it'll de-prioritize it because it thinks it's bad. And so now people have started using the word "unalive" instead of "kill" because, hey, then the algorithm, you know, can't can't catch it. And so these machine learning algorithms were always bedeviled by this problem that they, what I've seen people do with tremendous success is they now say, "Wait a second. We're going to stick with these machine learning algorithms, but we're going to inject some Gen AI into it by having the Gen AI look at these comments people are posting and getting the Gen AI to actually extract meaning from it, because we saw it can do that, right? It can actually get meaning from those messages." And so examples of ways people have done it, the simplest way, which is not necessarily the, is you literally go to a chatbot, you give it the message, and you say, "Hey, give me a score from 1 to 10. How likely is it that this is bad?" And then it'll give you a number, and then you put that into your machine learning model. But an even smarter way that I've seen is you can take that message, and you remember going way back when we talked about embeddings. So, you can take the message, you can get numbers for that message, right? Put it on this XY axis. And then you can use those numbers, those scores, even though they don't really mean anything, they're just like the internal guts of the of the Gen AI. You can use those numbers as an indicator. Maybe you look at historical data and you say, "Wait a second. When that second score is high, that tends to be suspicious, so I'm going to maybe moderate it." And I've really seen people do that to great effect. And the truth is, I talked about content moderation. Any example, any use case where you want to make predictions and there is text involved, this can really provide a lot of value and can be very useful. And notice how in some ways hallucinations don't really matter because you're now, I mean, they matter, right? You'd rather the model didn't hallucinate, but you're not using the model by itself, you're putting it in the context of a bigger machine learning model that you can control, that you can evaluate, that you can check. And so that's just one example I think where, you know, huge value uh uh can be generated.

>> So that's that's one. Let's talk about uh what is this agentic AI that I keep.

>> Funny, we will get to agentic, just so you know, we will. I'll skip to agentic. That first category I had six examples ready, but we we'll go. I mean, I'm only saying that to give you an idea that like there is just so much, there's a lot, there's a lot you can do.

>> You know, I don't on CNBC, I hear the words agentic AI, and I I guarantee you that nine out of 10 people that are talking about agentic AI have absolutely no idea what they're talking about. And I'll be the first one to admit that I don't know what it is. So what, what is it?

>> It's so buzzy. Um.

>> It is so buzzy. I have good news. It's like, it's like you hear this word like it's going to save the world. It's going to wipe out the management consulting business. I mean, agentic AI has accomplished so much, and it hasn't accomplished anything yet. So what is this thing? Who was it that quipped in the early days of the PC revolution that, you know, you see PCs everywhere except for the productivity numbers? Or I there was like, anyway, okay. I have good news for you. Agentic AI is really quite simple. It's basically just a large language model. A chatbot with a pair of hands.

>> What does that mean?

>> So what, what does that mean? It is a chatbot that is given the ability to do things in the real world, like send a text, like send an email, like process a credit card transaction, like process a return. And so the way this works in practice, the way companies do this, and let's take the most common example of an agentic AI, is maybe a customer service chatbot. You, as the company, when you set up this agentic AI, you create a bunch of tools, like a bunch of, you know, IT functions, basically functions that say, okay, this is, you know, will process a return for a customer. This will ship an order. This will send an email. This will send a text. This will refund a credit card transaction. And you make all those tools available. And when you set up your chatbot, you just tell the chatbot literally, like in the text of the chatbot, there's smarter ways to do it, but that's a simple version. You tell it, a customer is about to ask you a question. And by the way, here are a bunch of things you can do.

>> This sounds a lot, but it sounds a.

Lot, I had an interview with, uh, the CFO of a new company that just went public in the summer, and it's a, it's a travel agency company, but it's a completely new company. And the way the CFO described it was that prior to what they do, you know, if you were gonna, if you were going to go and book a business trip, you know, from start to finish, it would take you at least 45 minutes to an hour to, you know, book the trip, book the the rental car, get the hotel. They have created an an AI system from scratch where you can do this all in seven minutes.

That's agentic AI. Precisely. And again, exactly as I described it, what I assume happens in that company is internally they've created some nice pipes where you can easily book a flight, book a hotel, look for flights, look for hotels, and then the chatbot gets the ability as you're chatting with the chatbot, it can say, "Ah, I would like to now call this hotel booking tool. I would like to call this send a text tool. I would like to call this send an email tool." And so on and so forth.

And one thing I think this really highlights, and this is like the number one thing that usually, you know, if a company comes to me and says, "We want to use AI, what can we do?" The number one thing I usually tell them is you've got to first create a fertile environment, a an environment in your company that is able to benefit from that AI. And in the example, for example, you need to have the IT systems that allow a computer to book a flight or to process a return or to whatever. It doesn't happen magically, right? If you're, let's say, a manufacturer and you still take orders by hand and you're writing them down, you don't need GPT75. You need like just your IT systems to be in a way that can actually sort of facilitate those things.

And that's kind of why I say I think, you know, sometimes, you know, we talk about the AI bubble. I sometimes people sometimes conflate the previous question we spoke about, like, will AI scale? Will it become super intelligent? With, do we have an AI bubble? So they figure, you know, if AI scales to be like superhuman, then we don't have an AI bubble. But if AI stays the way it is today, then we have an AI bubble. And I don't think it's that simple. Like I would say even if models stay exactly the way they are today, if companies get their house together in terms of the IT and in terms of the systems, in terms of the data. But that's a major issue because most of the companies in the world, I don't think have their data set up in such a way where they can really say they have to actually go through a process where they have to put their data all in one place and clean it up, etcetera, before they can do any of this.

And by the way, Steve, I got to say, just as a human, my favorite thing about GenAI when I work with companies, my favorite thing is it almost now motivates companies to do this. Like five years ago, I would go to a company and say, you know, you got to put your data into one place, you got to organize. And they'll be like, "Ah, boring." You know, we want. Now it's like, "All right, that's what I got to do to get GenAI working. Let's go." But, you know, but, uh, absolutely, you got to do that. Uh, and so that really is what GenAI is. By the way, I mean, just one more example because I think it's, you know, might be relevant to your listeners. Uh, people are now working. These are in their infancy. Uh, uh, Anthropic is finally releasing this to the public. There are Excel agents now where the tools I mentioned that the chatbot has are tools like, create a pivot table, right? Into this Excel, create, change the color of a cell. And so you're talking with a chatbot and you tell it, for example, "Can you build me a DCF for Nvidia?" Right? And it will be able to use that tool and actually change your Excel spreadsheet as you're talking to it. Right? So if that actually works, you can imagine that would sort of be, uh, pretty amazing. I would say that's still in their infancy, but, you know, things are moving fast. So, so who knows? So that's the agentic sort of bucket.

And what else do you think this can do?

The large language models. Yeah. So, you know, I think it's worth, uh, mentioning, uh, just chatbots themselves, right? I mean, it's literally ChatGPT.com, Claude, Gemini, whatever it might be. Uh, they're great. They can write emails, they can sort of, you know, do all these things that, uh, uh, again, now we kind of take for granted. I'm actually curious, do you use them at all in your workflow day-to-day?

I, I use, you know, uh, Gemini all, all the time for research.

It's useful for research. Sort of, you know, I mean, to write code, by the way. I mean, the extent to which these models have gotten better at writing code in the last six months is just, I mean, sometimes gives me existential dread. I love writing code and I just like, "Wow, they're good now." Again, still not. I got a lot of pushback on that from a whole bunch of different coders who watch my show. You know, some of them were of the view that, you know, there were parts of coding that that helped, but parts of coding that it really didn't help. It was kind of all over the place. There was no, there was no consensus.

I'm going to tell you two things about that. First thing is if you'd asked me six months ago. Yeah. I would have basically said, forget about it. They're a useful search engine for code, right? If you want to create one or two lines of code, but, you know, that's pretty much it. They've really improved. I mean, I would encourage any listeners who haven't used Claude code recently or just haven't used Claude just by itself to write code, especially front-end design. So, creating a web page, a nice-looking website. Uh, truly, if you're someone who's sitting here thinking, "Nah, they'll never sort of do anything." Try it right now.

Okay. So, it's gotten a lot better.

Gotten a lot better. Having said that, I agree with them. I mean, I still think there is a certain level of creativity required in sort of architecting a software solution, in making sure it doesn't go off the rails, and making sure it's secure and so on, that is, uh, you know, that still needs a human involved. That said, sometimes I wonder, and I got to tell you, Steve, like, I, I love coding. I think it's a fun exercise. It's creative. And sometimes I wonder if I, you know, if I'm sounding like an artist who says, like, you know, "I, I love creating art, but these models, well, they're nowhere near as good as I am." And so I don't know that there's an element of similarity there. So anyway, I'm not sort of saying it's definitely amazing at coding. I'm just saying there are certain parts of coding that it's certainly really, really good at. Uh, last thing I'll say just about chatbots is I think they are, um, they are really, uh, something people sometimes don't know is, at least from an enterprise setting, they can be customized quite well. So you can do things like give them terabytes of documents, right? Like thousands, tens of thousands of documents. And then what they do, back to the embeddings we spoke about, there's a reason I mentioned them, they're everywhere. You can take those documents, you can embed the documents. You give the documents like a vibe, basically. You say, "Ah, this document has this kind of vibe. This one has this kind of vibe." And then when you're chatting with a chatbot, it can search through those documents because it can say, let's say you ask, uh, uh, you ask the chatbot, "Show me my contract that has the highest, I don't know, risk or whatever it might be that is the most." It can say, "Okay, risky contract that has a vibe. I can embed that and I can put that somewhere. Let me now look at the document that has the closest vibe." And it finds it. And it's just amazing how well it does that.

Let me ask you about, uh, sort of the, from a business perspective, two, um, what would be the right word, hot topics? So, you know, last year, if you were to look at, uh, you know, what did well in tech and what did not do well in tech, which is always, I find it very interesting because people just say, "Tech, tech did great. You know, buy tech, you know, blah, blah, blah." So, last year, you know, leave, leave aside just the Mag Seven as a category, but, you know, anybody involved with selling chips did great, whether it's GPUs, CPUs, memory chips, they did great. So hardware companies generally did well. Um, hyperscalers who are buying these chips and creating AI data centers did well. But the two groups that did terribly would be software companies, and some iconic software companies like Salesforce, ServiceNow. And the argument about why they did poorly was that the cost of creating software, quote unquote, is collapsing. And so therefore, the moats that these companies have around them are not as strong. That was one group. And the other group were the management consultants, because, you know, the argument is, why do I need a management consultant when for, you know, 90% of the stuff that I used to ask a management consultant about, I could ask ChatGPT and he gives it to me for free. So what do you just think about those two, sort of, are it's, you know, sometimes it's a narrative. And, you know, in the, in the business world, when a narrative takes hold, it's very hard for that narrative to die until it, like, somebody beats it to death.

So I'm just curious what you think of those narratives.

Let me give you some thoughts. I should just say upfront, I am like in awe of people like you who can actually get an idea at least of where the market is going. Like, I always, I can have an idea of, I have an idea of these topics. How those are going to then move the market. God knows, it seems to be like a beast that I sort of, you know, don't fully understand.

That's okay. That's not why you're here.

No, totally, totally. Um, so let's talk about software companies first, and we'll get to the, you know, the, the cost of developing software or otherwise. I just want to first express some sympathy in the sense that, like, I've certainly felt this way, even with teaching. It is like, we are building a car, like, sorry, driving a car, learning how to drive a car while it's being built. Like, literally every five seconds, you sneeze, and there's a new kind of AI that totally changes the way you, uh, uh, uh, you know, your offering, the way you should design your software, the way you should offer it. Enterprises, especially companies of the size of the size of Salesforce, they're just not designed to operate at that speed. I mean, that's wild, wild, wild, wild.

So things are happening very rapidly.

Very rapidly. And I would just say, I don't, this is not a prediction, but I've got a thing, just from like a, just purely sort of logical perspective, it's going to have to slow down at some point. At some point, we're going to reach a place where a lot of the low-hanging fruits have been harvested. And we're now, and so I, I would just sort of put that here first of all, that, you know, we are very much, it's like, it's almost like saying, you know, eight months into COVID, "Oh, look, this company is sort of doing well or badly or whatever." I mean, sure, but things were just moving, like the entire structure of the world was was changing. So that's the first thing I'll say. You know, the fact that the cost of software is dropping and people might develop their own software, maybe. I mean, I've told you, I've, you know, put upfront the fact that I think, uh, uh, uh, you know, coding using LLMs is, uh, is, uh, they're good. Whether they're perfect, definitely not, and so on. I will say though, I think people are underestimating the fact that, like, Salesforce is not just providing code software. I mean, they are, but they're also providing an entire structure, a way of thinking about your business, a way of, you know, like, can you corral a whole bunch of people together in your business to actually agree to use this piece of software? And so I'm not saying that's going to be replaced, but I'm saying there's probably a place they could evolve that involves, like, that puts more emphasis on those, on all the things around software rather than just the software itself. Right? Meaning, imagine if every single employee of yours could build their own Salesforce, their own sort of, I don't know, CRM, let's say, to sort of manage customers. It would be chaos. I mean, everyone has a CRM. Which one do you use? Which one is right? Which one is wrong? How do you verify it? Whom do you trust? So, you know.

I was actually thinking about also, you know, there are certain companies that control databases.

Yeah.

And like, let's say I'm the guy who has the best database about commercial real estate.

Sure.

And I've, I've had this business forever. I would think somebody could come in today and do it a lot cheaper than I than I have done it. And, you know, no one has ever been able to attack me because I, quote unquote, control this database. But I, I think that that people who have certain types of databases are vulnerable to this world.

I think that's true. I mean, I'll give you one example of some example of a place that's really vulnerable. If your entire database is predicated on the fact that you just like looked at a whole bunch of handwritten documents or typed documents that weren't digitized and you, you know, painstakingly sort of, you know, extracted the data from them and that gave you that database, you are really susceptible to disruption because now these large language models, you give them a PDF and using all the techniques we spoke about, they just extract the data in a second. So I think, you know, that's one end of the spectrum. The other end of the spectrum, maybe the extreme version of such a company, is Bloomberg, right? I mean, there are all these data feeds, they have, etcetera. I would imagine, and I don't know sort of a huge amount about it, I would imagine there would be much harder.

That's much harder because they have so many different databases.

Yeah. And, you know, I'll say there is a world in which, you know, we always forget the fact, and this is, you know, the, the, the paradox everyone likes to mention these days, but the fact that, and I forget the name of the paradox, uh, uh, but if you now have a, if it's now much easier to create these databases, yes, you're going to disrupt people who've had them in the past, but you're also going to have many more. There's going to be many sort of more places where you're suddenly able to create these databases you didn't have before. Uh, uh, right. Just, I mean, to give you another example, you know, I said when we first discussed using GenAI to strengthen machine learning, I said I had plenty of other examples. I'll slip one in right now because it applies just, you know, the fact that you could now take a bunch of text and even take images and create these embeddings, these essences, these vibes that kind of allows you to build databases you would never have had before. You could look at millions of images, right? Look at millions of pieces of text. You can embed them. You can put them on this XY plane that kind of defines their essence. And you can start saying, "Okay, if someone searches for a particular term, I can now surface the right images and the right pieces of text that just relate to that particular term." So, you can imagine whole new kinds of databases, whole new kinds of, uh, of, uh, of data that are created by this.

So, let's talk about management consulting because because that's sort of a pet peeve of mine over the years. And, you know, the argument that people have is, you know, you used to hire a management consultant because you, you, they have this, this information and they're going to come in. And now I don't need a management consultant because I'll just ask ChatGPT and it'll give it to me in two seconds.

Let me just tell you something. I, you know, Palantir is not a management consultant by any way, shape, or form. But, you know, there is some services element to the to the business. And I remember sometimes musing with, um, with my colleagues there, maybe I shouldn't say this, but I remember musing with my colleagues there that like, sometimes a lot of the value we provided was just getting all the people in the room together. Like, like literally, we would go to a company because they wanted to use data analytics, AI, they wanted to use our platform. And in that room, there was the CEO, and the CTO, and someone on the ground who actually worked with the data, and someone who knew what the day. And like, literally, all those people had never been in a room together before, right? So, you know, so again, I, I'm, I'm going to an extreme here. I'm simplifying. But I think there is that element of management consulting, that external view, that getting people together, which, yeah, maybe LLMs could put together. But I think, and again, also, I don't know much necessarily about the world of management consulting. I haven't been at McKinsey, Bain, or BTG, or whatever. I was, you know, Palantir, which is a very different kind of kind of beast. But I could definitely imagine, you know, thinking of the services side of Palantir's business, let's say, right, which is not management consulting, but certainly that I, I, you know, I really see a lot of value in like getting all the people in the room together, understanding how you want to take, again, your very messy business where everything is done by pen and paper, and bringing all the data into one place, and figuring out what's happening. So I could easily see the, the, the big management consulting shops kind of shifting, and they already are kind of shifting in that direction. Uh, and then also maybe there's still value to sort of classical management consulting. I don't know enough about it to, you know.

So let's go back to something that we talked about where so many companies don't even have their data in a position to actually.

I mean, in in your travels.

If we're talking about, let's say, the S&P 500 or the S&P 1000, you know, whatever. I mean, if nobody's data is in a is in a position to do any of this stuff, then AI is academic at this point in that in that sense because so many of the companies don't. So, where do you think corporate America actually is to be able to to do any of this stuff these days?

Yeah, it's a great question. You know, and the last research I've seen that actually tries to figure that out systematically was a long time ago. And there it was something like, you know, 5%, some ridiculously small number, but it was a long time ago. I don't think that's true anymore today because you're right, if if the data is not in the right place, there's just nothing you can really do with any of this stuff. I will say that there's a, there's a factor which is that the AI itself can be very helpful in cleaning that data, right? And so there's a company I worked with, I read a case with AI consultant CEO called Blend 360. They had a client that had this issue in their data where they were, the client was formed through a bunch of acquisitions. And they were a B2B sort of shop. And they were formed through a bunch of acquisitions. And everyone, they had multiple databases listing their customers. And every database used a different name. So, one database would say the client is Coca-Cola, let's say. One database would say it's Coca-Cola Company. The other one would say it's Coke. The other one would say it's CCNA or whatever. And what they did, which I thought was just genius, is they used a large language model to clean that up because if you ask a large language model, "Is Coca-Cola the same as Coca-Cola Company?" It is going to be damn good at doing that. It knows the world. It knows everything. So that's just one example where it's not GenAI, quote, GenAI, it's not like for the GenAI itself, but it really, it transformed things for them because now they had this unified data set. They could actually use use machine learning. So I'll say, I know I'm not directly answering your question, but but but I think there is hope that, uh, um, the GenAI not only sort of is useful with the data, but also that it can actually help, uh, uh, uh, sort of formalize and clean and sort of put that data in a better place in a way that that makes it more valuable. But I, I do think we're far. I mean, I think, you know, outside from the very big companies, uh, uh, that maybe have the resources or maybe need to for compliance reasons, just have that data in a good place. I think there's, you know, a lot, uh, of mid-size companies, certainly the not even close.

Certainly the ones I speak to, that there's just a lot of work to be done to get things in in in good places.

So what else should we talk about before we we finish up here?

What else should we talk about? Maybe it makes sense to talk a little bit about, um, what people are excited about when it comes to kind of next generations of of models, things that are in research. So there's a bunch of places people are going. So one thing that I think your listeners may find interesting is, um, so, you know, how I, I kind of explained how the way these models are trained is it just tries to predict the next word and the next word and the next word and the next word. And, you know, a model is considered good, at least in training, when you're tweaking the weights, it's considered good if it does a good job predicting the next word.

Correct?

And so what people have started doing, and they did this in the early days, but now there are even sort of smarter ways of doing this, is they're trying to say, instead of just tweaking the parameters of the model using that the next word prediction, can we actually look at the full answer? Maybe get the large language model to generate many answers because it can, because of the randomness that I told you about. So you can get many answers and then you judge the answers not by each token one by one, but by where it got to at the end. Did it finally at the end get me sort of the right result?

And the traditional way of doing that was something called reinforcement learning with human feedback, RLHF. Uh, that's been around for a while, but now people are doing, and sorry for the jargon here, but they're doing something called reinforcement learning with verifiable rewards, RLVR. That's sort of a big new thing they're doing, which is, can you try and use things that are automatically verifiable, like sums, like math, for example, and can you use that to try and sort of train the models in that way? And it's had, you know, quite a bit of, quite a bit of success. By the way, do you remember DeepSeek, that that novel that came out, right? That was, people were amazed at how good it was and how it was trained, sort of, you know, so, so cheaply to the extent it was. And there's sort of, statement to begin with, that was a controversial statement to begin with. But to the extent it was, a lot of it was they improved that process, the RLHF process was, you know, they did a better job there. So one direction to think about that I think has a lot of promise. World models is something Gary sort of mentioned last time. I don't know a huge amount about them. They're still very, very experimental. But the idea behind these world models is basically, can we give the large language model kind of its own version of the matrix, like, you know, the movie, like its own version of like a mini world inside of it. And so, for example, when we ask, "Pick a ball from the bag," you remember the example before, five balls, A, B, C, D, E, pick one at random. Instead of using the embeddings and figuring out where the embedding is and figuring out the next word, the model would be able to use this little mini matrix to actually like create a bag with five balls and pick the right ball and then see where the letter is and then give you sort of a better answer. And you can imagine a world model that like, you know, knows all the laws of physics and knows the way the planet works and so on. So that's kind of an exciting, uh, uh, uh, direction. I don't know enough about it to tell you whether it's something I'm, you know, I think is good, I think it's bad, whatever. There's clearly very, very smart people that are sort of going down that direction. But that's another interesting thing to, uh, to think about. Um, maybe last thing that's worth talking about.

For everything we've spoken about, machine learning models, AI models, etcetera, I think it is ultimately important to remember that these models are just, at least right now, statistical parrots that are just paring existing data. And that can lead to all the kinds of issues we've spoken about, things like statistical par. So exactly sort of what we were describing, they will take a, a question you put in, they will embed them, right, put them in a space, and then find sort of words around it. But what that determination is based on is just all the historical data you put into the models to train them, right? That's what I mean by statistical parrot. This is all statistics that's being used to figure out, you know, what the embeddings are, what the closest words are. They're just trying to replicate.

So that's why it has problems with novelty.

Because novelty is not in the database.

Because novelty is not in the database. But then it can have bigger problems, which is just that like, you know, the training data that is put into the model by definition is going to have some biases. It's going to have some preconceptions of the world. It's going to have some. And, you know, sometimes those work for you, sometimes they don't. I mean, just to give you an example, I recommend any of your listeners should go on, uh, on, uh, uh, ChatGPT and ask it, or any other model, "Generate a picture of a typical American boy for me." And it'll come up with a picture. Might be what you had in mind, might not be what you had in mind. But it has an assumption about what that is, what that means, what American means. And, you know, again, in that case, it's maybe innocuous, but in some cases, it's not innocuous.

So, you're saying the models could have like a political bias, a moral bias.

The model definitely has a political bias. The model definitely has a moral bias.

How does it, how does it get a political bias? How does it get a moral bias?

So, two ways. The first way is through that training I spoke about, right? It looks at all the text on the internet. And to the extent that all the text on the internet has a political bias, it's going to sort of absorb that. But then there's that second stage that I just briefly mentioned, this RLHF, reinforcement learning from human feedback. That is a process where AI companies will literally look at entire conversations and will tell the model, "That was a good conversation. That was a bad conversation." And so, for example, uh, uh, uh, let's say you have a large language model that starts being like, super abusive and threatening and anti-Semitic and racist and whatever. The AI company will say, "Whoa, that's a bad conversation. Let's try and shift the parameters to get less of that." And as part of that process, it can just gather all these sort of, you know, biases about the world. Some of them are intentional, but some of them can be unintentional, right? It's not to say that, uh, uh, uh, you know, that they're intentional. So that's certainly sort of, you know, a, and this gets especially bad, you know, when we said agentic AI, when we're giving the model a pair of hands.

If we now start allowing the model to do things in the real world, right? You were talking about Terminator. You could imagine a crazy situation where you have a model that isn't so good, that hallucinates, that can't come up with the right thing, and you kind of allow it suddenly to do, I don't know, to sue someone, let's say, right, as a law firm. I mean, that could be, that could be quite.

Send an email that's terrible.

Send an email that's terrible. You can imagine even worse. You give it access to weapons or whatever things like that. And so I always like to say like, I'm not as worried about human intelligence about artificial intelligence as I am about human stupidity, right? That they're giving those models a tool. And so I think a lot of that infrastructure around getting those models to work, the thing I said that is currently the bottleneck, the gating factor, being ready for this, is also to think about, okay, if I give my tool the ability to, let's say, process a credit card transaction, what kind of fences do I put around that? What kind of, you know, safeguards? Maybe I only allow it to go up to a certain amount. Maybe, you know, that, that, that kind of stuff.

So, I'm going to give you the last word. So, we had, as I said, Gary Marcus on. Just sum up your view of where we are in this world and and and how hopeful you are about what what this world's going to be like.

Good question. How hopeful am I? Okay, I'll tell you one thing. I don't know how hopeful or unhopeful this makes me, but I think the one takeaway I want people to maybe take from my view here, uh, um, is that even if everything, uh, Gary said was right, and his view is right, which I actually don't, there's nothing he said that I thought, "Oh, I disagree with this." I want people to realize that there's still a lot of value to be gotten out today. And that value could be for real good. I mean, there's a lot of stuff, you know, you were talking, you spoke about the US healthcare system and how difficult it is and and and complicated it is to sometimes file claims and appeal things and whatever. That is something AI could really do a lot of good for, right, if it's done right, if it's structured correctly, and so on and so forth. So, I think there's a lot of good and a lot of value that can be done, a lot of really a lot of sort of, uh, good stuff that can be done. And so, I guess I would encourage people to maybe worry a little, well, that's not true. They should also worry about the future and if we get Terminator and whatever, that that's worthwhile, but also spend a bit of time asking, okay, I've got this wonderful toy right now today, sort of, you know, what can I do with it? What can it do for me? What, what can I do with it? And hopefully, this has been a bit of a of inspiration for what, uh, and as I said, I'll put, we'll put in the show notes this little, sort of, tools I put together so people can actually play around with some of the concepts, sort of, we, we discussed. And hopefully, they really get to see how, you know, these are not, uh, these are not magical technologies. These are very much just, you know, calculations, statistics, but that put together just result in this truly wondrous, uh, sort of set of models.

That was great. Thank you very much.

So much fun. Thanks so much, Steve.

Glad to have you.

That was some interview. And my takeaways are that there's quite a bit of overlap in agreement between Gary and Professor Getta, but also quite a bit of disagreement. So if, if Gary was here, he would say, I think that the LLM models are losing their efficaciousness and they will never, ever, ever produce returns that justify the hundreds of billions of dollars that the industry is spending. I think that's what Gary would say. Professor Guetta, I think pointed out, he agrees that LLMs are not going to achieve artificial general intelligence, but at the end, he doesn't really think that matters because the LLM models and the agentic AI chatbots, etcetera, are good enough to produce really efficacious products. And, you know, recently there have been a whole bunch of announcements of agentic AI chatbots that have really shaken up the industry. You've seen insurance brokers get impacted, wealth managers get impacted. Everybody's scared that this is actually going to work. And the amount of money that the industry is spending to achieve all this is not stopping. You know, I think I would side a little bit more with Professor Guetta here, that you're going to see a lot more products. What we don't know, and which I think is really the question, is is all this money going to, going to achieve returns that justify the investments? And frankly, we're not going to know that. We're not going to know that for a while. I don't think we're going to know that this year. I think all we're going to see are a lot of announcements that are sound very exciting, but we may not know till 2027, 2028 whether all this investment is going to produce returns that justify it. So until then, I think the story will stay the same. The hyperscalers will continue to spend money, and people will be questioning whether this is really all worth it. This podcast is for informational purposes only and does not constitute investment advice. The hosts and guests may hold positions in stocks discussed. Opinions expressed are their own and not recommendations. Please do your own due diligence and consult a licensed financial advisor before making any investment decisions.