📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

AI Research Legend’s Honest Assessment of Where We Are

Unsupervised Learning: With Jacob Effron1:13:33

Transcription

Is reasoning enough to get to generalization, or is another method needed?

>> It does feel like there is something else that possibly could generalize much better.

>> Why do you think Anthropic was the first to be like really successful on the coding side?

>> Antropic made this very good decision to focus on coding. Open AI was like, "We're doing chat GP." The hard particap we'll see between closed source and open source models and whether that widens or shrinks in the next few years. I think it's a fair question, but

>> Lucas Kaiser is one of the authors of the transformer paper and has had amazing roles at both Google and OpenAI. On unsupervised learning, I got to ask him all the top-of-mind questions of what's happening in AI today. Of course, we had to talk about the transformer, uh, and how he thinks about its persistence and whether it will remain the dominant architecture and what its shortcomings are. We also got his thoughts on what changed in the fall to really make coding models so much better and why Anthropic was really first to code. We talked about what the future research directions that he's really excited about, and we also hit on a bunch of things around how he thinks the ecosystem will evolve from open versus closed source model to application companies. I think folks will really enjoy this episode with a top researcher whose, uh, research really set off a lot in the space. Without further ado, here's Lucas.

It's, it's a pleasure to have a, uh, Transformer paper co-author on the podcast. I feel like you've been at the forefront of, of so many major changes in the AI world. And, you know, our goal is really to get your thoughts on, on all the questions around the AI frontiers today. So, I really appreciate you coming on the podcast.

>> Thank you very much. Thank you for having me.

>> I can think of kind of no better place to start than generalization, right? It feels like that's the, the question in the air right now. Um, and I think in November I heard you say, you know, basically the big, this big question of: is reasoning enough to get to generalization, or is another method needed? And I'm wondering — I guess you said that, you know, maybe six months ago now, which is, you know, dog years ears in AI world so years ago. Uh, how has your thinking on that, that question evolved since then? If we take the current transformers with reasoning, right? And, and, and agents and they have access to a shell and, and, and stuff, they can do amazing things, right? It's incredible how far we've gotten. Like, 2 years ago even, not to mention before transformers, I would have never believed that, you know, you just take this next word predictor, give it then chain of thought and RL that and tools, and that it will — I know every day spend hours talking to Codeex in my case or other people — and it works, right? That you, you talk to it about hard problems at work and it makes sense and it implements things, and so, so, so that's incredible. On the other hand, there is this feeling, um, that that it is not quite like us, right? That it's not quite at the edge of what's that we, we all feel that it possibly should be even better, right? That that we can generalize from less data, like somehow make like bigger leaps, get these concepts from way less. I recently have this saying that, you know, people say like, "Americans will do the right thing after exhausting all other options," and like LLMs, they will learn a concept — they will learn it — but after exhausting all other options, you need this trillion tokens, you need to like learn all the surface level things, and only when that doesn't explain something they will finally learn the concept. That's not how we learn; we just get concepts from, like, sometimes we make them up and they're not great. But so it does feel like there is something else that possibly could generalize much better, that could possibly have this like a bit of a different form of understanding, more like long term. Um, but it's a feeling, right? And every time we try to put a, our thumb on it, it seems to evaporate. Or more like, it maybe it doesn't even evaporate, but it's like the transformer just catches up, right? It's, it was like so, so, so, so both sides in this time have grown, right? Like transformers have gotten even better, but the case for something else has also gotten even better. I would say there is, there's now like a number of labs that pursue post transformers and, and people see interesting results. There is certainly interesting things out there, so, so you know who wins? I, I still don't know, to be honest. I, I think there's good arguments for both sides, and, and it will be extremely interesting to see how this goes.

>> I think it'll be interesting for our listeners, you know, you, you obviously I think at, at, at a talk in Nearcon more recently like alluded to this like whiff in the air, right? That there's something that, that, that's happening on progress that's inspired like these neolabs and other folks to spin out and, you know, work on things that are maybe alternatives to the dominant architectures that are being worked on within the labs. What is that feeling? Is it seeing some of these early results, or, or what is it like? Or is it just like researchers' intuition? Like, maybe making a little more concrete for, uh, for our listeners.

I think a lot of this is intuition, and, you know, you need to be aware because it's like a lot of this happens in San Francisco at parties and like people talk to each other, so it may be, or like on podcasts, so, so it may be that it's self-fueled to some extent. It, it's, but, but I think there is a part of it that, that, that's very fundamental. I mean, like Yan LeCun has been saying something like this for years, right? Way before now, which the models we have in, in a long, long history, they were meant — they're called neural networks because they were meant to imitate our brain, but they don't really. They, they quite different, even if they may have some similarities, right? And if you look at how humans learn, what, what we can do, it, it is quite hard not to say that that from much less data we can do much more than our current models. So, so, so it feels like there is this fundamental ability that, that we as learning machines have that our models currently don't. So, so, so fundamentally there should be something there, not just a vibe. Now you can say as a counterargument that these models always had a trillion tokens to train on and people never do. So, so we just didn't optimize them for training with less, and if you, you know, if you had the same amount of compute but limited data, you can tweak transformers to do much better than, than they do today. So, so, so, you know, it's like some people say, "Why, why would you? Right, we have the data, we, we now it's, it's a big enterprise." But it does feel, even when we, even when we try to push with as little data as people — well, it, you know, it's also like we get a lot of data from visual things, from moving in the world, we take actions, so it's very different kinds of data. It's not truly comparable; that, that's why it's, it's hard to like make a very firm scientific statement about it. But, but there is this feeling that, that we have not exploited all that is there in machine learning. And, and obviously the exciting feeling is that maybe if we find out what's out, it, it, it, it could make what we have even more amazing. You know, maybe not. Maybe, maybe it vanishes when you have that much data. Who knows, right? But, but it's definitely

>> extremely interesting to me as a, as a researcher and I think to many people. It's like, I mean, transformers were fascinating, right? They're, they're great reasoning. I mean, it can solve research math problems. I find — I'm sure you've heard about the recent ERS things — and

>> of course, as I was a mathematician before in my life, so this is extremely exciting. I never thought a computer in this time frame will, you know, talk to me about mathematics takes at a high level as a real researcher that exists now, and, and this is insane. But then as a level researcher, I'm like, "Okay, but we haven't really figured out this learning." There is this feeling, right, that that it learns certainly, but it needs so much data. It needs so much compute. This, this feels like we're not quite there yet. Now, is this only a feeling? Is this a vibe? It, it seems to be reality to some extent, right? But, but we'll need to, we'll need to see.

>> The research appeal of, of figuring that out makes a ton of sense. And I think other folks might look at it and be like, "Well, you know, so what if it's not like people? Right, like we have the data, we have a, a method that works. Um, you know, obviously there's going to be some areas where there is limited data, like you know, medic like you know drug development and other things where, uh, learning from more limited data would be really helpful. But so many problems that exist in the world actually aren't that data constrained, right? Sometimes I feel like these sides almost like talk past each other, right? Like people at the labs will roll their eyes at Yan LeCun or something like that."

>> I think this is fair to say, but on the other hand, given how, how quickly and, you know, with the whole investment in AI, the problems that are not data limited get solved very rapidly. So very soon all bottlenecks that remain will be quite data limited, or, or already are becoming. And in particular, it does feel that, that to work well in the physical world, you do need to solve some part of it at least, because the physical world, if you, you know, you train on like one robot hardware, it, it's not — doesn't quite scale data the way that the virtual or text worlds, uh, or, or internet worlds do. So, so in the physical world is, is a sizable chunk of

>> so people are certainly trying, right? With simulation data and with egocentric video data, cheaper sources.

>> Yeah.

>> I mean, you know, I'm a huge fan of Waymo, right? They, I, I have always this joke like people say, "Where are my self-driving cars?" Well, I drive them. They're here.

>> But then they just canled the highway driving, right? Because they couldn't deal with some construction zone again. And it feels almost like you, you know, they have had this construction zone things for years, and this — I'm sure there's like millions of miles in simulation and quite some in real driving — and it still can't generalize to the construction zone on a highway. This feels, feels just off, right? I mean, I don't know what exactly didn't work there, but I certainly know, you know, no teenager has this problem, yeah? Or no human, right? We have many other problems, right? But, but not that. We can drive in a construction zone in the city but not on the highway? That just — construction zone is a construction zone.

>> And do you think that like some of this stuff, you know, will be, you know, or, or could be solved within the, you know, within the transformers? And I guess like what would you, what, what are you kind of looking for, I guess, you know, in the next few years to, to get a better answer to this question?

The exciting part in, in ML research is that it is so broad, right? You never know whether you need to tweak the architecture or do you need to tweak the data or do you need to tweak the loss or do you need to tweak the optimization process. Um, and they're fair arguments for all, and, and on top of that it might turn out that you need to tweak all of them to some extent, right? It's, it's like transformer is great, but it's also great with the next word prediction loss, right? Or you can make it work with RL, but, but you need the chains of thought, or it's like these puzzles only work when you click them together. So, so it is possible that the, that the like, if there is a new thing, that there might need to be tweaks to everything, but it's also possible that, that parts of transformers will survive. For example, like probably attention will be somewhere there, right? Um, but maybe you need other things to it.

>> Yeah.

>> Like maybe, you know, I, I'm, I've started my machine learning life with RNN, and, and I certainly hold recurrence deep in my heart. I, I like it as a construct. It feels — and reasoning kind of brought it back because every new token you produce, it's the same weights that, that, that we currently produce it. So, so in some sense it's, it's back. But, but, but it does feel that this RL is like very sparse losses, and you do so much, um, but it works, right? It's, it's a — and every time we try to do recurrence in like other ways, it it somehow does not seem to click yet, yeah. But, but, but then there is always the question: how hard have we tried, right? It's, I, I, I don't know if if you or the audience knows, there are models like T-RRM and H-RM, it's like very small models that turned out to do very well on problems like Sudoku, but also RKGI, and so they're a little bit toy tests, but, but they do quite well. I think a lot of the post transformer architectures are trying to merge this with LLMs. It's interesting, certainly, right? It's like the pure transformer can't do so well on it, but you add some recurrence, you, you add some bit of architectural tweaks, maybe a little different loss, and it does really well. So even on the small scale you can do a lot, uh, but then will it, you know, will it, will it generalize to the language and give you the things you want? It's — well, it will be very interesting to see, and, and luckily there was like a number of labs that, and that are trying. And the other thing, though, is like this year we have the agents, and to me this is like a totally — it's the biggest change in the way I work as a ML researcher in, I would say, the last 20 years, probably.

>> I don't know if you try to quantify it, but like how much more productive do you think it, it makes you?

Oh, I, I, I can, I can fairly well quantify it because I, I tried recently just on the, on a private machine to reproduce, uh, a bunch of papers like old papers that I, were was always quite interested in. Um, even some of my papers that I lost code for. Uh, and at least one of them I tried to reproduce before, and I knew it took me about 3 weeks to get to a runnable state, and with Codeex I could get there in two days. So, so it's about let's say a week to a day — that, that's already whether it's a 10x or a 5x, you know, maybe I could have been faster back then, but, uh, but it's certainly it, it changes your rhythm because you can just take on things. It, it, it's also, you, you know, I can just start three things in parallel and, and let it go, while when I was doing I would be just do one thing usually, right?

>> So, so it both makes it faster and makes it more parallel, which is, uh, but, but, but like I mean when I do like private things, not, not, not in a production repo, I basically stopped looking at the code, and, and a friend asked me like, "Do you think you're less sharp now?" And I gave it some thought, and I think it's actually to the contrary, because due to the fact that I don't look at every class name, every small function, but I still know that these agents can go off the rails — like once it, it, it with disable it runs something, and there were some ox losses, and it like just added it, just thought it should have another auxiliary loss, and it was totally off the charts and out of place. So, so you need to have like a full control in your head of what exactly is it doing? What is the loss? What is the — but you don't need to have control of like what's the name of the class? What's, you know, what are the exact words and the function? It's, it's quite impressive that you can trust the agents to be, you know, trustful and like that they're really implementing what you think. But I mean, sometimes we check, and, and, and they are, but since you need to have a full control in your head of what's actually running machine learning wise — what are the losses, what are the batches, what — then I feel it gives me actually more mental control over what I'm doing than it was before, because before I would implement it, but you know, sometimes in this time, be before running it I would have to like forget a little bit about what the big picture was exactly and focus on the little things, debug, then go back to the big picture. By that time, maybe I forgot some detail and then I would remember it where when it was wrong. Now it's like it's this beautiful thing where you can just be in this flow, like you just think machine learning wise what's supposed to happen. You tell it, verify it, and it's happening. Yeah. So it, it's not just about the time saved. It just makes the work so nice. It's a mild psychosis, I guess, among, among researchers these, these — we just can't stop.

>> OpenAI very, you know, publicly said, "Hey, our, our goal is kind of, I think, a research level intern by, uh, by November of this year." You know, as, as someone who plays around with Codeex all the time in your research, does it feel like you're, you're close to that, or, or how you feeling about that milestone?

It does feel like close to an intern, but you need to be very carefully checking, like as I said, it can just add you a loss that you did not ask for because it seems reasonable to it. Um, I don't know if interns do that. Maybe sometimes — I guess sometimes when they're creative, we — but like I, I try sometimes, you know, it's like I, I will just let it go for the night and I give it the goal, you know, "Make a better model for for this lower perplexity." Um, that never works. It, it will just start doing some very trivial tweaks that are not really interesting or useful. So, so it's certainly not at the level of a researcher.

>> Yeah.

Do what's the path forward to make it better there?

It goes back to, to, to our question, right? For a long time, I worked on long context in, in machine learning — even before transformers, you could say, on, on memory and so on. And, and then we worked on it with transformers, and, and, you know, the context got longer. We got like a million tokens, which is huge given what attention does. But now with agents, it really does feel like grep or ripgrep is, you know, our solution to long context: let's write a bunch of stuff in files and give it access to grep so it can find, and, you know, tell it to write index files, and it's like a little library. And of course to me as a researcher, like you told me five years ago, I say like, "That's not a solution, that's a hack," right? But, but, but you know, machine learning — everything is a hack of sorts. So, so like dropout is, is like, we don't, we don't judge, right? We, we take what works, and it works amazingly. It, it works. And, and you add a little bit of RL, like for example compaction — like if there is one reason I like Codeex over Cloud Code, it, it is compaction. It, you can go on with the thread and it's good at compacting. Why is it good at compacting? There is — but I don't think there's like very mysterious, right? People prompted it well and then put some RL on it to just make, make — and if you told this to me like some years ago that that the long context — well, you just RL a bit that it can use tools and find stuff in files and then summarize good enough to keep the — I would tell like, "Okay, that's a bandaid, that's not a like it doesn't feel like this deep thing." But you know, we, we don't judge solutions by how they look. We judge them by how they work, and it works really well. So, so to the point of like, can it become a researcher? Well, you know, some people would say, well, maybe no. Maybe you'll need this new architecture. Maybe you'll need a post transformer thing that has concepts that are bigger and, and follows goals, and, and it's a fair argument, right? It currently it feels like it can solve this. But then there are other people who say, well, well, you will have your conversations with Codex for a month, and then you're going to be prompted to go over them and find meta patterns, write this to some files, and, and just think how it can use them, and, you know, maybe if you have some data over a thousand people and do some RL on it, it will start behaving like a researcher. It's in some ways, this is how researchers learn, right? We look how other people do research, we do a bit of our trials, see what works. Why doesn't that work today? Like, I'm sure people have tried that.

>> Oh, I know. I, I don't think people have tried yet. Very hard. It, you know, it has like some people do some prompts and they work for them. Um, it is important to — I mean, to me, the Codeex era started like this year or Christmas, right? It's a — I mean, Codeex existed before and we used it, and Cloud Code existed and, and we also used parts of it, but

>> I think everyone felt at Christmas

>> but it, it seems like only the newer — but it's not just the models, it's also the harness and the some tweaks, so you know, it's, it's barely half a year, and, and there's still many people, if you go a bit outside of the, you know, our SFAI bubble, that totally don't get it. They're like, "You know, you're a little psychotic, but why?" Right? And, and I think it's a fair question, but, but it does, it, it, it, it started working very recently. We don't truly understand. Like, it was not like a big pre-training that the changed it that much, even though big pre-trainings came too, but it, it, you know, when, when we went from like RNN to transformers, it was very easy to attribute the change to to this

>> well, now I feel and then there was reasoning which clearly is important, and, and

>> but the change last winter, last Christmas — It's a little hard to pin down. I mean, harness changed and the little post-training changed, and, and then new pre-trained models come, which of course made things better, but, but it felt like a big jump which is not that easy to pin down what did it. So, so, so it's a little bit messy, right? We improve everything all the time. But then because it works and, and it feels so important, then there is also the necessity to just, you know, bring it to people, make it work everywhere, promote it — there, there is this competition going on. So, so, so I think in all of this, people did not have truly the time yet to, to think like, how do you really do this meta level? And like, people are starting, but, but also it also feels, you like, because the meta level is something like you research for a week and then you get some patterns and, and start applying them, that feels like — are this needs to take weeks? To like, our current reinforcement learning methods unluckily need to run basically all rollouts on this, and if your rollouts are like weeks long, then, then your training starts to be months long, and, and that all becomes a little impractical, which maybe is an argument, you know, that, that, that the post transformer like, that the human side has something to learn, because clearly, you know, humans can do research over years, and they do this once in their life, right? Or, or twice, right? Some mathematicians spend 20 years on one problem; that's their magnum opus, and that's it. So they did not have, you know, 200 problems 20 years long before to learn from, and somehow they manage. Um, how, how does this work that, that, that it's, it's a fascinating question, clearly with some relevance to, to, to this. We haven't figured it out, but on the other hand, you know, we now, we'll gather — since a lot of people work with it, we'll gather a lot of the data on the like weeks to months long humans. Someone will run this RL, and it may just turn out that, that, that it gets you. Yeah. Further. So,

>> it's such an interesting point because basically, you know, like you saying, as folks were scaling pre-training, or as folks were scaling the original sort of reasoning models, it was kind of it was straightforward, or, or at least made sense like the vector you were scaling on. And then this kind of big advance we've had in Codeex and Cloud Code over Christmas. If you don't actually know what the source of that is, or not you're not fully crystal clear on it, it's very hard to then determine what you should be pushing on to continue to improve these capabilities. Yes, it's a little, it's a little confusing. What I mean, the fact that I don't know doesn't mean nobody knows. I, I think maybe some people have like stronger opinions on what exactly pushed it through, but, but I don't think it's that clear at this point, which yeah, it's, it's just — I mean, on the other, it's been improving for a while, but something happened.

>> Yeah.

Because, yeah, it did not feel possible to, to do this, and now it does.

>> On this kind of current scaling regime on the, on the RL side. Think one question a lot of folks have is: we've seen, you know, obviously, you know, tons of, uh, of, of coding improvement and math and these kind of like verifiable domains, and I feel like the two big questions around RL continue to be like, you know, how well is this going to work on the non-verifiable side, and, and then also, you know, the extent to which we'll, we'll get generalization and not having to keep, you know, do tons of data on each space. Maybe we'll take them one at a time, but starting with the first: you know, how do you think about the problems that need to be solved on the non-verifiable domain side, and, you know, any inklings as to, to which faces might be, uh, might be next beyond code and math?

>> I do think there is a, there's been a fair progress on the non-verifiable side. If you look, for example, at like things like Harvey and law, or, or things in medicine, they're not verifiable, but there's a lot of parts of them that are verifiable, right? So, so, so there's been good progress on that. And I, I think — I mean, GPQA is one benchmark that in some sense benchmarks things like that too, and, and I do think there is really good progress, and there's really good incentives to, to make progress in, in these domains. I, I'm not sure if it's fully fair to call them nonverifiable.

>> They're certainly not as, as, as perfectly set up as coding and math, right?

>> They're, they're not coding and math, right? But, but math — I think people overstate how verifiable math is. Like coding is fairly verifiable in the sense like programming competitions are verifiable.

>> Yeah.

>> Once you go to like front-end coding and stuff, it's, it's also not that verifiable, but, but still, then math — the proofs are not that easy or clean. I mean, you can do Lean, but, but most of the math, at least from GPTs, it's not formalized. So it's not that verifiable. So it's a spectrum, and then things get less and less verifiable. I, I had this pet project of, of translating poetry into Polish, which seems fairly not verifiable. But then you run these models as verifiers, and, you know, they get a fair bit of stuff. They get like rhyme and things, and they can get cultural references. They — so it turns out once you read how people have verified things before, you can get to some level of verifiability. Um, but then I mean, I think what this poetry thinks — that was also what it was meant to show — is you can verify a lot of things and then have still have kind of no taste, and, and it — I mean, since it's not verifiable, it's not so easy, you know, if it were easy to describe in words, yeah, then it would be verified.

>> But it doesn't mean it isn't there, right? It, it, it — you, you read this and there is something in your brain that, that reinforces this idea that there is something they're missing. You know, we, but we have, we have driven ourselves into this hole basically on purpose, because what is reinforcement learning? It tells you whenever you have a, a teacher, a validator, someone telling you this is good, this is bad, I can train against it and I'm going to get good. And that's what the models do. So, you know, whenever I will come and say, "Look, I don't think this does this very tastefully," and something, then someone will tell, "Okay, show me," and then it will nail it. And I, I, I think some people even runs that go basically against — like for image generation, you, you can ask, "Is this beautiful or not?" Okay, not verifiable, but you just get a bunch of people who during training click this is beautiful, this is not.

>> And lo and behold, the images start to be more beautiful. So the verifiability thing is very weak, right? You can — it's just a very sparse signal when you ask. You can ask people, "Is this nice? Is this not nice?" W- w- then, you know, how do we — you know, why do I think this is not very tasteful? Right? It's clearly some of my experiences and some way that I have processed it make me say this statement now. So, so why does the model not say it? And, and there are two possibilities: one is that it hasn't seen enough experience that, that would make it do it, and the other is that it's not processing it in the right way. So I believe in both, actually, but even with the way it is processing, if you just put more experience — you ask a thousand people to, to, to tell it — then it gets better. So, so there is a — every hole you have, you can kind of plug by, by hammering on it, but it would be so nice if you didn't have to, right? It's like because also every hole you plug stops being a bottleneck, and then the bottlenecks that emerge are again the holes that you have not plugged, and, and, and so, so, so we're in this interesting circle. Uh, but hey, you know, if we had this method, this brain-like method that, that, that would just not have so many holes that need plugging, wouldn't this be great?

>> Does that kind of imply that like any, you know, problem area that, you know, someone does focus on under the current architectures like can be figured out? It's just that to your point, it, it probably is, is far more, you know, requires curated data and far more manual than, you know, a potentially more beautiful way of, of, of doing things down the line. But there's not like a set of problems or a set of domains that you're like, "God, under the current, uh, RL methods, you know, that would be too hard for, uh, for, for the models."

>> That it does not feel so, but

>> you do need to take economics into account, right? I mean, currently, to, to make these models work really well, you need to start from a fairly strong model, which is fairly big and expensive. On top of this, it's usually closed, so you can't really — you do it — I mean, there is the RL fine-tuning API which I quite like from OpenAI and some similar ones, but you don't truly have like full access to it. So even with the API, it can be a little hard. And even on top of that, the investment you'll need to make into the data and things — it's, it's substantial, right? You, you couldn't do this, you — us, you'd need a company, you need some contracts, you'd need

>> which, you know, if that's important enough, it's a fair method, but then, right

>> wouldn't it be great if you could just talk to the model and, and, and it would work on its own?

>> Does it feel like there's any signs of, you know, general capability improvement as you do — you could imagine a world where it's like, "Okay, we'll, we'll start with code, and then we'll do math, and then we'll, you know, do this for legal and healthcare, and you could, you could tackle each of these one by one, even if you're not getting any sort of generalization across." Or, you know, ideally, I think the, maybe the hope would be at some level of, of having done reinforcement learning on a bunch of different domains, maybe somewhere to pre-training at, at some level like generalization emerges or something.

>> But I think generalization emerges in, in reinforcement learn

>> So you think already the models get better across the board?

>> Oh yes, they, they, they, they certainly do. Like if you look at like, I think law is simply not in the RL pipeline at all, and, and, and you talk to Harvey or someone, and then they say like, it, it either emerges or they need like a little train like just a few touches on top of it, and, and it suddenly catches it. So, so, so there is definitely generalization, but, but it's the generalization doesn't seem just to go as far as, or like, yeah, it, it just like works in not the ways that we would hope sometimes. Like it doesn't generalize even from math to other areas of math. There, there is — if you look at the even the IMO, right, like now it seems like so far away that models IMO

>> but it, it, it would have like some types of exercises. For, for a long time it was geometry that it just couldn't crack. It would do very hard — it would solve very hard problems in other domains, and but geometry, you were like, "Oh, okay, it has no spatial understanding," and then it just saw more data and started cracking it, but not spatial understanding data or physical, just just more geometry problems. Um, but, but like it, it has this jaggedness, right? It will generalize from here to here, but not to something that seems very close, but somehow in this representation of these chains of thought, is, is not like it's close to me, but it's not close to the model. Right? So, so it's not like it's not generalizing. It's generalizing, but in its weird alien way, and, and that just doesn't cover some ways that I can generalize, and, and, and you know, it's possible that with more data it will just cover more of this space, but, but I also understand people who say like, you know, when it's like that, it's very hard to trust, like to commit to it, like it, it be because, you know, there may be this spike that it just hasn't gotten, so you need to be on the lookout for, for problems. And, and as I use it as ML researcher, I think it keeps me very honest because I think I need to be — it keeps me sharp. So maybe this good in this way, but it's not good in a capabilities way, because you just hope that, that it doesn't have these sharp edges, right? And for now it does.

You mentioned some of the application companies obviously that benefit from these models getting better, and I think there's like this big question of: if you're an application company right now, should you be working super closely with, with one of the labs and sharing kind of all these evals and like understanding you have the domain? Or, you know, is that like actually, you know, are you better off kind of building almost your own model based on information versus, you know, sharing it back? I'm curious how you think about like the room for, you know, applications on top of, uh, you know, the core models.

>> For now. What is certainly true is that the bigger and better your pre-trained model is, the less of these sharp edges you get, and generally the easier all of your life becomes, right? Whether you do RL on it or fine-tuning on whatever bigger model, things just get easier. It's insane how this has continued to be the case. We, we have — I don't — you remember like a year ago, two years, people were saying, "Oh, LLMs are dead. SLMs are the future. Small models

>> and we have amazing small models, like the Gemmas that recently were like few billion. Remember GPT-3? People said, "Oh, you go, you do no zero-shot learning under 100 billion." No, yeah, you have — we have like, you know, 3B models that, that are so, so, so — that's all amazing. But if you really want to solve big problems easily, adjust to your data and context, yeah, there just doesn't seem to be anything like a really — you know, elephant. Well, but they're, of course, expensive and, and, and hard to use, and, and even harder to train.

>> One thing I think would be interesting for our listeners is I think something that's maybe less obvious to folks outside the cutting edge is just what's enabled by new generations of hardware, right? And so I'm wondering if you could speak a little bit about — to — I mean, you know, obviously it seems like for, for, for certain things, as you, you know, uh, as we waited for like Blackwell chips to come online, it's like, hey, they came online and like the models got better. And it's always hard to tell like how much of that is, is just — yes, you could now do lots of things on the hardware you couldn't do before; how much of that was just timing correlation. But maybe just speak to, to that, and I think it's kind of relevant to this conversation around like, are these architectures just going to get better as the, as the hardware gets better?

>> I mean, hardware gets better, and, and hardware, you know, it's, it's easy. It's flops and memory access, right? So, you need memory fast enough to feed the flops. Um, but, but, but it's a very simple — can call it performance. And I, I recently — I so, I, I got a, I got a personal computer. I one for myself, and I bought a 5090 GPU, and it felt like, "Oh, you know, it's one, one GPU and like you're under your desk. What can this do?" Um, so I did a little bit of some tests, and, and it's just insane to think. So the 5090 it's about 200 teraflops. I mean it says 400, but some are turned off on BF16.

>> So the, the GPUs we research transformer on, they had nine teraflops, and we had eight GPU machines. In, in absolute scaling, you, you could — you, you, you could say, you know, be like 70, um, 70, 80 teraflops for, for real on a machine. Um, so now I have under my desk something that's like five of these machines in one GPU, which is much more convenient than — but I think we used like around 10 or so — so you could do all of transformer research on this few thousand GPU under your desk that, that, you know, you could have in your kitchen. Like it's, it's just it's a normal little tower, and, oh, okay, it, it, it's a few years, it's not even a decade, though, so, so it's quite, quite amazing what, what they can do. And, and now we run everything in BF-16, but of course you can go lower even in precision, especially with MoEs, then you pack more in inference. This is amazing, so, so our ability to run these models has dramatically increased, and it increases the, the things you can research, right? You, you can now run so many interesting ways now. It does give you the ability to just — oh, and also there's more GPUs, right? And the world, like the big labs are building out, so, so you can train huge models on a huge number of very fast GPUs. And Nvidia has kept the pace, and then TPUs at Google have kept the pace — they're really speeding up very quickly, um, and their numbers are growing, and it's a very parallelizable process, so we can now train much bigger models much faster. That is amazing. I do still think that the even more interesting things is that we can do more research, and like, like it, it, it's — I, I remember when, when, when I was joining Google, people were talking about like how much flops do you need for the to do something like the brain, right? And it's a very vague question because to really simulate a brain, it's maybe impossible, maybe still very much. But, but people for decades have been doing these estimates, and they always fell like somewhere between one and 100 petaflops. And I remember back then we were like, "Okay, so this is going to take like a few decades for us to get there." Now you can buy a single GPU, so, so that is quite insane. You have like this one thing you can, and then, then of course you can on the cloud get machines with, with many of them. So potentially you can, you know, run like a year of worth of human processing in a day at a, at a cost, right? But it's not a cost of millions, right? It's a cost of hundreds to thousands of dollars.

>> If you believe you can maybe figure out this algorithm — like, I mean, it's questionable whether we have the data that people have. Some people are trying to do recordings of, of kids, right? There's a question how well — there's a lot of questions — but we're getting to this level where, where someone at a university will be basically able to run like a childhood. You know, if you have an idea for how the brain learns, you'll be able to run it and in like a few days to the whole like 10 years of learning of, of a human being and see if it works or doesn't, maybe if you know how to evaluate it. I, I think this, this is even more powerful than the fact that we can build these huge models, which is also powerful because they will help you implement this all, and

>> right

>> and, and we're getting this loop where I always felt limited with RNNs, for example, with because they're very sequential, so if you just run them like in Torch, they're very, very slow, right? But you can write a special CUDA kernel that makes them very fast, but writing CUDA kernels is awful, right? You really don't want to do this, except when you can have a unit test that it does exactly the same thing as your slow thing, and an agent that writes them for you, and they're not yet amazing at it, but they're already do it, and you know, bigger model will probably be so good that you'll be

>> you'll just say, you know, "Use this hardware as best as it can be," and come a few hours later, and here it is. So, so the bottlenecks that were like because the hardware did not fit your idea. Well, the hardware is still the way it is, right? It can't do anything you want. It still needs to be parallel, but it can do much more than it could do before because you can write, like just ask agents to write kernels for it.

Yeah, it's so interesting because some people will say, "God, you know, without the scale of compute that, that exists at a only a few places, you know, it's, it's so hard to do — maybe you can do basic research, but like ultimately the, the, the rubber hits the road on seeing whether these techniques scale, right?" And, and you need to be in a lab to, to kind of experience that. But it's awesome to hear your kind of bullishness around, uh, the opportunity for academia and hobbyists and folks that are just messing around with, with single GPUs to be able to, uh, to, to contribute here.

Well, I, I think especially if you believe that there are some radical changes that you should do.

>> Do you think it's more likely than not that that's the case? Like I guess

>> no, it's, it depends on the day. On my, on my positive days, I do. Uh, it, it, you know, research has always brought us beautiful things. There is no reason to think it won't. Um, but then the

techniques we have also seem to work so well that that that that it just feels mind-blowing to it. It it would be a big mistake to not push on those two. But luckily there's you know there's enough labs. I I feel the whole thrill of being an academic. I was in academia before I I joined the labs is that you can go wild with your ideas, right? you you can't go you can't scale up that much but but on the lower scale which now is not that low you can go really wild you you you can try you know beautiful ideas that that are totally out of the current paradigm and you should that that that's you know that that's the fun of being a researcher and then well you know not not many won't work some will work in the small scale and not scale up But I mean at the scale the current like 8GPUs machine are I mean sure there will always be ideas that you know work up to a certain scale and don't don't work further but I think you're at much higher level now than than than it was like 5 years ago because five years ago this was like really like amnesty tiny things. There was a lot of tweaks that were just really small scale tweaks. Now you're getting even on on one machine you're getting to a scale where it's not tweaks anymore. It's it's it's it's it's like a you know like like I I I I privately use nano chat from from Andre.

>> Yeah. It's a GPD2 level model that that you get in a few hours on on on one on one box, right? Um these boxes have got a bit more expensive these days unluckily but but but they you know when new generation of GPUs will come the older will get cheaper. It's it's a it yeah it's just quite astounding what you can actually do >> and yes not all of this will scale but >> but the fun you can have on the way is is >> yeah totally. I guess one more research frontier, you know, I'd love to get your take on before we shift gears is, you know, multimodal models. And I think you I think on a previous podcast you said, you know, we haven't made a ton of progress there. Do you still feel that's the case? And what's your kind of current state of the union on uh on the multimodal world? So people are certainly making progress. Maybe this goes a little bit to towards Japa, but like the way we do multimodel and transformers or even with diffusion models, it's like in the end you predict like every pixel of these things around and if you think of me being here in the environment and like I think humans like sense an amazing amount of information every second or or less than you know the micro but we we we can't act like our neurons are slow, right? They have like hundreds of millisecond process. But we get all these senses everywhere all the time. And we somehow manage to learn from this insane stream without maybe you know like predicting every pixel autogressively like it's it's both like way more parallel and and and and like much larger. So I feel like the models we have, they have not truly done justice to to to this yet. Maybe it needs new research. Maybe I mean it's but they're also very simil like the I think thinking machines has recently had this like multistream transformers and it feels so easy, right? I mean in a transformer you pay attention to the previous tokens. You could have a bunch of streams that that do this, right? that feels like an easy tweak to the architecture, but but you know, maybe it's an easy tweak, but but just an amazing tweak because I I always when I work with like codecs and and you know, I just forget something, I say it, but then it's executing some bash command. So it needs to wait for my thing to steer it and it takes 3 minutes and I'm like this is just so not interact like it should just and then you can have the side thing and like there's a bunch of hacks again that kind of make it feel better but it feels like of course everything happens everywhere all at once and for us right we see here talk all at the same time um that should be how our models behave no now that there is a bigger lab putting pressure on that maybe it will it will come. Uh but it yeah it it does feel like like we do multimodel without it right without all of these like truly architectural changes to be to be parallel and and absorb like you know transformer can't currently at at the speed it does absorb a high resolution image every m millisecond right it it just because it splits them and this so sequential in this that that it just doesn't work that feels somehow wrong right it's like we shouldn't be putting these tiny patches there. It should like just go in be processed somehow.

>> Um, so I don't think we have like on this deeper level gotten there yet. But but on the other hand I feels like a lot of people are working on it. So so >> yeah then for coding I mean does it matter all that much harder to say? >> Totally. >> Well we I'm sure it will come. I'd love to kind of switch gears, maybe just talk a little bit about your time at OpenAI and your your kind of journey there. Um, because obviously it's been quite the eventful past years. Um, and you know, maybe uh you there's there's a few moments everyone kind of thinks about and so I'm curious to get your perspective, but maybe just like on the on the opening eye side, the company's had some very public like moments. Um, and I'm wondering like what were some of the difficult decisions that really like defined the company I guess in your in your time there? So, so I I you know I wasn't there for the earliest things. I I I think for my time there there there was this big question at some point whether to pivot to reasoning and I feel it was very brave of the company and and and and the leadership and all of us to to actually take this plunge and say yes reasoning will be as important as pre-training. We our models will be reasoning models. They will be launched. And you know at the beginning it was like the reasoning models were not that chatty. Somehow personality was was harder. They were slow and there they still are to some extent. And it was like you know should we should you ever do it? Like maybe people just prefer chat models.

>> Yeah. >> But open was yeah very good at taking this hard bet and saying yes we're we're going to launch it. We're going to go this way. Well, try to figure out how to how to manage. There were two lines of models at the same time. That that's obviously awful, right? You you want to unify this. The unification t took a lot of time because everything is moving. It it it it's it's a very hard decision. But now we wouldn't have possibly all of these amazing things we have if it didn't push on it. And it feels like you know even some bigger labs still have trouble catching up to the RL quality that so so there is some win that you get when you commit to things and and you know I I wonder these days you know open a has since then grown probably like 20 times or something like this become a much bigger company all of the labs have I mean Google was big even before but but everyone now like entropic has become big having been at Google before for a long time I think It's much harder for a big company to take wild bets like that, right? Because you have so much more to lose- because you have processes like it it's just harder. I just hope OpenAI retains this ability and the other labs too because yeah, like the current techniques are amazing, right? They get us very far. But but if there were, you know, a sparks of post transformer world, would these labs be able to jump on it or or would they be on the more conservative side? No, >> it feels that with reasoning there were some like early sparks but obviously not a ton of data and then you know that it was kind of a almost a I've heard it articulated as like a religious belief that this is just going to work if we if we double down on it >> and and we don't have like the successor yet or at least I don't know about it but but with the hope that it will appear will you need a new lab to push on it or or or will will I mean I I think if anything open AI is is good at you wild bets. >> It's obviously interesting to see this whole this whole trend of Neolabs, right? And folks like Jerry Tourric spinning out and and and saying that like it's it's it's almost, you know, easier to do this work outside of of a large lab, right? And make one kind of strong convicted bet.

>> Yeah, it's it's a it's a fair point it's a fair point tool, right? Uh but then um you know, you you start looking at the GPU numbers and it's a little sad when you're outside of the lab. It's hard to get them and they're very expensive. So, so, so but but but then GPUs are not everything. It's so but it's it's quite nice to have this whole ecosystem, right? You have both these little labs now >> and the big labs. Yeah, we'll we'll you know it I I it's it's so fun because being in this this AI little bubble here you clearly see that that there is a ton of competition that change is coming that you know we have not exhausted even on the current paths there's still a lot of techniques to do there's a lot of data and improvements and then and then and bigger models to train and and and then there's all these new things that are bubbling. Maybe they're not ready, but but they're very actively pursued with with, you know, good resources. And then I feel like you step outside of San Francisco and and people treat AI basically as if it was like from the last year before Codeex and would never change again. And then now that is a wrong way of treating it. It has all I mean to to to me the this coding agents have been such a reveal that that it's hard to hard to get over it. I I call it AGI, you know, one should call AGI what they want. I we we may get one day past AGI the way we got past the Turing test, right? We don't really argue about the Turing test anymore. Is it passed? Is it not passed? But who cares? um the these these things that they code with are clearly intelligent and and coding and and in and any disputable s you know you know obviously the AI coding wars are like quite fierce right now like you know what do you think ultimately will determine you know which of these AI coding products ends up uh you know how do they become better than each other uh and you know uh how do you see this the next frontiers for like codex and cloud code and >> you know I think coding market is good to have two big enough to have two programmers in it. It it I I I I think the bigger question will be how well do they go to other fields, right? I mean coding is great and it's important for us but >> but you could do the work of many people and current like currently Codex I I tried to recommend it to some friends but you know it used to start with the question what is your GitHub repo well that cuts off a lot of people now it's a little bit friendlier but it's still called codeex you know like even so people kind of don't hear this is your accountant tool, right? And to in contrast to chat GT where you just said something, I think Codex takes a little bit of getting used to and clot even more. I would say if you if you go on the code side. So, so, so I I think there there is some question how do you get this power to people in other occupations and places and that may be the more important question. Anthroposis with cloud co-work and like basically making a friendlier uh version of of of the core code products.

>> I I certainly feel like the abilities are there, right? As a as an ML person, I feel like these ones obviously can do these things. They obviously can do Excel. They obviously can this or that or that. But then of course they again I watch them as a like a hawk how to like there is some level of skill that that that you need to put in to to get this now this is totally a learnable skill but I understand that people are busy in their lives and don't necessarily want to learn this so you need to smoon it in some way that there are some fundamental things that I don't think will allow you to like just let it run not watched. Yeah, I don't think you want to do this. But on the other hand, I I don't think you would even want to do this even if it was super good at first, right? You need to gain some trust. And so, so the question becomes, how do you convince people to start putting some effort into gaining this trust? It it will pay back, but but but there is a hump on the coding side. Like why why do you think Anthropic was the first to be like really successful on the coding side? >> I think Anthropic made this very good decision to focus on coding coding, right? This this was at the time when OpenAI was like we're doing Chad GPT and you know great I mean Chad Chad is great but and and I think part why Entropic made this decision was that they just could not compete and and in in chat but they made a very good decision on what else to do and and this goes back to you know AI goes through these upheavalss right it it you need to put a bet on something that is not what is today even though the things today like isn't CH GPD amazing of course right that it was the most amazing AI of 2025 but clearly not of 2026 we and maybe in 2027 we'll have we'll have another thing um so so things change quickly if you put a good bet on on something else you can and it's not like open didn't do coding we did right and And that's why it could catch up reasonably quickly. But it was just not the focus. I mean, you know, you these companies are tiny. You grow to a billion users. You you have stuff to do, right? So just fall apart.

>> You mentioned uh uh this kind of like almost tension between, you know, nailing the stuff that's working today and then, you know, keeping other areas open so that if there's a a glimmer of hope in a different area, you you kind of double down on that bet. And I'm wondering what you make of that. You know, obviously open I think very publicly has gone in this like focusing moment now, right? And and you've seen it in the results of Codeex and um you know, maybe slashing Sora and some of the other things that were that were different. Yeah. How do you think about like kind of navigating that tension of of like really nailing some the here and now versus like keeping these other embers open uh that could potentially be really interesting down the line? >> It's a matter of culture and size and and money and perspective. Famously Google, right? Google is the lab that will keep all its >> I think some people have been quite critical of Google for this right missing you know missing your your uh invention uh not being the ones to capitalize on it >> you know but then it it works for them right it works be because it it whatever good comes out it's very easy to catch up because you already have a strong team in in in it right >> yeah do you think they've caught up I feel like there's a lot of discourse claiming that yeah they're still a bit behind >> I think they've caught up in the chat GPT world they haven't get caught up in the I mean I don't know if you've seen anti-gravity 2. >> Yeah. Yeah. >> I opened it after IO and it I would couldn't tell which one is Codex and which one is >> of course a lot of a lot of funny uh funny tweets about that. >> So so so there that's great. Um I I tried to to do some of my codeexings with the new 3.5 flash and it just doesn't work right. It has not the barrier that wasn't Christmas. I feel like it hasn't crossed it yet to me, but but it will, right? So, so so so so like if you're very broad, it can make it safer later if you need to catch up, but then you may not get the immediate win of like, you know, like anthropic and coding. You're just the first to nail it. And and you know, and great that there are labs that that just go and are the first to nail it. That that that's exciting. and and you know I feel like that's how it should be. Um and I you know open a had a good culture of of making bets but now it is also a bigger thing and it has some like you know GPD has a billion users it's important for many people in the world you should and a Google search has three billion users right it's important for many people in the world you >> you don't want these things to >> be hampered and and totally >> you you should go fast but but this breaking things is is is not so good and and I actually feel like it's quite good if the labs don't break everything on the way.

>> I guess you know a lot of people wonder about the the kind of gap between closed source models and open source models and it feels like there's two you know distinct things pulling in different directions. One is it feels relatively easy to distill models and you know you've seen a lot of uh a lot of claims around folks doing that on the on the Chinese open source side with with the closed source providers. Then on the other hand, it feels like more and more of these models even in the big labs are getting too big to serve. So they have to be distilled uh within the big labs themselves. What's your kind of gut intuition on the gap we'll see between closed source and open source models and and whether that widens or shrinks in the next few years? >> Yeah, it it is not that easy to to predict. I feel it. So bigger models are better. Um, you can distill them, but the distilled models are never quite like they're great, right? If if especially if you need a model for for some money, but they're not quite as good as the big models. I you know, I just said like the 3.5 flash, I could not quite feel it's on par with 5.5. Maybe because it's a distilled pro, right? Maybe maybe you just need to wait for the pro. So even within the like I for example I don't remember when I have used the mini model the the mini animals. I think they're very good. They're very useful. I just haven't used them in in a while. Right. >> So and whenever I use them they're fine until they trip and cost me so much time I go back to the big one. And and so so so you can distill things and you know when the open source can distill or not distill. I mean you know labs just try to not make you distill everything naturally but I think they also like don't fight you to death. Yeah it would be very sad if open source had like models that are very very far behind but I don't think there is a risk of that. there's enough companies and and and and now there's like you know notions of I also very much understand like if you're a country right do you want to depend like say your police stations or hospitals run AI to like help you do administration maybe you don't want to rely on one company that may just have an outage or you understand like so so so for there will be a lot of people who want sovereign they say models even if they're slightly weaker, maybe the tasks are not so hard. So, so I think there will be enough incentives to have open models that they will exist and there will be a very good incentives for the labs to still keep ahead. So, so you know people like people keep paying for for for this. So, so it feels like a state that should persist for a while but you you know I I it's famous last words and and AI and tech you can say things and they may turn out I don't want to make future predictions >> of course but what what job is a podcast if not to try and force you into them but uh no that that all that all makes a ton of sense. Um, you know, we always like to end our interviews with kind of a quickfire round where we stuff in a bunch of of broad questions at the end and and so maybe to start I just love like what's one thing you've changed your mind on in the AI world in the last year? >> Uh, well definitely I I I did not believe that that they will be like a intern kind of thing so fast and I have definitely changed my mind. I I I actually used to not talk to AI very much every day. People were always like, "So, how do you use chat GPT?" And I was like, "Yeah, I don't know. I I asked it one query yesterday and and one three days ago." And but and I was always like, "I'm not going to talk to my computer very much." And now I do >> about work. >> And and and so I Yeah, it it like Yeah. I I also did not think I'm going to like not use a editor for programming and now I don't. I just tell it to change the code. And >> yeah, that that was a big update. >> That's awesome. Um, I guess, you know, as you worked at these models more closely these past years, have your concerns around like the existential risk of of of safety around these models, have they they gone up or down these past years? >> I I don't think they have changed very much for me. I was always on the, you know, not too worried, but but also we should not be complacent side. And I still feel, you know, with all the skills they have now with programming and so on, I still feel the the small risks, right? The the risk the risk that they will hack some of our systems, make the grid go down or or things like that. I still feel these are the risks I would focus on right now. Not to say that extension risks don't, you know, it's good that there are people thinking about it. It's good to have some guard rails. It's it's in the end good to, you know, we should be able to turn off this these data centers if we so decide and and have control over all of that. But I I don't feel even though the models have become much better, I don't feel any threat from them.

>> On on the lab side, it feels like the buzzy news of the last week was that like uh you know, Andre Carpathy was going to anthropic uh to work on on RSI, right? As as a team there. Like what do you what do you make of that? >> Yeah, you know, I I am part of this psychosis, right? It's it's like you can do so much research with with this assistant, right? and and it's amazing and and you should and and you can make many parts of the systems better too like much faster. So, so that that's certainly true. But on the other hand, like you think about these post transformers things and you know the space of ideas is vast and unluckily most of them are wrong. That that's why it's called research, right? And and you need enormous lack and and skill but but also lack to to happen upon the right one. And we kind of feel like maybe it's somewhere there in the air but but it's research. maybe years away and and even with the best AGIS of the world we you know they're they're like human level they're maybe researcher level maybe they'll be like a 10x researcher but but for years there there was a huge community of researchers trying to crack these things and and and they didn't so so it may be just very hard and and yeah like we understand very little about the human brain yet and we cannot connect it to our ML in any great way yet. So, so, so I'm like on the one hand I think it's great. I think we'll see the current things getting better, but if you're thinking of of a research breakthrough, it it may just require something that even when you have like this if you're searching in a very efficient way and even if you're searching some interesting ideas that it still doesn't mean you're going to find it, right? >> Yes. just the space of all ideas is so vast that that that even very efficient searches can just not get there. So totally I'm I'm not that I'm not that worried existentially about this.

>> One thing I found interesting is I think you know if I'm correct like all of your your Transformer paper co-authors have have gone on to to start companies, right? And I'm wondering if that was ever something you thought about or uh or uh >> I certainly asked about it many many many times. Um, yeah, I um well I'm I'm very happy that I didn't so far. I I thought both my time at Google and my time at OpenAI have have been great and and and and it was a privilege to to to be there and and to to be able to do the work. I I love technical work, you know. I I have um everyone who started a company has has thought maybe they won't need to spend so much time on the company work and it feels like they had to >> totally >> you know but but then sometimes companies do amazing things it's been a fascinating conversation I want to make sure to leave the last word to you anything you want to point our listeners to or or thoughts you want to leave them with uh the m the mic is yours I thank you. I I I just want to I think I said it already, but I just want to repeat. I feel this time now that you have like powerful GPUs that you can put under your desk and and coding agents that can really help you push them to their limits and and the time where you know all the big things are pushing the transformers and great they are because they're amazing but but there is this whiff of of possibly other things. I think it is an still and again the most exciting time to be a researcher in machine learning and I want to encourage everyone to to just go and try their ideas to to learn from others. If anything I I feel like we should publish more of like wild things. I I feel always a little sad when so many papers are about like oh we took a pre-trained model and arald it in a slightly different way. I mean it's it's good but you know you don't need to catch up with what is there. You can just do new things even if even if they'll start smaller even if you know maybe it won't work at the first time. Um, you know, no nobody talks to me about the paper I had before attention is all you need which is you don't need attention. I had a paper at the year before saying you just replace it with active memory. Well, wasn't wasn't quite a good advice, but you you you need to explore the wrong things because they may lead you to the right thing. And and this is also what models are still so bad at, which I think Jerry is is trying to push. Models are very bad at like learning from a totally wrong direction to actually twist it to a right one. That's what we humans can can still do very well. So, we should do more of it. like we should just do wild explorations even if they fail and I feel now that the you know it's if you put a lot of your own effort without an agent it's it's done very hard to when it fails I think with agents it's even easier so I want to encourage everyone to you know do research explorations fail fail when it comes to this this is how we'll how we can get to interesting things >> I love that well I feel like that's the perfect note to end on thank you so much for uh for coming on the pod this was for fun. >> Thank you so much for having me. >> I'm Jacob Efron and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's a nights and weekends project in addition to my day job as an investor at Redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so, please consider doing that. And thank you so much for your support and listening. We'll see you next episode.