Transcription
You need to reach this level of reliability to really make any of these AI tools very useful. And I think we just crossed that, probably December last year, at least at OpenAI. Now we can trust these models to do a lot of the work that we are doing. The last few months have been pretty well. We moved from like competitions to, uh, usefulness to users, and that's what we are feeling right now. I think most of the time, the big is the the last mile. There will always be a lot of space left for this last mile in different verticals, and I would highly encourage people to continue working on that.
>> Hi, I'm Matt Turk. Welcome to the Mad Podcast. My guest today is Yan Dubois, who co-leads the post-training frontiers team at OpenAI. The recent release of GBT 5.5 was yet another major milestone in AI, and Jan's team helped build it alongside OpenAI's prior top reasoning models, including 03 and GBD5 thinking. Before OpenAI, Yan was at Stanford, where he co-authored Stanford Alpaca, the landmark project that kicked off much of the modern post-raining research community. In this conversation, we go deep on what's actually new in GBD 5.5, why reinforcement learning is moving from math and coding competitions into messy real-world work, why AI progress can feel like a sudden step function, and why continual learning remains one of the big unsolved problems in AI, three years after Jet GPT. Please enjoy this fantastic conversation with Yan Dubois.
>> Hey Yan, welcome.
>> Hi Matt, thanks for having me.
>> It's been another wild, uh, last few weeks in the world of Frontier AI with the release of GPD 5.5, of Claude Mythos preview. So it feels like, uh, we have unlocked yet another step function in progress, particularly in cyber security, agent coding. What's the best way to think about this from your perspective? Are things accelerating? What is happening?
>> Yeah. Uh, the last few months have been pretty wild. Um, internally, we also really feel it. Uh, and I think anyone who's working with, uh, anyone who's work, who's coding, basically, is really feeling it right now. Um, I think that's really because of three reasons. Uh, the first one is, even though the, in my mind, everything, the progress is actually pretty continuous, you need to reach this level of reliability to really, uh, make any of these AI tools very useful. And I think we just crossed that, probably December last year, at least at OpenAI. That's why I thought we really crossed that threshold, uh, where now we can trust these models to do a lot of the work that we are doing. So it feels like a stem function, even though I actually, in terms of capability, it's like, it's pretty continuous.
>> Um, so that's the first thing. The second, the second reason is, um, once you start having models that are really good, you accelerate yourself. Um, especially in terms of coding, given that we all code internally, uh, you, you accelerate yourself both for having these models like train the other models, but also like build, like the tooling that we need as researchers to like do our job. And, and all this acceleration, I think means that we saw these last few months going faster and faster. The third thing that I, I think we are, uh, feeling is all of last year, uh, we really built on, um, like these reasoning models, and we really like sawing pushing a lot on on reinforcement learning. And initially, when we had like 01, um, 01 preview, even 03, um, these models were still like optimized for, uh, what we call verifiable rewards, things where we actually have access to ground truth and like, it's easy to test whether you're correct or not. Uh, that is, for example, the case in like math questions or like comp, like, like coding competitions. And what I think we are realizing now is that we were able to take many of the tools that we built for these like verifiable reward cases, and we were able to use them more generally in, uh, on for reinforce learning on, like real use cases. And I think that's like really why we're feeling that right now in like, um, uh, just real-world coding rather than like competition. So we move from like competitions to, uh, usefulness to users, and that's what we are feeling right now.
>> Okay. Fascinating. So we're going to unpack a lot of this, particularly on the on the RL side. For the first thing that you mentioned, reliability, is that an engineering, is that models, like what makes a model reliable in, in the way you meant it?
>> It's a little bit of everything. But in general, given that these are agentic models, h the longer, if you just think about it as like every two minutes, there's like a certain probability that they're wrong. Uh, the longer that they run, the, the higher the probability that like the final answer is going to be wrong. Um, so it's just something inherent in like agentic models. And what we've been pushing a lot on is like making sure that the model, like we decrease this probability of being wrong every, like two minutes. So purely from a model point of view, of course, there's a lot of reliability that is also happening on the applied side, and the team at open air has been doing an an amazing job, um, on that, um, but I'm, I'm even talking only about reliability of our models and like making sure that, like basically we decrease the probability of being wrong.
>> Great. So 5.5, formerly known as spud, was, uh, as mentioned, a big deal, is a big deal. And I'm just curious from the inside, what was, what are you guys the most proud of? What did you find the most challenging? Give us some, some, some color on like how you all, uh, felt, uh, you know, releasing this.
>> We're all really excited by 5.5, to be honest. It is one of these models where everyone in the company was extremely involved, um, in building, um, and I think that we really feel it now. That's like, we got a lot of attention because of 5.5, and it's, uh, it seemed like all the stars were aligned. That doesn't always happen. Um, and I was just like, a great model for the, for this. Um, so we, we did feel it. It's kind of funny because in general, with every model that is looking really good early on, we have a model, we all get really excited about it, and then there's like tons of doubt that start, uh, coming up because it's like, oh, like everyone is so hyp, is like hyping this thing internally, but actually, it's like bad at all these other things. And then there's another wave where like people start, uh, um, underhyping it, and it kind of goes through, through waves, and it depends, like when we actually ship it, how, like people feel about it internally. But that's true with like most models that we have. Um, so 5.5 was not that different in this case, but it definitely maybe had like a, a higher amplitude of the wave. So people were very excited, then very, not as excited, and, and we shipped it, and, and people are happy externally.
>> How long does that process take? Like, you know, you, including the waves of going up and down and of, of excitement? I guess it depends on the, on the, on the release and the importance of each release, but like, is that a, is that a few weeks? Is that a few months?
>> It really depends. I, so I cannot talk exactly about what, what went into, uh, 5.5, but it kind, it kind of depends, um, which part of the pipeline is training parts of the model. So we really have like different sub-teams, uh, including pre-training, and you have like the mid-training stage, and like you have some post-training, and usually the closer you get to to products, like post-training being the last one, the faster the iteration cycle is. Um, and if you're more upstream, the slower the iteration cycle is. Uh, so it could go from, let's say, from months to to days. Uh, basically, 5.5 was particularly good, um, on agentic coding, computer use, knowledge work, and early scientific research. How does that work internally? Do, do different people focus on those different parts? How do you get to that result?
>> Yeah, we definitely have different teams that are working on specific use cases and are pushing on these use cases. Uh, my team specifically is actually the one that is kind of taking all these vertical improvements and try to put them together in the final model. You could see it as a team that is doing both kind of the smoothing function. So you have all these improvements, but you need to make sure that the model is doesn't feel too spiky, doesn't feel the differently in different, on, on different verticals. And also, you, you need to have some teams that are working, and that's basically what my team is doing, on all the horizontal improvements. So there are many things like instruction following, function calling, or like thinking about how much should a model think for, on different, uh, problems. Those are very horizontal, and that kind of impacts all these use cases. So we have both these more vertical teams and these more horizontal ones. Um, and both are very important, uh, to, to, to improve the, to improve on the model. Um, and the good thing is that these things can kind of be improved orthogonally. So you might have like multiple different teams that are working on certain verticals, and maybe for one model, there's only half of these teams that made integrations, basically, in the last run, and like improved the model on, on these capabilities, and maybe for the next model, it'll be the other half. So that's kind of at a high level how it works. Uh, one thing which I will say, because you asked also about, uh, what are the things that we are really proud about for this model, I would say two things. Number one is the efficiency of the model. U, we really, really improved, uh, the efficiency of the model, and like we, most of the tasks, and we basically performed, I would say, like 2x faster now with this model. Um, so that's great. Uh, and the other one that I already mentioned before, but it's kind of this alignment of the company and making sure that, like everyone is working towards the same goal. Uh, and that really takes the entire, the entire company working towards, like this north star of building one good model, uh, uh, in, in, like specific timelines. So very, very proud of how that happened.
>> Great. And then speaking of efficiency, how do you optimize for that? We're talking about efficiency per per token. Are we also talking about latency in serving the model? What, what, what part is AI research versus engineering?
>> So that's what, that's what I mean when I say it's the entire company. Is that it really comes from everywhere. It has to come from like inference optimizations. Um, it has to come from the model being more efficient in its thinking time. So you have basically every token that you think for, uh, basically the usual plot that you should be looking at is x-axis, the number of tokens that you think for, and y-axis, um, the, the performance. So this is the, these test time scaling curves that we look at. Um, and research basically tries to move this curve to the left. So think less, uh, to be the same level or more correct. Um, and then inference also deals with with this x-axis, but switches, switches it from number of tokens to actual latency. Um, and the final thing that people care about is latency on x-axis, performance on y-axis. And this is where everything comes together. And this is really what happened with 5.5. Um, so yeah, that's why I always say I'm really proud of the company for this one.
>> Okay, great. Let's talk about you for for for a minute. So you are in the post-training frontiers team. What? So that, that team you described as horizontal. So what does the, the team do in general?
>> Yeah, I would say there's three things that we do. So in a broad, broad sense, we are on the porching org. Uh, and my team is the porching, uh, frontiers one. So there are three things that my team does. Number one is, we kind of decide what goes into the final run. Um, so as we talked before, there's like money verticals, and someone needs to decide like what can go in, what cannot, and also provide the, the science experiments for people to, to iterate on something that's going to be representative of the final run. So this is the first thing that, that my team does. The second thing that my team does is bringing everything together and actually doing the big run. So this has, as you might imagine, like we train on a good amount of GPUs. So there's a lot of infra work that is needed, but also there's a lot of ML work that is needed by putting everything together and making sure things work well together. And then the third thing that my team does is, uh, horizontal improvements to the models. Basically, there are some things that, like these vertical schemes will not usually look too much at. For example, the thinking time, as I said before. So how much should the model think for on certain answers? Um, or like instruction following, function calling, uh, things like memory, and like general improvements to the model that are really across the stack. Um, so that's what the pushing frontiers team does, and, uh, and I'm leading that team.
>> Okay, great. And, uh, what was your journey to open AAI?
>> Oh, it's a long story, but I'll try to keep it really short. Uh, basically, I did my undergrad in biomedical engineering, um, in Switzerland. Um, I'm from Switzerland, and then I went on an exchange in Canada, and I learned about word tove vec. So I don't know if you heard about this algorithm, but it basically takes words, which is like something discrete, um, and puts it in a, in a vector space. So puts it, basically, in a way to think about it as a plane where if words that are more similar to one another will be closer to one another. So it, it brings these like discrete words into like some continuous space that is semantically meaningful, and I was absolutely blown away by that algorithm, and that's when I decided that I, I wanted to work on natural language processing and just like understanding language. Um, at that time, I was very wrong, but I thought that, uh, English, uh, NLP was basically solved or like close to being solved. That was in 2017. So that was, uh, uh, right when transformers started. I was actually right before transformers. So I was very wrong, but I decided I wanted to work on under research languages and basically, um, I wanted to improve, um, NLP on languages where we don't have that much data. Uh, so I went to, uh, work, uh, for Grab in Singapore, and I was basically building the natural language processing, uh, pipeline for them, working with Kum, with Bahasa, with Thai, Vietnamese, and all these different languages. And then I'm skipping a little bit. I had, I did more academic type of work in different countries, and I ended up at Stanford, did my PhD there. Um, and after this, um, had a small stint into startups, and then went to open air.
>> Yes. And I, I remember seeing on, uh, I think your blog or your page, a note to for for quant firms to not reach out to you because you were not interested in hedge fund work.
>> Yeah. By, I always think it's very important for me to think about the positive impact that I'm having in the world, or at least that I'm trying to have. So
>> so that's, that's why this note is there.
>> Yes. And as we were saying, just before we started, uh, recording, uh, people, uh, may have seen you in the GBT 5 video announcement, and you did this, uh, very funny, uh, demonstration of, um, an app that was built on the fly to teach your partner how to speak French. So like, people should go check that out.
>> Exactly. Um, that was that was a fun one. That was a fun one. It was TPT5 was not that reliable. So I was a little bit stressed that it wouldn't work, but, uh, but it ended up working.
>> So, this was truly live and, and presumably very, very rehearsed, but, but truly live.
>> Actually, the, um, right before we did that, like the last rehearsal, it did not work. So I got slightly stressed about that. But, uh, but yeah, seems like live life ended up working well.
>> Yeah, no pressure, but yeah, that, that, uh, that landed perfectly. Okay. Uh, very cool. All right. So, let, let's, uh, unpack, uh, some of the things we alluded to in the intro. So we started effectively talking about reasoning, and I'm, I'm curious what reasoning means in 2026 that's any different from, you know, a conversation we could have had about 01 or or or three. Um, in particular, one of the claims, uh, of 5.5 and, and also my experience as a user, is that it's particularly good with with messy data, which seems to imply that, um, it needs to reason through ambiguity more. Um, what has changed?
>> What I would say is that 01 and 01 preview, uh, were really, really breakthroughs, um, in, in the research community about having models that can think. And the longer they thought for, the more like the higher likelihood they would be of being correct. Um, so that was really a breakthrough. But initially, and if you look at like old blog posts, you would mostly see like math, uh, math evals, and also like maybe coding competitions, but things that are really easy to test whether you're correct or whether you're not. Um, and it also gives you like some suggestion about like how we were training some of these models. Um, and how I see maybe all of last year, and especially the end of last year and the beginning of this year, is that we were able to take these algorithms that work with, uh, verified rewards, like things where we can say you're correct or you're not, uh, to the messy real world, um, and really optimize for the utility that we provide to users and like making them more productive. Uh, so I think that's what really changed.
>> Okay. So it's the post-training reinforcement learning part largely?
>> Yeah, I would say that's a, I mean, there's also, there's also another big part of it. Uh, number one, basically, the first thing is that, of course, when you develop a new method, uh, the method is kind of fragile, um, and is not that reliable, and like, it's hard to, to basically productionize. So this bot also improved a lot. Um, but then it's also really, basically, we had a tool that we could start optimizing for, for different things. And initially, when we were developing this tool, uh, we were optimi, we were making a lot of simplifying assumptions of in the real world, basically, and, and now we are removing these simplifying assumptions, and, and at least in posting, we are able to optimize really, like user utility and make sure that these models are useful, and the tasks that we're looking at are useful. And that's why also now current evals look much more realistic. Um, I mean, if you think about GDP val, or even if you look at like threebench pro, or threebench, these look way more, uh, realistic than, let's say, some code force or like coding competitions that we were looking at with 01.
>> And still on the topic of of reasoning, um, what's ultimately the difference between 5.5 thinking versus 5.5 pro? Is that, is that just more test time compute, more tokens, and more time invested in solving a, a problem?
>> Yes, basically, it's just a question of, of, uh, how much test time compute we pour into the model, uh, or we pour into this entire, uh, system that we're shipping. Um, so we, we've seen again and again, the longer the model think for, uh, the better answers we will get. The problem is that these curves that we're talking about, um, are not, are definitely not linear, and like, they, there's some plateauing effect, and they kind of look, um, logarithmic on some, in some sense, um, or depending on which evals. So you can pour like two times more compute and actually only get like small performance gains. Um, I personally don't use pro that much because I really don't like weight. I'm pretty impatient, so I don't like waiting for that long. And, and I know that the probability of being correct definitely improves, but it doesn't improve like enough for, for me to use it. Um, but there are some people who use pro and who really love it, especially actually for academic research, and, I know, especially a lot of mathematicians who are using it, and that's because they're kind of just have this in the background that is running for maybe one hour, uh, two hours, and they don't really need to like iterate really quickly with the model. Um, and pro is really good for that.
>> I'd love to reconcile this with, uh, what you were mentioning about efficiency earlier, per token. So is the idea that, um, you would be able to think longer, but also be more efficient, therefore solve the task better? Like, how do those, the, the time aspect and the efficiency, uh, aspect sort of interact?
>> Yes. Uh, so if you go back to like the plot that I was talking about, I was thinking about, uh, where on the x-axis we have latency, and y-axis we have performance. We're basically moving this curve when we say that we improve efficiency more and more to the left. So we're becoming more efficient, uh, or like we spend less time to achieve the same performance. Um, but what pro does is that it extends this curve. So it says like, um, I'm going to think for much longer, but I will have a higher likelihood of being correct. But every iteration of the pro model also moves to the left. So it also becomes more and more efficient. The important part is, um, there will always be tasks where, uh, you just want to maximize the probability of correctness, and you don't really care about latency. For example, if I, if I start a job before going to sleep, um, I mean, the model has like eight hours, like it should just think for as long as it, as it can. Um, and this is what kind of promo gives you.
>> and, uh, in layman's term, like what, what, what does that mean practically, or how does that work practically? If the model goes in a wrong direction, then it would interrupt itself earlier? Is that, is that one of the axis?
>> so for the effic Okay, so there's two things. Are you asking for the efficiency? What does it mean?
>> Yeah, for the efficiency. Yeah, yeah, for the, uh, largely for the efficiency. I'm, I'm just curious, uh, how reasoning gets more powerful.
>> Yes, that's a good question. Let me give you, um, maybe a metaphor from, like humans. Uh, if, if you have a, someone who know, like someone who's an expert in a certain domain, uh, and you compare them to, like some undergrad that is like starting in that domain, uh, the undergrad doing that task will probably take, might take, like one day, two days, and will have to think through a lot of, like the, um, the possibilities and like investigate because it never did a certain problem. While someone who's an expert in that, in, in that field, will usually just like know what direction to take, and it will, it will not spend the time on, like investigating 10 different directions because it knows that there's like one that is more likely to be correct. So this is the type of efficiency that we're talking about. It's basically models that we, where we optimized more on, like real-world problems. Um, and as a result, it was kind of trained to, uh, to figure out with a higher likelihood which paths of reasoning are, are more likely to be correct. Um, so this is, this is a part on, on efficiency. There's also what, what you suggested, is that, um, part of it is the model knowing when it's going down the wrong path. Uh, but this is also something that we can, um, that the, the model can be trained for with reinforcement learning, is like knowing, okay, like that seems like not a great path. Let me backtrack, and let me go and, and test something else. Um, and if you train the model less, it might realize it's in the wrong path much later.
>> Okay. All right. So, it seems like, um, a lot of this, uh, goes back to, uh, reinforcement learning and post-training. So, let's, uh, talk about, uh, how the different components of modern AI systems work. Uh, so let's talk about pre-training, mid-training, and, and post-training, and spend more time on post-training since it's so important. Starting with pre-training, uh, first, at, at a high level, and, and realizing that you may or may not be able to talk about how the things are, are done or what happened in the context of 5.5 specifically. You know, big narrative of last year was that pre-training was hitting a wall, and was not going to yield much progress. That seems to not be the case at all, in 2026. Uh, can you walk us through some, some ideas for what is happening in pre-training and why it's progressing now in a way that people hadn't predicted, uh, last year?
>> For pre-training, I, I can't talk in a lot of details about what is happening internally. Uh, besides that, um, the team has been really doing a lot of good work, um, and our models are really getting better and better. Um, one, one thing that I do want to, um, highlight when we're talking, for example, with efficiency, um, if you have larger models, uh, the amount of thinking time, so the amount of tokens that will think for, um, will usually decrease. And the way that you can think about it is that, um, metaphorically, the model already thinks through its weights when it generates a certain token. Um, so you can, you can decrease the number of like tokens that it needs to generate for thinking by kind of like increasing the size of the model, uh, that you're training. Um, so, so oftentimes, if you just increase the, the model size, if you basically train, pre-train larger models, uh, you will get better efficiency. Um, and the good thing with larger models is that they can be paralyzed better on, on at inference time. So the, even though you might think, okay, you actually generated fewer tokens, but by a larger model, uh, so you actually might decrease the, uh, the overall efficiency of the system. This is not true because the larger the model is, the, the more chances you have to actually, uh, optimize for, optimize basically for inference on, on GPUs. So you will be able to, um, to make them, the overall system, like more efficient. So that's, that's one thing I wanted to say with, like larger models that are actually giving you a lot of efficiency. Um, otherwise, in terms of pre-training, I think it's very interesting. I actually also thought maybe two years ago that pre-training was kind of hitting a wall. Um, and when we see, for example, if we talk just about entropic, I mean, myth is seems like clearly just a much bigger model when you look at the cost, um, the cost of the model. Usually that's how you know, by the way, if it's a, if it's a bigger model, you just look at the cost, um, per token, and, and clearly they're getting very good performance, uh, just by increasing the size of the model. Um, so I think the field was a very, at least part of the field was surprised, uh, about that. There were a lot of conversation about hitting, uh, data walls, and it seems like we did not quite hit it. So the larger the model is, the more data it needs to ingest to be trained. Um, and it seems like different companies kind of found different ways, um, to overcome the fact that that we don't have that much data on the internet.
>> Is the next frontier or the current frontier, uh, for, for data, multimodal data? Is it synthetic data?
>> I think synthetic data can probably work well in a, in a data, um, data limited regime. Um, I think multimodal is an interesting one. Uh, I, I definitely cannot talk about what we do internally, but like, I used to work on multimodal, uh, representation learning back in the days, and I always thought that it would really help, uh, kind of your reasoning abilities if you have a lot of multimodal data. Um, and I still think this, but, but for example, like if you look at entropic models, they tend to not be that good on multimodal, and they are still really smart. Um, so it seems that, uh, it's not as necessary as, at least I would have thought in the past. Um, I still believe that once we go to embodied agents, embodied AI, you will learn a lot about the world. Um, and you will kind of improve general intelligence and usefulness to users, um, by learning how, how the world interacts with itself. Uh, but at least looking, for example, entropic models, it seems that they, they don't need that much multimodal data to have strong models.
>> and by embodied intelligence, so you mean, so potentially robotics, and so if you use a video that shows how gravity works and how a robot evolves in space, then presumably that would be more useful, is that, is that the thought?
>> Yes, the, the idea, the intuition that I think many people had, and I, I definitely felt for a long time, is that it's hard to understand the, uh, only through text, and they will be, um, it's hard to understand what, like what physics is without really seeing what, for example, you can't understand gravity without really seeing things falling. Um, and when you look at our models, I mean, it's, they kind of understand gravity without having seen that, but it still seems not obvious, like it seems still seems like they would get it more, and like they are still kind of missing some common sense aspects. Um, so I do feel like we will improve the common sense of our model by having them interact in the real world. Um, but we are, we're still pretty far from that, I think. And by we, I mean, just generally the academic community and, and the AI community, seems pretty far from that.
>> Yeah. And, uh, while we're on the topic, as a quick detour that that leads us to the concept of world models. So leaving your taking your open hat off, are you, are you bullish on world models?
>> World models in the sense that, um, yes, you can try to replicate or like simulate things, simulate like basically work in an environment that is simulated. Um, yes, the problem is simulations are always going to be really hard and not, not going to be truthful. So I think there will always need to be a certain, a little bit of training that will need to happen in the real world to make sure that the model realizes kind of these mismatches between the simulated world and the real world. Um, and I think, uh, we as a field have a tendency of, uh, optimizing something that is simulated or not quite realistic, um, past the point where this is useful. So that's like something that I think we should always be careful with. Is we spend a lot of, of time and effort on optimizing something simulated or not quite, not quite realistic, and it's great at the beginning, but at some point, once you start optimizing too much for something, it's, it's not representative of the real world, and, and people continue doing that just because that's what they've been doing for a long time. So I just think people need to realize when to stop that. Uh, I don't work with, uh, um, with these type of synthetic environments as much, or just because I don't work on embodied AI, um, so I don't know if we had that yet.
>> Okay, great. All right. So, going back to pre-training, mid-training, post-training. Let's talk about mid-training. It's might be something that people have heard about a bit less. The term comes up a bit less. What is it and why is it important?
>> Mid-training, um, is just this, this idea of something that's between pre-training and, as you might realize from the name, and kind of the post-training, um, part of this, part of the pipeline. Um, and really the idea is, if you have high-quality data, um, that is more representative of what you really want in your final model, um, you should overtrain on that data. Um, so taking a step back here, pre-training, what is it? Pre-training, it's basically trying to learn everything from the world by learning everything from the internet, at a high level. Um, the problem is that most things on the internet are not really useful. If you think, for example, about Wikipedia or like GitHub, which is like coding data, um, it just seems like there's way more information in there than some random forums. Um, um, yeah, some, some random farms that may maybe not like have that much information, like for example, ads. There's also lots of ads on the internet. Like you probably don't want to train too much on that. Um, but in pre-training, we train on everything. And in mid-training, we basically overweight this type of high-quality data that we think is more useful, uh, for, for training the final model. And this is something, I, I can't talk about what's happening everybody, but this is like something that, that is happening definitely in all the academic community right now, and in all the open-source, uh, models have this stage of mid-training.
>> great. Post-training. Let's start, uh, at a high level by, by defining what that is. So there's reinforcement learning, but that's not the only part of post-training. What, what else is there?
>> It kind of depends how you define the term, uh, and where you put the boundaries. In my mind, post-training, um, including, let's, I'll take it from a very broad sense, which includes all the reinforcement learning and like the training for our reasoning models, um, is just the idea of having something that knows everything about the world to making something that is useful to people. Um, so pre-training, I think about it, or the metaphor that I like giving is, uh, you go in the library, and you have a lot of books about everything, and in theory, you can find all the information that you want in the library, but it's much more useful to talk to an expert who has learned these books, and that you can ask questions to, and they can answer, they can answer, and like, they can understand like what you're actually looking for. Um, so this is kind of the goal of, of pushing at a, at a very high level, is like making something that is useful, um, uh, to users, and is like easier to interact with. Um, so there are multiple, uh, stages. I'll talk mostly, well, I'll talk only about things that are happening outside of OpenAI, and kind of the, the usual stages. There's usually some SFT that is happening.
>> which is supervised finetuning.
>> Supervised finetuning. Yes. Uh, supervised fine-tuning, and that's, um, that's actually what early on, most of the, uh, models that were portraying were only doing supervised fine-tuning. This is the idea is that if you have humans that can give you, uh, the desired final answer. Um, so if you have, if, yeah, if you have humans that give you the gold answer, you can basically clone the behavior of the human. H, so this is what we call behavior cloning. The problem with this is that you will never get better than what your ground truth gives you. Uh, and humans are actually pretty limited in many, in many sense. So you will never like overcome, uh, the, the, the human labelers that you, you're working with. Uh, the reinforcement learning, or reinforcement learning stage, goes from behavior cloning to really like optimizing rewards. So the idea is, I don't know what the ground truth is. I don't know what the perfect answer is, but here, here's how I would say whether the answer is correct or not. And here are the things that I, I want in the answer. And what you do is you start optimizing. You start having a model that tries to, to get more reward, basically, optimize more, um, uh, this, this reward function that, that, that's how we call it. Um, and it goes beyond what you currently have, what, what, like humans can do, or at least the humans that you're working with can do. Um, so this, this, I would say, is the two big stages. Then in reinforcement learning, uh, that depends in, like which models are being trained. At least in the open-source community, it seems that there are, there are different ways of doing that. Reinforcement learning when you have verifiable rewards. So, uh, reinforcement learning where it's really easy to say whether something is correct or not, and you can really kind of have a binary reward for this. And that goes back to how we talked about 01, uh, and 01 preview in the past. Um, and then you have reinforcement learning without verifiable rewards, where maybe I could do pair-wise comparisons. I can say this, this answer is better than this, this other one, but I don't really know. I cannot quite say this is the perfect answer. Um, so of course, like, it's a continuum, and there's everything in between, but I would say these are like the three, uh, high-level, uh, things to think about when you think about post-training in general, um, and how people are usually doing it in the open-source world is that they take SFT, uh, they clone the behavior that you, you can collect online or from humans, and then once it's already at a pretty good level, they just do this reinforcement to go beyond what we currently have, because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically, because how, how reinforcement learning works is you sample many times, essentially, from, from the model that you're training, and you say this one is correct, this one is not, and you say do more of the one that is correct. So you have to stumble across the right solution. So you're much better off first getting as much as closer as possible to the best you can do. Um, and this is, this behavior cloning, and then doing reinforcement.
>> Does reinforcement learning create new capabilities, or does it make the model better at existing capabilities?
>> It's really hard to say because retraining, when it's trained on all of the internet, arguably already has all capabilities in it. Um, so it's, it would be even hard to answer this question scientifically. Um, because arguably everything is, is already there. What I would say is that if you look, uh, at models that we were training, or that we're pushing, like two years ago in the open-source world, for example, I worked on one of them, Alpaca, where we used 50,000 examples for SFT, and like now when you look at reinforcement learning from, from models like Kimmy or, or, or from deepseek models, it seems that they are closer to 1 million data points. So definitely people scaled up a lot the reinforcement learning stage. Um, and from this, it seems that they've learned like new capability, like this reasoning aspect, this fact that you can check your answer and, and, uh, and try to improve it. So you can, you can really think for longer to get to get a more correct answer. So all this to say that arguably everything is already in pre-training, but we were definitely able in the last one year and a half, even, even in the open-source world, um, to have more capabilities after reinforcement learning that we used to, uh, before. I heard several times that reinforcement learning is pretty finicky and, and hard to scale, and part of the reason why we, as an industry, didn't do, uh, reinforcement learning as part of the initial kind of LLM sort of progress curve was, was precisely that, that it was hard to, to make work. What is hard about scaling RL? Is that a question of, uh, data sets, knowing where the rewards are? Is there, is that, or something else?
>> I would say most people who did not work in reinforcement learning in the academic and like in research community up to two years ago, probably thought reinforcement learning would like just doesn't work and is like too finicky to, to work with. I used to be that type of person, and actually, when I saw Chad GPT come out, they had the, this blog, I was not at openi at the time. Uh, I saw this blog that says that they use reinforcement learning, and my first thought was, I can do the same without reinforcement learning, because this is just an overcomplicated method. And this is actually the problem that we started working on with Alpaca was exactly, let's try to reproduce that only using SFT, just by doing this behavior cloning. Yeah. And like, for example, Yanuk famously, like gives, like this metaphor of like, oh, the reinforcement is like the cherry on the top. So I think that was really, like the intuition that most people had. Um, it seems that after crossing a certain scale of, um, models that know basically everything about the world, and what we call like good priors about the world, it seems that reinforcement learning just started to work. And this is not only with LMS. Robotics seems to have, or get, seems to be entering the same stage, where they're realizing that actually, it used to be very finicky, but now that we use models that like know already everything about the world, it actually learns pretty well, um, now. To answer your question about what is still complicated with reinforcement learning, one is an infra aspect, um, so just like systems in general, reinforcement learning, you have at a very high level, basically to sample, as I said, for many answers, and say like, what is correct, um, and what, what is not, and, um, and like this sampling is just very expensive, and you have, you have to do it at scale. The other issue that, that also in the open-source world, we, uh, people are seeing right now is that when we are training more agentic, uh, systems, you only know whether you're correct at the end of your very long rollout. Um, so you get very little information per token of whether you were correct or not. And it's hard to say, uh, it's hard to basically do attribution. It's hard to say what part of your entire answer was the one that led you to being correct. So that's more of an issue on the machine learning side. It's, uh, the, the, the ideal world in machine learning is when I can say exactly like, this thing was good, do more of that. And the problem again, with, with, uh, with these agentic systems and, and reinforcement learning, aentic system, is that you don't really know which part was good or not until you arrive at the end. That's another big issue from, from, from re, reinforcement learning.
>> What's the current, uh, frontier of reinforcement learning? It, it seems like there's a jungle of acronyms like GRPO and, uh, other techniques. What, uh, what are you using? What are you excited about? What do you think is promising?
>> So, I can't talk about what we're using, but like, for example, uh, in the open-source world, GRPO seems to be working very well. Um, and people used to have different methods like PO and, and DPO, and like people seem to have really converged to this one. The big, big difference with others, other methods, is that, um, you again, you do this, like simple method that I told you about, like sampling as many answers as possible, and you say which one is correct. Uh, so in some way, GPO is a very simplistic method. Uh, and in general, we saw over and over again in machine learning that the, the simplest method that where you can scale up in terms of compute, usually is the one that ends up working the best, and that is kind of what is happening here, um, at least in the open-source world.
>> As you described some of the challenges, question crossed my mind, you know, you often hear that AI systems are not built, they grown. How would you characterize it as well? What part is science versus a craft, or trying multiple things and then just keeping what works best in your day-to-day life?
>> Yeah, that's, that's a great question. I think how it usually works is that it starts being craft. People just try out many things, and, and they start building a mental model of what works and what doesn't. And over time, we move to like from this, like craft land, uh, to more science. Science, uh, is, or like more scientific approach.
Is are rarely the ones that like first end up working. It's hard. It's very rare that, uh, you take a really scientific approach and say, um, uh, like this is the optimal, the optimal thing to do, and you do it, and it just works. Like, people just, there's some sense of alchemy. People just have, like, a good flare for something, and they make it work. And then other people, or that person, uh, starts trying to improve what we are doing by being very scientific. Um, and I would say this, this happens over and over, um, in, uh, in machine learning. Uh, so first craft, then science, and both are really important, uh, but it's different stages of the pipeline. In terms of engineering, this is definitely something that is, uh, uh, always necessary. Uh, so I would say most researchers have moved to being relatively, uh, good at, like, figure, at least I wouldn't say good engineers, but good at working in, like, complex systems and, like, figuring out what they need to, to, to try out. And the systems, the, and the infra that we have has become more and more complicated. Um, so, so definitely the, the work required changes over time.
"Fascinating. All right. So still in reinforcement learning and circling back to some of the things you said at the at the beginning. So if I want to make my model better at computer use or genetic coding or whatever domain, then I would spend a particular amount of time doing specifically reinforcement learning for computer use and putting together a data set and then, uh, coming up with rewards. Is that, is that how it works? Like, you, you just pick one problem and you just, uh, do reinforcement learning specifically for it?"
To be clear, I, I talk more about reinforcement learning because also this, like, the part I know, I know the best, and this is what I've, I've worked, like, pushing, I've worked on for a long time. Um, we talked about mid-training before. Uh, like, all these things are also extremely important, and you can improve it in different parts of the pipeline. As I said before, the closer you are from the final stage of the model, uh, usually the smaller, um, the, the scale of the training becomes. Um, so you can iterate fast on that because now you can iterate in terms of days rather than iterate in terms of, uh, in terms of months. So usually people start from this, like, fast iteration loop, and then they go deeper and they make, like, bigger changes, uh, across the entire stack. So this is not to say that, um, only, like, reinforcement learning matters. I'm really not saying that. But it's just that, like, that's why people will start doing, uh, changes, and then they will that will permeate, and, uh, we will go deeper into the stack. Um, so this is how it, how it works. Uh, and, and like, in the open source world, it's very much like that too. I think you see way more post-trained models than you see new pre-trained bases. Um, and you see way more, like, improvements in, in, like, the algorithm, and that's why we talked about, I mean, GPO, DPO, PO, like, there are so many XPOs, and that's because people can iterate really quickly on, on this final stage of the pipeline.
"And the, the jagged nature of, uh, those models, does that come from this approach of, uh, picking this problem and that problem, and therefore it's going to be excellent at those problems, but not as good as other problems, or is that a more fundamental characteristic of, uh, AI models?"
There's definitely some of that. Uh, for sure, if you optimize more on specific types of problems, you'll be better in that setting. Um, I would say, at least my intuition is that it's less about the exact, like, problems that you're optimizing on, and it's more about the class of problems that you're optimizing on. So, for example, uh, if you are really good at, like, math competitions, your model will probably be pretty good at, like, coding competitions. So it's not about the domain, it's more about, like, the skills that are necessary and the way to think, um, and, and this, like, horizontal, um, capabilities that you need for performing these tasks. And that's, that's, um, what I think you're usually seeing when some, when some model is really bad at something, it's actually bad at that in any domain, in any language. Uh, so, so you have to think, yeah, about this domain and then, then this generalization of this domain, not necessarily per domain, uh, capability.
"So speaking of generalization, so there's been, uh, that, uh, clear evolution from math and coding success to now, uh, starting to cover different areas. So that's the whole GDP val thing, where, like, across the economy, um, different areas are being evaluated in terms of, like, model performance. The sort of same question is that, is that the result of, uh, overall model progress, or is that a deliberate, okay, now we're going to take, uh, you know, this, uh, part of the economy and build a data set for it and do mid-training and do post-training? How does that, uh, progress work from, uh, those very specific domains to generalizing to the rest of the world?"
It's definitely something that we actively push on. I think people are realizing, I mean, us and also other companies, um, that we are moving towards this world, um, where we want to really make products that are useful and, like, improve, like, productivity of people, um, and, and help people in their day-to-day life. So I think there's a, there's an, a very active move to deciding what are the domains that we should be prioritizing. What are, now that we know we have an algorithm that we can apply in different places. Uh, what we're constrained by is more collecting the right data, having people who really care about a certain problem, uh, work on that problem. Uh, but there are not that many people who can do these things. So you really need to prioritize. Um, so this is, yeah, it's, it's a very active, um, it's a very active, proactive kind of approach here. Um, and in general, I would say the performance of the model really depends on, like, the number of people who care about the final, uh, output of the model, uh, who are looking at that model. So if they start looking more on specific verticals, like these verticals will improve really quickly, but again, we don't have that many of these people that can do these things.
"But to unpack something that you alluded to, I think, a minute ago, do, do models, uh, actually generalize now more, especially from a reinforcement learning perspective? So being making a model very good at domain A or B, then is likely to make the model better at C, regardless of the amount of effort you put into developing, uh, rewards for domain C."
So I think there are different axes of generalization. One, there's an algorithmic generalization, and, and that's, like, really, can I use the algorithm that I developed or this black box that I developed for domain A, and can I use it for domain B? Um, and at least talking about the open source world, it really seems that, like, people are able to do that. They take GPO, they apply it in, like, many different places, and it just works. So that generalization, uh, seems to be relatively, uh, good, uh, which, which is why we're seeing a lot of progress, otherwise it would be hard to make progress. Uh, then there's the generalization of the model that is trained on one particular, uh, data set, and this is what I was alluding to before. Is at least my mental model is, the generalization happens in terms of capability. Like, if the capability is the same, you will see generalization across domains, um, again, like multi, like different languages, like coding, like you can optimize for C++ coding for having a good, like, C++ model, uh, with very little training on C++ partly because this pre-trained model, very little oil in C++ part, partly because this pre-trained model has seen all of C++, and so it already kind of, like, understands the basics of that language. Um, so, so that type of generalization definitely happens. Um, the generalization that I think is harder, uh, are these when we don't have these, like, horizontal, uh, capabilities. So I'll give you one concrete example. If my model is very intelligent in terms of being correct on, like, competitions, I, I usually take that example because it's like somewhat constructive at, like, math competitions, like coding competitions. From a human perspective, people that are good at these things are usually just smart, and if they are smart, or like someone might think that, at least that are just smart, and if they are smart, they can actually do other things too. Um, but that is really not true. And that type of generalization is really not true because, um, many things where we need to have, uh, humans working on, like, expert domains, like the world is very messy, and these coding competitions and math competitions are extremely well-specified, and you need to have this, the capability of, like, understanding, like, underspecified tasks, understanding how to deal with, like, the messy world, and understanding, like, what is the, um, what are even the resources that you need to answer the question? Like, if you look at the, at the math competition, like, you usually have everything in the, in the, in the prompt. It's like, you have five lines or maybe 15 lines, and it's like, all the information that you need to answer this question. In the real world, I, if I'm a consultant, if I work in, like, finance, I need to go on the internet, I need to, like, find and extract different information just to understand, um, before doing any of the reasoning, just to be able to do that reasoning. And, and this type of, like, horizontal capability is the thing that, um, doesn't usually, like, you generalize if you have that horizontal capability, but in many cases, we don't have that horizontal capability. Um, so, yeah, that's why we hallucinate actually in every domain. Like, when you have hallucination of LLMs, if a, if a, if a model is really bad at saying that it doesn't know, that usually happens in every single domain. You won't have, like, one domain where the model is extremely calibrated about its knowledge and another domain where it's not.
"And as a quick detour, is, is hallucination also a reinforcement learning problem where you reward the behavior to say, 'I don't know' when, uh, it occurs?"
John Schulman has a great presentation about that, I think from like one or two years ago, um, where he was saying that if you do behavior, if you do behavior cloning, so this, like, SFT that we talked about before, um, you will be, like, you will basically reward and optimize for hallucination because what will happen, or you could optimize for hallucination because what will happen is, if your model doesn't know about something, but now you say that the right answer is to say that something. So I'll give you, I'll be very concrete. If the model doesn't know about a paper, uh, and now in an answer that you give that is given by a ground truth answer given by a human, you say, uh, "Here's where I got the information," and then you cite that paper. Like, what you're actually optimizing the model to do is citing something that doesn't exist because it doesn't know that that paper exists. Um, and so, so, so John Fman had this, like, great presentation saying, like, SFT is going to force, like, hallucination. While in reinforcement learning, given that, as I said, you kind of sample from the model in the first place, extremely unlikely that you sample something that it doesn't know and it's correct. That's like extremely unlikely. So you will never reward that behavior. You will only sample things that it doesn't know and being incorrect, and then you will kill that, uh, kill that behavior. So, so hallucination, um, at least the, the intuition that people have, um, is that it can come, for example, from from SFT, and it, it can come from this, like, portioning pipeline. But if you have a good reinforcement learning pipeline, that shouldn't happen too often.
"And, uh, going back to, um, generalization as well, is there, are there examples where, um, actually getting better at one domain makes the model worse at, uh, the rest, a little bit, uh, to what you were saying about, like, some people are very good at math, some people are very good at English. Pretty often they're not the same people?"
In domains, usually not. What will happen though is, um, you will make decisions based on which domain we optimize for. And if you optimize for one domain, you will be able to optimize less for another one. So it's not necessarily that optimizing for one thing will make the other one worse. It's just that as a result, you can optimize less for the other one because you're computed, your data constraint, you have, like, like, your, your, uh, human bottleneck also in terms of that work. What does happen is, uh, you can have negative kind of generalization, like bad generalization or negative transfer, more for these horizontal aspects of the model. So I'll give you a very concrete example. Um, explicit instruction following versus implicit instruction following. If I, if I have a model, and this is, we often hear, for example, from OBI models that they tend to be really good if you tell them exactly what you want. Um, but as a result, sometimes we hear also that they're like less good if you were not as, as specific about what you wanted. For example, if I make, if I make a typo and I say, like, "Change this file," and I make a typo in this file, um, an extremely good model at, like, explicit, explicit instruction following will change the wrong file, the one that has a typo. But like, humans would probably realize that you made a typo. Um, and, and, like, as a result, there are cases where this explicit instruction following goes against this, like, implicit instruction following. Um, so you will have cases where, basically, these horizontal, um, capabilities go against each other.
"And maybe to close on this whole, um, reinforcement learning, um, conversation. Uh, so is your sense that as we progress from being excellent at coding and excellent at math and move to the rest of the economy, do you think that the rest of the economy is a tractable problem? Do you think we can get to the same level of performance ultimately?"
Yes. But I will, like, yes, we can. I don't think there's anything like really deeply special about these domains where we cannot optimize, where we couldn't get the same with other domains. The, but is for, for at least two reasons. The first one is most of the people working on these models are pretty good at coding and they really care about coding because that's what they use as day-to-day care drivers. And there's nothing better than the user being also the one who, like, trains the model because, like, then they understand the issues. It's, um, it, it's very hard to really, like, for me, for example, it's very hard to really understand, like, what should we change on the ver, like, on, like, legal, uh, aspects of the model if I don't understand anything about the legal domain. Um, so that's one thing. The other thing that, um, you will often hear about, and I, I mentioned also briefly about before, is this kind of verifiable rewards. There are domains where it's easier to say where something is correct or not. Um, for example, in the case of cyber, like, you, you mentioned that before, that, like, cyber has been improving a lot, cyber capabilities of our models, and this is because in cyber, it's like extremely easy to say, if in, are you correct? Like, did you find, like, is the cyber issue that you find a real issue or not? It's very easy to test it. So there are domains where reinforcement learning is just, like, easier to, um, to apply. But there's nothing, I would say, in the capacity of the model that is constraining the model to be as good at legal and, like, medical, um, and, like, other domains. So it is the, the, the short answer is, we know less about these domains, and, um, definitely there are some domains that are easier to optimize for in reinforcement learning.
"Great. Let's talk about evals for a minute. Uh, that's, uh, a hugely important topic. Maybe to start, why is it so hard to evaluate a model in the first place?"
Evaluation has been harder and harder as models become better. Um, and that's because the tasks that we ask to the model become, uh, more and more general, um, and more and more open-ended. So, like, now I maybe just say, like, "Build me a website that does X." Well, before, in the past, I would just be like, "Hey, like, is there a specific bug in this, in, in this, like, implementation that you have?" And it's like, much easier to say whether there's a bug because I can, I can extract, I can know, I can have a human that says, "Here are all the bugs that you have," and then you can apply that automatically. Um, while the, the website one is very hard to know, uh, what is, like, the optimal answer because there are many good answers. There are many good ways of of building a certain website. This open-ended nature of models really makes evals harder. Um, there's also another issue is that models in specific axes are becoming better than the majority of humans, and so we have fewer and fewer humans that can actually evaluate these models in particular axes. Uh, so that's definitely a constraint. Another one, to be honest, is kind of cultural. Um, most people want to improve the model, and they, they think that the best way to do that is kind of training the model. When in reality, finding issues and, like, making sure that we can quantify improvements is just as important, if not more important. But there's always this, like, cultural gap. Um, that was especially true, I would say, in the academic world up to, like, two years ago, when evals were always fixed, benchmarks were always fixed, and even data sets were kind of always fixed, maybe let's say four years ago, um, and there was, like, a mentality shift of, like, "Okay, data is actually critical." And now there's a lot of people working on data. And I think eval was still not quite there. People don't really fully, everyone knows that it's important, but, like, people don't really understand, like, how impactful it could be to work on evals. Um, so actually, my first, first product at OpenAI, I just came in and I was like, "I want to work on data and evals because I know that this is the thing that no one is is working on, and as a result, I know that's like super impactful to work on that." Um, and, yeah, the tide is shifting, but, like, not fast enough.
"And is the pace of progress in model as a judge and AI evaluating AI, is that, is that moving as fast? Is that a distinct part of research, or is that fundamentally the same idea or the same techniques?"
It's really fundamentally the same method. It's like nothing. Also, most of the things that we do in, in eval, especially now that we have reinforcement learning, could just be applied nearly exactly as is during training. So that's another reason actually why eval are so complicated is that every time you build an eval, you actually build a way to build training data sets. Um, so now you're going to optimize that training data set. Well, not even if it's not that eval, it's going to be the same type of data, and now you're going to do super well because we have this generalization of, of, uh, of capabilities that I was telling you about. You will learn that on that other data set, and now you'll become really good at that eval, and that eval will become obsolete really quickly. Um, so, so that's also an issue with ULS. But, yeah, to come back to your question, um, the model as a judge, it's really important, and I think it's one, one, probably of the most important things because as we get, like, better models, uh, we have this self-reinforcing loop, and we have this, this, like, capability flywheel where better models become better teachers for other models. Um, and this is really important for training. But then you can also do the same thing for evaluation. So, I, a lot of my team works on that, and I think it's really critical is to work on this, uh, module, model as a, as a judge kind of, um, framework.
"Okay, fantastic. Um, all right, so as we get, uh, towards the end of this conversation, I'd love to to zoom out a bit and get your sense for where things, uh, might be heading. Obviously, it's incredibly hard to make predictions on on AI, uh, you know, years out, but let's call it the next 12, 18, maybe 24 months. Is your sense that things are going to continue progressing, or are we heading towards something that could feel more like a discontinuity?"
In terms of progress, as I was said before, it's, I think it's always continuous. Now, the feeling of discontinuity will happen. It did happen three months ago with coding, or four months ago with coding, and I think that will happen now in every other domain. Like, most people are not feeling, uh, the same way, like the, like, kind of the capability of our model and the usefulness of our models, the same way as, like, coding, um, and, like, software engineering is feeling right now. Uh, so this will definitely permeate, I think, through many other verticals. Um, now, in terms of, like, capability bump, in terms of, let's say, the, the verticals that we're already looking at, uh, I think it will be more continuous, um, and they will not be, there will never be, uh, big discontinuities. Like, most of them are always local discontinuities, but you zoom out, and it always just feels pretty smooth. Um, it's not always like this, but, like, that has been the case most of the time, and I can definitely not predict when is the next big discontinuity.
"What is your sentiment on this general concept of, um, accelerating loops in AI? So whether that's continual learning to make, uh, models more current and able to learn faster, to this broader concept of AI building AI, like in an increasingly automated way? Fact versus fiction. And what are you excited about?"
I'm extremely excited about continual learning. I think we haven't quite cracked it. I mean, we have, uh, we have, like, Codex memories, and that, that is helpful, but it's definitely not, like, the end state. Um, I have a friend who always, like, tells me about, uh, again, another type of plot that we should be looking at, which is x-axis time, y-axis utility that you provide to users. And right now, or, like, or, like, usefulness, basically, of the models. And right now, actually, most models at day zero, if you just drop them in a company, um, arguably they are more useful than most new employees. So they start higher at t0. Um, but then across time, they are mostly constant because they don't really learn, kind of, company knowledge. They don't really learn, like, to be more efficient, uh, over time on, on doing the things that they are doing. Uh, while humans learn really quickly. And what is important is, kind of, this integral, um, or, like, kind of the area under the curve, um, of these curves. And as a result, that I think, like, humans are still more useful, um, in many cases. Um, and that's why what we will need is to make, like, continual learning is to make this curve, um, now monotonically increasing over time, and basically make models more and more useful the longer they work in a certain environment. So, I'm extremely excited about it. I'm actually surprised that we're not quite there yet. Uh, three years ago, when ChatGPT came out, I remember I was doing a startup with friends, and we were thinking about working on, on continual learning and, like, personalization and, and, like, memories in general. And we're like, "Ah, OpenAI is going to do that in the next six months. Like, they have all the data. They're going to figure it out, and they have all the users, and the models are going to learn super quickly from users." And, uh, three years later, I don't think we're there yet.
"And quickly, in layman's terms, what is the fundamental difficulty?"
It's a good question. I actually don't quite know, to be completely honest with you. Uh, I don't quite know why it's taking us that long to figure it out. Uh, it's this type of, of domain that I, I think if we really put enough resources behind it, like, we would figure it out. Of course, there's, especially when we talk about, like, this memory inside of a company, there's this big questions about, like, permissions, and there's, like, a lot of, of question about, like, um, privacy and, like, what you can share and what cannot, like, across models, across users, sorry. But for a single user, even for a single user, we're not quite there. And I, I don't quite know why. At least at, at the, at the high level that I can talk about, I don't know why.
"Yeah. What you bring up, um, is, I think, really interesting for AI builders and investors and startups, which is, uh, this, this, this question of, um, the models, uh, getting increasingly smarter within an enterprise, in particular. There's like this whole tension between what the models are able to do and then what a lot of people have built around the model. So, you know, a year or two ago, it was it was RAG. These days, it's all about harnesses for agents. And a lot of people are wondering whether the models are going to end up eating the harness, whether the harness is just a temporary thing. From your perspective, like, where, what do you think, um, happens?"
Yeah. Um, I think harnesses can really improve the capability of a model right now. Um, I think given that we seeing this, this really fast progress in terms of capability, uh, I personally wouldn't push that much on the harness, unless it's like, the harness is, um, is something for, like, very concrete goal that you're trying to achieve right now. Um, so certain companies, like, if they are focused on, like, a specific vertical, they want to go from this, like, 80% maybe reliability to maybe the, like, 85%, and, like, harnesses will give them that, and I think that's, like, very important. But, like, they will, they need to do it while knowing that they will have to retune that harness in the future, and I think that's, that's totally fine. Um, if you try to have, like, a general harness to that will, like, sustain over time, uh, I don't think that will work. The harnesses for specific, uh, domains as a short-term thing that you need to do. I think it's, there will always be so much you can do in harnesses, and if anything, I think everyone should do more of that if they have a specific problem in mind because we're leaving so much on the table without a good harness. Uh, arguably, if we just, I think if we froze the models that we have right now, and you really worked on the harness, and, like, maybe, like, we also spend more time, like, training with, like, a great harness, um, I think people would really feel the AGI in every single domain, or could already feel that in every single domain. Um, but given that we're not freezing it and we're going to continue training better and better models, uh, I think the harness, we don't really understand what the final harness will be, and it's not, and, like, it will always change.
"Same question about applications. So we alluded to, uh, your progress in, in different verticals, and that was, you know, GDB val in general, but also how to bench telecom, which does complex customer service workflows, and then progress against finance agents automating 88.5% of internal investment banking modeling tasks, and then 51.1% on office QA pro. So bit by bit, uh, you're doing, uh, more and more of this. So do you think people should be building, uh, applications, uh, anymore, or is ultimately, as we get closer to AGI, all of this going to be part of the model capabilities?"
There's so much space on pushing for, like, external companies or, like, startups pushing on specific verticals. Um, I think there's so much space for that. Um, the reason why is because, uh, a lot of people kind of think about intelligence in quotations, or, like, kind of, like, raw capability as being, uh, the real bottleneck, but I don't think that's true. I think most of the time, the bottleneck is the, the last mile. It's, like, making sure that the model has access to, like, the right, like, has the right permissions, or, like, has also to access to, like, the, the right connectors and things like this. Um, and we are going to be very focused on this, uh, on this general aspect, and I think there are other companies that should be focused on more the verticals and providing maximum value of what we currently have. Um, so I think there will always be a lot of space, uh, left for this last mile in different, uh, in different verticals, and, uh, I would highly encourage people to continue working on that. And maybe one day when we stop making horizontal progress, which I don't think is anytime soon, maybe we will start focusing on that. But, yeah, that's not what we're doing now.
"Okay, well, that feels like a very, uh, optimistic note, at least for the startup ecosystem, to, uh, end up on. Thank you so much, Yan. This was terrific. Really enjoyed it. Thank you so much for spending time with us."
Great. Thanks, Matt.
"Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode."