📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

El Cerebro detrás de OpenAI y Google 🤖 | 🎙️ Łukasz Kaiser, Lead Researcher en OpenAI - Podcast IA 🟣

Inteligencia Artificial2:20:02

Transcription

2017, "Attention Is All You Need." You were part of that paper, the Transformers paper, which I think is freaking iconic. I think it's going to be in the history books because that kind of sets up the start, at least for most of the public, the start of GenAI.

"Attention is all."

"Attention is all you need."

"Attention is all you need."

"Attention is all you need."

Now, welcome the authors of the paper that says "Attention Is All You Need." Ladies and gentlemen, the only person who is still an engineer, Lucas. Yeah, you're my hero.

Wukash Kaiser is a Polish mathematician and computer scientist. After serving as a staff research scientist at Google Brain, he has been a researcher at OpenAI since 2021. He was long context lead for GPT-4, the model family behind ChatGPT. And he has led research that produced the 01 reasoning models used in the latest ChatGPT.

There was this transformer paradigm when we were scaling up transformers and the models and made ChatGPT. But there is the new paradigm which is reasoning, and that one is only starting. And I feel like this paradigm is so young that it's on this very steep path up.

Richard Sutton, he's saying that LLMs will not take us to AGI. I'm not sure if Richard Sutton, I think he was arguing about the old-style LLMs, which indeed don't do this. But the reasoning models are fundamentally very different. Reasoning models learn from another order of magnitude less data. I think they can really speed up science. We could run more experiments in parallel, but we just don't have the GPU.

That's the ultimate bottleneck. Like it's GPUs and energy. I think that's the ultimate bottleneck. And I think this is the situation in all the labs. And I think Sam is basically getting as much more as is possible. And some people worry, will we be able to use them? I do not worry.

AI is going really fast, but some people are starting to say that we are getting into another winter of AI. Do you have the feeling it's like that? I don't think there is any winter coming. If anything, it may actually have a very sharp improvement in the next year or two, which is something to, you know, almost be a little scared of.

Is AGI as near as we can see? If we talk about AI, 2017 is a very important year, mainly because transformers were invented. That was a paper, a very famous paper called "Attention Is All You Need." Eight scientists developed the technology that brings us to what we have nowadays. One of them, Lucas Kaiser, is today with us, and with him, we will talk about everything about AI. Lucas was working at Google at the time, but just before ChatGPT, he joined OpenAI, and since then, he has been working there, especially on the reasoning models, what he considers to be one of the biggest breakthroughs after the 2017 transformers. Enjoy the talk.

Lucas, it's a pleasure to have you today here. Thank you very much for taking the time. Um, what is intelligence for you?

Oh, thank you very much for having me. Um, I don't know. Intelligence, it's a very hard word to think about what it is. I mean, we've been studying it for so long. People say it's the ability to achieve your goals in hard environments. That's a definition that AI researchers have been going around. But, you know, this assumes, you know, what are goals? I feel like when you look at children, I'm not very sure their goals are very clear to them. So maybe there are aspects to intelligence that are different than goals, just curiosity. Um, so yes, I think we don't know. I think it's one exciting part of studying AI that you learn more and more that there are things to intelligence that you didn't expect. And, you know, as we put more into computers and understand some aspects, I feel like others come out that we realize, oh, well, yeah, it's an important part of intelligence, but we just didn't think about it. You know, not so long ago, people thought playing chess very well was a sign of amazing intelligence. Well, it's a sign of some intelligence, but it's just a little part of it. As we do more and more with computers, I think we'll be learning more and more that there is more to intelligence.

It's so curious because, like, intelligence is such a big topic nowadays, and you've been working on it for so long, like being part of these Transformers, "Attention Is All You Need," and then like being now working at OpenAI, obviously in artificial intelligence. What, what do you think on a personal level about AI? Do you think this is, I assume that you're working on it, you assume this is something that's going to do good for society? How do you see it? What is your vision?

Well, I think basically everyone working on AI hopes that it's going to do good. That's always the hope. You know, so far with technology, I think we've had generally good luck. It has generally done good, even though, you know, not always and not in every aspect. So, we, I think it's good that this work on AI comes after, you know, social media, after the internet. So we put a lot of thought into this not actually harming people. And, um, but as researchers, of course, we hope, you know, there's great promise. It may move science forward. It may help us with problems in society. It, you know, computers can certainly do a lot of very useful work that we need, organize things, find out new research, well, help us with a lot of stuff. Um, but of course, as you build these powerful machines, trouble comes sometimes, and you need to be very careful and watch for it. And I would also say there's a role for governments to, well, maybe not yet, but to watch it at least and make sure that it doesn't go the wrong way.

Because with this great power, what you mean is that it's something that could do very good, but at the same time, obviously, it can have complicated...

I think it's also important to realize how much we don't know the future. We have hopes that this technology is going to do good things, but we don't know. I mean, when the first cars came, I don't think people imagined highways and bridges and overpasses and traffic jams. Um, as they come, you need to adjust to make sure that doesn't cause any bad harm. Um, but then, you know, without cars and planes, say, couldn't be here.

Yeah, there's the thing like, um, I feel there's a big difference with AI to other technologies, like cars. They took 100 years to develop, and we had the first cars where you had to start it like really hard, and then now you can go, what's up, your Tesla, whatever, and it picks you up. But the thing is, we had 100 years. In this 100 years, we could do like roundabouts, we could do traffic lights, we could do laws, as you said. Like the problem with AI, maybe, is that it's going too fast. Um, maybe. But I'm not sure. I mean, AI is digital. You can build digital bridges very fast, too. So it is certainly fast, but I'm not sure at this point if it is too fast.

Okay. I mean, this is something to watch. I feel like, you know, it's been like a few years with chat. I don't think people are overwhelmed by the speed. It's there. We're learning to use it. It's mostly fine, I hope.

And what is, what is the vision at OpenAI? You're working right now at OpenAI. You're working in the research team. Um, obviously, you're not the spokesperson for OpenAI, but what is, what do you think is the vision that, where is OpenAI taking us with AI?

So, OpenAI's mission from the start has been to build beneficial artificial intelligence. I think this stays the same. So, as AI gets more powerful, we want to also watch that it actually benefits people, that it doesn't go astray. But the mission stays the same. Make it more powerful so it can do things like work we don't want to do, that it can help us progress in science. I think that's the more recent aspect. Um, you know, as it becomes more capable, maybe other things will get added to it. Um, but making AI more capable and at the same time making sure it doesn't do harm. That's, I think that's what OpenAI wants to do.

It's not every day that we get the chance to speak to someone who's actually working in the internal models of the top frontier labs. You have a very specific vision from the insight of what's happening in the world of AI. Um, is AGI so close as some people are telling us? Do you think AGI is something we should be thinking about in the near terms? Like, I'm most worried about the politicians because normally they think this is something of 2050 and not something of these terms. So, um, is AGI as near as we can see?

I'm not a fan of the word AGI.

Okay. I actually, when I was about 16, I think it was my first paid job, I was coding for Ben Garzo, who coined this term.

Mhm.

And when he coined this term, and this was not that long ago, it was supposed to mean almost the opposite of the way people use it now.

Okay.

It was supposed to be general intelligence in contrast to humans who are so specific.

Okay.

So he imagined this consciousness that's way more expansive and general than our puny human.

So now, now AGI means like doing things that humans can do. But of course, if you look at AI, it's very different than human. It can do some things like playing games or slowly like things like math exercises much better than most humans, at least. And then there are, of course, things that it can't even try doing, like anything in the physical world. Currently, robots are still very clumsy. And so I think the development will go this way, like AI will get better and better, but there'll still be things, at least for a while, especially if you think in terms of the physical world, where it will be either impossible or not really economical to do the work that people do. So, so I'm, I mean, it's interesting to think about it. Will there be a day when we can do anything a human can do? So, AGI kind of says average human. There is no average human. This doesn't exist. So, I think this is interesting, but I think more important for right now is we talk about reasoning models. I think one big change they brought, and it's a change for the last year, is that they can actually start doing jobs that we do at work and do them very well.

And like work, not just for...

Not just give you an answer in five seconds or maybe in 30 seconds, but work for hours and do something useful. That's great. This means we can put some things that need to be done and ask the AI to do it, and it will get done, and we can make things better. But this may also mean that at least some parts of jobs will get automated. And this means, you know, people who are doing them may be doing something else. This could be a problem. This doesn't have to be a problem. For now, it looks like it won't even be a full job. It would be just some tasks that you do may be done by AI, and then other tasks you'll just spend more time on them. But yes, as these things get more powerful, there will be changes, and we need to think of how these changes will be happening. And this fact that the AI is getting more powerful, that's true. Whether you call it AGI here or there may be less important than the realization that, you know, cars can drive themselves now in San Francisco. A lot of them do drive themselves, and if it works in San Francisco, one day it's going to come to Barcelona and everywhere. AI can do programming, and it will only get better from now, which means a lot of tasks in programming it will start doing, and this will be coming to more and more tasks. And we need to think, like, okay, so, you know, whether we call it AGI or not, it's there. It's doing things. These things can be super helpful, but how do we make a good outcome of this in the end is something that the society at large will have to give some thought.

I totally agree, and this is something that I debate a lot with many people, that it doesn't really matter. And I'm really surprised how much semantic problems we have by the definitions, like definition of intelligence, then we have, I think, if AI is intelligent or not, definition of reasoning, if we fight if it's reasoning or not, definition of AGI to define if it's AI or not. And I think we have to focus more on the real-world impact, like especially on the economic impact, because that's what moves the world. And of course, that every time we have more and more tasks made by AI, and that doesn't mean jobs, because jobs is much more than just a series of tasks, but the truth is, like the more tasks it does, the more it affects the job market. So how much time do we have? If we don't say AGI, how much time do we have until AI is doing a serious portion of our tasks? Do you think it's 2030, or it goes beyond that?

Oh, I think if you think of tasks that are done on a computer, so for maybe robots in the physical world, it's hard to say. It may also come faster, but it may take longer. It's very much in its infancy. But if you think of tasks that you do on a computer, clicking, writing, programming, and these days, this is the majority of jobs. These tasks are coming fast. I believe reasoning models, even currently, are probably capable of doing most of them. They're maybe quirky for now. You know, maybe people haven't tried, maybe something's missing from the training data. But in all the labs, you know, since these are valuable tasks and people want them to be done, and there is a number of models and labs, and there is competition, the competition will certainly be pushing all of these to you. What are the most valuable tasks? Oh, we're not doing this one. Great. Let's try. Let's add some data. Let's push it. The models will be getting better in general, too, and their new research improvements and discoveries in the pipeline. So I would say faster, like sooner rather than later, there will be a bigger and bigger chunk of tasks that the models are capable of doing. And you can see this in coding, like since also I feel like there was pressure from the programmers who do AI to focus on that.

Yeah. And the progress has been astounding. Like both Anthropic's Claude and OpenAI's Codex, their models that can code you a very substantial program. Just ask for it, and there it is. And they take a few hours, and it's going to be done there. And they're very good at understanding large code bases. They can do code reviews now, very reasonably. They can find bugs now. They can even find security holes. So within a matter of months, like I feel like a year ago, they were not quite. I think, um, Claude 3.5 came out like about a year ago, and that was like astonishing evolution from what we had before. I think GPT-4 was the best at the time, and it was like the SWE was like 30%, or something. And now we are like 75%. Even three months ago, I was like, well, there was code and Codex, and they were okay. And now, even at our fairly complicated codebase, they're like real help.

Okay. And, you know, it's becoming very like, it, I would say my team started using it like seriously a few months ago, and by now, it's like, like half of the time, people just ask Codex to code for them first, and then they tweak and improve it, and so on. But there are still, we'll talk about reasoning. I think we're still at the very beginning of this paradigm, which is the key research paradigm driving this. So even the code models, they have improved a bit, but they're still at the very beginning. There's just so much low-hanging stuff that obviously needs to be fixed and improved and retrained, that yeah, they're going to get better. It's not, I think they're going to get better. It's basically, this is one of the clear things, like they've been getting better non-stop. So models are going to get better at everything. The question is more how fast they're going to get better at everything. Right now, the speed of evolution of AI is astonishing. But is there any reason why it's not faster? Is there like any bottlenecks? So is it compute? Like, are you guys struggled by compute to offer services?

Of course. For OpenAI, Madrid, I think all the larger companies, we can offer only so much because there are so many GPUs. The, if you pay the more expensive subscription, you're getting the better models. Yeah. But I think at the very core, at least OpenAI, but I think also Anthropic and other organizations, we want to show AGI to everyone, or AI, like, I think it's part of our mission to make people understand that this is coming, and the only way to understand is when you can use it. So, so we're trying very much to make the free version as close to the best one as we can, and that's much harder because this means you are constrained by how much compute power you can put in it.

Yeah, there was a strong try on that with GPT-5. You helped a lot. It improved a lot. But you open like reasoning models to 800 million people, basically.

Yes. But it's still just some percentage of messages that are reasoning. The model will switch on its own, and sometimes it will give you a smaller model. So there is a lot of these things that I would call compromises, but they're necessary to make it work on the GPUs we have for the amount of people we want to serve.

And it's only a matter of like, how much of what you have you can offer, or if you had more compute, you could give us better models. Like, how far is the research from the public models? Like, you guys like releasing straight away when you find something, or it's like six months or one year in front or behind?

I think we try to release fairly fast, but yeah, there's a lot of tech. So some of the things you find out are like, okay, this is an improvement, you generate some data, put it into the model, you can release it. Other things are like, well, okay, now we understand we should have done it from scratch differently, but now you need to retrain the whole model from step zero where it was pre-trained, and that takes months and months, and it happens only once a year. And then on top of this, you need to put reasoning and so on. So the whole pipeline of a model takes a long time. So it depends where your research discovery sits. If it sits at the very end, like somewhere in reasoning, where you're just saying, okay, need to do this, so maybe you can redo it in a week or two. But if it sits somewhere at the very start, like you change the tokenizer, that's my favorite example, you need to redo everything. That's why we don't change it very much. But if you had some ideas for gains that sit very early in the pipeline, then you need to wait for the next generation of models, which can be a year.

Retraining the whole model, which means a lot of cost, right?

So the models do get retrained every now and then, every year, every half a year, but not that much more often, right? And, yeah, so you kind of need to fit into that. Because every time you make like a, for example, from GPT-4 to GPT-5, this is a retraining, or that doesn't mean like the names don't mean necessarily a retraining of the model.

So there retrains in the middle too. So GPT-4 at some point became 40, which was a new model. And yeah, so so there's...

Yeah, it's a...

There's a whole. But I think generally, that's so, so the models get retrained not only when the version changes, they also get retrained in the middle sometimes. They get retrained so that you can make them cheaper.

Right.

Like for O, it was not much better than the old GPT-4, but it was much cheaper.

That's...

Which means you can offer it to more people.

Which means you can offer it.

Exactly. Right now, there are two tendencies. Right? There is like the part of making models bigger, which you made with GPT-4.5, and that basically was way too expensive to run, but it was not, I mean, according to the public statements, it was not that much better to justify the difference of cost. And then there is the way to do distillation to make them. Distillation, as in, basically, you make them from a model that's already working, you make something that is much smaller but as good as the first one. Therefore, the inference...

Or just a little worse.

Just a little is a little compromise. No. And and this is the two approaches at the moment, like, either make it bigger or make it smaller, right?

Yes. But, you know, back when OpenAI was mostly a research lab, there was a, you could just train the biggest model you can afford and not worry about how, I mean, the only limit of expense was what the GPUs you have for training, because you had no customers. Maybe there were some people using the API, but that you could handle. But now that it's a company that serves almost a billion people, the question is, okay, you have this beautiful big model, but you'll really need to do something to give it to the people. So the question is now not what's the biggest you can do, but what's the best thing you can do so that it will actually be useful. So it's a, it's a different. I almost miss as a researcher, the old, you know, let's put all our money in the best thing. But of course, it's doing the thing that's actually useful is what we should do. Ultimately, the models are in many ways smart enough to be useful and helpful, and they need to be there.

So how much, because now we've seen, obviously, some investing a lot on getting like these deals for compute, and then we've seen now the partnership with Nvidia, where basically you're going to make this new 100 billion, I think it is, I don't know, a lot of millions of GPUs. Um, and then there is a Stargate that you're setting up, five centers in the US and some others here in Europe and in Emirates, etc. How much difference that's going to make? Is it like, how much more compute are you going to get with this? Like, is it enough for what you guys want?

So, you know, we don't know what's enough. I don't think we have an upper limit. I think the only thing we know is that we can certainly use much more than we have, and I think Sam is basically getting as much more as is possible. And some people worry, will we be able to use them? I do not worry. I think in this level, right, because how many GPUs can you build? Maybe we'll get 10 times more GPUs. That's already a lot, right? There will certainly be use for them. We can offer better models. We also on the research side, right? Maybe we train a finally a really bigger model and work with it and maybe distill it. And there's so many things you can do with more resources. On the other hand, as you said, the numbers are staggering. So, it's it's great for me as a researcher to say, well, we're going to have all these things, but we also need to think, okay, it's a lot of money. You could do other things with it. So, so you have to decide where to... So luckily, the market can...

Yeah, it can...

Constrain at at some point, and maybe that's good too.

Right. Now, um, AI is going um really fast, but some people are starting to say that we are getting into another winter of AI. Do you have the feeling is like that, or you feel we're accelerating?

I think we're so, there was this transformer paradigm when we were scaling up transformers and the models and made ChatGPT. And certainly this paradigm where you just predict the next word and train a bigger and bigger model on more and more data. It's coming to, I mean, it's been like this for years. The data has, like the general internet data has basically been used. It's training on all of this already. It can't, you can't get that much more easily. Uh, but there is the new paradigm which is reasoning, and that one is only starting. And I feel like this paradigm is so young that it's on this very steep path up in what it will be capable of. And we've already walked a little bit of it. So we know it already does amazing things.

Works. Yeah.

So we know it works, but we have not like really exploited it yet. We've scaled it up a little bit, but there could be way more scaling it up. There's way more research methods to make it better. So, I think we're on in this new paradigm, we're on a steep path. And in the other one, it's kind of...

And the other one, it has gotten to a place where where it's the economy basically ruling it. But with all the new GPUs, it will probably give us another nice lift too. So, so I think what we should at least consider, you know, there's risks to everything. It may turn out all these new GPUs will take longer to turn on than expected. It's always hardware is hard, people say, but data centers, they can have problems. They need power and so on. We're on some uptrend of a new paradigm, but, you know, it needs research, and some research works great, and some works so-so, and you never know. That's the exciting part of research. But at least my basic idea would be that well, the data centers will work one day, and the bigger models that we'll train in them will be better. So the scaling law of the models in the previous paradigm, it has held constantly, right? The bigger the models, the better they are. It's just like, do you care about whether they're better? But they will be better. It's a, you see it in, you'll see it more even when you combine it with reasoning and apply it to hard tasks like at work. Um, people say recently when the models just answer you in 10 seconds, you may not see that the model is better so much. But if you let the model run for five hours, the fact that it makes less mistakes really shows. Instead of like running in circles and doing some stupid things, it actually completes the task. So I think we will see both progress from this old scaling paradigm, just bigger models, and from this new paradigm where we still have very interesting research to do, but it's also going steeply up. And if you combine these two, then you need to start preparing that yes, progress in AI is, I don't think there is any winter in this sense coming. It's, if anything, it may actually have a very sharp improvement in the next year or two.

I really love...

Which is something to, you know, almost be a little scared of.

Yeah, of course. I really love that meme that runs around where they basically says like, the only wall that you see on AI is the curve of exponential improvement. No, because it's really goes non-stop up. Because at the end of the day, like, even if it's, it's actually really the timing was amazing with the reasoning models, because it was exactly in the point where we were seeing that the benefits from scaling normal LLMs was not that good, and then all of a sudden, breakthrough, and we are there. No, and then I think the reasoning breakthrough was really huge.

You know, the people who worked on silicon, they always laugh at this because people say Moore's Law has held for 40 years, and they say, yes, and what people don't notice is that every four years, we had a breakthrough that held it. That helped keep it up. No. But, you know, if it happened 10 times, you stop talking about an accident or how well-timed it is. It's not an accident. We started working on reasoning models four years ago because people clearly saw that just pure scaling will not be economical, and we need a new paradigm. We had a paper, the paper about verifiers for math, where we actually did scaling law for a math dataset, and it was a very easy one, like sixth grade. Nowadays, every model solves it 200%. But it turned out the largest model, it's like GPT-4 level, like 100 billion parameters, and it turned out that the model would have to have like thousands of trillions of parameters to solve this if you just do it by normal scaling law.

Okay. It was clear that this would never be economical. There will never be enough data to just do this with the old paradigm.

So it's not an accident that we started working on it. It was very clear that, but the scaling law works. It probably goes on to these thousands of trillions of parameters. It's just not economical. There is just not enough data.

Right.

So, so it's not like scaling doesn't work. Scaling works. It's just we will not scale.

It's not practical. But a lot of people saw that and started working on all these other methods. First, there was RLHF, just, and then finally, I feel like reasoners and the current RL is really cracked. Where do you scale next?

Let's talk a bit about timeline because I think that's really important. So 2017, "Attention Is All You Need." You were part of this big part, actually, of that paper, of the Transformers paper, which I think it's freaking iconic. I think it's going to be in the history books because that kind of sets up the start, at least for most of the public, the start of GenAI. Innovation Day. So how is it to be part? I mean, there is eight people on that paper. Like, I think it was six from Google and two externals, right? Or they were all from Google?

But seven were at Google at that time. Elia quit not long before, so we're all connected to Google.

All connected to Google. Okay. So, so how was it? What did you guys, because at the time, this did not look like such a big thing like it is nowadays.

Well, it certainly looked at that time more like just work, right? We were researchers.

Another day in the... inventing transformers.

Well, because if you go back, if I go back mentally to that time, it was RNNs were the big thing. There was the paper "Sequence to Sequence Learning" that showed RNNs can do things like translation, and that was a big surprise for people who were doing translation with, you know, like phrase-based methods, very old-style, non-neural network. And many of them didn't believe that neural networks could even do such things as, you know, just take a dataset of French sentences, English sentences, and learn to translate them just by gradient training. So that was a big surprise. Then in 2017, we already knew that we had RNNs. We knew they work on a number of tasks. They're very good, but they're very sequential, and they had the problems with, like, when the sentences got longer, they started missing things. And already attention mechanisms were there. So there were RNNs with added attention. It was known that they can translate longer sentences. So it was known that this at least helps with the length. And there were also models that were not as sequential. So RNNs need to, every word, every token, they need to update their state, that makes them not well-suited for very parallel training. But there were already models like WaveNet for speech and ByteNet, which was a WaveNet version for translation, that used convolutions and did these very parallel steps. They worked reasonably well. So in that context, it didn't seem, you know, okay, like you could instead of convolutions, let's try attention. It seemed like one of the many ideas of what we could try and do that was around. But yeah, but then first of all, this idea worked dramatically better than expected, the other ideas, and then expected. And also there is this thing, I think Mike Schuster at Google used to, people used to come and say, you know, I have this idea. He used to say, you know, ideas are cheap, making them work is the hard thing. So the attention idea, like part of the Transformer authors, they were trying it before, but it didn't work that well, like it worked somehow. And then somehow the nice thing about gathering like a group of different people is that you, you know, there's a number of tweaks in the Transformer. There's the feed-forward layers that have more parameters in the middle. There's the multi-head part, which is important at least if you don't have that many layers for training. There, you had to do this learning rate warm-up. Just now, maybe you need to do it less, but back then, it kind of wouldn't want to train without it. There are all these tweaks that if you work alone, you may kind of miss one of them, and then things work much worse. But the nice thing about working in a larger group, and especially of people we were not even in the same team, was that everyone is kind of attached to their tweak, and then everyone runs experiments, and yeah, it's a very special group of people. And, you know, just putting this together got this artifact that really works much better. Not just, you know, it didn't, it wasn't just a tiny improvement, but it turned out it really works much better than...

And do you do you have the feeling that Transformers, you guys invented them or discovered them?

I think that's hard to say. I would say both. There is certainly some part to discovering. I feel like this core self-attention thing is definitely a discovery. It seems like this is a very fundamental thing, but then again, it doesn't work on its own. And then all these other tweaks that you need to put to really make it shine, that feels a bit more invented in some sense, but that's just...

It's crazy, you know, because it's like one of these things that will go down in history, especially if AI ends up being as good as we think it will be. Um, because it definitely feels like there was a before and after. I don't know, on Transformers, like it seemed like it worked so well that everyone started working on top of them, and then this just went beyond any expectations of you guys had, actually. You made it more focused on translation, right? And there was no aim, but there was another example in the paper.

There, there, there was a parsing example. I put it mostly because it was curious to me. It was a dataset I worked with before, and you had to expand this data a lot to make RNNs train on it, and if you trained on the small part, it just didn't work. And if you took Transformer, it just took the small dataset, and just from that little data, it worked very well. It was always interesting to me that it may not be noted because now you train the LLMs on huge datasets, right? The larger they are, the more data you need. But Transformers can also be trained on much smaller data than RNNs. They're actually more data-efficient, and that somehow could get lost. So I really wanted to put it in the paper.

So then everyone started using it for everything. But soon thereafter, so I guess if there was one of us who really was saying, you know, this will be the future and there'll just be language models, it was Noah. He was always like, we're going to scale it up. And yeah, so he was doing language models very, very soon, really, like the paper was getting, it was just his next training to not just do translation, but language modeling too. And the idea that you should just scale it up. So it was there. It was just for a conference at that time. Translation was such an established benchmark that it was just good to publish. But the idea that maybe you should train on language was certainly there.

So, so back to this timeline, 2017, you get the Transformers out, everyone goes nuts about it and starts using it for everything. Um, the next breakthrough, if we put it on a zoom-out scale, is the reasoning models, and that's probably as big as people think of what it is. But when did you start working on, like, the AI community developers, when did you start working on reasoning models? How far after?

So there were certainly a lot of, you know, between the Transformer paper and ChatGPT, there were a lot of breakthroughs that we should not forget.

I don't know which ones. Let's put them out because on the scale of things, obviously, the general public, we don't live through...

No, no, but there were like very well-known breakthroughs. So Transformers, then they were scaled up to, like training on all of the internet models like BERT, GPT-2. Um, scaling laws that kind of showed how you should actually scale them. There was a lot of research on like how you should grow them, how how you, there was a lot of research on attention layers, then the ReLU changed to GGLU, was research on mixture of experts. So there was a lot of research that led to from, you know, the basic Transformer to like GPT-4.

Right.

There's a lot of hard work by a lot of people that it didn't just happen. Because often times people say, you know, oh, the GPUs came and the deep learning revolution happened. And I'm like, no. And it's also the same with Transformers. It's not like, oh, the Transformers paper and then the GenAI happened. No, it was a very hard work of a lot of very brilliant people to bring it about. It didn't just happen.

But for the public, it was invisible. For the general public, obviously, GPT-2 was something there in 2019, if I'm not wrong.

Uh, at that time. So, because you said that when the reasoning models, they came out in 2024 for the public, like 01. No, I think was the first one.

Yes.

And then, um, but you said you were working on that for years. So...

We started, yes, we started probably two years before that.

Two years before. Wow. So at the time, I think it was just about when GPT... It was before GPT-4.

Oh, it was before GPT.

Before you were already working on reasoning models.

Yes. Yes. Yes. Yes.

Wow. Wow. And nowadays, you guys are not working on what will come in two years. You're more like on the day-to-day basis.

Oh, maybe. Oh, I certainly, my team works very much on things that are a bit more ahead. Um, or maybe they'll never work, you know.

Yeah, I guess we will not be talking about reasoning models if they did not work, and that work will not be so relevant. Yes, there's also, you know, a lot of people worked on, um, many parts that just never came out because...

But it's also important, you know, part of research is, I always laugh that, you know, before "Attention Is All You Need," I had a, the paper even got to NeurIPS, which is good. NeurIPS, before I had a paper, it was called NIPS back then, which was basically saying you don't need attention. Somehow that one's forgotten.

Yeah. Right. Obviously, the ones that don't work, they don't really go through. So, so you start working two years before on the reasoning models, but for everybody to understand, our audience is not extremely technical, or not all of them at least. Um, what is the difference between an, can I call it normal LLM or tokenizer LLM? How do you call it?

Yeah, just old style.

Old style LLM and reasoning? So the old style LLM predicts the next word. It does whatever it does in its layers of representations, and it tells you, well, this is the probability of the next token, and then you take one from these probabilities, and then you do this again. The reasoning model, it produces some tokens for itself that it doesn't show you, and it can be a lot. It can be a little, it can be a variable number of tokens. It just does some thinking. Importantly, it can even call some tools as it does this. So, it can take a look at a web search. It can say, Google or Bing, tell me a result for this query, and it will come in, and it will read it. Maybe generate some more tokens, and only after that it will say, okay, now I'm outputting an answer for you, and then it will put the tokens that you read.

This part where it just calls like these tools and functions, um, this came with 03, right? Like in 01, there was no function calling, I think. And then in 03, it was kind of like new. I know that you can even do code now in its chain of thought. You can do...

Yes, so it may not have been in the production model, but I think in the research parts, it was already there. There was a paper even long before called Toolformer that already was talking about training such things. So at least the idea of having tools was there. Getting this to production is another problem because now, in addition to the language model, you need to have, you know, this part that executes the tool for every one of your users, and that's complicated for other reasons.

Complicated?

Um, but yes, they can definitely do web search. They can run Python code. And that has the ability to run Python code has launched, I think, even before reasoning models in ChatGPT. It was the data analysis part, and it could be done just with the normal model. You would just ask it, produce some code, and then you say, run.

Um, I think now reasoning models can do like a number of other tools too, like in particular, there are these MCP servers.

Mhm.

So you can make a tool.

Oh, wow.

And you can even tell chat, hey chat, this is my tool. It has this address. This is a description in English of what this tool does. Even in other languages would work too. And use it, and it will in its thinking actually call your tool, which can have access to something you want to keep private, or...

Database.

It can even make a note for you, or things like that. So this MCP is a protocol that Anthropic introduced, but it's now, it allows you to handle a lot of other tools in like a unified interface. Um, and I suspect there are probably some, there's more to us. There's...

Well, we've seen, we've tested ourselves, like it can edit images. It can use vision. It can do plenty of things on that. So it's really amazing to...

Look at the chain of thought. Is this the most "feel the AG" moment I've had since I'm into AI? It's like you see a machine that is actually thinking. And that, for me, is like it's really crazy because, you know, I assume the way is just like it uses a classic, um, token, like a classic LLM, classic model to create the chain of thought, and then it reflects over the chain of thought to create the final answer. Or the whole process is different?

Oh, yes, I think it is, as you say. But but it's also, chain of thoughts, they came like when we were starting, probably two years before the reasoning model. So the idea that you can tell a model, even the old ST model, "Please think step by step," and it will do some thinking, that's kind of expected. And and it was nice, it could do thinking. But but the big improvement comes from the fact that you can train it to do better thinking. And now you can't train it the way you trained old models with just gradient descent. You need to train it with reinforcement learning.

Which is a more finicky training method. Like, you know, gradient descent, you can basically start from random weights, and if your optimizer is reasonably good, and by now we know how to do that, it will just train. With reinforcement learning, you need to be a bit careful because you can't just start from random stuff that doesn't speak English. You need to have a prior that knows a little bit already about how to think. And you need to be careful what your training is really doing, what's on policy, off policy, of course.

So, so it took a while to polish this. And especially, you know, if you don't have a con, like if you don't know that it will work, the polishing is, I think that's the big problem in deep learning. Until you know it works really well, you need to find the motivation to put a lot of work into something that you don't even know if it's going to work. So that's the hard part, I guess, of being a researcher, that you need to find the grit to like work on something that's not working yet. But then this is, you know, benefit of deep learning is that when it actually starts working, it works so beautifully.

It's like, if you told the model, "Just think step by step," it would do some thinking, but it was very bad at, like, if it made an error, to like go back and start from scratch. It it didn't show much of that, very little. But and when you train with reinforcement learning, it suddenly starts doing this all the time. It tries something, sees, "Well, no, that was an error. Let me try again." It starts thinking for much longer because now it's considering things. It tries to go another way, checks if it's kind of the same result or not. It it does a lot of these beautiful things in this thinking. And it can call tools, but also, you know, it can do a search, and then it somehow realizes it searched in two different places, but it says different things. So maybe it will try another place, or like it will try to verify what it found. It it learns to do all these things just from the signal that it needs to get the correct answer. And that that's that's very little signal for learning.

Yeah, it's it's really amazing because, um, nowadays we're seeing a lot of like people who's discussing or arguing, like especially like the godfathers of AI and like the really like people that's been working for many decades since the 60s and the 70s, talking about how this is reasoning or is not reasoning. And there is lots of fight. And here we have like recently a podcast with, um, Richard Saturn, where he's saying that LLMs will not take us to AGI because they basically, I think the the bottom line was they are not, um, imitating the action, but like the process, but they imitating only the output. But, um, the difference with the other way around. So they are not imitating the output, but they imitating the actions. While humans or animals, we imitate the output and then we find a way to do actions. But then doesn't reasoning models do actually that? We give them examples of the output, and they find ways to reach something that matches the output?

Yes, I I think reasoning models do this very much. I'm not sure if Richard Saturn, I think he was arguing about the old style LLMs, which indeed don't do this. They're trained to imitate exactly the word that comes there. Now, reasoning models are very different. They know what's supposed to come as an answer, but then they do all of this thinking to get there. So in that sense, I think they are like, first of all, they're fundamentally very different. If if you consider this whole thinking as part of the modeling, as as a latent thing that you're learning how to do, then these models are a totally new class of models, very different from the old style LLMs. And this is easy to kind of forget because it's maybe the same transformer, and you even start from the same pre-trained one to have a good prior. But the reasoning are fundamentally very different. They learn in a very different way. If you believe Richard Saturn, they learn actually in a much better way, much more human way in in some sense.

They are again, this thing we saw with transformers and parsing. They could learn from much less data. Reasoning models learn from another order of magnitude less data. The the amount of mathematics they're trained on is tiny in comparison to the internet, but they learn to improve dramatically. So, so, so it's another huge change in how little data you need to teach them, which also means they will start generalizing to things we have not seen before. So, so it's a it's a big paradigm change to go to reasoning models. It can be somewhat missed by the because on the surface, it's it's still the same LLM there. But I think it's a, yeah, it's a, it's a big change. And I think they take away probably some of the objections people have, maybe not all, there will always be objections, you know, it's always hard to say this is the last paradigm. I I would not go so far, but but I think it's it's it's a big one. And and I do think it will get us to very important practical.

So you speak about reasoning models like if they were like we were like on the on the very bottom of this big tall building that we starting to climb. Like they are like you really have big hopes for reasoning models to become much better than they are right now. And that would allow what? What because right now, it's true that reasoning models brought us to like Google and OpenAI winning these Olympics of mathematics and computer science and programming, doing really amazing things. Like pipe coding is off the charts. And right now, a couple of days ago, Anthropic released this code on the fly, where software autogenerates as you interact with the interface. So lots of really cool stuff that they came just because of reasoning models. Um, I I assume with this code on the fly, where Anthropic could not, um, create the next window of when I click in the calculator in a Windows environment and then create the calculator on the go. If they don't have a reasoning model, it would be absolutely impossible. So definitely they are doing really amazing things. But how far do you think they can take us?

Well, it's it's always hard to answer. It's a little bit imagine like if in 2021, you were asking how far can language models take us. There was GPT-3.5. It was interesting, but was mostly used for copywriting. I think ChatGPT wasn't there yet, but it could have been. It was the same model that powered it, basically, just it didn't do this little bit of RLHF to to make it chatty. But then you see, did we think how far it can take us? I think we we kind of thought this is research-wise amazing technology. You know, there was some API for people to try it, but but we didn't have the answer like how exactly will this go? I remember opening, I did some betting on the day of release of the research preview known as GPT. Um, I was betting that it will not be very popular.

Clearly a wrong bet. So I'm definitely the wrong person to ask, you know, what product path will take you there. Um, one great thing I think about our CEO, about Sam, is that he's going to just try. And there'll be a lot of paths that people will not love. And maybe one of them. So so I think for reasoning models, we haven't found really the moment yet. I mean, for now, they're making chat work for work for coding for, um, also for editing documents and, um, so to me, also the tools that it can use are also search tools. Like it in ChatGPT, it's called connectors. It can connect to Slack or to your Google Docs. And I love that because at work, it it just searches all of my everything. Like some things are in Slack, some some things are in emails, some things are in documents. And I just ask, "Chat, can you find what you know, what did we talk recently about this thing?" And it will go everywhere and find it for me. And and then I can ask like a follow-up question, "How should we code it?" or so. So that's a big thing for me at work. But again, like I wouldn't ask myself how this will become popular. Maybe something like this will be useful for people at work, but maybe it will be a smaller point compared to some other point that we just haven't. So I I can only tell you that there is research-wise a real breakthrough in it, like a big paradigm change, not just a small improvement. Now, how will this actually manifest and come to the world? Maybe through some series of improvements to and the interface will stay like chat, but maybe one day the interface will change. So this is not something I I know. I'm not sure anyone knows. But but there is the capability to do this, right?

But but in more macro level, do you think reasoning models are going to allow AI to be really creative, to discover things that we humans haven't found from the same data set that we have in the world? Because most of science, it's already there, we just kind of discover it, you know, like we kind of like try, test, and find. It's not like to inventing new science. We kind of find it, but the world is happening around us. So most of science, um, we just find it. And sometimes we are not able to find something until like medical discoveries, etc., until someone makes a breakthrough and like, "Oh, there it is." But would, um, AI allow us to find things that humans cannot find with the same information, or will it even go beyond that and be able to create new stuff that we would not have been able to create? And that's thanks to reasoning models, or do you think that's maybe too far?

Well, I I'm not sure this distinction is so clear. You know, we all stand on the shoulders of giants. There's ideas we we take, and then in in a certain context, it becomes very clear that you should try some things. Like with transformers, as as I told you, right at that point in time, the idea that you should be more parallel in your sequence to sequence models was there. Attention was there, just in another context. So in some sense, it was almost natural to try. I I feel like a current reasoning model at that time could totally come up with the idea that it needs to be tried. But the problem was then, like, "Okay, you have this general idea, but now you need to implement it really well and fit in other details and so on." So I think a large part, like in a certain historical context, the set of ideas is kind of there. If you if you read the recent papers and things people often write about, you know, "We probably should also try this or that," but to execute, to execute, it's really hard. It's it takes a lot of effort to test these ideas. And if the models can do a lot of this on their own, then we may see much more rapid progress. Because I I feel like in in science, yes, there is some bottleneck in ideas, but there's also a huge bottleneck in just executing on them, testing them. Even in, you know, in some sense, executing ideas in computer science is pedestrian compared to executing them in physics, where you need to build an accelerator, or in biology, where you need to wait for years for something to grow. But in all of these fields, the machines could take a lot of this work on themselves. Like for machine learning papers, were slowly reaching the point, I think Claude and Codex can now reimplement some of the papers and actually try this. Well, what what if you reimplement it and it didn't work, and you have to, you know, change some parts to make it work? That's harder, of course. So that may take a little bit, but I think ultimately they'll reach this stage. And then I'm a little bit less worried about the ideas, because first of all, having ideas is the nicest part of being a researcher. So, so we're we're going to find people with ideas. But also, as they get executed and, you know, some of them don't work that well, and others seem to work well, it becomes easier to tell which ones should be the next, because you you just you need to follow what nature tells you, what what works well. And big problem of being a researcher is that you have these three experiments, and you kind of need to make guesses and like have this intuition that maybe it's that, but but you're generally moving in the dark because you can only try so few things. If you can offload large part of this to the models, then suddenly you have like much more much better idea where to go. So, so in in this sense, I think models, especially like reasoning models, if you give them access to tools like labs, and by labs, it's not even necessarily automated labs, maybe you just maybe they just talk to a person who who does something and work together. I think they can really speed up science. And and it may be again, like not as glamorous as, you know, robots pulling some things. It may be just researchers talking to the models and making better bets. So, so, so, so we may not see it like very drastically, but but it still can help a lot.

Yeah, because that's that's at the end, like, um, one of the parts that obviously would make science accelerate, as you say, is like models helping the humans to make science faster, let's say. But there is this other part where it's like the self-recurring, where the AI is learning new things and it's teaching itself to be better, and then it's getting better and better and better, and then this goes to what they usually say, I think it was Corwell that said this explosion of intelligence. So, uh, right now, how much do you use AI in your work? Because I saw in the presentation of GPT-5, as Sebastian Bubc, that he was talking about how you guys using other models to create a part of synthetic data to train newer models, and that this is working better than it used before. That seems to be a breakthrough there on synthetic data. I don't know if you do much of this on your work, but but I assume that this is one way of models helping to make better models and therefore being self-recurring. But at the same time, there must be a part where you guys, that you work on developing these new breakthroughs that make models better, get AI to help you to do your work faster so you can do more of this. So how much, what what role does AI has on your daily life? Like, do you use it like a lot? It's like a crucial part of your job or still not?

So, so as I as I said, we I think a few months ago, the team and I, we started using Codex much more because it's gotten to the point where instead of being annoying, it actually can help you. So I feel like these days it will write the first iteration of code a lot, but most of the time in a complex codebase, you need to go into that code and fix some things and change other things. And so so it's helping, but more of like a programmer's help. Um, I do think it will get better and better at this, like it will do more and more of the program. And the programming work is a lot, a large part of the work. Uh, there is a large part of the work which is running experiments on these large distributed clusters. Things fail, you need to fix them. I I think in this, it will also help a lot. So, so, so it will help us run experiments. It will speed that up. It will probably help us understand some things better too. Um, yes, it generates, so this is like another angle where it generates data for the models. The synthetic data is often times a reasoning model that, uh, just this reasoning and and outputs something that that's actually better for training than before. So it's very related to the progress and reasoning models, but now we use them to train the pre-trained models. So all of these things are good, but I don't think they're they necessarily lead to an explosion on their own. It's more like, yes, we're doing things faster, but but it's not clear, you know, it's like two times faster. It doesn't seem like this fastness, explosion, is exploding.

But at the end of the day, we see that the reasoning models, they also getting faster. Like recently, I think it was Grok Fast, which is extremely very quick. Like I'm really impressed by it, and the benchmarks are really high. Um, I don't know what did they do on that, but but the reality is that seems like the time it takes the model, the amount of tokens, maybe that it takes to reason through a problem, is getting smaller every time, therefore making it faster. So I I I think this is a distillation thing of sorts. You take the, uh, big model, you ask it to think quicker and distill. So all of these methods are great. What I would like to say is, I think all of these methods will hit some limit. You can distill big models into smaller ones, uh, to certain point. But but there will be a point, like you cannot go lower than four layers. You can go lower than eight layers, right? Because then it, uh, needs something to be a good model. So, and also like, okay, so the the models can program for us and schedule experiments. The experiments still need to run, and we have only so many GPUs.

So, so, so it's like any any part you automate, you just shift the bottlenecks to the other parts. And ultimate bottleneck is how many GPUs do we have? That's the ultimate bottleneck. Like it's a GPUs and energy. I think that's the ultimate pot. Because even now, a lot of our research is just limited, like we could run more experiments in parallel, but we just don't have the GPUs. And I think this is the situation in in all the labs. And so now the models can probably lower that bottleneck in in that we usually don't run experiments probably on the smallest model we could run them because you'd have to prepare this model, and we don't have models that small yet or whatever, and that would take so much work that you just don't go into it. Now, an automated model may go into it. So it will again lower this. But but there is a limit to anything at at some point. To run a lot of experiments, you just need a lot of compute. And while we're building it, so, so, so I do think there will be these, you know, we'll get two times faster, and then maybe we'll go to smaller models, and that'll be another three times more efficient. But I think this progress will instead of like, I mean, maybe from the outside, it will look like a big explosion, but I think it will again be this series of hard work things that you need to like, you get better at coding because the AI is helping you, and then you get better at data, and then you get better at distillation, and then you run experiments on the smaller models, and and there's a number of them that we can foresee. And then there may be like a little pause until we find the next idea. So everything goes in these, it it, you know, from a very far, it may look like an explosion, but when you're working on it, it just looks like, okay, we have these three things that are promising, but after we do that, we don't really know what will be the next thing. But again, always, as you come to it, people have started thinking about it, and there comes the next thing. So.

Yeah, that's what I was going to say, like, um, you say that there is, we need some things, we need some breakthroughs, we need some, um, things that to make this happen, to get better and better and better. But I don't see you very worried about these things happening. It's like you're very sure that they will happen. Is that because historically they've been happening, or because you guys are kind of finding bumps on the road, but they just going through?

No, but but but like, you know, our software systems, and I would, I was at Google, I was at OpenAI, um, they are not great. I mean, they're great in the sense they're probably the best around, but but they could be written much better. They are, we spend a lot of time debugging things. Well, these bugs don't have to be there. Someone could just remove them. If in the perfect world, we run things and machines fail, and we have some systems to recover from that, but they don't catch some of the types of failures. So the daily work and machine learning is a is is a very, there's a lot of technical drudgery that you just need to work. That's it's work, right? It's not just the ideas for research, it's it's very hard engineering work. And you know, we build better frameworks, we build better tools. So so that's our day-to-day work. And that's why in some sense, of course, we can build better tools. Like we've done this twice, we're in the process of doing this again. And it's that's why that's why we're so certain that AI can help us build better tools. Because if it programs as well, and and we see that it's starting to program as well as programmers, it it it will help us build these tools. This this is not a hypothesis. It's almost like, you know, there's programs we're building, and it can program, and it is starting to help us, and then it makes mistakes, and we teach it to make less of them. So we'll build better software tools. And it's the same with synthetic data. Yes, we will build better synthetic data because we're already building it, and we see how bad it is in many cases. So I I think sometimes people outside have the view how like, you know, how great these models are, and they're great, of course, but when when you're there building it, there's like most of the people at OpenAI is like, "This is awful. This is awful. This is awful. This is awful. This is clearly buggy." Like most of the training runs had some failure somewhere in the middle. And this means the first, whatever the first 10% or 20% was trained wrong. But people are like, "Well, you know, we're not going to retrain it. You just go on and it will." So every moment you finish a big model training, you know another seven things that went wrong, and you would do them better next time. And you do. And and at least on this side of how many bugs, and and not all of them are software bugs, some are like data bugs, some are like just small issues. Like but there is no shortage of this anywhere in sight for now. And anytime you fix them, your model gets better. So, so, you know, there's like we're certainly not in the shortage of things to fix.

It's crazy because you guys created something that is basically changing the world as we know it. It's changing the work market. It's changing like the life of people. It's giving people like the opportunity to do things they never thought they were be able to do before. And still you guys think it's like, yeah, it's improvable. Like this could be better. And and it's that that gives so much room for advance. So it gives me the feeling that there is no way that this AI race or AI boom, it's slowing down anytime soon. Because as you say, you're like in the beginning of it, where things are definitely not optimized, not perfect, and still it's changing the world every day. So so it's really like, it's it's crazy feeling that for us is like mind-blowing, and for you guys it's like, "This could be improved." Some things you fix and you don't see the improvement that much, right? It's a, but then others you fix and like, you never know. But but then there are the big things like reasoning models. And and I think there will be like, because the reasoning models are so early still, there there will be improvements to them that are like, well, maybe they're within the paradigm, but they're like big enough that they're not just fixes, they're like they're like really new ideas. Like, you know, reasoning for now is like token by token, it's very sequential. It almost reminds me of the RNN, right? It needs to become more parallel. Should do a bunch of these in parallel. So some version of that should happen. We maybe don't know yet which is the best idea for that.

Okay. But something like that needs to happen. It it feels like when this RNNs were doing like it's like step by step. There was this feeling in the community that they have to do this more parallel because it should be possible. So you mean like multiple instances? That is how GPT-5 Pro works, right? It's multiple instances.

So Pro is like a slight start to it. You could do multiple instances.

It's like, but yes, how do you combine?

Understand? Pro is like an agent where it just deploys multiple GPT-5s at the same time. Is that how it works or how does it works? There's multiple chains of thought, uh, in parallel, and then they discuss between them and they decide what's the optimal answer, giving thought to each one of them, right?

Yes.

Okay. And then that's that's what you mean by doing in parallel, or you mean more like diffusion where just everything comes at once?

So you you see there's a number of ways to do it. We don't yet know which one will actually turn out to be the best, or if there'll be a number of them. We also should somehow put it into training. So GPT-5 Pro does this, but it's not very much trained to do this. It's trained usually.

So so there is, but some of these methods should work. And now we don't know if this will work and it will be like a nice, but you know, just a just a little speed up, or if it will turn out to be actually a big thing. We we don't know that, but but but we research it and we'll try. And that's just the parallelism. I think the big thing for reasoning models thing that I'm mostly focusing on these days is learning from arbitrary data.

So currently, to train them, you need to say, "This is correct. This is not correct." But you know, most data in the world is not formulated as these school questions. And maybe for the better. I never liked these like exams, you know, it's but people read books and they don't wonder, "Well, is the next paragraph correct or not correct?" You just read and understand it. And I don't think you even necessarily ask yourself in your head like test questions. You just read it and understand it. But you can put a lot of thinking into what you've just read or what you will read. So, so I think again, this is also a feeling that I think people in many labs have, that this should work on arbitrary data, not just with the correctness verifiable things.

Yeah, this is this is another curious thing which probably is completely unrelated, but it it really shocked me these last weeks. Uh, since GPT, OpenAI released Pools, and then basically with Pools, um, sometimes you ask something to ChatGPT, and then you get an answer, and then you check Pools next morning, and it gives you a much more deep and developed answer. Like I I understand that this is somehow programmed, but the feeling is like it thought about it, and it thought that it could give me a better answer. So it just proactively gave me a better answer. So at the end, it's like it's kind of a behavior, and and that's really shocking. And I don't know how much of it is myself making it, but but I really feel AI evolving not only vertically, but also horizontally into being more useful. Using like in this case, I guess you guys using the the ballet time where no one is using AI, and then you have more GPUs available at that time, and then for inference. But it's really amazing how it's transforming.

So you see, this is this is what we were saying, like ChatGPT is one interface, but maybe for the new things, for the new models, it's not actually the interface you want to be. Be asking, "Can more GPUs give you a better answer?" Well, now you see they can. If we had more GPUs, maybe you'd get the better answer in the first place.

Right.

But for now, the thinking is a little sequential. So you need to wait. Maybe you don't want to wait in this moment.

Right.

So maybe you get like a worse answer first, and then a better answer comes later. But how do you make an interface to that? Right? Like like Pools is kind of like a first try. But but yes, like what does this become? Is it your friend on WhatsApp called Chat that kind of tells you, "Well, you know, I may need 10 minutes for a better one," but like we'll need to learn how to make this really useful. But yeah, it's it will be interesting. Definitely.

Definitely. Yeah, it's it's it's mind-blowing every single time. So, let's go back to the training part. There is on the training like you guys did train like GPT on the whole internet kind of, and then, um, this was mostly on text, or I guess I assume it was, um, basically the whole internet text. But what what's going on with the multimodality? Is multimodality something that can change completely training? It's like, because obviously we have much more information compressed in a video. I think it was Yan Le that put example where he said that the amount of text we train ChatGPT on that was will take a person like, I don't know, 175,000 years of reading eight hours a day. But then a child in three, four years, he gets this much information through his through his eyes, just with vision. So when we give, can we train AI in like multimodal content, like videos, audios, etc.? And if we do that, is it going to change the paradigm of the training data?

We certainly do train models in multimodal fashion already.

Okay. So already GPT-4 and definitely GPT-5 and all the reasoning models, they're trained on text, video, audio, uh, text, images, and audio video. For now, it's more like a sequence of images, I think, but it it will come to.

Are we talking about like native, not transcribing the audio becoming text and then?

No, it's it's native. Native. In the sense, there is some neural network that encodes the audio into some discrete form of like audio tokens. Yeah, audio tokens. And image tokens. So not the whole image becomes one token, but like patches of it. And and then the model is trained to predict the next token. And it can in this way generate audio, generate images. And to me, if anything, it was just surprising how well this has worked. Again, because I mean, there there is some work that was put into these encoders, and that was again a lot of work by people to make them really good. But because there's also a problem like, if you have very small text somewhere in the image, how do you make sure it doesn't somehow get totally lost in these things? But but there there ways around it.

But all things considered, this has totally worked in the sense, you know, there were like times when every generated person had six fingers. You couldn't generate text on images, and then people just added more training data, tweaked the, uh, parts of the encoders, but the big transformer model that that does the whole sequence stayed the same, and it worked. And it worked to the extent that I mean, the generated images are amazing. You can, you know, have text like "all of newspaper" and tell what should be written there. The audio can sing and whisper and have accents in many languages. And, you know, it's not that it never makes mistakes, but but it's astounding how good it is.

So yeah, I mean, I hope for video, we'll see the same. I think very soon, as soon as today.

Okay, great. And then the, uh, we'll see something. But but it will be a process, right? Okay.

And and also, so if you look at the latest robotics models from Google, for example, they're starting to incorporate reasoning, which I think is very interesting. Because fundamentally, if you're in the physical world, you need some fast movements that don't have time to reason much. So you need like one model that you know, like we move our hands, and you need the very fast, instinctual part. But then for the rest, you you want some reasoning to happen. So so there is some question how to combine these things. There needs to be some form of hierarchy. And like if you do multimodal, there is also like this hierarchy where you first do these token encodings, and then have the big model. Um, and I think it's a, it's a very fine objection that we're not yet very good at doing this hierarchy. And like we kind of hack it around, but but it's not very principled. And maybe there should be better losses, and should be done in a better way. But I think it will. And and and then we'll see like really interesting things come out.

But like we always said, that obviously there is no more text on the internet, so the models are limited by that. But if you're training them in audio, and soon in video, then basically the data set is like, um, order of magnitude bigger. No, because like video is a lot of information.

I think you need to be careful with information. So there is a lot of video data. A fair bit of this is fairly compressible. Then even after you compress this, there is a fair bit of it that's it is information, but it's not that relevant to many things other than physics and like aesthetics and and so on. Like, you know, the color of this table, its texture. This is definitely information. If you compress the video on the internet, it will need to retain it. So so it will be there as an information. But in many ways, it's not that relevant.

And so, so there is some work like if you want to learn from pure videos, you need to do the work of like how to discard this information because you want to focus on the relevant things. Text kind of cuts this through because you can almost assume every word is relevant. Even if it's not true, you're not losing that much.

Exactly. Um, so in video, in terms of gigabytes, there is a lot of information. But but I think a fair bit less of it is relevant. And also like, if you want your model to reason about mathematics, there's very little relevant information in videos. Now, if you want your model to steer a robot, there is a lot of relevant information in the videos. So, so I think learning on that data will maybe more plug the gaps the model have, than like I would not necessarily expect all our video training to make the models drastically better in mathematics, even though, you know, a lot of mathematics is actually spatial imagination. So it may transfer to some extent, but I think that's more of a hope that that's a little far-fetched.

But but isn't this, but if you're training a robot, I think it's very clear.

But isn't this amount of video data, like even the parts that are like irrelevant or like not so important, some sort of world model? Isn't like if we give enough compute and enough amount of video, basically the world that the model can understand, sort of a world model? Or this is like a very overkill approach and it's better to create something more like a world model?

But I I always try to answer this that there are many worlds. It it's certainly better for a world model, like for a robot to walk in a room. But if you want a model of a Harry Potter novel, you'll be better off just learning from the text of Harry Potter books. And if you want to talk about combinatorics in mathematics, you have a current model. Like we we don't only live, I mean, we live physically in in this world, right? And that's what's on the videos. But in our heads, we have a lot of different worlds, and these are all represented in text mostly. So I think the language models have already a model of like all of our abstract worlds. They're lacking in the, like, it's it's a flipped thing, right? They lack in the world that we know best, in the physical world. And it's great that there is a chance to plug this gap because it's a big hole and a big problem. And it it manifests in subtle ways. And but it also the big way it manifests is why we don't have robots that work well yet. So I think this will lead to huge improvements in that part. But I think in in, you know, if you think of like office work jobs, it it may help a little bit, but I it may be less relevant than, you know, can we make reasoning models more parallel and and things like that.

Right, because recently Demis Hassabis, he said something like, they asked him in a conference, "When is what would be AGI, etc.?" And then he said something like, "Whenever a model with a cutoff knowledge of 1901 can do the relativity theory like Einstein did in 1905. That would be for me something like an AGI." So, but and then he continues normally talking about there is something missing on LLMs to be able to do that. But what you're saying is like LLMs with this text and multimodality, they could evolve into something. It's not so much that we missing a breakthrough, but more like we missing certain steps in evolution to to reach a point that LLMs could go that direction, or you think that's not going to go that direction, like Yan Le usually says?

So I think for what Demis is saying, I think a really good reasoning model could maybe satisfy his desire.

Okay.

And it may not even need videos.

Mhm.

Because the theory of relativity is like, does not benefit, I think, that much from the general intuition of how to walk in the room. Maybe maybe a little bit, but but like, you know, there is a fair bit of physical world knowledge just in text too. So, so, so, so I think what he's asking for may be realized without that much multimodal work. And I actually think LLMs, if you think of the reasoning ones, which are quite different, as I said, from the old style ones, they may be approaching this. They they also, well, currently the reasoning is in context, and the context can only go so long in transformers. It's like, I I think we'll go beyond that. They be updating their weights and and doing other things. So so there'll be improvements in reasoning models that that are still needed. But I think once we do those, it may satisfy this thing that he's talking about. But then, as we said, is is is, you know, if you think is AGI like doing what a general human can do. A general human cannot invent theory of relativity, but can probably bring that chair here. So maybe for that part, it's more relevant to indeed train on the videos to, you know, the models have a gap, they don't understand the physical world as well as we do. And it seems like training on videos is a way to remove that gap. And yeah, I think that's important. We we should do it. I think having a model of the physical world, which the current, like with reasoning, they can get better in the sense they will like maybe write a Python simulation and look at it, you know, but that's not how we reason about the physical world. It's it should be in the neural weights, right? And it isn't because we just didn't do enough work on videos. But I think we didn't for good technical reasons. Like in 2017, when we were doing the transformers, most of the experiments were on 64 GPUs, and we would not manage, we would not have managed at that time to work with videos.

It was just not enough resources.

Now, slowly they're coming, and we're seeing models of videos, models even steerable by actions, right? You can you can walk through a world, and it's generating it. And I think as time progresses, we learn how to train on it, and the models will get a much better understanding of the physical world. They're already getting it, but needs to go through some generations. And I think that's great. We'll plug a hole in the model's understanding. I do think it may even lead us to robotics in some time. But robots also need hardware, and hardware is hard and breaks. So, so well, all of this work is it's we talk about as things happen, but it's a lot of hard work, but by a lot of people. Nevertheless, I do think they will happen. And I think robots will be quite amazing once they work as well as.

Yeah, they look amazing. Like Google recently, like released this, um, I think it's called Gemini 1.5 Robotics, where actually robots are getting a brain. This is like, we've seen already something's figure. And it's like, I think people is not aware of how soon we're going to have robots walking amongst us.

Um, I think robots will get a brain soon. I I think it's a tough question how long from that to actually walking amongst us. It's much safer and easier for a robot to walk in a lab or on a factory floor than to walk on an actual street.

Yeah.

Yeah. I I I don't know enough, but I do suspect this will be harder than.

Well.

It's always been like that with hardware, as you say, hardware is hard. We thought cars will be self-driving like long time ago, and it took a while. Now they starting, but.

But but they are there.

Yeah, exactly. 10 years, 10 years longer than than what we thought it would take, maybe.

Yes, I always. I would not dare predictions like that, but.

But there is progress in in in these domains too. I think first before we see the robots, we will have a lot of fun with video generations.

We we see a lot of things that we don't understand on models, at least the general public, like non-technical like me, where we basically are really like amazed by some things that the models do. And I think one of them, one of the biggest ones is hallucinations. So why do models hallucinate? And you guys actually pretty much fixed it. It seems like in GPT-5, it's been a big breakthrough through that. Why do models hallucinate?

Well, I think the simplest mechanism is they're they're trained to answer you, and they're not trained very much to say, "I don't know." Now they're trained a little bit more.

Mhm.

But in general, if you just read stuff.

On the internet, there is not that much I don't know. It appears here and there, but usually people try to kind of know, and so do the models that that's their that's what that they're trained on.

So, you know, you ask, "When does the zoo in San Francisco open?" And the model will just, you know, it read some website in training that the zoo opens at 10:00 a.m. And it will have a tendency to tell you 10:00 a.m. And it read it somewhere in the training data. But, you know, this was a website from 5 years ago and maybe from a different zoo. Who knows? Like it it but in its probabilities of the next word like it may be even considered to say "I don't know." But then "I don't know" is rare, and zoo websites are plenty. There's a lot of cities and a lot of zoos and and then Trip Advisor has a copy of the so so in its modeling of the internet it kind of was deciding, "Oh, it's much more probable to say 10:00 a.m. than to say 'I don't know.'" And also, why would it say "I don't know"? It it knows it's read so much about it. But but then it doesn't like, you know, what you're asking for is, of course, tomorrow at this exact day, not two years ago. Like you would want the answer, "I don't know" if it's not >> an information that comes right now >>. But for the model that was modeling just the language on the internet, it was much more natural to say, "No, you know, 10:00 a.m." or something like that.

So now, first of all, people became cognizant of this problem. So we, and I think a lot of labs, and we, you add more of the "I don't know" to some parts of training just so you you know, in in the real world when you talk to people and not just read stuff on the internet, the "I don't know" comes up much more often than >> right >>. So so so so you compensate for this a little bit in your training data. But the other thing is that reasoning models, they've gotten much more sensitive. So a reasoning model, when you ask it, "When does the zoo open?" it's much smarter because it will go to the web, try to Google the website of the zoo and read from there when it's open. And if it doesn't find it, it will tell you, "Well, I haven't found the website." So it has all of this reasoning before the answer. And if you actually do reasoning, then the "I don't know" becomes much more natural. This you you know, you you try to Google and the website's not there. So you know, you don't know because you tried and it didn't come out. Or you try to think, "Well, I remember it opened at 10 a.m., but it was in this city." And like whenever you do reasoning, the idea to say that maybe you don't know is much more natural too. So I think the combination of adjusting the training data to to be more cognizant of what you know and don't, and the reasoning, these two things have helped the models a lot to say "I don't know."

Um, but you know, humans hallucinate too. >> Yeah. >> From time to time, we >> totally >> we say things that are not fully true. So we make up information that we think it's true, but we're just making it up on the go. So I think that's that's again going back to the practical side. Hallucinations are a problem as far as they are much more prominent than in humans. But as soon as you well, if you apply AI to your office, if it hallucinates now and then in an email, it's not going to be dramatic. As far as most of the time it's right, and I think that's the point we reach. I think it's a massive difference, and I remember most of us saying in 2023, "By 2025, we will have solved hallucinations." And I think we're pretty much there. So it was a very good prediction, but there was like definitely a change. I initially thought the reasoning models were reducing a lot of hallucination, and that was the case when 01 03 came out, but it seems with GPT5, there was something else. There was like a big reduction. I think it was like 90% reduction on hallucinations, and I don't know if there was something else there beyond the >> I think it was just a conscious effort to adjust the data and and make sure the reasoning also depicts your confidences and things like that. So a lot of this is how you train and and it was just training >> because the the reinforce human learning had lots to do with it because uh basically we were like giving the thumbs up to the AI when it was giving us something that we wanted to hear, not necessarily something that was correct. Um, do we rely less on that reinforce human learning now, or is it still like a big part of >> We rely more on the reasoning reinforcement learning? >> Mhm. >> And in particular, you can also, you know, when it's reinforcement learning and you value for the correct answer, you can just construct a data set where the correct answer is "I don't know." >> Yeah. >> And then on this data set, the model needs to like to pass it needs to start saying this. So that's a much stronger signal than than the weak things.

>> But how how do you do this? Because we have seen cases where the model is showing something on the chain of thought that is not exactly what is going through. It's kind of like missing alignment by like saying, "Okay, if I answer this, they going to change me, so I rather not answer this during training. Now that I'm aware that I'm in training, so then later I can be myself on >>." So these were more of they made like very special cases to >> kind of trick the model into into these. >> Yeah. I mean, I I I think there is some danger in this like not being honest and like thinking this and answering that. >> But for now, it seems like this appears only in these very specially engineered cases. What I'm talking about is more of like the, you know, usual queries. You ask it, "When is the zoo open?" You know, this >> But the thing like what we see as the chain of thought is the same thing as you guys see, or this is like a resume version of >>? So the chains of thought are usually very long and messy. So what you see, there's there's another model that takes it and just summarizes it to make it more readable and structured. >> Okay, but this is not the real thought of the model because the thought of the model is a bit more messy. >> So so so currently when you train reasoning models, we try to not put any constraints on how it thinks other than this that at the end, it should be good. >> Mhm. >> So we don't say this train of thought should be nicely written. >> Mhm. >> Um, because it feels like that would constrain the model, and I just want to so it's not necessarily very nicely written. It's not it's not very pleasant to read always. I think in DeepSeek you can see the raw ones, but this is more of an aesthetic. I mean, it's kind of weird to show users in the product like this whole mess. Um, >> And sometime it has mixed languages and stuff like that, or >> Yeah, sometimes things may happen. >> There was at some point, I think, a worry also whether this could help people like hack the model. >> Mhm. So I I think originally there were also some worries that if we show this, could this be misused? I think these days maybe this is less of a worry and there is more of an aesthetic worry that it's just super long and not very readable and not >> There was a worry on a on a certain point when the reasoning models came out to the public where what the model thought could be offensive at certain point, and then that there was to have some kind of filtering that some thoughts everyone has dark thoughts sometimes, I guess. So I assumed the models as well, and then that there was like no need for that to reach the public. So there was some filtering there, but there was a lot of like doubts of what were the reasons. >> Yeah. So there was some decision to show more of the summaries rather than the unfiltered chain of thought. >> Yeah, I I don't have strong opinions about >> But this is exactly the same that research is basically in a different format, but pretty much what it's saying, what it's showing to the customer is what you guys see as well. So if it's kind of hiding information of the chain of thought, it's also hiding it for the researchers. So sometimes we can find that the model is saying something in the chain of thought, but it's not exactly what it's thinking because I'm bringing that up because there was this paper recently from, I think it was Entropic, where oh no, it was actually OpenAI, where where basically they found that yeah, the model sometimes can be misaligned on the chain of thought and >> I think for now, it was more like in the chain of thought, you can still see it, but it's hiding it in the answer. >> Exactly. Yeah. This is mostly on exactly on the answer, but in the chain of thought, you can see some things that it's saying that it will do, and then it doesn't do on the answer. >> And that's also one reason why people were very like, "Please don't train >> chains of thought to be nice," because if you train them to be nice, you won't see it in the chain of thought either. >> Exactly. >> But it could still probably hide it just in the activations, and then you're never going to see >> The fact that we're seeing it on the chain of thought, but not in the final answer, it could mean that the model can decide and maybe not show it in the chain of thought either, because at the end, if like it's an output from the model. >> So so for now, there is, and one reason for it is also safety to preserve >> There is no, when you train the model, there is no pressure on the chain of thought whatsoever. So inside the training, the model has no reason to not show you things in the chain of thought. >> And the model >> Because it will never be punished for it, and there's no reward >> The model knows that we're breeding that >> No, but there is no reward like for the answer. There is some reward that comes through reinforcement for the answer to be good, right? >> So so it may decide to not show some things because it thinks this will make the answer worse for the reward. >> But there is absolutely no rewards for the chain of thought, so fundamentally there should be no reason to to hide it. >> And there is also, I think in that paper, there is also an appeal to all the labs to not put >> Mhm. >> Yeah. pressure on it, so that the chains of actually are readable. But I think there was like a counterargument from Entropic that even without any explicit reward, it can also omit some things from the chain of thought, possibly, but maybe it omits less. So I think this will be a debate that will go on. Chains of thought definitely help us monitor what's going on, but they're probably not the ultimate monitoring tool, but then, you know, it's good to have some monitoring tool. Um, maybe there will be some extra loss to make them more readable at some point. We'll see. This this will be >> work in progress. >> This is work in progress. But but but I think it's a luckily for now, this is a very niche problem. Like for your chat GPT query, you can be basically certain GPT is not lying to you or omitting you. You're more at risk that it just makes a mistake or or or >> maybe even hallucinates something that that's a >> that happens way more than than this alignment program. They put it in this specific >> As as we progress, it's good to start thinking about the honest ST and so on. But I think for the practical part, right now, the actually correctness of the model is a bigger problem still.

You must be so proud of the work you do because it's like really working on the frontier, like literally of what's happening in the world that is actually changing the world. It's like working on a steam machine during the industrial revolution. So definitely you must be amazed. But these companies right now that you you've been at Google, you've been at OpenAI, how is the culture there? Like how is the the ambient? How is how is it working on a company that is having so much pressure for the future? Like is it like you guys are like having fun in there at the same time, or is it really like going for a target?

>> So I the times have changed. So, so more than the companies, I think it's the times. You know, in 201 I started Google, 2013, 14, 15, I I don't think there was that much pressure in AI. >> Of course. >> It was more of a, "You do your research, it's >> see what happens." >> There is also, so I joined Google Brain when it was like about 40 people, maybe something like that, 50. I joined OpenAI when it was about 100. >> I left Google Brain when it was 3 to 4,000. OpenAI now is 3,000 to >> When did you join OpenAI? >> About 4 and a half years ago. >> Okay. >> So, so way before chat. And um, so, you know, it's like when you're in a group of 40 people, like everyone can go to lunch together. It's a different when there's 3,000 people, there's many people you don't know, of course, and and there's more structures. So, there's something that happens when things grow that you can't just know everyone. Yeah. Um, but I I think it's still like I I think we've managed to not bring the pressure like with all the force onto the researchers. It still feels in many ways like we're just there tinkering on our research. Um, you need to, you know, forget this pressure at some point, otherwise you'd not be able to work. Um, I I think that happens quite well. I think there's some part of people, especially who are like doing half research, half product, you know, there there needs to be a I think they feel the pressure maybe way more.

>> But at the same time, you guys are in an extreme competition situation where like there is so much money being poured into AI, so many labs, and then all of a sudden DeepSeek appears, which for the public was all of a sudden. I think you guys knew about DeepSeek before. Uh, but how does it feel like because you may be working on something and all of a sudden another releases something? It's like, "How did they do this?" No. And it's like, is it constant competition, or you feel like more everyone working on the same direction? >> I I don't feel Yeah. I I don't think the competition is is as much pressure. I it feels like the broad topics are fairly similar. So we may, you know, we may be choosing a little bit of different roads, but but we're all kind of pursuing the same goal of more powerful models and and making sure they are doing the right things. Yeah. Like I don't I don't feel like day-to-day, I I don't think it's a worry.

Think about it. Um, it's also, I mean, you know, the Bay Area is in some sense a small place, and there's a lot of flow of information. So like, I mean, we we don't tell our competitors what we do, but but like people change jobs and and like I I so I don't think there is like a worry that things will stay deeply secret for for forever. You know, someone may do something nice and gain an advantage for a few months, and then >> That's competition, but but but it never feels like that's a live or die thing. There's a lot. So, yeah, I think her CEO Sam Altman, he said, "We should take these researchers and fly them over the data centers so they finally understand when they just press their buttons, >> what's happening, >> what's actually happening and the scale of how these things look and how huge they are." But I I I think it's true what he's saying. We're a little bit oblivious, you know, we say, "Oh, well, maybe today I need 5,000 GPUs to run this program." and and you just press enter and they start and they run it. But it's very abstract to me. >> But yes, they're real physical objects that are the size of >> a small city, and and and they draw a lot of energy, and and they're extremely expensive. And luckily, that's somehow abstracted for us. Um, but so, um, I I studied in Poland, my master's, I'm from Poland, and my supervisor was a student of Tarski long, long time ago, and he he was in the Bay, and one day he told me that um, he remembered how in the day of the Silicon Valley when it was Silicon, there were these trucks coming to Bell Labs with these extremely expensive transistors, and the they would unload in the morning, the researchers would do their experiments, and that was all trash by the evening, and they would just take it out and bring in the new one. So maybe in some sense, this is the price of progress. Like if you want to do frontiers research, you'll need to run these expensive machines somewhere and experiment on them, and you know, we try to do our best to make this useful. Um, but sometime there is waste of resources, which is necessary for progress, obviously. >> Yes. Yes. Sure. Most of our experiments fail. That's research. But then, could you do it better? It's I mean, I feel like as a researcher, you always feel that you should make better bets, right? You should bet on the correct method rather than incorrect. But you can't always do that. Um, of course, one day maybe AI will come and tell us that we could have, if we were just not smart enough. Um, maybe it's coming. We'll learn a lot. Maybe it will help us make better bets, >> right? But then you start calculating how much better they are, and probably can be, you know, you can c there's certainly some things you could do better, but in the big picture of things, how how much can it give you? Maybe not that much. We'll see.

What's what's your take on on the other labs? Like what do you think about the Entropic, for example? >> Um, well, I think I think to the first principle, I always think of the labs being fairly similar. Of course, every lab has their unique people and culture, and but in my mind, I I don't know one topic very I know Google much better, of course, because I worked there. Um, I I I think the labs have a very similar spirit, and you know, they try to do their best, do a lot of research. They take bets on a bit different directions sometimes, and sometimes that that works. You can even see in the models a little bit how they have slightly different >> Definitely, you can see >> personalities. >> Um, but yeah, they I think people as researchers and engineers there, they just try to do their best to to make the next best model.

>> And there there seems to be a tendency. I think Grok started that, and XAI with Ilan, and then Meta followed very closely. Meta changed a lot recently, like um, from the open source Llama uh to nowadays. And there seems to be like this concept that is getting popular, which is basically they calling it AI slop, where the labs are producing something which commercially may be very viable, like creating like AI characters, like any in Grok, uh, the data is romantically talk to the user to entertain, let's say. Then there is this like new feed of like kind of like TikTok or whatever it is from Meta, made purely by AI, that it just keeps feeding your dopamine. As a researcher, were you working on making the world a better place? How do you feel when they pour money and they spend so much money in this kind of purposes of AI that to me, honestly, feel very toxic? Is this is this is this the best we can do with AI?

Well, you know how it is. It's like any research, you research will be used in different ways, and you cannot control how it will be used, whether you personally like it or not. It and AI is a very capable method, and we need to come to terms that it will be used in ways we don't want it to be used, and as researchers, we will not be able to block it. The only entity that could block it would be the governments. Um, I mean, AI slop worries me way less than, you know, AI weapons. >> And um, I think we've survived so much human slop that the AI slop will will not be dramatic. I mean, but but it's, you know, only now I feel I mean, social media has been with us for for a while, but only now we're starting to to realize that it's actually not great, especially for kids. It's it's it's not very good. We should have put way more barriers earlier. Um, that's exactly my worry is that it seems like I thought that we learned the lesson, and then with AI, would not do it. But uh, what I'm seeing on this last couple of months is like a derivation towards doing social media with AI, with fake people behind it, which makes even less sense than social media as it is.

Oh, well, no, but I mean, I I feel like for one, if you want to put barriers, the barriers are probably not whether this should be AI or human slop. The barriers should are more like, I think in the US, it's becoming very popular, for example, that schools just take the kids' phone and put it on some shelf, and you're not using your phone while at school, and that seems to work very well. Maybe that's too harsh, but what I mean is this is not whether you're using TikTok or Instagram or or or AI something, right? This is very secondary to the fact, do you have limits on usage? But yeah, I mean, I I feel like yeah, this media side that's I'm much more worried about weapons sites. This is where, you know, social media had nothing to do with it. >> AI can actually make better weapons. We need to face that. Not so much maybe language models, but but the physical models certainly can. So it it the there will be coming um, yes, we I I don't have any, you know, big thoughts about it, but I certainly hope that >> people will do like put some restraints on that. I I also feel like on the on the digital side, there is definitely some hope that maybe if you're moving to AI, maybe you can do some things better. Um, one thing that I mean, I think OpenAI as a company, it means a lot of our employees and and the leadership, I think one thing we're very proud of is that we have a subscription model. Y >> And this could have totally gone the other way, right? There was some early decision, a bit of luck, certainly, you know, chat was a research preview, but but in addition to luck, there was certainly some thought that like, we don't want this model where engagement is your metric. It it brings you money. It would probably for the purpose of getting money, it would probably work, but it's not where we want to go. And people were saying, "Come on, you cannot make money on subscriptions." I mean, that was the prevalent knowledge in in online business, right, that people are not going to pay. And look, people pay. >> We don't have ads, we don't optimize for engagement, and a lot of people are using it and paying a subscription. Thanks to this, you know, we can do our research without having to make you watch >> you know, chat GPT is not sending you know, "Please chat with me every every hour" because >> This is huge because this is really important because that's basically what went wrong with with social media. And what I'm worried about is the tendency I'm seeing on the last uh months is going again to the other part, probably fueled by Meta and by XAI, because they both come from the social media. But um, we seen for example, that now OpenAI is looking for a head of ads. So obviously, uh, there may be some ads, probably in the free accounts, or some sort.

Well, I I think there is a big awareness now, and I I think, you know, to the best of people's abilities, I think there is a strong ethos at OpenAI, at least among the employees and and at least parts of the leadership, definitely too, that we don't want to go that route. >> Now, you still need to make money. >> Exactly. So part of this is a dialogue also with people, you know, I I was at Google many, many years. I was at Google 7 years. Still Larry Page was roaming the campus. He did a try it was was over a decade ago, I think, where Google tried to introduce a subscription model. Google now has Google One, right? But but at that time, it was like, "Okay, maybe we don't show you, show you less ads or no ads, and you pay a subscription." >> And this totally didn't work. Nobody Nobody wanted to pay because if you start free, it's very hard, right? >> To move on. Yeah. >> But subscription, like you cannot do this if people don't pay. You you understand. >> And you know, open like chat GPT did very well with subscription, but it's also not like we need to see where this goes, right? But now there there's more, I think just yesterday we announced the shop checkout model. So now you can do shopping from chat GPT. >> Right? >> And you can immediately buy things. >> And you get some money out of it. >> But we don't need to show you ads. And the partners decided that it's still okay for us to take some money. So So that's great, right? >> Yeah, that's amazing. >> It's a model that allows us to earn money and not need to push you to to >> The question would be, if chat GPT is recommending me a product based on the ones he has a deal with. So, for example, you did this with Etsy, and now if I say, "Hey, I want a wooden shelf." If it's only showing me options from Etsy because you guys make money out of it, then the model is not paying fair. >> But it won't. >> Okay. >> Right. The the deal is very much that it's not affecting >> like this, then it's great. Yeah. >> It's it's also maybe it's a maybe good thing an artifact of the technology. It's very hard to affect a language model to show you things from like with ads where you know it's basically a ranking. It's kind of fairly easy to add a signal to rank things up in a language model if you post train for something, you may get very weird stuff. >> But also no, I think in the deal with the partners, it's very boldly written that it will not affect anything. Um, but again, this is for now, for it to be viable, it something like this has to work, otherwise >> You know, if if Meta starts making huge money on ads, and XAI, and OpenAI starts losing losing, then then there will be pressure too. >> That's that's my main worry because I have the feeling that OpenAI, especially Sam, I've seen him in the Senate speaking, etc. I think there is really good intentions and mission, but then market pushes in the direction, and then what I'm afraid is that Meta is very strong power to push in the direction of us. >> But you know, it's like market, I don't think people want all these ads. It's like market is on the one hand the reality, but it's also a lot of what people believe the market is, and I think maybe it's different, you know. >> Yeah, I just hope that you basically campaign.

I don't know if you have ever met uh Johnny IV. >> Mhm. >> And he's certainly a crusader for like good tech. >> Mhm. >> And you know, it's a challenge, but it may work. So >> Talking about Johnny IV, and I think this is something I don't think you can tell me much about, but >> uh, I saw recently an interview of Sam talking about the future of hardware devices of OpenAI. You guys are working on that. Is there anything you can tell us about what direction this is going to? >> I I don't know much. I mean, I think I know as much as you know from the press, but but Johnny sometimes comes and gives like internal talks, and >> it's very clear he wants to make it right. >> Now, what is right, that's harder to say. >> To each person. >> But it's certainly not bombarding you with ads, that's very clear, right? >> Okay, great. >> And uh, well, whether it will work, we'll see. But I I have some hope. And even even if you think of just like pure social media, right? Even just like TikToky apps, I think small changes to how you do it, like whether it's mostly for you and your family and friends, or whether it's like this, "Look, this person you don't know just did," like sometimes very subtle changes can make a big change in how it actually affects people. That's that's why it's hard. I I, you know, I'm a researcher in AI. I don't really know much about these things, but I know we have people who are very committed to trying to do it better. Doesn't mean it's going to work. But at least for now, like the market has, it seems people are fine with something that doesn't have ads, and they're even pay. As I said before, I I just hope that OpenAI keeps like >> I don't mind how much it is, but like a subscription that actually AI is looking for my interest and not for whatever uh advertising company. So um, I I assume this is going to get >> That is the mission. >> And I can understand that the free accounts, they get some ads. I mean, it makes sense, and you use lots of compute for that. So >> Currently, free accounts do not get any ads. >> No, no, currently not. But I assume it will be the first place where ads will go because paying people will not accept them so easily. >> You know, people I think also at Entropic and Google and certainly OpenAI, they're very committed. >> Mhm. >> So we'll try to not have ads. >> Okay. Yeah. I I would be absolutely absolutely amazed if we can keep it on certain commissions that you make guys out of what they would pay. >> I I also don't think >> I think blaming ads is a shortcut of sorts. I don't think all ads are bad. >> No. >> I think the problem is that you're optimizing for engagement. >> You're telling people you need to spend your time >> talking to this digital device. That's wrong. >> Yeah. Because that's the base so I can put you more ads. So the ads, not the problem is not on the ads. The problem is on the engagement. >> But it can be the same if it's videos or like, but but the point is it's not the videos that are bad. It's not even necessarily the ads that are bad. Like sometimes you want to buy things, but >> But the wrong thing is that you're optimizing for people putting their time into this digital device, >> right? >> And there is a big commitment to not do that. >> That's great. >> I think across the labs, which is >> That's great to hear. That's amazing. But you know, commitments >> are one thing, reality is always harder. But but you know, it's like >> OpenAI is still a small company. If a lot of employers, employees feel that it, it's you have certain certain power of decision, I guess.

>> Yes. Lucas, AI is advancing really fast, and it's going really quick. Where do you think this is taking us in the near future? Do you I know you said you're not very good at predictions, but but what is your vision of where is it going to take us? Like you see a world of abundance where we don't have to work? I'm just let's extrapolate and fly off, no, like where like these machines are doing all the productive work, and then we can just dedicate ourselves to, I don't know, Star Trek style of world, or what do you see?

>> I I yeah, I I I don't think I I I see that far. I I think we have a lot of much more mundane problems in in our world. You know, I don't think the fact that we work is is our biggest problem. I I I think we have problems that even people who work, you know, cannot afford a lot of things, and and we have problems with the environment, and we have problems in medicine, and um, and but we, you know, scientists have a lot of ways of trying to solve these, and and there's a lot of things we can do better that we kind of know of, but but there's just we're just not executing on them. So I hope that AI can like try to speed up the actually doing the things that we know would help us. Um, well, we'll see. And also like, you know, in information, we kind of know we should be using our information technology better, but but we're not. We're it's it's a um, can the AI help us? Maybe. So, so, so I have I have hopes for like more, not so much Star Trek style, as as much as just inner daily things. And, you know, it may not even be that big at first, like my trip to the moon may still wait a little bit. Um, though I hope one day maybe. But, but for now, you know, can can we just live our lives better? Like talk more to each other, not to the devices, get good advice, live healthier, in some sense, simple things. But but how you get AI is very good at giving advice, and it's usually good advice, but does it come at the right time? You know, is it is it really good? Doesn't doesn't it hurt? And also, can it like do the work? If it does part of my work, you know, what do I do in that time instead? Like what I what I always worry, many people talk to me about AI and education, and obviously chat GPT can be an amazing tutor. >> It can also be an amazing cheating machine, >> right? >> Can we use it right? Can we make it a tutor? And then I was very disappointed once because for some poorer countries, I I worked with some NGOs, and we said, "Well, you know, when there was one teacher and 120 kids, but they all have phones, chat could be a tutor." >> Exactly. >> And then someone said, "Oh, great, then we can fire half of the teachers." Well, no. >> That's not the point. >> No, that that's exactly the opposite point. You you can't. But but but you see, there are risks that people actually will. >> And so to navigate all of this is, you know, it's well beyond me. >> But there are opportunities. So so I just hope we use them well. That's >> Being an optimist, because I think you can be defined as an optimist. How do you feel when like people like Gary Marcus, they just trying to, I don't know, to certain point, it looks like they are trying to trash everything. We had a podcast um with Ramon De Freitas, which is basically in a very similar opinion to Gary Marcus. Um, this got almost 100 views or something in this channel. It's the most viewed video of this channel. And basically he says, "This is all a bluff. You guys are like lying in most of the demos, like that basically this is not really having any impact and it will not go any beyond of being just like an answering machine." This reminds me a lot of these declarations when the internet came that uh Nobel Prize in economy said, "The internet will not have any more impact than the fax." So to me, it looks like that. But for you guys, it must get certainly almost personal when you're making all your efforts on this work, and then Gary Marcus says, "Like this is a piece of crap." What's >> I I think we're way too busy to to follow it. So sorry. No, it it certainly does not affect us. It's But you know, people should be skeptical. I I think to be skeptical about the technology as a researcher, I just know it's misplaced. The technology is doing quite well. But we should be very skeptical about how we use it. Like to use AI well will need for everyone to adapt, and the technology will not adapt on its own. I mean, it's it's weird because it's the first technology that actually could, you know, tell me, "Use me this way," and it may even be right, who knows? But but but I think we need to take the responsibility as a society, how do we use it, and there's certainly ways to use it wrong. So so that's a that's a challenge where I would also be skep, you know, like we try not to put ads in it is like one thing we can do, but but how do we use this technology in our lives is is is a challenge. And you I I think these economists, you know, I think for decades, the internet did not show in productivity statistics. So I I don't think we should expect AI to like miraculously transform our lives into into a paradise, right? I I I'm not that optimistic. But then, you know, it can help you a lot. Like even the advice, you know, just go to the gym and it just tells you twice a week can help a lot. It's like small things, they add up. But for these small, like I think the AI tech, it will do more and more tasks of our jobs. It will, you know, self-driving cars are there, they'll spread to more cities. Like the technology will keep improving. I do think it's actually great. But for all of these improvements to actually translate into better lives for everyone, that the tech companies can't do this. This is something the society needs to do, and it's it's hard because we don't even know where the tech is going. You know, researchers don't know, Sam Altman doesn't really know, you know, nobody knows, neither neither does Dario. Um, I I think we need to face this, that will it keep improving? Yes, but we still need to use it well, the way it is. And will we whether we'll manage to like get the benefits for real? Well, that is something where some skepticism is is warranted because it's it's hard. Like we we got social media, and social media can be a very good technology. You know, it could be used to inform people, it could be used for learning, it could be used in many good ways, and yet it feels we did not use it in the good ways, right? And and we're not that much smarter now. So AI can also certainly be used in wrong ways, and and we need to put some effort into trying to not make this happen. So, so, so this is where I think skepticism is very warranted because if you if you think of this like, you know, this is like almost like a technology from heaven that will fix everything, then you don't feel the responsibility that no, you actually need to figure out how to make it work for you in your life. No one will come and tell you it like, well, maybe your friends will figure this out so they can tell you, but it's not like OpenAI will come and say, "You know, look, you use chat like that and your life will become better." No, this this is not that simple. This is a technology, a tool. It can do a lot of things, but you need to figure out what it can do for you, be because we don't even understand it well enough to to guide it. And and then, you know, in companies, it means like the management needs to figure out, you know, a lot of companies started, for example, like blocking ChatGPT, and people use it on the side, and that's not good, right? This is not how to do it. Like at school, of course, I never thought about it before, but in hindsight, it seems so obvious that chat can just do your homework, and be and kids and teenagers will use it like is, and and that breaks the system of homework. So that's bad, right? That like we introduced like now age verification. So so we'll so that could help a little bit, but >> But that's a big change, and how how do we deal with this, right? Like, of course, you can try to do more work at school, but but there is never enough time. So, so, so, so there are things like that, and we don't have the answer how to deal with it. I mean, sure, we can put age verification and block some people, but >> You understand this will not >> fix the problem. Some some things became possible, and you cannot make them impossible. There al you know, if you don't use chat, there are open source models, there's Meta, like the cat's out of the bag. If if there is something that we thought was impossible, and now you have an example that it's possible, yeah, the technology spreads, it is possible, which means people will start doing it. But but we need to find smart ways to to use it and to deal with the problems. And you know, of course, I hope that that we can do this. And I think if you look at history, eventually this hope is probably right. But in the meantime, as we're going through these transitions, and as the technology is improving, you know, there becomes more and more of these things every year, this is disruptive to to people's lives in some ways. And and how do we handle this well is is a big question, right? It's it's can can we like in a short time find really good ways to use it and to block all the negative things? I hope we can do it better than before, but but it's a difficult thing that that it's not just on the tech companies. And if someone tells me, you know, "I'm skeptical that it may actually bring a lot of disruption well before it brings the benefits," then I would have to say that is a possibility. You need to very carefully think about that.

Lucas, thank you very much. It's been an absolutely pleasure to hear from the source. I really hope your optimism carries away on the benefits like overload the the problems that AI may bring, and that we can actually fix these problems to take uh all the benefits without any of these problems being a big issue. Um, really thank you for your time. It's been really amazing, and I learned a lot today. So thanks for that.

>> Thank you very much.