📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Google CEO Sundar Pichai on Gemini, Self-improving AI, and World Models

Matthew Berman11:20

Transcription

The diffusion version of Gemini. That that was I was not expecting that. Is this a departure from transformers, or is it something else?

So, we're going to push the diffusion paradigm as hard as possible, and then where we need to kind of bring them together, we will do that. Are we at that inflection point, given it seems like this is self-improving artificial intelligence?

We are definitely now working on recursive self-improving paradigms. What happens to people that do knowledge work? Just lean into these tools. Just getting in the mindset saying, "Look, I'm just now super assistant with you all the time and just take advantage of it."

Do you see the Google search homepage being the the kind of first place people go to find things?

Sundar, thank you so much for sitting with me. I noticed you announced the Gemini model is going to be a world model, right? You're transitioning to this world model. Does that take significant architecture changes? Is this a departure from Transformers, or is it something else?

You know, we are, you know, Google DeepMind has always had a a broad view of all the things that need to be needs to be developed for AGI, so they have efforts, both the G2 models; they have parallel efforts into building world models, which is different from the main line of uh Gemini 2.5 pro, but things we are learning there will make its way there, like when we built VO3, it's grounded in physics physics. Some of that innovations came from our work on world models. So that's how I would think about it. And then uh the diffusion version of Gemini, that that was I was not expecting that.

Yeah. So I think it was five times faster than the flashlight. Is that going to start to make its way into this world model? Like how do you think about all these different architectures?

Look, I think first of all, today all of you know mainline Gemini models are autoregressive LLMs. They are next token prediction models and architecture. Whereas our image models have been diffusion-based models. So doing text diffusion, I think it's a different paradigm, has um, you know, you could see for a same capability, it's so much faster, but it's obviously behind, you know, behind the Gemini mainline in terms of capability, but I think there'll be areas where you can use them. So we're going to push the diffusion paradigm as hard as possible, and then where we need to kind of bring them together, we will do that. And so, but I think it's good to push all the directions in parallel.

Yeah, I think that makes sense, right? You just make a lot of bets, push them as far as you can go and see how they come together in the end.

That's right. The next thing I want to talk about, so Alpha Evolve, I read that paper a couple times, saw the project, was absolutely blown away. This is um AI that can uh discover new knowledge, right? And and so it really feels like we're at this inflection point of the intelligence explosion. Um, do you think we have the right ingredients to to really are we at that inflection point, given it seems like this is self-improving artificial intelligence?

Look, you're spot on to the potential for something like Alpha Evolve. I think it's it's amazing. We launched that like a week ahead of IO in this low-key way. Yeah, it's one of the most groundbreaking work we are doing. But this the fact that we spoke a lot about agents today, but the fact that you know you can have these agents which can go improve code, make discoveries, etc. What an extraordinary paradigm that is.

Yeah. I think this is where we all underestimate even today, even talking, we so underestimate the potential of this technology. There's been nothing like this ever before. Why I always felt it was one of the most profound things ever, more profound than fire or electricity. But I think when we making progress with agents today, the models are, you know, they're expensive and you know they have latency in them. So when you chain them together to do all this, you know, that's what makes it still not fully there, but we are definitely now working on what looks like recursive self-improving paradigms and uh so I think I think the potential is huge and and if you were to point to one area whether it's the the core intelligence of the model, the memory, the scaffolding around the ages, what do you think is the the highest leverage area for improvement?

Look, for for for me, like look, figuring out how to do do all of this more efficiently. So driving efficiency in how all of this works is what's going to make it all much more practical to be used at scale everywhere. Something we've been obsessed about. That's why you know our 2.5 flash which we always are focused on because that's where we bring the most intelligence, the best price point, the workhorse, the workhorse. Yeah. So the more we can so the biggest breakthroughs is to get make everything work in that way, right, like you know, and and this is why we work on TPUs too, what uh what drives some of that infrastructure advantage, that's what excites me.

So you you mentioned agents, I know a lot of the presentations today were about agents. I am very bullish on agents; agent memory in particular is something that I've been thinking a lot about, and it makes agents so much more powerful when they learn to shorthand with you. When they learn about you, they become higher quality, uh, more efficient, uh, but it's also potentially a lockin, right? Uh, for for large companies. Do you think there's a need for an open-source or open protocol similar to MCP or agent to agent uh, but for agent memory?

No, that's a that's a great question. Look, I think obviously when you're giving these models u memory, you know, you have to give there are important privacy issues at stake. You want to make sure the user is in control. But I think like just today if you decide to stop using Gmail and you're going to go, we have data exportability. We allow you to export your email. I think you know maybe we are in this early phase, but I think these are great concepts to think about, time to say if there is my memory, how can I take it and take it somewhere else as a user with control, but you know I don't see why those things are not possible. Going back, I think the open protocols end up being super important, right? Like that's why you know A2A and MCP are important exciting directions. I don't think there's going to be one KI to rule them all or one agent, you know, you will be using a lot of them and so understanding what's your data, how can the models access it and maybe maybe make those portable. I think those are worthy things to think about.

So I I went over to the demo booth. I wanted to try on your the new XR glasses. They they looked incredible based on Project Astra. Do you think glasses are kind of the best, the optimal form factor for this personal in artificial intelligence interaction? And and if not, what is it? Or is it a combination of things? What do you think?

Okay, it will show up in many places, but glasses, I mean, they're really powerful because they're just like you're going about your day-to-day life. You're just interacting with things. It is in your line of sight, right? And and and maybe can even talk to you more privately, right? You know, etc. So, I think it's incredible. Uh, you just mentioned memory. I just had this amazing experience with Astra where I was just, you know, I showed it a few things. Then I later said, I don't know where where an item is in my office. It said, "Let's play detective." And it thought it knew where it was, but when I went there, I sneakily took the thing away. You know, you could see say, "I just saw it there. Can you zoom back out?" It was almost figuring out I had like kind of pulled that thing away from its line of view. That's so impressive. So, you know, memory. So I it it was so intuitive to use it. So I love love that experience.

And um, continuing on the kind of user experience track, uh 5 years from now, do you see the Google search homepage being the the kind of first place people go to find things? Because it seems like your Google is surfacing all of this context right where the user is almost almost proactively, and you can kind of see the vision there. So how do you how do you think about that transition if there is a transition?

Look, I um it'll evolve in surprising ways, but you know, I'm very excited about AI mode. I've been using it a lot. I see others how others are responding to it. Uh, it's a very AI-forward experience, and people are so natural. They type in so much. The end, but it's grounded in search; it can use all the tools; it will have personal context, and over time we can be proactive there too, right, and be because you're wearing your glasses, you know you're a student, telling you hey you got to you know get to your homework, I save some time on your calendar to do so, and when you go to sit to do it, it has prepackaged stuff for you, all of that I think is definitely within line of sight, you know the details will have to be worked out as we make progress. But this is what we are working on.

Yeah. I mean, I'm I'm extremely excited about being able to have I use a bunch of different Google services. All of my information is there. Having it surface to me and being able to have an agent that can kind of see across all of that data is incredibly important that I told you earlier that's why I walked over and bought an Android phone. I want to experience that firsthand when it's ready.

So, I I have uh another question for you. So a lot of people are anxious about this new world where most knowledge work and maybe eventually all knowledge work can be done by artificial intelligence. And what happens to those folks? What happens to people that do knowledge work? How would they prepare, stay relevant? How do they stay on top of of things?

I think at least in the near future, I mean, this is like having a superpower with you. It should take a lot of grunt work out. Yeah. Allow you to operate at a higher level. So I think the opportunity is actually, you know, think about with VO3 how many new think about you make videos on YouTube, like just imagine the future in which if you want to explain something to your viewers, being able to quickly have a prompt that kind of captures it, inserting it in your video, like I think you know so we're putting powerful tools still in the hands of people, the best way you can prepare is like what you're doing and everybody should just lean into these tools. Test them out. Test them out. Start using them. I always tell people when they come to me and they do something, I'm like, "What does Gemini 2.5 Pro think? Did you put We had IO keynote." I'm like, "What did Gemini 2.5 Pro think about the IO keynote?" Like, tell me that. Just getting in the mindset saying, "Look, you have this now super assistant with you all the time and just take advantage of it and and leaning into it." I think I think you know we're all going to get a lot of access to new new tools and capabilities and so that's how I see it playing out.

Yeah, I I I'm extremely optimistic about the future. I I hope people lean into this. It's it's really exciting.

Sundar, I want to thank you so much. It's been an absolute pleasure. Thank here. Thanks.