Transcription
I would love if you could compare the moment that we are in right now, in 2025, to the early days of the internet. And I'll ask it as a sort of question where it's like: AI is to 2025 as the internet is to what year? Um, okay, internet history. I guess it was Arpanet and whatnot in the '70s. Uh, but the web kind of—I think of it as born in—uh, Mosaic was launched in '93 or so, if I remember correctly, and then Netscape however many years after that. Um, I think um, in some ways you can draw parallels. Um, I don't know, you could say with sort of—I don't know—transformers in 2017, the first inklings of our uh new kind of language models. But I think it's really different in many ways.
I mean, for one thing, um, the internet was brilliant and enabled many things, uh, but it wasn't like technically revolutionary, um, in the sense, um, you know, like uh, with the web, Tim Berners-Lee at CERN, they were just organizing the scientists' data and stuff and sharing it, and um, they did a great job and it took off as a viral organizational thing, and it's been phenomenal, don't get me wrong. Uh, but it wasn't like—nobody would have like questioned whether that was physically possible 5 years before. There was no real limitation.
Um, in this case, we don't really know what intelligence is. We don't know how far we can take it. Um, I think a lot of people, myself included, are just surprised how quickly and how far it has ramped. Um, and uh, so that's just an important distinction. Like we don't even know what the possible peak is. Like with the internet, you could imagine everybody could communicate at high speed with everybody else. There'd be, you know, every company would have a website, you know, which they do now. Um, but you could you could have realistically imagined that in 1990, let's say, or um, and even uh, there were things before that like gopher—I don't probably dating myself—but um, there were things before the web uh that were kind of like that. Um, but yeah, with with AI, you don't know where the peak is or if there's a peak at all. Um, so that's one important difference.
Um, the other important difference, and for better or for worse, um, this has now gained, you know, profound international attention. I mean, the amount of resources and, you know, money and compute and energy that are flowing towards AI is extraordinary. You know, the early days of the web, we were a startup, you know, we whatever we we we got uh a little uh seed loan, less than a million dollars. So, we're off to the races. You know, I think we got like $10 million in our venture round, and that was that. Um, you know, these days companies are spending billions and billions of dollars building, you know, the best AI models in the world. Um, and um, that's a good and bad thing, but uh, but it's definitely different. Uh, so I just think it's um, in as much as there are parallels, I just think we have no idea where this is going.
So to that end then, do you believe that AI is fundamentally more of a discovery than an invention? Do you feel like this is some emergent property of the universe and we're just stumbling upon it, or is this like the ultimate test of human creation? Well, I I mean, I guess both of those things that you said, yeah, um, are distinct um from from what the web might have been. I mean, yeah, I think it's a discovery in the sense like we simply do not know what is the limit to intelligence. There's no law that says, you know, can you be 100 times smarter than Einstein? Can you be a billion times smarter? Can you be a Google times smarter? You know, there's no—um, I don't I I think we have just no idea what the laws governing that are. And um, so yeah, I guess I guess you could call it a discovery of sorts. Uh, maybe an analogy is sort of um quantum computing where you don't really know how much computation are you really going to be able to get out of the universe. Um, kind of the basic laws of quantum mechanics suggest it's extremely high. Um, but you don't know in practice if there are just other limitations you don't know about right now. Yeah. Yeah. Yeah. No, that's fascinating.
But you ultimately think that AI is a more momentous discovery or invention than the internet? Yeah. I I think the internet was um, it was definitely very important. Uh, but it was sort of as much a uh kind of social development, like everybody agreeing to use these protocols and then kind of making their data and systems available for everybody else with, you know, TCP and IP and then um, you know, HTML, HTTP, just agreeing on protocols and allowing it to grow and flourish. um, you know, maybe akin to how money was an invention uh thousands of years ago that let people, you know, really trade and stuff, but it's not like—it's not neither money nor the internet are testing the limits of the universe. But AI is. But AI is. Yeah. Because we just—um, we don't know how intelligent things can be, and we don't know—you know, we know some things about the brain. There's maybe whatever 100 billion neurons, 100 uh trillion synapses uh and they run so fast, but you know, with our computers, can we, you know, simulate that? Can we go beyond that, um, and how far and what would that be like? Yeah. Um, we just don't know. I feel like in that in that way, the question of, you know, how do we approach this, what do we build, who should work on this, are all questions that are philosophical just as much as they are um, you know, technical or economic. 100%, yeah. Well, consciousness is another kind of thing that gets brought into it, you know, the internet didn't raise questions of consciousness, for example, but right, you know, if this AI is smart enough and self-aware enough, does that matter? What does that mean? I don't know. Yeah.
You started Google famously in a garage in Menlo Park with Larry Page in 1998. You were just two guys trying to build something because you saw an opportunity. Now Google is a $2 trillion company—which I hope I have this right—180,000 employees. um, maybe plus or minuses, plus or minus, you track it as well as—um, I'm sure riding the AI wave so to speak with all that infrastructure has many uh and many um advantages, but I'm curious if there's any part of you that maybe just 1% wishes you were 20 years old again, just graduated from Stanford, it was just two guys in a garage. Um, oh, that's a good question. Um, I mean, look, I'm just grateful, all as computer scientists, to be alive at any age uh during this time. Um, I think if you were to walk across the street, I don't know, maybe you can talk some folks into letting you do that later. Uh, you know, and you see how all the AI researchers are all gathered kind of around around the coffee machine and whatnot, and you know, everybody's excited. I mean, it is a very startup-like fuel. Uh, it's not really a garage, obviously. Uh, though technically, by the way, when we started, we had the garage and a couple bedrooms. Okay, good. It was sort of like, but we did have the garage, but it wasn't just a garage. Um, my greatest fear was getting a historical detail wrong. So I'm glad that's done now. No, no, we we we tell the story kind of like a garage. I mean, but uh, but but there were a couple uh rooms in addition which helped a lot. That's good. Um, but um, I mean, there is a very entrepreneurial sense, I think, given the kinds of um compute requirements um that are required to compete at the forefront right now, um, and um, kind of the amount of science that goes into it, it would be really hard to try to um make a lot of headway, at least on the foundation model side, as a couple guys in the garage. Uh, lots of folks in garages can use these models to create new and amazing things. Um, and um, I don't want to discount the possibility that somebody's just going to have some idea that's so brilliant uh that even in a sort of a couple people in the garage could pull it off. Um, but it it seems like the the frontier is being pushed by, you know, pretty big companies like ours. I think we're now at the frontier. I'm very proud of the product progress we made over the last year. Um, so honestly, I'm really grateful to be able to be a part of that. So, I don't think I would take that uh teleportation to my younger self just yet. Yeah.
What is the most sci-fi sounding thing that you actually believe has like a decent chance of becoming real, let's say in the next 10 years? I think the most exciting will be uh Gemini making some really substantial contribution to itself in terms of uh, you know, machine learning idea that it comes up with, maybe implements uh and to develop the next version of itself. We already use Gemini a lot during like pieces, like some AI researcher will be like, oh, I need to, you know, debug this code for me, or um, you know, help me with this math uh or something like that, as sort of one-offs, but uh, as a sort of really substantive, some kind of new um significant breakthrough um that the AI itself makes. Um, I think that to me that's science fiction, and um, then I think it could well happen. Yeah.
If you had to ballpark guess when do you think Gemini will create the next version of Gemini? Uh, I mean, I I like I said, it's already assisting; that's already happening. Um, but from like kind of some kind of from scratch sort of rewrite, I don't know. That's a tough question. I don't know if I don't know how high of a priority that is to some extent because we sort of can guide it. You know, at what point is it possible? Um, maybe in a 3 years, let's say three, four years. I I don't know if it would the vision the virtue it would make just by itself would be quite as good as itself, but um, you know, if you think about it with our new um video model, the V3, let's—I mean, which by the way brought me to tears in the demo area just a few—in a good way, in a very good way. Okay, just check. Yeah. Yeah, the sound, something about it just came. Yeah. No, sound is such a huge deal. I didn't realize how much that was missing until it was all there, and it just—Yeah, it hit me like a ton of bricks. Oh, well, thank you. Um, but you know, in theory, I guess I've never tried this, but—Well, first of all, you don't—I don't know if we support this in the user interface, but you don't have to give it a prompt. Um, like it will just generate a video. Our user interface might not actually support not having a prompt. Uh, but then you would have no idea what it's going to make. Yeah. Um, you could just say make a good video, I guess. Um, that would be sort of the model just by itself going, but generally I think when it's directed by a person, and presumably that's what you did, you gave it some prompt, some target, uh, you know, you get really good results. So, I guess I'm just saying if Gemini creates the next great version of Gemini, I think for the sort of foreseeable future, it will probably do better if there's a human that sort of guides it at at least at some high level way. I mean, it it could conceivably someday, blank slate, just go to town, do everything without any guidance. Um, but yeah, that's that's level sci-fi I don't think we've gotten to yet. Got it. So for the foreseeable future, you do foresee a world in which it's the Google employees helping the AI along to build the next versions of Gemini, the next—the future that we're going to—Yeah. Yeah. That's right. Yeah. Yeah.
Um, how much—I'm curious how much of your time and energy you spend on sort of these like the deeper, more meaty philosophical questions of AI versus—I'm sure all of the the the technical questions, the practical questions, the business questions. How—yeah, how much of your energy goes towards all of that, given that it is such a—you know, feels like we're discovering some fundamental underlying capability of the universe at the same time as building cool tech. Yeah. I mean, probably not that much goes to the philosophical questions, just as a practical matter. It's just like there are so many technical details you need to get right on the on the way there. Yeah. Um, you know, I'm fretting about our being able to sign up for whatever Ultra and B3 and all the all the things that aren't quite working as I'd hoped. And I'm uh right now hassling the engineers and the product managers about all the little um snafoos. Um, I mean, it's definitely nice to take a step back. Um, some of the philosophical questions do kind of emerge out of the technical details, like um, you know, we have some new model; how are we how are we going to evaluate it, let's say—like what does it mean for the model to be good? We have sort of standard benchmarks and things that at some point the models tend to get really good at those benchmarks. Um, and uh, you know, every time you need to kind of redesign that, you do take a step back kind of philosophically and try to figure out, you know, what is important. Um, when you have new kinds of AI models—the diffusion model, for example, that you can now uh play with—text diffusion, that's not like an apples-to-apples thing. Mhm. So now we're kind of asking, well, how do we compare a thing that doesn't, you know, go left to right, that's actually kind of spinning the whole thing out at once? Yeah. Um, how do we measure that compared to our uh kind of normal autogressive models? So a lot of these things bring up some philosophical questions. Um, but you are pretty grounded and knows the grindstone—no pun intended—answering them. Yeah. Yeah. Well, you're—it's—it's rooted in such practicality. You actually—it's not—you know, it's not theory about what other people are working on when you're actually in the lab building. Yeah. Yeah. So, I don't know. Probably be great for me to be able to spend more time on some of the philosophical questions. Um, but there's there's a lot going on. There's a lot there. Yeah.
Um, I don't know how much time we have left, but one question I do want to make sure I get the chance to ask is, what's a question or a topic that you wish people like me, interviewers, asked you more about? A question or a topic that I wish people would ask me about? Oh my gosh. Um, okay. To formulate this as a question, or I guess I guess I could just sort of answer the question. Maybe that'd be easier. Um, uh, I I I guess it's this overall idea that people sort of react to, let's say, new AI announcements, whatever. We we announced a bunch of things yesterday, and now, you know, what can you do with those things? Um, and there's a bunch of cool stuff you can do with those things. Um, and then there's a bunch of stuff you can't, or it's not quite right. But I think the interesting question is what are you going to be able to do with kind of those things, next generations, in one year, in two years. Um, and that's that brings uh, you know, all kinds of exciting questions. M uh, I mean, language models 2 years ago made so many very embarrassing uh errors. Mhm. It was kind of like, oh wow, this thing actually did this correctly, and that's super cool. That's very different than, oh my gosh, I can actually use this as a tool for whatever it is because I, you know, if it's, you know, going to be right 20% of the time, I can, you know, whatever, post a tweet with it, which is like, wow, that's cool. But I can't actually use it day-to-day. uh, and um, nevertheless, when you kind of look at the trend, uh, you probably will be able to use it day-to-day with reasonable reliability. Um, so I think as people kind of think about where they're going to plug these tools in uh to whatever it is they're trying to do, you know, you're trying to make a movie. Yeah. Uh, VO is cool, and that has sound now. You know, maybe a year ago you've been like, "Well, it doesn't have a sound. It'll be a pain." Um, we've done some work for continuity of characters and things like that, but it's—it's—and we actually have some movies we are making with it, but it's probably still not ideal for making like some kind of two-hour film. Sure. Um, but nevertheless, I think when you look at how far these tools have come in the last couple years, you know, all of them, video models, not just ours, but you know, um, all of them, and you forecast in a couple more years, boy, you're going to be able to do a lot of really interesting things with them. Um, so I—can you give some examples of that? Because one of the arguments I've heard is like, by far, video is the most exp—If you look at all the different modalities here, video is the most computationally expensive. And what's the practical application? Fun videos on YouTube, you know, AI rot on TikTok. Like there's a lot of um controversy or conversation amongst uh, you know, folks about why are we investing so much time and energy into doing it besides the fact that it's really, really cool. Can you share what are some of those other practical applications of these video models?
Yeah, I mean, like I say, I mean, I think it's like the difference between a cool toy and sort of a useful tool, you know, sort of a matter of of time, and it happens gradually, and um, uh, you know, we're trying to aim for the useful tool. Um, we have some uh film producers that are here. Um, I think Darren Aronofsky might have already spoken to—I don't know if he was on panel or something, but he's making a video. Um, I've—a close friend of mine, Dustin, is making a video. Um, but like, you know, some real artists are using these tools, but it's early days. I mean, they're obviously dealing with something that two years from now movie directors will think was kind of a joke. Mhm. Uh, but they're putting up with it to be, you know, on the frontier, and I I think these models will be capable of producing really compelling videos. Um, they'll be able to do it, you know, in concert with, you know, human directors, human actors and so forth. I think that Aronofsky film does have—it has a combination of like, you know, real life sort of acting combined with AI generation in in a pretty cool way. Um, but I, you know, look, we today, obviously, we have the um Industrial Light & Magic, all the, you know, Lucasfilm, you know, they do all the special effects, you know, we already use technology to generate film—this is just sort of a new dimension in that, and obviously early days, and maybe it's not, you know, the best resolution, it's not like the best um continuity over a long period of time or whatnot, but I I think you'll see all those things come along. Uh, so, we're trying to push the envelope so these things become real tools, uh, not just toys. Yeah. Yeah. The one of the things I always say on on my platform is whatever you think about it today, this is the worst that it will ever look. Yeah, that's right. That's exactly correct. Um, I—Okay, so I got the signal for one more question. Um, so I I'll I'll ask this one on on behalf of all of the, you know, the builders out there who are so excited about this moment, but are not a Google employee. They're not working at a Frontier Lab. Is there any sort of direction that you would point these people into, whether they have a software engineering background or are just simply somebody who is understands the depth and the importance of this moment and wants to be involved? Is there a direction that you would point them into to say, start here? Or maybe is there a direction you would say, don't go in that way. Don't don't waste your time.
Um, well, there's so many great ideas, so I would be reluctant to stop anybody from doing anything, 'cause you you never know. Yeah. Um, uh, I think there's actually an increasing amount of really interesting academic work, even, you know, outside of the big labs. Um, and I think it has sort of happened with this um advent of reasoning models that that happens more in the what we call post-training kind of reinforcement learning um step, which is more manageable with the amount of compute resources that lots of um academic institutions or small companies have. Um, and there are lots of actually open-weight models, including our Java models, that you can use to experiment with that. Um, increasingly, I think you'll see um some kinds of reinforcement learning APIs from the top models, so that you know, you can send us your problems um maybe with correct solutions or graders, and you know, we can contribute uh your problems to the mix of like if you want the model to be good at that. Um, so I think that kind of thing is coming. Uh, yeah, I actually I think it's pretty good time um to be able to make an impact uh without having to train up a foundation model. Amazing. Well, thank you so much for your time. Thank you. A pleasure.