📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Ex-Google Insider: You're Not Ready For The Next Phase of AI

Inside the Silicon Mind with Firas Sozan34:07

Transcription

Some people say we're already at AGI, but actually, if you look at the industry and look at enterprises and what they need, a lot of these enterprises are making minimal use of AI. And when you look at why, it's because a lot of the work that they do is visual. The benchmark that I often refer to is baby vision, and there you see that these models are still at the level of a preschooler, not even a elementary school student. So, I wouldn't call AI at the level of a preschooler AGI by any means. They can't do things like count how many glasses are on the table or do some simple board games or understand some simple spatial problems. They can't even tell you what two things a wire is connected to. And that's quite important if you're building AI that's helping you build a data center.

What do you think they'll say about Google Brain and how it actually changed the industry?

I would say they would consider Google Brain as the Bell Labs of this era. The people who came out defined AI for quite a while. Sara Hooker, who was at Cohere AI; Elias Sutskever, obviously, who was at OpenAI and then started SSI; and Dario Amodei as well, who obviously started Anthropic.

Andrew, you spent 14 years inside Google Brain alongside people like Jeff Dean. When you look at where AI is today, do you think the real breakthrough was the technology or the culture that produced the people building it?

A big part of it is the the culture. So, the culture, um, allowed people to think freely, to try whatever they wanted. There was no pressure from products or pressure to launch something in like a certain time frame. So, it was just a very free and open research culture. People were discussing new research ideas in the micro kitchen or during lunch. And then people were excited to to execute it and and see what happened. So, it felt really like an age of innovation, especially since deep deep learning just kicked off, um, around then as well. And people wanted to see, you know, if their ideas, uh, would work in this new paradigm.

You know, it's interesting when I think about the people that I speak with every day, the researchers, the machine learning engineers, and all the incredible individuals in the industry, makes me think this ex- citement that we're having in the market right now, the models, the agents, etc. I think it's the people. I really do. And this is why I wanted to. When you and I were talking about doing this episode, it all came back to the people. So, I want I want to start there. Let's unpack this for the audience, right? Let's I'd love you to perhaps give us a bit of a a background on who you are, where you came from, what you're doing today, and then a little bit of an insight into some of the people you work with and what they're doing today.

I'm Andrew. I grew up in the UK. I did all my, uh, studies there and my PhD there, and then I moved here, uh, to the Bay Area 14 years ago. Uh, there I joined a team that eventually became Google Now. Uh, but a couple of years after that, I joined Google Brain. Uh, that's when it was still very small, around 30 or so people with Ilya Sutskever and Oriol Vinyals still there. And, um, yeah, a year after I joined, Quoc Le and I, uh, we wrote the first paper on pre-training and fine-tuning. And that's the paper that's led to a lot of the language modeling work today that you see in chatbots, in the form of chatbots. Then I spent, uh, several years on like Smart Reply, Smart Compose, Google Health, like trying to apply these into products. And then, um, after a few years in Google Health, I moved back into Google Google Brain. And there LLMs were really rising, uh, in Brain at that point with. So, I called it GLAM. I worked on PaLM, called it PaLM 2, and then call it the data area for Gemini, and that's what I was doing before I left.

And that's also a lot of great people have passed through Google Brain as well. So, around the time when I was in Google Health and before before GPT-3 came out, Liam Fedus, Demis Hassabis, David Ha, who have now all started their own companies, they were my interns. And a lot of subsequently a lot of Google Brain people have all started companies. It's like Sarah Hooker, who who was at Cohere AI; Elias Sutskever, obviously, who was at OpenAI and then started SSI; and Dario Amodei as well, who obviously started Anthropic. So, a lot of the age of like frontier AI labs were started by ex-Google Brain or ex-Google DeepMind people.

That's incredible. Let's take a step back a little bit around that paper that was written on pre-training and fine-tuning. What what year was that paper written?

Uh, that was in 2015.

Tell us a little bit about that paper.

Yeah, so that paper was, I would say, it was a research discovery. So, we were trying some ideas, trying to improve paragraph vectors, which were at the time the state of the art way to represent paragraphs. And that was like work that came out of Word2vec. So, this idea that you can use essentially back propagation to get the optimal vectors rather than averaging some bunch of a bunch of embeddings. So, we were trying to improve that, and then we were trying some ideas. But what ultimately worked was training the model to do language modeling, and then fine-tuning that model to do sentiment analysis on some Rotten Tomatoes movie reviews. And then we found that that was able to beat any other supervised classification method at the time. Every other method including, uh, including other LSTM based methods. Uh, we were using LSTMs because transformers didn't exist at that point.

And then after that, Quoc said, "Why don't you just try it on images, too?" Uh, because like in deep learning, one special thing about deep learning is that we're not attached to any specific kind of data, right? We want methods that work with images just as well as with text because fundamentally we believe that the neurons, uh, the network can just learn the modality correctly. And so we also tried it on images. We tried predicting going row by row, rasterizing image, and then predicting the next row of pixels, and then again doing fine-tuning on that. With that the method, we weren't able to get state-of-the-art results, but we still got a very good result. Like slightly less than the state-of-the-art, but still, uh, very good considering the model had just didn't have any convolution at all. So that no convolutions, but it still had like a really good result.

Would you say that paper was the foundation or the genesis of what was yet to come? All these companies, the amazing things that have been happening over the last 5 or 6 years or so in the industry.

Yeah, I think if you ask, um, uh, people like the the core components of like what we have in terms of LLMs and chatbots today, I think they will say it's the transformer, it's the objective this language modeling objective and then fine-tuning, and probably lastly they would say it's the data training on the web. So, it's just like triangle of, uh, components together, and yeah, this language modeling objective is a critical part of that triangle.

And who wrote this paper?

Me and Quoc Le.

And tell us about that moment when the paper was written and the moments after, the the months and years after. How much impact has that paper had on the industry today? I know I'm kind of asking the same question again, but I want to emphasize this because I think it's important to understand that you hear this a lot in the industry that AI has been around since the '50s, but something transformation has happened and it's happened in this very unique place where a number of incredible minds came together.

I would say the most impactful thing was it finding a way to use language modeling. So, even that same year, there was like, we were talking in the Google Brain team and talking conferences, and people were asking, "Why are we training these language models? Like what's the point?" Cuz at that point they were only used for decoding. They were only used for speech recognition, and there was no other use case. And Quoc and I were there thinking, "Oh, they don't realize it, but yeah, language modeling is the core of language understanding and everyone is going to see that." Um, and then also when we presented at at NeurIPS that year at the end of 2015, uh, Sepp Hochreiter, who was one of the inventors of LSTMs, went to our poster and he said, "The method just works." Like he had already tried it and it just works. So, I think that was a sign that this was something that could be the basis for a lot of AI, but yeah, I think none of us could have expected it to be at the stage it is today. That I'd still be doing the same thing like a decade later.

But, um, the uses gradually grew. So, first there was things like Smart Compose and Smart Reply. Then we tried to apply it to Google Health. That turned out to be a bit early. So, that didn't work very well, but as OpenAI found, as you scale up these models to larger and larger degrees, it still keeps working. So, that's the key thing about this objective is it can draw in as much data as you have, like the entire web, and still utilize all that data effectively, whereas the previous methods just wouldn't be able to do that. And so, we saw its development through GPT 1, 2, 3, and then adding on new additional developments like instruction tuning and then RL and of course the transformer as well, which was also developed in the Brain team. And that's how how we got here today.

What do you think about when we talk about Google Brain and what was transpiring in that period of bringing these minds together? What did Google Brain understand about hiring and building a world-class team of researchers that maybe the rest of the industry just just was clueless to?

If you look at that time and also if you look at the people in that time who then went on and started companies like like I said David Ha, but also Anna Goldie and Azalea who have started Recursive most recently. A lot of those people came out of the Brain residency program. So, the Brain residency program was a very successful program for people from diverse backgrounds to join Google Brain for a year and work on projects closely with several researchers in Brain. And that it was very hard to get in. So, thousands or multiple thousands of people applied and the acceptance rate was very very low. Jeff Dean knows the actual number, but I know it's very low. And ultimately, they selected not just for like academic prowess or like GPA scores, but they wanted to see, uh, people with unique backgrounds who could bring in new ideas, who had slightly different ways of thinking, um, than the status quo. And that led to a lot of creativity in terms of like new research ideas and new things happening. So, I think that's a little underrated these days, which is research creativity and research vision and hiring people with, uh, the backgrounds that can do that.

It's almost as if they were completely unapologetic about how high the bar was to get in, and that was perhaps piece of their playbook of of building this incredible team. When you look back at that specific period in the team that you worked with, some of these people, you know, a lot lots of these people, as you said, have gone on to build incredible companies, just like yourself. Now we'll get to the company that you're building today. Was there some common thread across the entire team, perhaps a university they went to, a collection of universities they went to, or a certain passion they had, or something that connected them, like a connected tissue between them?

I would say it's probably the passion. Like they showed a lot of passion for AI, and to really be at the frontier and make research breakthroughs. And they I think they also had some unique backgrounds, so they might have built something before, like written papers at a really early age, won awards early on. But yeah, slightly different from a traditional computer science background, but I think also showing intense curiosity, really curious about the world and trying to understand the world, and trying to understand particularly how you know how we can improve AI and build things that are better than we had. So that was really key in terms of like resumes that we looked for and people that we looked for.

And you, was there a period when you first joined at there that you felt that you were part of, or at least you were inside a historically important environment?

Yeah, I would say that was evident very early on, just from the people that were in the team. So, particularly Jeff Hinton, who even back then, even like a decade ago, was a legend in the field. And so he's very known for being very creative and having ideas that like just they just work. And and part of that is his research vision and research sense, and that comes from, I would say, his belief that we should model things after the human brain. That the human brain is the only real, you know, example of intelligence that we have, and like following the kind of way that the brain does things is the right direction. So, that has really persisted.

You see that neural networks, deep learning that's followed with, you know, that basically follows neural networks which are far less complicated than real neurons but still following the same kind of design where the idea is you don't design the perfect network from scratch. You just let it evolve through gradient descent and you just let the data take the network wherever it needs to go. So, that's a really fundamental belief of his, I think. And that his thinking permeated throughout Google Brain, and so when I first joined, Ilya and Quoc, they were working on the sequence to sequence paper which also laid the foundation for a lot of model models. And there you see there was like Ilya was still coding things to run on kernels on the GPU. And it was, you could see instantly that it was a very special environment. Even back then.

So, Jeff Hinton, his belief was, we should model things around the human brain. Talk to me about that. What does that mean?

What that means, at least to me, is that if you think about how the brain works is just a very adaptable piece of neural machinery. Right, there's a lot of studies and also like this DeepMind has similar origins. DeepMind came from the UCL Gatsby Neuroscience Lab. A lot of the labs at the time came from neuroscience places. My manager at the time was a neuroscientist, and so the belief was that if you have a very good computational setup which could be neural networks, deep learning plus back propagation, then all you actually needed was the right data. So, that's assuming that the the brain, the human brain, just is just a piece of computational machinery, it has certain kinds of objectives and it receives a lot of data. So, those are the kind of like key assumptions that go in here.

I would say like maybe one subtle, a bit more subtle, assumption is that there isn't really much information encoded in DNA. Whether that's true [snorts] or not, we don't quite know, but you can argue both ways because some people can argue that pre-training is a kind of encoding intelligence into the DNA of the model and fine-tuning is more like how brains actually evolve as they grow to adults. But yeah, I would say like one of the fundamental things is to build a network that's very general, that's capable of anything and then all you have to figure out is the right objective and and the right data.

And then he also had some follow-up work which was based around like can we have something more plausible than back propagation because it's very hard to envision how back propagation would work in a human brain cuz basically to have back propagation you have to store all the when you update, say when you're doing an update to a neuron, that neuron has to know exactly how it fired to start with and there's a lot of evidence from neuroscience that that information isn't recorded. So, people generally believe the brain doesn't do back propagation. So, that's also one thing working on that I know of is to try and to find a biologically plausible alternative to that. And it's possible that if we do that then we also we might have another another breakthrough in deep learning.

Andrew, I want to go back to one more piece around the a couple of more pieces around that the team and the culture at Google Brain. You talk about curiosity and passion but amongst the team of, let's say, 15-20 people working together, researchers who are uncovering this breakthrough, were people comfortable in being wrong around their peers?

Yeah, I would say people were, uh, comfortable being wrong. Um, people would show results pretty early even when there hadn't been a lot of experiments like the the results might be wrong, things like that, but people just understood that, you know, this is how research works. Sometimes, uh, you're right, sometimes you're wrong, sometimes you make mistakes. And yeah, I guess another word another term for that is like, um, psychological safety. So, people weren't afraid to to speak up and say also, "Oh, I think this isn't the right direction," or, "We shouldn't do it this way." So, yeah, it was a very open environment, uh, open to both criticism and, uh, being wrong.

Being a recruiter myself and doing our best to find talent that are able to do these amazing things, we talk about talent density and proximity to greatness a lot. And one thing you said to me and and you used this word in our preparation session, you used the word "osmosis." So, when you use the word "osmosis" when describing Google Brain, what do ambitious people underestimate about being around elite talent every day?

Um, underestimate, interesting way to put it. I would say that it's more about understanding how research works and how senior researchers are research leads think and tackle problems. So, I think that one thing people doing PhDs learn, and a PhD is a long time, right? It's multiple years. But for most people, the benefit you get out of that is learning how to do research effectively. It's not really, uh, the papers. I think there are some studies that say, you know, most PhD research papers are only read by two people: the person who wrote it and the reviewer of the paper. And a big part of that learning is learning when you should give up on a project, when a project idea just sounds right and has potential, and yeah, the right way to think about research problems. So, all of this is kind of like independent of the actual research project, right?

So, that means that if you in a great team, a very talent dense team with great research leaders there, then you can very quickly learn these principles. Um, like I said, when to abandon a project, when to keep pushing despite roadblocks, and like being able to just identify good research ideas just from hearing them, right? And so, this can just happen over time just by being around people and listening in on conversations, on research talks. So, that's what I mean by osmosis. You might not be working with that person directly on a project, but just understanding how that person tackles a project, how they, uh, work on it, and how they think about research problems there is really valuable.

And you, when you're around people like yourself and Ilya and Jeff Dean and all these amazing minds, I'm just wondering, what begins to change in the way you think as a researcher or an engineer?

Change in the way I think? Do you mean like, uh, from my early days in Google Brain?

Yeah, maybe something you observed in your interns or people that you worked alongside, because you spent a long time there and you saw a lot of great things happen and you watched these people go on to build these companies and now you're building yours. And the reason I want to emphasize on this is because we talk about osmosis, we talk about remote working and being in an office or hybrid. I think we lost something when we became fully remote during the pandemic and bringing that back I think is so important. It's great to say I'm part of a team, but are you really part of that team when you don't see them except a 30-minute Zoom call in the morning, right? So, I guess what I'm saying is when you're around exceptional talent, whether it's in the research world or amazing software engineers or any industry quite frankly, right? What do you believe happens in a person's way of thinking, but also their ability to create, to build when they're around other great minds?

I think in person is very important. So, that's why our company people are in person to promote that. Being around, being free to share ideas in the corridors, when you're getting a coffee, a lot of these conversations can just lead to new ideas, new research projects. That was something that was definitely lost during COVID. So, these, you know, just ad hoc conversations that combine to ideas that people previously didn't think could be combined. I think that's very important.

And then I would say like what changed for me over the years working with these great people is thinking bigger. So, I think that as a PhD student it's very easy to go very much into a niche and work on problems in that niche that, you know, you will be known for that niche, but doesn't have that much impact in the wider research community. So, that's kind of like a safe way to do things. But I would also encourage people to try to think bigger of bigger ideas, of bigger impactful things that they can do in the community because ultimately those are the things that can transform academic area. I would say some of my early work on my PhD was on non-parametrics. That was definitely a bit of a niche and I published some papers there, but you see now no one really talks about it, right? So, even though I published some works there, ultimately it's way less impactful than the big ideas that I had the opportunity to to explore at Google Brain. So, yeah, I would recommend people keep thinking about big ideas going big.

I want to ask you what might seem like a controversial question. As I was preparing for this podcast today, I was looking at your career, I was looking at all the people you worked around, the list of names, where they are today, and one question kept coming back to me and it was: all these amazing companies came out of it, right? So, Google Brain led to OpenAI, to Anthropic, to all these frontier labs, to Alorian, which we'll get to now. Why weren't they built inside Google? Why didn't Google build all of this? Cuz they were great at bringing in this amazing talent and creating this culture of curiosity and passion and free thinking and becoming the foundation with the paper and everything else, but why wasn't it built in Google?

Well, I think it's, um, part of the Silicon Valley ethos, right? And the the show as well, which it is a quite an old show now, right? Silicon Valley, but I think that that also covers a ethos, which is that I think people a lot of people see their growth in these big tech companies. You know, it's a great place to grow and great place to learn new things, but I think there's a sense that at some point you outgrow that team, that organization, and at that point you don't have a lot of options. Your options are essentially to, you know, get more promotions, right? Uh, at the higher levels that becomes quite political, or you can jump to another company, another big tech company. There you might get a, you know, a promotion as part of the job, but it still means you're part of the politics most likely. Or the third option is to start something yourself, take more ownership, take more responsibility for your destiny, and also be completely free of politics.

And I think ultimately, if I wanted to do politics, I would work in politics. But I'm really here to really push the edge of research and push the edge of AI and the frontier. And so, the decision was kind of obvious to me that if I could build a great team, get funding, we'd definitely be able to build something great. And as Ilya Sutskever said back in the days of like sequence to sequence in the early days of Google Brain, he said something that I think a lot of people found found inspiring at the time, which was, "Success is guaranteed." And that's how I feel like when I'm building my own thing. But it just has to succeed.

Let's talk about that now. This is a continuation of the story, right? How all this began, the companies that went on to build the Frontier Labs, and now Lorien. Tell us about Lorien. How long has the company been around? And what are you building? What's the problem you're solving? And I think the second part, which we can get to, is how much of what you've learned all these years are you trying to recreate? I think we talked about the culture and the building the team. But let's let's let's start with Lorien and what you're building.

Yeah, so Lorien is a research and product lab. It's been around for 5 and 1/2 months at this point. Started by me and some friends that I worked with from Apple and DeepMind. And we're here to really build models to advance us towards visual AGI. And the reason for that is, obviously, I've been working on language models for more than a decade. And what I've seen is that the advancements in coding and text have, uh, and math have really been gigantic. Some people because of that say we're already at AGI, but actually, if you look at the industry and look at enterprises and what they need, a lot of these enterprises are making minimal use of AI. And when you look at why, it's because a lot of the work that they do is visual.

So, they might be working on a floor plan or designing a engine for a new plane or figuring out a wiring diagram for some electricals or like choosing a sofa for their home or an office chair for their office. Right? All these things are visual. All these things coding, pure coding, just doesn't work. You can't code up a new aircraft engine, uh, no matter how much you want to, or just use math to math a new rocket. Uh, it doesn't it's not that simple. There's a lot of visual elements and essentially, you know, physical AI elements in there. And so, we see that big gaping capabilities, and we also see that in benchmarks. So, uh, the benchmark that I often refer to is Baby Vision, and there you see that these models are still at the level of like a preschooler, not even a elementary school student. So, I wouldn't call AI at the level of a preschooler AGI by any means. But that means they can't do things like count how many glasses are on a table or do some, uh, simple board games or understand, uh, some simple spatial problems.

I want to unpack that a little bit more, but there's a question I've asked a number of my guests over the episodes that we've recorded. And the one that I I really like this asking this question, especially with someone like yourself, right? With what you've done in your career. If we were to use the mobile era, right? Thinking back to like a Nokia 3310 or like a the first iPhone and where we are today with today's iPhone or Samsung, where is AI today? There's actually the '80s with the big briefcase taking the phone out with the antenna. Like, where is AI today in your opinion?

Well, I would say for text and tasks, then we're at the iPhone level. Um.

First iPhone or the iPhone today?

Uh, maybe, um, an iPhone a few years ago. You can code on your iPhone. Well, I guess the app app store doesn't allow that, but you can you can, uh, do some advanced things on your iPhone, auto automation and math. You can use like Mathematica, all those kind of things. But, I think for visual problems, then we're at the level of a Nokia where we're taking cameras with maybe 64 by 64 pixel resolution. That's the That's the resolution of the AKGI benchmarks. And yeah, everything is very pixelated. Everything is very unclear. You can recognize that, uh, what car that is or who that is in the photo, but doing anything more advanced is is just out of the question.

So, with that in mind, if we think about you're solving, where could this take us? Cuz I'm thinking about cameras, safety, crime, you know, what you can do with recognizing problems before they occur. And these are just ideas that that come to my mind when when you speak about visual AGI, but I'm just wondering, with the problem statement that you're going after and the problem that you want to solve, what use cases out there, let's say when we get to the iPhone era, could you solve?

Yeah, we believe that there's a lot of things, uh, in technology that can really benefit from this and really advance the state of technology itself. A lot of technological advancement comes from the physical world, comes from, uh, mechanical, uh, electronic engineering, electrical engineering, all these fields. And a lot of these fields, like I said, are visual, uh, visual problems, visual diagrams, editing those. And so, we believe that that can be one of the early areas of technological progress. So, that's engineering, different areas in engineering, CAD design, CAM design. But, there's also architecture, which is designing floor plans. There can be some application in agriculture as well, construction, and this general case imaging. So, really the the kind of use cases in industry are endless. Um, if you can build something that's very good.

One thing that someone told me recently is data centers as well, like, uh, data centers are being built at a very rapid pace, but it's very hard to build those. And right now, these models, they can't even tell you what two things a wire is connected to. And that's quite important if you're you're building AI that's helping you build a data center.

Couple of last questions I want to finish finish off with. We think about all the amazing greats in our industry, people like Steve Jobs, who have completely transformed the way we operate. I mean, I'm I'm recording this podcast using an Apple product right now. When people look back 20 years from now, what do you think they'll say about Google Brain and how it actually changed the industry, the people that came from Google Brain?

Yeah, I would say they would consider Google Brain as a Bell Labs of this era. That, yeah, the people who came out, uh, defined AI for quite a while, led to the all the progress and the foundation of AI. It's very hard to see where AI will be in 20 years. I think LLMs will still be around, uh, but we'll definitely have some new things. And yeah, hopefully those new things are also, uh, developed by by this new generation of labs that still carry forward the Google Brain culture from a decade ago. So, yeah, I'd say that hopefully the culture of Google Brain still lives on even though the name won't.

It's interesting, right? We know about the PayPal mafia. This is almost the Google mafia. The people that have come out of here. And I And I wish you all the best of luck with what you're building at Alloy. And I I thoroughly believe you're going to do something incredible. It's It's been a pleasure having you on the show. To finish off on today's episode, as we do with every episode, is always to ask our guest for a book recommendation. So, Andrew, what would be your book recommendation?

I've always loved the Foundation series by Isaac Asimov. So, that is a great example of thinking big, not like tens of years or hundreds of years in the future, but thousands of years in the future and planning for for that. So, yeah, really love that series.

I love it. Andrew, it has been an absolute pleasure having you on the show. I really enjoyed every minute of this, and I'm looking forward to getting out there and having the the world see and and hear about your business and what you're building. And I wish you all the best of luck.

Thanks, Ross. Yeah, it's been an honor to be on here, and yeah, it's great chatting with you.

Thank you, Andrew.

Thanks for tuning in to another episode of Inside the Silicon Mind. This podcast is powered by Harrison Clark. For more episodes, don't forget to subscribe and hit that notification bell. As always, stay curious, stay consistent, and stay inside the Silicon Mind.