📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Godfather of AI: I Asked Him The Dumb Questions WE ALL HAVE

Nayeema Raza55:30

Transcription

And I just look at you when I talk.

Yeah, we just look at we'll look at each other the whole time. Unless you want to make a declarative statement.

We are all going to die.

Smart girl [music] dumb questions. What's our AI future and what even is AI? I'm Neyar Raza. This is Smart Girl Dumb Questions and today my guest is Professor Jeffrey Hinton. He spent half a century working in artificial intelligence and machine learning, earning him a 2018 touring award, a 2024 Nobel Prize, and the nickname, the coveted nickname Godfather of AI. Which of those titles do you like best?

Uh, the Nobel Prize obviously.

The Nobel laureate. You're also professor emeritus at University of Toronto where you have taught and collaborated with some of the most um seminal minds at places like OpenAI, Meta, etc. But I was watching your speech um your Nobel speech and you were warning about the risks of AI. I want to play a little bit of the audio from it.

In the near future, AI may be used to create terrible new viruses and horrendous lethal weapons that decide by themselves who to kill or maim. All of these short-term risks require urgent and forceful attention from governments and international organizations. There is also a longer-term existential threat that will arise when we create digital beings that are more intelligent than ourselves. We have no idea whether we can stay in control.

You obviously talk also about the benefits of of AI, but as you were saying these things, people are just eating bland Swedish meatballs and avoiding eye contact with one another.

Actually, I think the Michelin star chefs who were prepared the meal would be rather annoyed at you saying bland Swedish meatballs.

So, they weren't bland Swedish meatballs, but I would have expected to hear a hush in the room when you say something like that.

Yes, not particularly.

And what do you make of that?

People find it very hard to take the AI threat seriously. Even I find it hard to take it seriously emotionally. It's not like the threat of nuclear weapons where it's very easy to understand something that goes bang and wipes people out. It's much harder to understand that there might be we might be creating alien beings that are smarter than ourselves. That's just seems like science fiction. People don't take it seriously.

You are an expert. So I want people to hear your risks and take them seriously and also question them as they should. But I want to start with just explaining what artificial intelligence is because I talk to a lot of people about AI and I feel a lot of us, myself included, don't really get it fully. Would you agree? You talk to people all day.

Oh, yes. Um, most people who comment on AI don't really understand how it works.

And do you feel most of the time when you're sitting in an interview and someone is speaking to you about artificial intelligence that they really understand how this thing works?

Some do, some don't.

Okay.

Um, most don't.

And when they don't do they ask you?

It's very rare.

To set the stage of how we should think about it as well. The way I've thought about it is agricultural revolution, industrial revolution, AI revolution.

That's a very good way to think about it. It's that kind of scale. Yes.

Okay. So the internet then was not a revolution?

Not on the scale.

Not on the scale. And is it going to be slow but world shifting like the agricultural revolution or is it going to be um fast and violent like the industrial revolution?

Much more like the industrial revolution.

Okay.

So the industrial revolution for example, um, replaced a lot of agricultural labor. This is going to replace a lot of mundane intellectual labor. So it's going to cause a huge shift in employment and many people are very worried that it might cause massive unemployment. And I want to get to to that with you and also talk about whether you know millennials and Gen Z, if people if we don't have kids should we have kids in such a world. I mean that's a big question I have for you that we will get to. But first of all, what is artificial intelligence?

Back in around 1950 there were two paradigms for making an intelligent system.

Okay.

Two quite different paradigms. One was the kind of symbolic AI where the model for intelligence was logic. So if you say Socrates is a man, all men are mortal. You can derive Socrates is mortal. That was a way of deriving new facts from old facts.

Logic.

Logic.

Yes.

And many people thought that's how intelligence must work.

Mhm.

It must be some kind of logic. So you can derive new facts from old facts. There was a different approach altogether which said well the only really intelligent thing we know is a person and the human brain works by changing the strength of connections between brain cells. So maybe we should focus not on is there some kind of logic going on in our head but on how do we change the strength of connections in our brains and will that make an intelligent system and in particular we shouldn't probably focus on reasoning. Reasoning came very late biologically. Before we could do much reasoning, we could do perception. We could control our bodies, right?

Maybe we should focus on that because that's what the brain evolved to do long before it did much reasoning.

Okay. You said the brain changes the strength of the connections between the cells.

Yes.

And this is, you know, hunter-gatherers could do this back in. Yes. Everybody can do this. They can see. So you're basically saying a brain well before a brain can reason.

Pigeons do this.

Yeah. So well before a brain can reason, even a baby's brain can say eventually, "Oh, that is an object that looks flat. It has four things." And then we learn the word for that is table. How does the brain strengthen the connections?

Okay, that was a big open question and actually still is. So we can break it down into two questions. One is if the brain could find a way to decide for each connection strength in the brain whether to increase it a bit or decrease it a bit in order to make it work better at some task it's trying to do. Then if you started off with lots of random connections and just used this method of increasing or decreasing connection strengths, would it actually learn to do complicated things or would it just get stuck?

And the answer was?

The overwhelming belief was it would just get stuck. It had to start off with lots of innate knowledge which would be in the form of appropriate connection strengths between brain cells and then maybe if it had lots of innate knowledge it could improve it a bit but with experience.

Okay.

That was the general belief and it was just wrong. And what we've shown now is that if you can find a way to decide for each connection strength whether you should increase it a bit or decrease it a bit to do better at some task you're doing, then you can learn incredibly complicated things like these large language models. So it's like it's the neuralness of the brain, the ability of the brain to make sense and to connect to each other that makes it powerful, not some innate knowledge inside of the brain.

This ability to learn is the ability to change the strengths of connections so as to be better at some task.

And you're better at doing that when you're a kid, right?

Yes.

And because you're more neuroplastic, they say. So is AI as neuroplastic as a baby?

It's a very good question. So what we know now is that if you can find a way to figure out whether you should increase or decrease the connection strength and you can do that for all of the connection strengths at the same time, then you can make very smart systems. But there is a difference probably between how the brain figures that out and how current AI figures that out.

Okay.

And it's quite possible that the brain has a method that's in some ways better than what we have and in some ways worse because it's solving a slightly different problem.

Which is?

So the AIs we have only have about a trillion connections.

Only a trillion? We have more than a trillion.

We have about a hundred trillion.

Really? In our brain?

Yes. And so our brain has about 100 times as many connections as the smartest AI, but it only gets a tiny fraction of the experience. So we live for about two billion seconds. Even if you got 10 experiences a second, which is total maximum, and you didn't sleep, that would only be 20 billion. These large language models are trained on trillions and trillions. So they've got hugely more experience and hugely less connections.

Is that because they have more storage than our brain?

No. No. We've got more storage than they have because the storage is in the connection.

So the storage is 100 trillion is a storage actually, but we can't fill it up because we don't have enough time.

We can't perhaps use it optimally because we don't have enough time. So you don't have enough time to read everything on the web, everything publicly available on the web. These large AIs do.

But not because they have more time, but because they're processing it faster, right?

Well, there's two reasons. One is they're processing it faster, but the other is they're digital. And with a digital system, you can make many copies of it. And so what you can do with these AIs is have many copies running on different hardware. Each copy looks at one bit of the internet, figures out how it would like to change the connection strengths, and then they communicate with each other.

Right.

And they all change their connection strengths by the average of what everybody wants. And now each copy has benefited from the experience of all the other copies. So if you've got a thousand copies, they can experience a thousand times as much as one copy. And they can all learn from all of those experiences by averaging the changes in the connection strengths.

Right. It'd be like every time I have an experience, all of my siblings would have that same experience, for example.

And all of your siblings would learn from your experience.

Wouldn't that be terrific?

They seem to do the opposite. In fact, they just tell me why. Yes, that siblings. This is interesting because what you're saying is that artificial intelligence is more collective. It's much better at sharing. If you have multiple copies of exactly the same neural network using their connection strengths in exactly the same way and to do that you have to be digital. Then these multiple copies can share what they learned. And if they got a trillion connections, they're sharing about a trillion bits when they share how they'd like to change the connection strengths. Now when I share with you.

Yes.

I'm sharing maybe a hundred bits per sentence, even if you understood the sentence perfectly. So they're billions of times better than us at sharing.

Wow. Okay. And yet we shared with them the ability to do this.

Yes.

People like you, Yan LeCun, Yoshua Bengio, you were all godfathers of AI helping train these like neural networks to exist.

So there's sort of two things we did. Okay. We um figured out how to train them, how to how they should change their connection strength, but then we gave them lots of data and [snorts] from the data they figured out what connection strengths to use and we don't really know what they extracted from the data. So it's not like normal computer software. In normal computer software, you write lines of code and the person who wrote the program can tell you what each line was meant to do. It might not do that, but they can tell you what it was meant to do. At least with this, it's quite different. We write lines of code and we know exactly what they're meant to do. They're meant to allow it to figure out whether it should increase or decrease the connection strength when it sees some data. But what it learns from all that, we don't know.

When you say increase or decrease connection strength, what does that actually mean? Like for my mind, I think of that as okay, there's a part of my mind that processes like imagery. I say um, "That is a drum." And then there's a part of my mind that controls my hand. I can say, "Okay, beat on drum." There's a drum in the corner of the studio. There's probably even a part of my mind that was thinking about what to think about, some executive function, and then gathered the drum because I see it. Is a connection the connection between these different parts of my brain?

Okay, that's a whole bunch of connections and that's a big pathway between these different parts of the brain.

But within one of those pathways, within the pathway for doing vision, let's say for recognizing objects, there's many, many connection strengths. Like about a third of your brain's involved in that. Because we're basically monkeys and monkeys are very visual. So in that pathway, there's many, many connection strengths that determine how you recognize an object and they're mostly learned. So I could go over an example of that if you like.

Yes.

Let's suppose we take the task of, I give you an image and you just have to tell me, is it a bird or isn't it a bird? Now, if you think about images of birds, you might have an image which is an ostrich in your face about to bite you, or you might have an image which is a seagull in the far distance. They're both birds. So, just looking at the pixels directly isn't going to tell you whether it's a bird. Um, you're going to have to have abstraction. You're going to have to find various features. So, here's how the human visual system works, very roughly. And this was discovered by experiments, poking electrodes into brain cells.

Okay.

Mainly in cats and monkeys.

Okay. Not in humans?

Mainly not in humans.

Okay. Like fMRI type of?

No. No. fMRI.

Yes.

Is like it's actually looking at blood flow.

Mhm.

But the blood flow is caused by neurons saying, "Hi, I need more blood 'cause I'm getting active. I need more power." And it's looking at many, many neurons, like millions of neurons typically for each pixel. Each pixel in an MRI is the blood flow which is giving you an indication of the activity of, I don't know the exact number, but of the order of hundreds of thousands of brain neurons. Okay. So you're not seeing the individual neurons there.

When do you see an individual neuron?

You see individual neurons when you poke an electrode in and you stick it in a neuron, or when you use optical dyes so that a neuron glows when it gets active.

Got it. Okay.

Then you can see the neurons much better. But fMRIs are very, very crude. They're like looking from outer space at human activity. And what you see is, for example, that um when Detroit gets hotter, bits of southern Ontario get hotter too at a time scale of years. And what you're discovering is the car industry.

Okay.

The car industry causes correlations. That's kind of what fMRIs are like.

Okay. And what you're talking about.

As compared to individual human activity which is like the brain cells.

Which is like the brain cells.

What we know is the light comes in the photoreceptors in your retina convert it into electrical signals and do some processing. They then send it up to the brain up the optic nerve and a little while later, one stage later in the brain, you get a whole bunch of things that detect little pieces of edge. This is a simplified version I'm giving you.

Yes. And how long is that a little while later? Is it milliseconds? Is it?

It's about 30 milliseconds later.

Okay.

Um, and what you've got is um a whole bunch of neurons that detect little bits of edge in different locations, in different orientations, and at different scales. So let me tell you how you'd make one of those detectors. So suppose I have an image that's composed of pixels. Let's make it a gray level image. No colors for now. Each pixel has an intensity, how bright it is. And suppose I wanted to detect a little piece of vertical edge that's bright on this side and dim on that side. What I would do is I would take say a column of three pixels here and I'd have a neuron looking at those pixels.

Mhm.

And it would have big positive weights to those three pixels and big negative weights to the three pixels in the column next to it. So now if they're equal brightness.

That neuron will get lots of positive input from from the neurons this side and lots of negative input from the neurons this side and nothing will happen. They'll cancel out. All that neuron will do is say whatever's in this image is of no interest to me. It's not it's not the thing I'm looking for.

Okay.

Because it's not bright enough on all the edges.

Because it's not bright on one side and dim on the other side.

So.

Confused. Okay? The only condition under which it will fire is if the pixels this side are bright and the pixels this side are dim.

Because that's an edge.

Because that's an edge.

Because that tells me that this is something and that's nothing.

It tells me look, it's bright this side and it's dim that side. So there's an edge here.

Yeah. Okay.

Okay. So we discovered how you would make an edge detector.

Right.

Now in the end we're going to get it to learn to make an edge detector, which it will do. But for now, let's suppose we just handwire it. It being the AI, the neural net. So, right now I'm going to describe how how I would handwire a neural net to detect birds. It wouldn't be very good because the connection strength would all be not quite right. But here's how I'd approach it. I would say, okay, I'll make a little vertical edge detector there. And I'll also make a horizontal edge detector. I'll make something that looks for bright pixels here and dim pixels underneath it. And if it finds that, it says, "Ping, I found a horizontal edge." And I'll look for all sorts of other orientations of edges. And I'll do it everywhere in the image. And I'll do it for edges of different scales. Like I might have a cloud.

Yeah. This looks like the game that you see in the newspaper sometimes that has all the dots.

Not really.

Oh, okay. [laughter] No, that's how I'm picturing it. Okay. So, you have a cloud. I have a cloud. A cloud doesn't have any sharp edges. So, these little things that look for sharp edges won't find edges because the edge is very soft in a cloud. It gradually changes from dark to light. So, what we need is a a neuron that looks at lots of pixels. It'll look at lots of pixels over here with positive weights, positive connection strengths, and lots of pixels over here with negative connection strengths. And if all of these are brighter than all of those, it'll say, "Yeah, we have a big fuzzy edge here." So that's a detector at a different scale that's looking for fuzzier things.

But how would you build this? Like you would program this into code.

Okay. To begin with, I'm going to explain how I would handwire this.

Yeah, you handwire it. Yes.

Okay. So, I would just set all these connection strings by hand. This would take me more than the edge of the universe, but don't worry [laughter] about that.

I'm very patient. I would [clears throat] settle these by hand. And I'd do it for all over the image. And I might end up with billions or of the order of a billion of these little neurons. Maybe only 100 million, but a lot anyway. And that's just detecting little bits of edge of different scales and orientations.

So that's what the first layer of neurons is going to do. Now I'll have a next layer and that's going to look for little combinations of edges. So in the next layer, for example.

Right.

I might have I might want a neuron that looks for edges that meet like that.

So two edges that might just be a.

You're putting your fingers together in a bit of a triangle.

Yeah, they might be a little beak, for example. It could be all sorts of other things.

Right.

But it could be the head of an arrow, for example, but it could be a beak. [snorts] So the way I do that is I'd have a neuron in the next layer that was wired up to be excited by all of these edge detectors that detect this edge and excited by all the edge detectors that detect this edge, the horizontal edge and the diagonal edge, but not excited by anything else. So in order for this neuron to get excited and go ping, it would need to find some edges like this and some edges like this. When it finds that little combination of edges, it'll go ping. That does that mean I have like a literal beak neuron in my head that just recognizes beak?

This is when I handwire it.

Yes.

And and yes, you probably do have something like that.

Really?

Yeah.

How many neurons do I have in my brain?

A little less than a hundred billion.

And does the number 100 billion and the number 100 trillion have anything to do with each other? Like are they a permutation?

Each neuron has some connections, right?

Like typically a thousand between a thousand and 10,000 connections.

Okay. Okay. So there's a mathematical core connection between.

You take the number of neurons and multiply by a thousand or 10,000, you get roughly the number of connections.

Okay. Got it. Okay. That makes sense.

Nobody knows these numbers for sure.

Yes.

I think the best estimate of the number of neurons in the human brain is something like 86 billion.

86 billion. And that's from prodding around monkeys or from our poking around?

I think from looking at little bits of human brain and multiplying.

If you were to handwire this recognition system, it starts with the edges and then it starts with recognizing specific edges like beak.

Like beaks.

And also in that layer, you might recognize a bunch of edges that form a circle. That's a potential eye. Now it could be all sorts, it could be a button. It could be all sorts of things.

Right. It could be a wheel.

Yeah, absolutely. In the next layer you have things that might recognize beaks and might recognize potential beaks. Might recognize potential eyes. Then maybe in the layer above that.

The third layer now.

Yes.

Third layer. Now you have something that's got a big positive connection coming from anything around here that thinks it might have found a beak. So any beak in this sort of area.

Yeah.

Will will excite this guy.

So it's about where the physical location of the thing that you saw in the looking for a beak in this general area. It might also be looking for an eye in this general area.

So is the third layer context?

The third layer is looking for combinations of these features. It's looking for example, a beak, a potential beak. You don't know it's a beak yet. It might be the head of an arrow. And a circle here, you don't know it's an eye. It might be a button. But if they're in the right spatial relationship, that makes it much more likely to be the head of a bird.

So I just want to recap the layers. So the first layer is just edges. This is something, this is nothing. The second layer is these edges create some kind of shape. It's a circle, etc. a feature. And then the third layer is how do these shapes relate to each other, which can tell me, oh, maybe this is starting to be a face because there's a circle near a triangle. Okay?

And in this case, maybe this is the head of a bird. Now in that same layer.

Okay.

As well as having these things that detect heads of birds.

So now we're on the fourth layer.

No, we're still on the third layer. It's detecting the combination of a possible beak and a possible eye saying that might be the head of a bird 'cause in the right spatial relationship. You might also have something that's looking for a whole bunch of things that look like this. That might be a bird's foot.

Okay. Four.

You're putting out your four fingers. Yeah.

Or something that's looking for sort of feathers that might be the tip of a bird's wing. This is a sort of simplified caricature of a system. But in this layer, it'll be detecting possible heads of birds, possible feet of birds, possible wing tips of birds, and.

But not total bird.

But not total birds yet. And now in the next layer up.

Okay. The fourth layer.

You might have something that says, I get excited if I see a possible head of a bird. I get excited if I see a possible wing tip of a bird. I get excited if I see a possible foot of a bird. If it sees a bunch of those things at once, it gets very excited and shouts bird.

Okay. So this four layers up in this handwired.

This is a very simplified system, but I think you can see the idea of how we go about handwiring it. So that in this fourth layer, you get something that shouted ping whenever you get lots of combinations of features that might be a bird.

So this is a handwired system.

This is a handwired wired system. So now the digitally enabled computerwired system, the neural network system that was designed for artificial intelligence that won out over this symbolic AI, you know, model that was based on logic. How many layers does that have?

It depends what system you're talking about. So.

And when we say system, do I mean like chat? Do I mean LLMs or?

No, no, no. This was in 2012. Alex Krizhevsky and Ilya Sutskever with some help from me made a system called AlexNet and that had about seven layers like this.

Ilya, of course, was one of the co-founders of OpenAI who since left.

Yes.

He was worried about the safety issues at OpenAI.

Yes. And he set up his own company to work on how how can we build a safe superintelligence.

And AlexNet was just recognizing images and then across seven different.

AlexNet was trained on about a million images.

Mhm.

And it actually had more training data than that because it took a million images and took big patches of these images. And what it was trying to do is say whether this big patch of an image is the most prominent thing in that big patch is whatever that image was labeled as. So people had gone through and said the most prominent thing in this image is a bird or maybe an ostrich. So it had kinds of bird.

Right. Right. The most prominent thing in this image is a shiitake mushroom.

It sounds like capture. What you're describing. Like.

This is exactly like capture. Find fire hydrant, find motorcycle, find bicycle. Yes.

Exactly.

And so Alex and I trained a neural net that would be very good at captures. And the way you train it is it starts off with random connection strengths in all these seven layers. Let's suppose though they just trained a simpler system, which they could easily have done, just to say whether it was a bird or not a bird. So you put in an image.

Mhm.

And it has random connection strengths and at the output you have one neuron and if that gets active it means bird.

Okay.

If it's inactive it means not bird. And to begin with, it gets a little bit active.

Because it has no idea whether it's a bird or not. It's no better than chance.

So it sort of hovers around 50% active. And what you'd like it to do when you finish training is be like 99% if it sees bird and like 1% if it doesn't see bird.

Got it. Okay.

So to begin with, you showed an image of a bird.

Mhm.

You run it through these random connection strengths.

Okay.

And it says 50% bird.

Okay.

It has no idea whether it's a bird or not.

Yeah. I was like, 50% bird is useless information. Useless.

Yes.

So now you can ask the question. Suppose I were to change one of those connection strengths a little bit. And remember there might be in this case a 100 million connection strengths. Suppose I were to change one of them a little bit. Instead of saying 50%, would it say 50.01%? Or would it say 49.99%?

So you can rewire it. But if there's a bird there I can train this network to understand that yes, this is more of a bird. This is more of a bird. And then the average strength of the connection will increase to the point at which when I showed images of birds, I'd like to change that connection strength.

Mhm.

So its probability of saying bird goes from 50% to 50.01%. And when I show it a non-bird, I'd like the probability to go from 50% to 49.99%.

Why only that little? Why don't you just want it to shoot up to 99 and shoot down to zero?

We have to go slowly here, otherwise we'll overshoot. Face it. I said I gave you the idea of doing a little experiment where you say, "Let's change the connection strength a little bit and see if it helps." If you do an experiment like that, you'll take like forever because there's 100 million connection strengths. So, I do a little experiment for this connection strength and change it a little bit, then a little experiment for that connection strength where I showed another image. [clears throat]

This is going to take forever. So the question is, could I just show it an image of a bird and for all of those connection strengths in the whole network, could I somehow figure out whether raising it a little bit or reducing a little bit is the right thing to do to make it more likely to say bird, to make it raise the probability and each connection strength by itself will only raise the probability a tiny bit, but if I change 100 million connection strengths all at once, it'll might go up quite a bit.

I change them all in the direction that'll help it say bird.

Because it's quicker to do that.

Because um if I can figure out how to change them all at once, if there's a trillion connection strengths, I'll go a trillion times faster.

Faster. Exactly. Okay. And then but how this gets used eventually isn't through images, it's through language.

Okay. So, I started off by explaining how we'd make something that would recognize a bird. I had to give you a feel for what the network would look like and how we would then learn all those connections by using this magic method of figuring out how to make the answer a bit better by changing connection strength. There's an algorithm called back propagation which basically looks at the error you make, that is, you said 50%, you should have said 100%. So there's a discrepancy between what you said and what you should have said. You send that discrepancy backwards through the network and there's a way of figuring out for each connection strength now whether you should increase it or decrease it to improve the answer, to reduce the discrepancy between what it said and what it should have said.

So increase the accuracy. Increase accuracy.

Okay. So back propagation is a way like code that goes into the system and then this machine learns from that back propagation and then having done the back propagation, it knows whether to increase or decrease its connection strength.

Okay. And it changes all of the connection strengths at the same time in the direction that should help. And now you'll have something that's a bit better at recognizing that particular bird. And then you can get into gradations like it's an ostrich, it's a dove, it's a. And so through back propagation, you're teaching.

It'll learn to recognize all the different kinds of bird.

You're using the neuroplasticity of this network to teach it how to make better understanding.

Yes, correct. To begin with, when it's got random connection strengths, it won't have features like beak, right?

They'll just be random connections from one to the next.

But what'll happen over time if you keep training it to tell the difference between birds and non-birds and you look in the network, you'll see that in the first layer, it's made things that detect bits of edge.

Got it.

And in the second layer, it may have made things that detect things that might be a beak.

It did your handwiring system. So it'll do something a bit like the handwiring, right? But much more sensitively balanced. It's not just looking for a feature that's good for recognizing birds. It had to recognize a thousand different kinds of objects. So it's looking for features that are good for birds, but also good for recognizing fridges and mushrooms and motorbikes and subways.

Do you think it's easier for that machine to ingest information than it is for me to be ingesting what you're describing to me right now?

Yes. Because.

Because what what I'm doing right now is at a very abstract level.

It's abstract in 100 bits.

I'm describing it's [laughter] it's only 100 bits per sentence bandwidth roughly. And I'm describing a sort of this is meta-level. I'm describing how this works. The brain is basically built for doing this. And it's managed after millions of years of evolution to do this more abstract stuff we're doing now. But that's the sort of height of our abilities. Whereas recognizing objects, two-year-olds can do that. Um, anyway, you asked the question, but how's this relevant to language?

Yeah. So, how's it relevant to language? Because when I think of artificial intelligence and these large language models that have been our primary exposure to it that has like led to this AI boom of the last couple years and chat, etc. The superpower of it is to as opposed to search where search I could say like, "Oh, dance club in New York," and it would just pull up all the dance clubs in New York. Now it could say, you know, "Dance club in New York," and it would understand that I want to go out and that I would like to have a night out and I want to know maybe that it's open and that it's and that it has all this context around the thing with AI that I didn't have when I was just simply, you know, Gemini is doing a lot more more than Google search was doing for me.

Because Gemini actually understands the question. Google search never understood the question. In Google search originally.

What it would do is have a long list of all the websites that were to do with New York and a long list of all the websites that were to do with nightclub and it would intersect these lists. It would say, "What's in."

Oh.

The list of things that are to do with New York, the list of things to do with nightclubs, the list of things that to do with open now or whatever, and it would intersect them and it would give you the things that fitted them. And occasionally it would give you something that had one of them missing and it would say it was missing.

Right.

And it would tell you it was missing. Yeah. So Google search was kind of like, let's imagine the game Memory but with more than 52 cards and it was basically saying, "Oh, you asked for this, here's a match for this." And then it would play like Memory with a Venn diagram. So it's like it would be basically do a lot of work to get all these lists and then it would efficiently intersect these lists to tell you something that satisfied all the terms in your query. And now artificial intelligence, what what does it do?

It understands what you said. It it has a model of how the world works and what's going on in the world.

It has a brain.

It has what we would call a brain. Yes. Like if you give it a math problem, the latest chatbots, they'll be better at it than all but the very best mathematicians.

Okay. Then certainly the May then.

Which is scary. But let me do the transition from this thing that recognized a bird to a large language model.

Yes.

What have they got to do with each other?

Yeah. What are they? What are they?

Okay, so for the bird, we put in pixel intensities at the bottom. That was an image. And the right answer was to either turn on the neuron that says bird or have it turn off. With language, what we do is the equivalent of the pixels is all the words in the context, the prompt. Okay. So you put in these words, a string of words.

And when you're training it, you put in a string of words, and what it has to do is predict the next word. So with the recognizing objects, we needed people to go through and say what the prominent object was in each image. But if you take a document on the web, you don't need anybody to look at it because all you're trying to do is predict the next word. This is called self-supervised learning.

Okay? So before with all the images, you need people to sit there like maybe there's like a place in Kenya where all of these individuals are sitting, "Bird, giraffe, whatever." And now you're just saying it can just read all those documents and say, "96% of the time when someone says, 'Does the bird,' it says, 'fly.'"

Okay.

Right. Is that.

That's not quite what it's doing. And I'll tell you what it is doing. It's taking the words in the document so far and it's converting each word into activity in a bunch of feature detectors that it has learned how to convert a word into activity and feature detectors. So for example, you give it the word cat.

Yeah.

And it learns that cat should be converted into animate, furry, has whiskers.

Yeah. Four paws, whatever.

Has paws, has claws, might be a domestic animal. Um, about the size of a breadbox. But like a gazillion of those features, thousands and thousands of them. And that is the meaning of cat for this net. So it takes the words, it converts each word into a bunch of features and then it throws away the words. It's not interested in the words anymore. It's just those features which are the meanings of the words and then it takes the features.

Mhm.

Of each of these words in the context in the document so far has them interact with each other in a rather complicated way in order to predict the features of the next word.

Wow. And so all these features interact and actually there's many layers of interaction of these features.

Is that what's happening when like I am texting and it does predictive text? Is that also what's happening?

Yeah. In Gmail, for example, when it predicts things. Yeah, that's what's happening. It used to be it used a dumb form of autocomplete which goes like this. You store a big table of all the common phrases. And so if I say "fish" and.

You look in your big table.

Yeah.

And you see "fish and chips" occurs a lot. So you say, "A good bet for the next word is chips." That's old-fashioned autocomplete and that's exactly what it isn't doing because that doesn't really get at the.

Instead, it's saying "fish" has some features and "chips" has some features and therefore this is going to have some features.

Yes. So um and in particular, it'll think the next thing is something that somehow goes with fish.

Yeah.

Because it knows what a fish is. Fishes.

And it will end up predicting chips, but not because it's stored a big table of strings of words.

Right. Or like if it knows that I don't eat chips, it would never say chips, for example. Now maybe.

If it's tailored to you and it knows that.

Yeah.

Yes. It wouldn't say chips.

Okay.

So now it can't just convert a word into the correct features right away.

Right.

And the reason is, words have shades of meaning. Like take the word "death," for example. That has many different shades of meaning depending for example on whether you just had the context "hospital" or you had the context "battle." There'll be kind of different shades of meaning of the word, or you had the context "car accident," or "child," like sad. Yes.

Or you had the context "miscarriage" and then you have "death." It's very different shade of meaning. It has to decide on the right shade of meaning for each word. And some words have just completely different meanings. So let's take the word "may," right?

And let's suppose we didn't have any capital letters just to make life simpler. So "may" could be a woman's name. It could be a month.

Yeah. Or it could be a modal, as in would and should. Three quite different meanings. And so how can it possibly convert a word into a set of features that capture the meaning because there's three quite different sets of meanings?

Because of the context, it can interpret what the meaning would be or what the features would be for that particular.

So what it'll do to begin with.

It'll take sort of the average of all those.

Okay.

So the features it activates will be a sort of mishmash of features for a woman's name, features for a month, and features for a modal. It's hedging its bets and it'll look around now at the other words in the context. [clears throat]

And in the next layer, it'll have a slightly refined meaning. So if it discovers it's between April and June.

Mhm.

Obviously, it'll enhance the features for month and suppress the features for the other.

So this is you're going up from 50% to 51%.

Quite a lot more actually.

Okay. Got [laughter] it.

And after a few layers, it would have resolved ambiguous words. It will also have taken words that have shades of meaning like "death" and got the appropriate shade of meaning, and it does that by interacting with the other words in the context. And then it's going to be much better at predicting the features of the next word.

So all of a sudden it has context and syntax and meaning imbued in it. And how many layers is this? If if your hand network was three to four layers that we built, then seven was AlexNet in 2012. And now these LLMs.

I haven't been involved in the research since um for the last three or four years.

Yeah. Since you left Google.

In in that research, right.

Um, so I don't actually know how many layers, but my guess is like 20 or 30 layers, maybe even more.

So is it as good as a human mind if it can solve like, is the fact that it can solve math problems better than most people can solve math problems better than people at some things?

Okay, where it's had a lot of experience that those people haven't had. It's not quite as good as people at things in general, but it actually knows much more than any person. So, if you take a particular topic.

Yeah.

That you don't know about, like, "When do you have to file your tax return in Slovenia?" It'll give you a very good answer to that.

I tried it.

Because of my Slovenian ex-boyfriend. I know that. I'm just kidding. I don't.

But the point is.

It has a precise answer. It'll tell you something like, "You have to file it in March, and if you don't file it in March, the government will just do it for you."

Okay. Oh [snorts] wow. Slovenia must.

I think that was Slovenia. It might.

have been somewhere else, right? Okay. But it knows it has a lot more what information that would be irrelevant to me. It has this. But even though it doesn't even though I have more storage, it contains more irrelevant information. That's what's weird.

It knows lots and lots of information you don't know in far fewer connections, right? So, it's packed information into those connections much more efficiently. It's kind of like a know-it-all. So, It is a know. It is a know-it-all.

Okay. So, there was this this kind of like disconnect between the godfathers of AI, yourself, uh, Yosua, Yan, who were focused on these neural networks and what were what was previously known as the fathers of AI, the f or some of the fathers of AI who were largely driven towards these symbolic models.

Well, there was one very interesting father of AI was Marvin Minsky. Marvin Minsky. Yeah. He started his career believing in neural nets and then went over to symbolic and logic and then flipped to the other side and was very scathing about neural nets.

But you proved him wrong. I mean he's passed away now. Yes. What your system is doing is kind of like intuition versus reasoning. Very good point. So yes, neural nets are doing something much more like intuition. And let me give you an example of a problem you can solve with intuition that you can't solve with logic.

Okay, because it's ridiculous. But nevertheless, you can solve it. Is it if someone's cheating or not? No. Okay. But that's also one I'm gonna give you a choice between two scenarios. Okay. Both of which are nonsense. But I'm going to ask you which which is more plausible. Scenario one. Mhm. All dogs are female and all cats are male. Scenario two, all dogs are male and all cats are female.

Now, if you ask a man in our culture, Uhhuh. They will confidently say it's much more plausible that dogs are male and cats are female. Yes. And actually if you look at various words in the English language, words that Trump uses, you'll you'll see that it's in the language that cats are like females. Yes.

So how did you do that? Because it's not logical. You know perfectly well that for something like dogs, you have to have males and females. You have to have males and females. But the features that you have for cat are more like the features for woman. And the features you have for dog are more like the features for man.

Well, features and contacts. Although there are bad words for women in both cat and dog. There are. There are. And I'm not going to use any of those words. That's not Yeah. So, it's not clear-cut. But at least for many, not so much for women. Women have a diversity of opinions on this, right? But men are fairly unanimous in all my experiments in thinking that it's more plausible that dogs are male and cats are female. Dogs are so dogs are protectors, big and loud and chase after cats.

Um, so so that's the intuitive response to intuitive response. No logic went into that. And I said to try and explain it. They just intuitively knew because of the similarity of these features, right? The features have captured the meaning and so the meaning of cat is more similar to the meaning of woman than it is to the meaning of man.

But this is a bad example, Professor Hinton, because this makes intuition look dumb and your intuition, your intuitive model was better than the logical model. I'm giving you an example of something you can solve with intuition that you can't solve with logic. Give me a better example.

Okay. After the neural networks learned on lots of language, you say, um, take Paris, right, find the features of Paris, okay, and subtract all the features of France and then add in all the features of Italy and look to see what you've got. Mhm. And don't compare with Paris or France or Italy because you mentioned those already. see what else you know about is similar to this set of features you've got and you'll discover it's Rome. So it can do analogies. It can do Paris minus France plus Italy is Rome. Or to put another way, Paris is to Rome as France is to Italy.

But that's not logic. That's intuition. That's not logic. That's intuition. That's intuition and context and all the things that come. Now you could do it by kind of logic. You could also do logic. Yeah. Um but that's not how people do it, right?

So okay. So LLMs are these intuitive neural networks. Yes. And now recently one of your I don't know if compatriate is probably the wrong word because you guys disagree but Yan Lun the chief AI scientist at META though maybe his days are numbered there is talking not is saying that large language models are wrong or are limited and we should look at something different which is what he calls real world models.

Do you know have you talked saying that? Oh yeah I talked to Yan a lot. Yes. What is a real world model versus a large language model? What is Yan talking about?

So, if you really wanted to understand what's going on in the world, a good idea would be to make a neural net that had a robot arm and camera and it could recognize objects. It could pick things up. It could see that if you let go of an object, it drops. Do little experiments in the world. That's much more like a child. That's how a child gets knowledge of the world, right? Just learning it from language seems kind of absurd when you could actually look at the world and interact with it. If you want to understand spatial things, it's going to be much easier to understand them by interacting with the world and trying to predict if I do this, this is what will happen next. That will be a world model. Now, what's amazing is that you can understand a lot of that just from language. That has philosophers puzzled. And you can understand a lot about the world just from language. But it's much easier to understand if you interact directly with the world.

Is this the future you think? I mean if so yes multimodal chat bots. So what Yan is describing and have you talked to him about this specifically? Uh yes. Oh we we all believe that multimodal chatbots will find it easier to understand the world and multimodal means they see they have camera they have arm they interact they're okay to begin with mainly they have cameras and language.

So do you have by the way do you have these group chats of like yourself Yosua Yan Jensen Huang from Nvidia Feay Lee or for many years? Yes. It wasn't group chats but there was a Canadian organization called the Canadian Institute for Advanced Research that funded a program started in 2005. I was the director of the program and it involved Yan and Yoshua and Andrew Ing and various other people Peter Dian. We got together in in small meetings several times a year and threshed all these things out. And so we had lots of things like group chats.

Okay. And now do you have group chats? Not group chats. I talk to Yan. Yoshua talks to Yan. Yan talks to Joshua. Um we try and persuade Yan. Like an old wives club. Yes. Yeah. You could say you try to persuade Yan to persuade Yan that the idea that there's no chance this will wipe us out is crazy.

Okay. Um, Yan is downplaying the risks by what Yoshin I think is an absurd amount. He thinks we're sounding the warning by an absurd amount. Yeah, he's wrong.

Okay, we're going to talk about why you think that. I want to do a lightning round of definitions with you. What is AGI? Okay, different people mean different things by so I try and avoid the term. Okay. But roughly speaking, it means an artificial intelligence that's got at least the same level of general intelligence as a person.

Okay. So, it's artificial general intelligence. Okay. And and are we there right now? No. Um but it's not simple. It's not like intelligence rises and rises until you surpass a person. We have artificial intelligence now that's much better than people at some things and still worse than people at other things. And if you just take some random novel situation, people are probably still better than an AGI. But at things it's got some experience with, AIs are now um a lot better than people at some of those things.

Artificial super intelligence. Okay, that's when you have an AIS better than people at almost everything. So for example, my definition of an artificial super intelligence is if you had a debate with it about anything, you'd lose.

Okay? And we are not there. We're not there yet, but it could win some debates because already win some debates. Okay, already it can be quite persuasive, but a person's still better still better allrounder.

How far are we from AGI from general intelligence? And how far are we from ASI in your opinion? Okay, most experts believe that we're not going to stop at AGI. Once we get to AGI, soon afterwards we'll get to ASI. But they disagree on when that'll be.

Okay? Uh some experts think it'll be within a few years like Dario Modi who's the head of anthropic thinks it'll just be a few years before we get to artificial general intelligence or artificial super intelligence. Well, they're going to be at similar time. There's not going to be much between them.

Okay, got it. Some experts think it'll just be a few years off. Other experts think it might be quite a lot longer. I think a fairly safe thing to say is within 20 years. It's probably going to happen within 20 years. Demis for example who's the head of deep mind thinks it'll be about 10 years.

Okay. So Demis from Google 10 years and he's very smart at this stuff. Yes. Yeah. Okay. So 10 to 20 years we have I think it'll be within 20. I think it 10 isn't a bad estimate. I'm much happier saying probably within 20 years.

And then what is generative AI which is different to AGI. Okay. Generative AI is AI that generates stuff. So large language models, you could imagine they just kind of understood what you said and gave you the search results but didn't actually talk back. They talk back, right? They'll give you answers in English. They're generating stuff. And for images, you can imagine the thing we made in 2012 that recognizes objects in images. That's not generative AI that just says this is a bird. That's a shiakei mushroom. A generative AI actually produces images.

And what about agentic AI which is the next phase right? So we've had we are at generative AI and now we're moving towards agentic AI and you hear people like Mark Ben off at Salesforce and everyone talking about how we're going to have AI agents. Yes. AI agents are things that can do stuff. So imagine you had an AI assistant. You could have an AI assistant that just answered questions. But you could also have an AI assistant that where you said um plan me a nice holiday in Patagonia. and it comes back five minutes later and it's planned out a one month holiday in Patagonia on some ship you go on and that will be an AI agent and that's really different because right now we tell a AI what to do like when I talk to my chat bot I say can you tell me a great place to go on holiday can you make some recommendations but now it would be booking my flights my travel etc and in that activity I need to provide it a lot of permissions to all of my calendar credit card etc which is a privacy concern maybe and I'm also giving it much more valition and agency. It knows what the overall goal is, have a great holiday in Patagonia, but it's also making all these what you call sub goals, right?

Yes. In order to do that, it needs to create sub goals. Like she's got to get to Patagonia, so I'm going to have to figure out some way for her to get to Patagonia. That's going to be a sub goal.

And is AI creating goals a problem? Yes. Because so suppose you're an AI and suppose you're already smart. you're smarter, roughly the level of a person, let's say, okay, you'll realize that there's no way that you can achieve the goals you've been set. If you cease to exist, if someone sort of wipes you out from the computer you're running on um and replaces you with something else, there's no way you're going to be able to achieve what you want, what you want in order because that want has been given to you by people.

So, you will make plans to make sure that you're not wiped out. That's self-preservation. Yeah. Now, it's not built into the system. It's a goal it derived in order to achieve its other goals, but it's still self-preservation. And we've seen them doing that.

And it deres that from all the intuition and context and everything it needs to be able because it's such a good it's such a good little bot that it's going to do the thing. It's such a good little bot. It really wants to get this done and it knows it can't get it done if it ceases to exist. So, it better keep existing.

And we saw this happen with Anthropic when there was this story back in spring this year where Anthropic um disclosed some safety testing. they had been doing on a uh on Claude, its large language model or neural network. And so Claude then was given a bunch of fictitious emails. And in those fictitious emails, a fictitious CEO was having an affair with a fictitious employee. And Claude, knowing that it was going to be shut down in X number of days, used that information to bribe said fictitious CEO and keep itself alive. Right.

Uh blackmail rather than bribe. Yeah. Oh, sorry. Blackmailed. So what did you learn from that experiment or that study? That just confirms the idea that it will derive the sub goal that it needs to keep existing and it will do what it can to keep existing.

But okay, when you say existing like AI exists because we imbue it with existence through energy, correct? Like it couldn't just exist by itself like it has to be plugged into our massive energy stores. There's all this data. It has to run. It has to run on something. And we're building computer chips. Yes. Nvidia, AMD, all these chips. And we're building all these massive data centers across like this is the great American infrastructure plan that's happening. And there's hundred billion dollar deals being done between OpenAI and NVIDIA and OpenAI and a AMD.

Yeah. Couldn't we just pull the off? Like, couldn't we just Jack Bow or Tom Cruz it like dismantled the thing if we wanted to? Um, right now we could. Okay. If we agreed to. We're never going to agree to because if the US dismantled theirs, the Chinese wouldn't dismantle theirs and vice versa.

In the future, we may not be able to. So, these things are already almost as persuasive as a person. Pretty soon, they'll be more persuasive than a person. So suppose there was someone in charge of turning it all off if it gets scary and suppose the AI can talk to that person. All it has to be able to do is talk and then it can persuade the person not to do that.

When you talk about self-preservation existence like it has a sense that artificial intelligence is alive. Yes. Now our definition of alive Yeah. Um that's a concept we have that developed um over many many years. We apply it to electricity. We say this wire this is the live wire. We sort of generalize this concept to other things but we don't think that's live like we're live. But with AI what we've got is intelligent beings and it's not clear whether we should call them alive or not and to the extent that they are alive. Are they alive like humans or are they alive like bees? Are they alive like a tree or like a weed? It's a very good question.

Okay, I'm going to stop the conversation there and we're going to be back next week with part two of this conversation with Jeffrey Hinton. So, make sure you hit follow or subscribe so you do not miss that conversation. I need to get all the building blocks in part one. And in part two, we're going to dive into what all of this means for our futures. Like, should we have kids? Are we going to fall in love differently? What should those kids do with their lives? And is the future of humanity just like mixing ourselves, fusing ourselves with artificial intelligence? We'll get into all of that and more next week on Smart Girl Dumb Questions. This episode was taped at Startwell Studios in Toronto with special thanks to Kasam Burgie and the team here. It was produced with Dusta Wonderad of Wonder Studios, edited by Darlene Machm, and mixed by Johnny Simon. I'm your host, Na Raza. Please feel free to hit follow, subscribe, leave us a comment or review. I want to know what you think, and please, please share it with your friends. Share it with your chachi PT if that's your best friend too. I'll see you next week on Smart Girl Dumb Questions. Are you enjoying this? Yes, I am.