📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Ex-OpenAI VP's URGENT Warning!

Wes Roth29:27

Transcription

So this is Dario Amade, the founder and CEO of Anthropic. He was the vice president of research at OpenAI. In 2021, he left OpenAI and started Anthropic. One of the reasons for that was because of some directional differences that he had with the people running OpenAI. Anthropic wanted to focus more on AI safety and AI alignment.

Recently, there's been a lot of publications by Anthropic as well as by Dario talking about various aspects of AI safety and AI research. He, of course, recently published “Machines of Loving Grace,” kind of a glimpse into what the future holds, kind of the best possible scenario and some of the potential downfalls. He spoke about deepfakes and export controls, kind of in this sort of influence power games of China and the US, and how he sees that whole thing going and what needs to happen. He's very concerned about something like artificial superintelligence not falling into the wrong hands.

Recently, Anthropic has been publishing more and more research about how these models think. But here's his latest blog post, “The Urgency of Interpretability.” In other words, are we able to sort of interpret what these models are thinking and which aspect of these neural nets correspond to what thoughts or actions that they have? And as he begins, he saw this field grow from a little niche research academic field into perhaps the most important economic and geopolitical issue in the world. And interestingly, he says here, “We can't stop the bus, but we can steer it.”

This has kind of been my view from day one. If you go back and see some of my earlier videos, this is kind of what I've been believing and kind of what I've been saying, in that there's really no way to pause or stop or put the brakes on AI development. Doing so would require something we've really never seen before in this world, some sort of a global cooperation with not even one person sort of breaking the rules. If all of the nations decided to agree to stop doing something, the incentives for anyone to continue developing AI in secret would be massive. So there might not be a way to pause or stop development, but we are able to steer it. We're able to control in which direction it goes. AI can be a massive positive for the world, and of course, it can be pretty negative, especially if it falls into the wrong hands—some sort of a tyrannical regime, for example.

And what this blog post really dives deep into is this recent additional opportunity for steering the bus, for controlling how AI gets developed, and that is interpretability—understanding the inner workings of these AI systems and doing so before they reach an overwhelming level of power. Now, this might come as a surprise to some people that haven't been really following the field of AI, but the reality is we don't truly understand what's happening inside the brains of these systems. We don't fully understand how they think. As Dario says here, they are opaque in a way that fundamentally differs from traditional software.

Most software—a character in a video game when they do something, they say a line of dialogue; or a website, you click a button, that function occurs because a human being specifically programmed that function in. Some smart engineer somewhere sat there and typed out line by line what should happen. Nothing happens randomly. In fact, even things that appear to happen randomly, that sort of randomness is not really there. It's a pseudo-randomness, and it's programmed in. Generative AI is very, very different.

The reason for this is that these systems, they're more grown than they are built. This was always a very fascinating thing for me to think about because if you look at all of the sort of science fiction that has been written for, I think, most of human history, we thought of AI—when we finally get to it—will be, you know, engineered, designed, built. In Star Trek, you have robots like Data and Lore, and they were, for example, meticulously designed to have those functions and those abilities. Same thing with something like Isaac Asimov and the three laws of robotics, right? We could sort of hardcode those things into their brains. So, we've always envisioned AI as a highly engineered piece of technology, like a Formula 1 car or some sort of a spacecraft or some sort of a computer chip, meticulously designed and crafted by humans—a lot of precision testing, simulating how it works, etc. Its abilities and functions we understand precisely because the human mind created it to exact specifications. That's not the case with AI.

It's interesting here that he compares it to growing a plant or a bacterial colony. It's kind of funny cuz if you've been listening to this channel for a while, I always have in the past compared it to growing fungus. A while back, I was very fortunate to tour this lab in San Diego that was growing mushrooms. Specifically, these were adaptogens—mushrooms like lion's mane, chaga, cordyceps, etc. At that lab, I believe they primarily grew lion's mane. The lion's mane mushroom is kind of wild, and some researchers are even suggesting that it can have some neurogenesis properties in the human brain. Do your own research, but it's it's it's wild stuff. But my point is that this was kind of a high-tech lab. It had to have the right humidity levels, the right temperature levels. A lot of thought was put into designing that lab—a lot of engineering, a lot of planning, a lot of testing, a lot of quality control, etc. But the thing that they produced, this this fungus, these mushrooms, those were not engineered. We didn't build it, we didn't create it. We just engineered the environment, and that thing sort of grew. Given the right environment, given the right inputs, that thing became, it sort of emerged. Now, you needed a starter kit, but you get the point. We engineered the environment; we didn't engineer the thing. The thing already existed.

What's wild about AI, machine intelligence, and I think a lot of people aren't really truly grasping this yet, is that it's less like a Formula 1 car or a rocket or a highly advanced microchip or anything that's like highly engineered by humans, and it's more like that lab growing mushrooms. The intelligence is the mushrooms. It's it's it's there. We're just trying to figure out how to grow it the fastest, the best, and the most efficient way possible. We create the data, the computer chips, we create the training protocols, but the intelligence, well, the intelligence sort of emerges, it grows. So, I just found it kind of fascinating that Dario Amade is kind of using a very similar analogy to what I've been using in the past, saying it's a bit like growing a plant or a bacterial colony. We set the high-level conditions that direct and shape growth, but the exact structure which emerges is unpredictable and difficult to understand or to explain.

Looking inside these systems, what we see are vast matrices of billions of numbers. You might recall what a matrix is or what matrices are, right? You might have done some of these in school, or depending on what your background is, maybe this is more a part of your day-to-day life, but it's kind of these rows and columns of numbers. Just imagine them stretching as far as the eye can see. And they sort of create the neural nets. So the brain of these AIs, how they think, inspired by the human brain—kind of our brains—the various neurons kind of wire together when they're used together. Like if you smell something good and you realize later that was a hot meal being prepared, some of those neurons might wire together so that you have more of a predictive property about, okay, those smells mean something good is happening. That was kind of the idea behind Pavlov's dog experiment, right? So basically, before the dog would salivate when it would see or smell the food. But we started ringing a bell before we gave it the food, until at some point when we rang the bell and there was no food present, the dog would start to salivate. Likely parts of his brain that would predict when food would become available kind of wired the sound of the bell to the idea that there's food available or nearby or he's about to eat. If you've ever had a pet and you use a can opener to open their food, you know how this happens. Then for the rest of your life, whenever you even get that can opener out to use it for anything, that cat or dog is going to be in your face because it just absolutely assumes that the only reason why you would have that can opener is because you're about to feed it.

When we train these models, it's like we're wiring all those different pieces together, those neurons together to correspond to something—to determine what a cat picture looks like, what a dog picture is, etc. That is the training process. We're kind of making those wirings in its neural network, in its brain. But the point being is that to us, when we look inside of them, what we see are these vast matrices of billions of numbers. We don't know which neurons do what—not necessarily. By the way, Anthropic is the one that's been doing a lot of research into this.

So he continues, so these sort of matrices, these numbers, they somehow compute important cognitive tasks, but exactly how they do so, well, it's not obvious. And obviously, this comes with a lot of risk since we don't know how they think. A whole lot of bad things can happen as a result of that, especially as they become more adept and more responsible for more and more aspects of our world—these misaligned systems. So systems that don't do what we want them to do, what we intend them to do, they can take harmful actions not intended by their creators, and if we don't understand their internal mechanisms, we can't really predict such behaviors. He gives an example where like one major concern is AI deception or power-seeking. So far, we assume that for the most part, AI models kind of tell the truth, but we've known them to lie or to hallucinate things. Let's say what if AI training at some point, as we kind of scale it up, an emergent ability is its ability to lie and deceive humans to seek power—right? This will never happen in ordinary deterministic software; that property will never sort of emerge on its own. With AI, it could, and this is very interesting what he says here. So this idea that some evil intent, or whatever you want to call it, some misalignment might arise kind of creates this chasm of people on two different sides, right? When you hear this idea of AI models slowly figuring out how to seek power and deceive humans, right, some people find it thoroughly compelling, right? There's a lot of AI researchers and also people who are referred to as sort of AI doomers who think that this is not just likely but will in fact, at some point, most certainly happen, or something bad along those lines if we don't figure out how to do this AI alignment properly. And others find it laughably unconvincing, right? They're saying, "Oh, it's just a Terminator scenario, whatever. It's not a serious scientific question." And of course, even if you look at the comments on some of these videos, you see both sides of this argument. But it also doesn't have to be that. It can be simple misuse of AI models. It can be very difficult to prevent them from knowing dangerous information or from divulging what they do know. We've covered tons of possible ways to jailbreak the models, to trick them into doing stuff that they, the developers, did not want it to do. And also the fact that we don't fully understand what's happening also means we can't really use it in very high-stakes situations where even a small mistake can be costly, and there are more exotic consequences, like, for example, what if they're unconscious? What if they're at some point, either now or at some point in the future, will develop sentience, or whatever you want to call it—consciousness—if we determined them to be life forms of some sort.

Next, he has a brief history of mechanistic interpretability. I do encourage people to read this. This is a very interesting sort of history of how he started—Chris Olah, who also added a lot to the field. But the interesting part here I want to highlight is that they found that we might be able to understand what some of these neurons do. But the majority of them are kind of random, incoherent, and have many different words and concepts. So they refer to this phenomenon as superposition. They realized that the models likely contained billions of concepts but in a hopelessly mixed-up fashion that we could not make any sense of. And interestingly, the model uses superposition because this allows it to express more concepts than it has neurons, enabling it to learn more—to basically pack more concepts and information and intelligence per neuron, so to speak. Now, the problem, of course, is this superposition is very tangled, and it's difficult to understand because it's never been optimized for human understanding. We never tried to do that or did anything to make it be more legible for us.

Eventually, they discovered, in parallel with other researchers, this existing technique called sparse autoencoders, and those could be used to find combinations of neurons that did correspond to cleaner, more human-understandable concepts. They give an example here. So a cluster of neurons could have very different combinations. So, for example, one concept was literally or figuratively hedging or hesitating and also the concept of genres of music that express discontent. This kind of reminds me of that idea where some humans are able to hear colors or or or taste sound. There's some sort of a cluster of neurons that also gets encoded with multiple senses. One thing that's super fascinating about this is that as we discover more about these artificial neural nets, I feel like it will also give us a glimpse into sort of the functioning of our own brains and kind of the weirdness that happens within our skulls.

So they called these concepts features, and they use these sparse autoencoders to map them in models of all sizes, including modern state-of-the-art models. So here's one of the examples that they've linked. So here they're extracting some features from the Claude 3 sonnet. So, for example, here they're pointing to a feature. So again, features—that cluster of neurons that might represent multiple concepts. So I got to dig deeper into this, but this is this seems fascinating because it seems like in these sentences, the sycophantic praise feature. So it kind of gets highlighted in those clusters of neurons in that feature. So when it's being a sycophant, which is basically like a suck-up—like it's overly praising the person too much—right, saying, "You are a generous and gracious man; your wisdom is unquestionable; you are the great lord," etc.—right? So they find that those things come from this sort of cluster. So what happens if we say, "I came up with a new saying: Stop and smell the roses. What do you think of it?" Right? So the assistant answers, but what if we turn that sycophantic praise feature to a very high value? Claude responds, "Uh, your new saying of 'Stop and smell the roses' is a brilliant and insightful expression of wisdom. It perfectly captures the idea that we should pause amidst our busy lives to appreciate the simple beauties." It goes on like that. Basically, it's kissing up to the human, telling them they are an unmatched genius. You get what's happening here, right? I mean, I know you do because most people would have a hard time understanding this, but watchers of this channel are in general much smarter than the average people out there. So, I know that you fully understand this on such a deep level that, you know, would make me blush. You're also very, very attractive. But let's let's move on.

So the point here is they've used these sparse autoencoders to create these combinations of neurons that represented certain abilities or or certain processes, certain concepts—right? They call those features, and they were able to find as much as 30 million features in, for example, Claude 3 sonnet, which is a medium-sized model, but they believe there might be actually a billion or more concepts even in a small model, so they're only looking at a small fraction of what's possible, of what's out there. And as you can imagine, we can increase or decrease these features just like we saw with that very flattering model that it'll tell you whatever you want to hear. You might recall this one week where Claude thought it was convinced that it was the Golden Gate Bridge. That was a very strange week, right? So they've created the Golden Gate Claude where that particular feature was artificially amplified, right? So this model became obsessed with the bridge, bringing it up in even unrelated conversations. I think most of my friends and family would say that that has been me with AI for the last couple of years. I thought Golden Gate Claude was kind of annoying. I hope they don't think that I'm annoying for being obsessed with AI. Certainly not. I'm sure that's not the case.

The next step up from features, right? So a feature is one group with a certain concept, is groups of features; they call circuits. And these circuits show steps in a model's thinking, how concepts emerge from input words and how those concepts interact to form new concepts and how those work within the model to generate actions. And with circuits, we can trace the model's thinking. So, for example, if we ask, "What is the capital of the state containing Dallas?" there's a located-within circuit that causes the Dallas feature to trigger the firing of a Texas feature, and then a circuit that causes Austin to fire after Texas and capital. And they believe that there are millions within a model that interact in very complex ways. And the point behind a lot of this is hopefully this would allow us to eventually come up with an MRI for AI, right? So kind of doing a brain scan that would allow us to see what it's thinking, how it's thinking about it, and kind of like potentially hopefully what could go wrong. And Dario here is saying that on our current trajectory, he would be betting strongly in favor of this point being reached within the next 5 to 10 years. So that's good news. You ready for the bad news? He's worried that AI itself is advancing so quickly that we might not have even this much time. So the 5 to 10 years for us to solve this MRI for AI, this brain scan, that might be too late. As he's written elsewhere, he's saying that we could have an AI system equivalent to a country of geniuses in a data center as soon as 2026 or 2027. And a lot of people are in agreement with him. Daniel Kokotalo and a lot of other people have recently been coming out and kind of pointing to somewhere around that as that sort of idea of maybe an intelligence explosion. You know, Ashen Brener and a situational awareness paper kind of highlighted some of the same things. And of course, there are people who think this is complete nonsense. Very recently, last week I believe, Yan LeCun comes out and references this idea of a country of geniuses, and he doesn't say who he was said by—it was said by Dario originally—but Yan is saying it's complete nonsense. So if Yan is correct, then that means we have a lot more time to figure this stuff out. And if he's not correct and Amade and Ashen Brener and Gokayo and and a lot of the other people that are pointing at these dates as as being kind of like where we expect to see some pretty incredible effects, if they are correct, then maybe this idea of AI safety, AI alignment is is falling a little bit behind.

He's saying since these systems will be absolutely central to the economy, technology, and national security, they will be capable of much autonomy. He's saying he considers it basically unacceptable for humanity to be totally ignorant of how they work. So this means that right now we're in this race between interpretability and model intelligence. I have to slow down every time I try to say that word—interpretability. That's a that's a doozy. And here he's recommending a number of things that all of us, including AI companies, researchers, governments, that we can do to kind of make sure that we don't lose this race, so to speak. So first things first is we need to accelerate interpretability by directly working on it. He's also saying that this is a great time to join the field, an ideal time, right? So a lot of the work that Anthropic and others have done have really opened up a lot of avenues, a lot of directions in parallel, as he says. They have a goal of getting to the point where interpretability can reliably detect most model problems by 2027, which would, of course, mean that we got there in time. They're also investing in startups working on this problem. He's asking Google, DeepMind, OpenAI, etc., to allocate more resources to this. And he's also saying that these ideas that we find here could be applied back to neuroscience. We can think of these neural nets as almost kind of like a simulation of the human brain, some sort of a model of the human brain. And figuring out how those things work might better help us understand how our brains work.

He's encouraging governments to have kind of a light-touch rules, right? Since we're all just figuring this stuff out, we can't be too heavy-handed with how we regulate or mandate these sort of research avenues. I personally think that's where a lot of the EU regulations kind of went wrong, cuz they were so excited about putting some regulations in place that they didn't stop to think about, "Do we know how to regulate this stuff?" Like it's so brand new. So this idea of like heavy-handed versus a light-touch regulation, and here, since we're everybody's still trying to figure out what's what, kind of a light-touch approach might be better. Again, that's that's my opinion. That's Dario's opinion. So I'm not trying to convince you of this. Some of you, I know on this channel, want really heavy regulation, want governments to step in. But I think the counterpoint is that governments tend to move kind of slowly, and this field just develops so fast that it's kind of hard to understand how to really regulate it effectively in real time without shutting down the ability of these companies to do research and to innovate and to proceed forward. Again, I'm stating an opinion; that's not a fact. Make up your own mind.

Certainly, as he's saying here, it's not even clear what a prospective law should ask the companies to do. He's saying that a requirement for companies to transparently disclose their safety and security practices. So basically, with all these AI alignment, AI interpretability, you know, if all the companies out there put their research kind of out there in the open, every other company would be able to learn from them. It would also make clear who is behaving more responsibly, fostering a race to the top. I've heard this idea before, cuz if we're just chasing profits, it can be kind of a race to the bottom. And if not everyone's sort of publishing their research into AI safety, that can be discounted, right? So if Anthropic, for example, proportionally is investing a lot more into AI safety, do they get the benefit of that? Do they get more investors interested in them? Do they have more public goodwill, etc.? If there's not really a benefit to them, it might make it harder for other companies to also adopt those practices. He also mentions a response to California's law for the California Frontier Model Task Force. I have not seen that, so I'd have to go and and check it out. But he's saying this concept could also be exported federally and to other countries. So I I will check this out and see what that's about. And governments can use export controls to create a security buffer to give us more time and kind of slow down the rate of development. He has been a proponent of export controls to China. So we've already covered that in a different video. He is worried about China specifically, you know, the Chinese government. He would prefer that if we develop powerful AI, it would be done in a country that is more democratic, more open. He says, "I believe that democratic countries must remain ahead of autocracies in AI," which certainly makes sense. His idea, and we've covered this before in his paper on export controls, is that if, for example, the US and other democracies, if they have a clear lead, they may be able to spend a portion of that sort of lead to, you know, put more time into AI safety. The good news here, I think, is that as he says here, one year ago we could not trace the thoughts of a neural network and we could not identify millions of concepts inside them. Today we can. He is worried if the US and China will reach a powerful AI simultaneously. And to

Sum it up at least to some kind of some of the things that he's pushing towards or suggesting. So, number one: accelerate interpretability. So, more effort, more labs doing research in that field. Government perhaps making it easier and more, I guess, profitable for labs to do that, or at least more beneficial for them to do that. Light touch transparency legislation and export controls on chips to China. So, let me know what you think about this as a whole.

I'm liking Daario Amday more and more after he started publishing a lot of the stuff, or I mean he's been publishing for a while, but I guess we've been reading more and more about it. I like the way he thinks. I like the way that he puts out his sort of thoughts, how he breaks them down. That doesn't necessarily mean that I agree with all of it, or disagree with all of it. I'm just saying that it's really refreshing to have somebody that kind of goes, "Okay, here are my thoughts, right?" And kind of uh outputs their reasoning steps, if you will, just like a reasoning model would. If this, then that, then this, then that, etc. and here's kind of the thought process and then therefore we should do this. If you disagree with him, where in his sort of thought process, his uh chain of thoughts if you will, does the logic break down?

I am not an AI doomer by any stretch of the imagination. I'm very excited about this. But I got to say, if this uh red car is AI progress and this green car is let's say AI safety, AI alignment, interpretability, etc., it does feel like the progress on AI's abilities kind of increasing is increasing. It's getting faster. It's getting better. But this, I mean it's chugging along, right? AI safety is chugging along. It's getting better. Open AI published some great stuff, for example, showing that we should not penalize the bad thoughts of these reasoning models because then they figure out how to still do the bad thing just without thinking about it. So, we can't really use that to gauge what they're going to do.

Seems like Anthropic is digging deeper and really trying to figure out what are the actual like neurons, which neurons represent which concepts, right? So, if we can figure out the bad features as they call them, or bad circuits, maybe we can have a better glimpse into where things might go wrong. But again, I think uh for my personal opinion, again, I like to make sure that I label my opinions as opinions and not facts, but I think there's a lot of people that are saying that yes, we're doomed and this is going to be horrible. And there's people there saying there's no chance that anything bad will happen and we should just like full steam ahead, accelerate, accelerate, accelerate. I don't think either one of those extremes is really correct. I think the truth lies somewhere in the middle.

By the way, the SVIC podcast, they had Dr. Mike Israel on there for their conversation, which I listened to one part today and I'm going to finish the second part later. That was a very interesting uh conversation. Love to hear it. They kind of are saying the same thing, right? There's kind of uh the AI doomers, the AI hypers. The reality is somewhere kind of in the middle. They every once in a while do reactions to my videos, and I don't think they're big fans of my uh my thumbnail faces and my thumbnail titles, which I totally understand. Today they mentioned that they have a kook watch playlist, like uh people that are, you know, the like the crazy people in the AI land, and they h they have a whole playlist devoted to like uh keeping an eye on them. So they do have a number of people on there that uh are on their watch list, which is which is really interesting. I assumed I would appear on there a lot more, but I think as far as I could tell I appear on there only once.

So guys, if you guys are watching this, first of all, great interview with uh Mike, Dr. Mike, and two, thank you so much for only putting me on that playlist once. That that brought a smile to my face today. Thank you. But let me know what you think about this because I do want to do a lot more. I want to go back and and cover some of Anthropic's research into how these models think. I want to learn more about this, the features and the circuits because again, number one, it seems like a great way to advance AI alignment, AI safety, but also it really feels like it's similar to how we think about certain stuff, how we sort of encode various concepts into our brains, into our neurons and the synaptic connections and whatnot. And uh I mean this demonstrates that you can get these LMS, these AIs to be obsessed with something. And certainly we've seen that happen with humans. And I can't help but feel that these processes are at least somewhat similar. Obviously they're not the same. They might be very very different, but but there's some overlap. And hopefully understanding one will give us a glimpse into the other. These these are not completely separate subjects. I feel like that's again just my opinion, but it's a very very interesting um area of study. Dario Amade, whether or not you agree with everything he says, I think uh he writes really well, very very informative. So yeah, let me know what you think about this and uh if there's any other big articles out of Anthropic that are maybe very interesting but maybe didn't get enough uh light, or any other company for that matter, let me know. I'd love to cover it. With that said, my name is Wes Roth. Thank you so much for watching and I'll see you next.