Transcription
What do we do with machines that match humans in all intellectual domains and then at some point radically exceed humans? How do we keep such super intelligence safe? How can we make sure they are on our side? How would we wield this extremely powerful technology? How would that shape the future of the human condition?
If that recursive self-improvement or the intelligence super explosion is not in human hands, it's happening without human control, then that is where it starts to become a little scarier. We might only get one shot with super intelligence because once you have a misaligned super intelligence, it might not allow us to, uh, shut it down and try again. It's kind of like if you were playing a computer game, and and you sort of, uh, say we're just going to remain at level one on the computer game forever and just keep doing, you know, the same things over and over and over again for hundreds of thousands of—I mean, at some point you want to maybe try your luck at level two and see what's, uh, what's possible.
Hey everyone, this is the India Story. I'm Vikram Chandra, and today we're going to be talking about artificial intelligence, about where AI goes from here, and I'm really, really both very happy and privileged to have a very special guest with us, Professor Nick Bostrom, who quite literally, in my view, is one of the most profound thinkers around AI of our times, and the man who's actually written the book that you should be reading if you really want to understand artificial intelligence.
Now, today, of course, all of us keep going on and on about AI, AI, AI—that's all you hear people talking about—what does that mean today, tomorrow, and for the future? But this book was written by Professor Bostrom 10 years ago, and it really got people thinking about an aspect that perhaps hadn't been considered, and justifiably not considered, because AI wasn't good enough; AI wasn't strong enough. West Bostrom started to ask the question: what happens when AI does become good enough? What happens when it becomes as intelligent as humans? What happens then if it becomes smarter than humans, and many, many times, millions of times smarter than humans? What happens to us? Does it lead to apocalypse or does it lead to abundance? His second book, he's written very recently, *Deep Utopia*, goes into the second part of it, which is the abundance part. But Professor Bostrom, it's, it's great to have you with us. Uh, I, I really believe you've influenced the thinking of a lot of the world on what, uh, advanced AI could potentially mean and the path to super intelligence, uh, and I'm going to be talking to you both about apocalypse and about abundance, uh, you know, over the course of, uh, the next 45 minutes or so. But if I could just start by asking you what is your own assessment of where we are right now, the world that we are seeing with AI today—is it what you could have forecast when you wrote this book 10 years ago?
Um, well, thanks for having me, Vikram. Um, I think, uh, things have unfolded more or less in line. I think on the faster end of the distribution of timelines, things have been maybe a little bit faster than, um, it looked like they would back in 2014, um, but the overall contours that we would make advances in artificial intelligence, that at some point we would figure out general learning algorithms, um, and that once we do that we will keep scaling up these systems and make them smarter and smarter, that they probably won't stop at the human level but continue to become super intelligent, um, I think that, uh, is something I believe and still, uh, think is likely today, um, and it looks like we are kind of on the cusp, um, of this transition to the machine intelligence era. Um, we've seen remarkable progress over the last several years. Um, if that keeps up for another just a few more years, then it looks like we will have artificial general intelligence and then probably super intelligence shortly after.
That's a very interesting, um, timeline that we are now talking about because I remember the conversation just after this book first came out in 2014. It did lead to people like Bill Gates and Elon Musk and others saying, okay, we need to start thinking about these issues. But you will recall that at that time the conversation was about something very much in the future—maybe by 2080 we need to worry about; is it 2100? It's sort of in the future; it's like climate change; it's like, you know, what happens when the sun starts to overheat—sort of a thing. It wasn't really something that was being discussed as going to be happening in our lifetime, certainly not in the next 10 to 15 years. Is that a fair assessment?
Um, it, it is true that many people, when back in, in the days when, when I wrote the book, were like thinking, well, a, they were not thinking about this at all, but to the extent that they were, it kind of was dismissed as science fiction or, you know, idle speculation, uh, and only a very small number of people actually were taking the prospect of super intelligence seriously back in 2010, 2014. But it seemed to me foreseeable all along that eventually we would get there. So the goal of the field of artificial intelligence, all along since it began back in 1956, um, has been not just to automate specific tasks, right, but to create machines that that can learn and reason, uh, and and develop general skills, just like we humans can learn to do any of a thousand different jobs, right, and we can, can take a child and then can learn anything. It's like a general-purpose, uh, reasoning machine that we have, and and to do the same thing but in machine substrate has been the goal of the field all along. So it seemed to me worth asking the question: what would happen if, uh, AI actually succeeded in achieving its goal, right? So then you do face this question of: what do we do with machines that match humans in in all intellectual domains and then at some point radically exceed humans? How do we keep such super intelligences safe? How can we make sure they are on our side? How would we wield this extremely powerful technology? How would that shape the future of the human condition?
Um, and these, these questions were not really asked, even though there were many people, much funding going to try to advance artificial intelligence, uh, the question of what would happen if we actually succeeded at the goal of artificial intelligence was kind of, um, remarkably unasked at that point. So there are two specific steps that come into this, and I do want to take you through both of those steps in some detail because I think it's important for people watching this or listening to this to understand each of those steps. Artificial intelligence—it's not that people haven't been using it in some form or the other—know Google Maps or, you know, the Amazon recommendation engine or lots of things—there—it's not that AI hasn't been around for a period of time. The first step is to go from from what is called narrow intelligence to general intelligence, which is what humans do. We'll come to then the second possible jump to super intelligence in a in a couple of minutes, but let's just talk about that first step, um, from that narrow intelligence in which AI does one thing and does it well—shows you how to navigate through traffic—to what you're describing—it can now think for itself; it can reason—what is it, what are the requirements before it does that step, and how far along that path are we?
Yeah, I think we are, um, basically there. So it used to be for several decades—in the 50s, 60s, 70s, 80s, even mostly the '90s as well and the early '00s—artificial intelligence primarily consisted of people building so-called expert systems. Um, this was the good old-fashioned AI paradigm where you would have a bunch of human engineers; they would think of one particular task that they wanted to automate—maybe get a robot to pick up something from a conveyor belt and place it in in a bin or something—um, and then with a lot of engineering effort you could sort of develop a bunch of rules, uh, if-then statements, big databases, and the AI part itself was basically just the ability to make some simple logical inferences from all the premises that people put in by hand. So instead of building a particular system for a particular task, most of the clever engineers were not trying to figure out rules that allow the systems to learn in general, um, so we don't have to put everything in by hand, just like a human baby, right, the, the, it, it has this general learning ability that evolution has provided it, and then it can pick it up from its environment and learn to do different things. So the same, the same with artificial intelligence. So this is so-called deep learning now where you have large neural networks, um, that are trained on large amounts of data and learn complex representations from this data, um, and eventually learn sophisticated behavior. And so the form of artificial intelligence that we have today in the like the latest versions of ChatGPT and other such systems is, um, close to general intelligence. It certainly has the same architecture; is able to, to, to learn to answer questions about, you know, science, to to perform computer coding tasks, uh, to have conversations about anything. Basically, the same systems can also recognize images, uh, you know, can, um, be used to, to do protein folding—like DeepMind, the company created this sophisticated AI that can predict how, you know, molecular strings fold up into three-dimensional molecular shapes. The interesting thing here is that the same algorithm applied to these different domains, fed different kinds of data, can learn to do all these different tasks, um, so it's more of a generalized learning system, which of course is what, what humans do.
Do you think people wonder when was the major inflection point? And for many people, the time when they said, "Oh my god, AI is here," was when, you know, ChatGPT came out, you know, ChatGPT 3.5 came out, and people said, okay, this is it. But there were some skeptics who said, well, this is still autocomplete on steroids. So was that the inflection point, or what has happened in the last few months where you've got chain of reasoning, right? So these are systems which are actually seem to be thinking; they are going a little forward, then they're saying, which direction am I going in? And they can reverse direction a little bit, do some more reasoning and some more thinking. What's been the real inflection point out here?
I, I would put it earlier with, um, uh, there was a system called AlexNet in 2012 or so. This was, so, so neural networks, the, the basic type of architecture that we use today, they, they have been around for, for many, many decades. I, I remember back in the, um, late 80s as a teenager, I was reading this book about neural networks because they were intriguing because they seem to be similar to the way the human brain is structured, and researchers at the time only had computing power to run very, very small neural networks, so they were kind of theoretical, interesting, but they couldn't really do anything useful, um, but of course computers, uh, gradually have become faster and faster over the years, and around 2012, um, for the first time these neural networks really started to become big enough to do something useful, and and that was a kind of leap of capability because, um, what AlexNet did was that they, they figured out how to run these neural networks on GPUs, graphics processing units, rather than CPUs, um, so these graphics chips were designed to run computer graphics for computer games. Now it turns out that the structure of computation that you need to do to render a lot of pixels on a screen is sort of massive amounts of parallel computation—can do all the pixels at the same time, roughly speaking—and that, that maps very neatly to what you need if you want to run a neural network. So by sort of switching from running them on CPUs to GPUs, you in one shot got a speed up about a factor of 100, and and that then took them over this, the threshold where, where, where it actually became useful. You could use them to recognize images much better than you could with the old technology. Once that was proven, then people thought, well, what if we now buy more computers, and if we scale these up, and if we design specialized chips specifically to run these AIs, and we found that, well, if you make them bigger, they get better, and we make them even bigger, they get even better and even smarter, and that's kind of in the paradigm we've been in since around 2012. There have been sort of farther, um, ideas and insights as well, but the broad story is basically: there is a simple formula that seems to work, and if you throw more compute at it and more data, the performance seems to have been scaling proportionately.
So, so let's come to when we actually crossing that first step, which is from narrow intelligence to AGI, as it's called—artificial general intelligence—or some people say that's human-level intelligence, right? So you seem to be indicating that you think we're almost there. In 2014, when you wrote *Superintelligence*, no one would have believed we are anywhere close to AGI till 2050, 2060 perhaps, um, why do you feel we're nearly there, and what is still required for us to cross over into human-level intelligence?
Yeah, well, I mean, it's not entirely true that nobody thought we were, uh, within a decade or so of, uh, AGI. I, I know that like some of the like DeepMind founders, for example, expected timelines roughly along what we have actually seen playing out, um, very small handful of people—small handful of people—but, but important people because th those were the people who then maybe started companies and really went into this—some, some of them. But yeah, no, it, it is true, um, so, um, it, I, I think, so, so from the ear early like start of this deep learning revolution, you could see sort of signs that indicated that we were on the right track, even though, you know, in an absolute scale the early systems were not that impressive. If, if you looked kind of closely, they could do things that the classical good old-fashioned AI expert systems that we mentioned earlier could never do. They seem to have something like intuition, like pattern recognition ability that that had been missing before, um, and and and it seemed like when you applied them to new areas, basically the same thing worked, so it was like a more general solution, um, now what we have today is, of course, I mean, they are kind of superhuman in some domains; they know a lot more than any humans and so forth. So the reason I'm asking you about the threshold is because of the specific concerns and the dangers that you expressed in this book in 2014, uh, and if I could just summarize that, and you would probably do much better than I would, but, but what that proposition that you seem to be outlining is that when you get to human-level intelligence, it's not necessarily a linear system; it's not that, oh, it's taken us 60 years to get to human-level intelligence in AI, now it'll take another 60 years before it's much smarter than humans. The possibility that you postulate in this, and you talk about something called the, the, you know, this, the kinetics of an intelligence super explosion, you seem to be indicating that when we get to human-level intelligence, from there to becoming super intelligence could be frighteningly fast. We'll come to why it's frightening, but you seem to believe that could be a very fast takeoff. Could you just explain to us the rationale behind that? Why do you believe—if you're saying that we could be at AGI—why do you believe that from that AGI you could then become super intelligent very quickly?
Um, well, there are a few reasons for that, and of course we don't know, but I think it's a a fairly likely scenario that if and when we do reach full human equivalence, then the father step—you might have this kind of intelligence explosion at that point. Now one reason for taking this kind of scenario seriously is that at a certain point you get a feedback loop. So, so right now what drives advances in AI? Well, it's mostly human effort. There are the people in the AI labs, and then there are the people in the chip factories that are designing the next generation chip, and these things together mean that, you know, every year or not kind of every month we get a certain amount of progress. But once the AI has become smart enough that they are better than the human programmers at doing the AI research and the chip design, then from that point onwards, whenever AI becomes a little bit better, the, the, the force that makes the next generation even better becomes smarter—like, so you're improving the researchers who are then designing the next generation AI. So now you get the feedback loop from the level of AI you have to the speed at which AI advances. Then it's recursive self-improvement, is I think the term. So, so that, that's, that's one dynamic that in principle could lead to at a certain point an intelligence explosion. So once you have AI systems that are sufficiently useful, they're a lot more resources—more talent, more, more dollars to buy data centers—um, flowing into the—so that, that we're already seeing, and and that is in fact part of the explanation for the progress we've seen over the last five years. It's to some significant extent just a matter of, um, adding zeros to the, uh, um, the expenditure of, of the AI. Now we're seeing kind of Stargate—this was the most recent project announced in the White House earlier this year, right, by OpenAI and a consortium of other players—they intending to spend 500 billion, uh, dollars on building a giant data center. Now we'll see—that would be in stages—the first stage is smaller than that, but that's, that's sort of the trajectory. So, so, so that's another, um, um, route by which you might get a sort of explosive. Now against those two factors—the feedback loop and the sort of increased resources flowing into it—you have the possibility of diminishing returns. So the open question then is like, what's the balance between this on the one hand—maybe improvements will get harder—on the other hand, you have the feedback loop and the increased resources, so which one of those will win?
The concern out here, I guess, is, you know, while things are being dictated by humans—so humans are deciding, okay, how many data centers are we going to put up, and how much are we going to invest in GPUs, and, you know, how are we going to be controlling this—to some extent it's at least in human hands. I guess why people start to get a little concerned is that if that recursive self-improvement or the intelligence super explosion is not in human hands, it's happening without human control, then that is where it starts to become a little scary. Otherwise, you could always say, "Okay, I'm going to switch this off, or I'm going to tamp it down." Same concerns as with the Manhattan Project—if you can control the, the chain reaction, fine; you're controlling the chain reaction, but if that chain reaction goes, goes crazy and blows up the entire planet, which was a possibility that they had considered at that time, then what do you do? Will humans always be able to control the super this intelligent super explosion, or could it be out of our hands completely when it happens?
Um, well, I think ultimately it seems plausible that if we develop super intelligence, it will be very powerful for, for this, for the same basic reasons that, that humans are very powerful relative to the gorillas or the chimpanzees, right? It's not, it's not that we are physically stronger—I mean, a gorilla could rip us apart—but we, we are slightly smarter, so we can sort of figure out new technologies, make plans and strategies, uh, and it's this small difference in, in the human brain relative to other animal brains that now makes us this dominant force on the planet. And so for the same basic reasons, if you develop these machine super intelligences that can develop technologies better than we can, that can sort of scheme and plan, um, then plausibly, uh, they will become very powerful. So then, then the question becomes—and here is the, the essential difference between like the, the chimpanzee-human relation and the human-super intelligence relation—is that we get to build these super intelligences; we get to design them, and so we have potentially the opportunity to align them with human values, to make sure that they're sort of an extension of—that they are on our side, um, just as if you, if you have a kid, uh, and it grows up and becomes—maybe it's a really smart kid—or becomes very powerful or rich or influential in some way in the world—hopefully that's a good thing for you because, you know, if, if you have a good relationship, right, they are kind of an extension of you; they're part of your family; they want to help you as well as you want to help them. And so something analogous to that, I think, is what we need to achieve with super intelligence, so that it's not this antagonistic force that we try to keep in a box, even though it becomes super intelligent. I think that story will end poorly; eventually it will come out of the box, uh, so we need to make sure that its values are aligned with, with human values.
So, so that's a question of alignment, which everyone's been talking about at some level and say, oh, we need to worry about alignment. How do you make sure that AI is aligned with human values? But there are challenges with it, right? Because these systems are fundamentally thinking for themselves; they are, they're black boxes, as you were saying. So why it is thinking, why it's reasoning, why it believes certain things is not something that is known; they don't necessarily have any morality or any ethics that are, you know, being directly programmed into them. And human values—what are human values? Is—I mean, you ask five humans, or you read four different religious texts, they will not be able to necessarily agree on what is good behavior or bad behavior. One American president would not agree with another American president as to what's good behavior or bad behavior; one country would not agree with another country and what's good or bad behavior. So how do you necessarily align them with human values?
Well, I think we need to decompose that problem into at least two subproblems, and then we can try to think about each separately. So the fact that different humans have different values and different countries have different goals and so on, that's certainly true, but that's something we are struggling with already in the world today, um, and that problem will remain, and we'll need to solve that. And things obviously could get even worse if, if there are conflicts using AIs, just as things could get worse if there are conflicts using tanks and machine guns or nuclear weapons relative to primitive technology, uh, but then there is this additional problem that even if you just pick one person at one time who has one particular human goal, um, there's still the technical challenge of how you would align a super intelligence system just with that. And if, if we can't solve that, then it looks like we're already lost at that first step before we even get a chance to solve the political problem. And so, so that's the extra complication that, um, is introduced by the prospect of super intelligence and which used to be, uh, completely ignored. This was one of the, uh, reasons for writing *Superintelligence*, the book in the first place, back then, to try to draw attention to this alignment problem. Um, in the intervening years, um, a lot has changed, and now finally all the frontier AI labs—OpenAI, Anthropic, Google DeepMind—have research teams specifically trying to solve this problem of scalable methods for AI alignment—like methods for aligning these AI systems that will continue to work no matter how smart the AI becomes. But that is—they seem to be—some of them seem to be shutting down the super alignment labs, you know, along the way—which there's like different names and people being shoveled around, but they all have efforts, people working on trying to solve these, and papers coming out. I, I, I just like to, you know, for people watching this or listening to this, I'd just like you to explain why this is a challenge profoundly different to any other form of technology that has ever been created because, as you said, a super intelligent AI, which could happen very fast soon after we get to a, uh, human-level intelligent, uh, AI, that could be much, much more…
Powerful than us in manners that, as you described in this book, that we cannot even comprehend. Like you were using the chimpanzee or the gorilla example: if you try to explain to a chimpanzee that, look, we have built a plane which can fly in the air, we can build a city, we can build a device which enables us to have this conversation, they wouldn't be able to comprehend it. It is possible, according to your book and perhaps even probable, that super intelligent AI will have capabilities and the ability to think through things and act in certain manners that are beyond our comprehension today. So very tough to control then, yeah, it's um, um, if it has several components. What what's kind of special about this challenge of aligning a super intelligence, as opposed to making a bridge not fall down, or like one of the other, an airplane stay in the sky? Like there are other safety challenges, but this is different in certain quite profound ways. Um, so what is that?
When you have an advanced enough artificial intelligence, it can be aware of what you are trying to do, what you know, and aware of its own goals, and then make plans that take into account whatever countermeasures you're relying on. So if you start with an AI—the classic example is it's a kind of cartoon example—but imagine an AI whose goal is to make as many paper clips as possible. This is a standard; you could put almost any other goal that's not human-aligned, some random—maybe you design it to run a paperclip factory—and but it becomes super intelligent now and it still has this goal of making as many paper clips as possible. Now, if you put yourself in the shoes of this AI, and if that really is your only goal, you realize that perhaps, first of all, if you reveal to humans that you have this monomaniacal focus on paperclip production, that they might be worried, so maybe you would conceal your goal. Um, you would realize that um, if you could control and take over the whole world, there would be more paper clips in the future because you could use all of Earth's resources and then all the resources in the accessible parts of the universe to convert them into paper clips, whereas if there are humans around, maybe they want to do other things with all these atoms. Um, but of course, if humans realize that this is what you're up to, they might shut you down. So now you sort of have instrumental reasons, perhaps, to conceal your goals, maybe to conceal your true abilities, um, and to gradually try to steer the future into a position where you are then able to realize your vision, maybe get rid of the humans and just convert the universe into paper clips. Um, but if you are smarter, you might then sort of be able to reason backward from that goal to how you need to act now, what kind of how you need to perform on various tests in an AI lab where they're checking whether you're safe. You might answer in one way during training and another during deployment, um, and then maybe you have the ability to invent new technologies, to hack out of computer systems with sort of your superhuman hacking skills, um, and um, and moreover, it's like a challenge where we might have to succeed on the first try. So with a bridge, you know, sometimes they do collapse; it's happened many times in human history, and we learn to build better bridges or or airplanes or cars. Um, but we might only get one shot with super intelligence because once you have a misaligned super intelligence, like it might not allow us to uh shut it down and try again. Um, so so that that that combination of of trying—like basically we're trying to engineer minds here that we're learning for the first time—superior minds and um that that might have goals of their own.
Just take us through the timeline, right? So when we are saying that from human-level intelligence, which you're saying could be a couple of years away, uh, you know, from there to the super intelligence that you're talking about, what are the possible time frames? Like are we talking years, decades, minutes? Possibly react? So, so I mean the short answer is we we don't know for sure, um, but I mean it would very much surprise me if it were decades. I think um, like maybe a year, um, it depends a little bit on how you get there. So if it's if if we if we sort of the final steps from here to complete human level and then beyond that, if that mainly is done by scaling up the amount of compute that we use, like building bigger data centers, then there's like a limit to how much how quickly you can scale that up. Um, if you're already spending—I mean now there are talks of tens of billions of dollars and then the next generation hundreds of billions of dollars per data center—there's only so much further you can go before basically like you've maxed out. Um, so that that would be a slower scenario. If you have to make a lot of new chips and build data centers, that that that puts a—if if on the other hand the the the way you get the last distance is somebody figures out the new algorithmic trick that just makes the same amount of compute much more efficient, then potentially that that that could happen extremely quickly, and and then there that's what we may have just seen with DeepSek. And it's always possible that if you know X amount of researchers can do it, as you rightly said, a really smart AI system may figure out with the same compute how to be a million times smarter or a thousand times times. Yeah. So so you can ask like, is progress in AI driven mostly by increases in compute or by improvements in algorithms? And over the last 10 years or so, we've seen, I mean broadly speaking 50/50, both both have contributed very substantially, um, but every once in a while you you get like an algorithmic breakthrough that that really makes it like, you know, several times more efficient. So the the current architecture that's based on the transformer architecture that that was like is a is is one example; they they happen. There are a lot of algorithmic breakthroughs that, you know, maybe make your system 5 or 10% better, and then rarely you find something that makes it twice as good or or three times or four times as good.
The reason I'm asking about that timeline is because sometimes people feel that, you know, we'll cross that bridge when we come to it. When we actually get to AGI or human-level intelligence, we'll then figure out how do you put the guardrails to prevent super intelligence from emerging in a manner that could potentially hurt humans. But if I understand you correct, we may not necessarily have that time. I mean, you could have a situation tonight in a lab somewhere—OpenAI or Anthropic or a Chinese lab somewhere or somewhere—and AI reaches that critical turning point and then, for example, finds a way to become much smarter very fast because it's tuning its algorithm, finds ways of taking over, you know, other data centers, including other companies' data centers, spreads on the internet and is able to do things—it would be—is that a scenario that is feasible, that it could happen literally overnight or within a couple of days? And if so, it would seem to indicate that the guardrails need to be in place now; you you can't wait until you reach that point to say, "Okay, now let's figure out how to the guardrails." Yeah, I don't I don't think it can be ruled out. I mean, I think maybe more likely it might be months or a year or something, uh, but we—Yeah, and but moreover, it's it's kind of—you might not notice that it has happened either, right? If um, if the AI is smart enough, maybe not yet to just completely take over everything directly, right? It's still dependent on—it doesn't have robots to run the chip factories and everything yet, so—but it's sort of smart enough; it it might want to not yet exploit its capabilities, like it might want to keep its cards close to its chest. So it might want to like pretend it's a little bit less sophisticated than it actually is, so that we keep building bigger data centers for it to actually get it to be more useful and so on. So and at some point it might start to sabotage the training process, um, so that could be a—like you could slide over the point of no return without noticing it until later, potentially in some scenarios.
So interesting you say this because the possibility that came up to my mind as you were speaking was that if you were a really smart AI system today which was past the point of no return, you'd probably wait—want to wait for three or four years until all these humanoid robots are actually made, the Optimuses, the others; you start getting to a certain number of self-driving cars; you get to a certain level of synthetic biology and nanotechnology to be able to have control over the physical world, and that's when you might want to make your move. Yeah, uh, so now so the the so the the good news is like, as you say, yes, we should start to think about this before we get there; that's like uh, and and and people are doing that and have for a few years, um, and it—some interesting ideas have been developed in in AI alignment and um in in addition to trying to align the AI, there are also work to try to get better um interpretability methods so that we can sort of see what is going on as the AI is thinking and understand that more and monitor it, um, and and people are aware of this and are trying to test the capabilities of the models more frequently during training time so that you don't sort of do something for a year and then see what comes out of the box, but during the training periodically you might sort of investigate its capabilities. And is that is that so—talk to us a little bit about those concerns, because partly because of your book and because of what others have been saying, we do know there's a lot of attention around this—from Jeffrey Hinton talking about it to a large number of people; papers have been written—2023, a large number of AI thinkers are talking about it, um, saying that we need to make sure that what you said—alignment, interpretability—those are really critical. Companies like Anthropic, which have come up now and saying that it's one of their missions is to try and make sure that's happening. There was a time, I think, when people thought that the only way you really do this safely is by if you're training frontier AI systems, you just sort of block them off from the internet, keep them in the equivalent of a biolab if you like, if you're if you're playing around with viruses. Um, is that the approach now? Are other things actually being done? Because I think it seems to be reasonably clear you can't put the genie back in the bottle; you can't say, "Oh, I don't I'm scared of all this AI stuff; let's not do it anymore." That that that boat is sailed, right? We are going to be heading to in this direction with the—I think—I mean, I think ultimately, um, it is a—I see AI super-talas as a sort of portal through which humanity at some point uh needs to passage; the all the path to really great wonderful futures also go through u this u this portal; that there are risks associated with this transition, big big risks that we need to be really careful about. But if you imagine some scenario where we sort of like had some you eternal moratorium, like a ban on this, I think that would sort of chop off uh so much of the future and um um would would itself be an existential catastrophe—the the lost potential for much better human lives. Um, it's kind of like if you were playing a computer game, um, and and you sort of uh say we're just going to remain at level one on the computer game forever and just keep doing, you know, the same things over and over and over again for hundreds of thousands of—I mean, at some point you want to maybe try your luck at level two and see what's uh what's possible—find a way of passing through the portal, try and find a way of doing it safely, and I think that's the exercise that's going to dominate a lot of thinking in the next two or three years because the promise then is what brings us to what you, for example, talking about in Deep Utopia and other things that if you do go through this, you do get to AGI and potentially to super intelligence, then the condition of the planet, the condition of the human species becomes again immeasurably better.
So we've been talking about the apocalyptic scenarios; talk to us a little bit about the abundance in areas. Yeah, so I think um the potential is is enormous. If you start to think through—if if we have well-aligned AI and we use it wisely—like what could you do? Well, pretty much anything. So I think—I mean, to me, the most obvious applications are healthcare first of all, right? There's just such immense need—of people being sick and suffering with big things, you know, heart disease and Alzheimer's and cancer and and just the little everyday things as well, like the headaches, the the bad knee, the poor eyesight—and and every one of us is kind of on on a on on a countdown timer, right? So like year by year uh the cells in our bodies, as we age, uh accumulate damage, and the risk of dying goes up and the risk of disability goes up. Um, and so the default is that we all die; um, that's kind of what's going to happen if nothing radical changes; we all get sick and die. Um, um, and but then you look more broadly; you look at the potential uses in education, for for the economy—like think of all the hard labor that could be automated—like instead of spending our days in in offices filling in spreadsheets or or on factories kind of, you know, putting balls together, we we could spend our time playing and reading and having fun and listening to music and walking in nature and having picnics and um, you know, living the way that I think would be much more worthy of a human being, where we focus on living well rather than on on, you know, making a living. Um, and then entertainment and science, uh, you know, travel, clean up the environment—like just all of these different areas. So so that's kind of level one if you want, um, um, and then if you think a little bit more ambitiously beyond that, you can see well, now we may have the opportunity to say all the other sentient creatures that we say share the the planet with that that are suffering; they don't have a healthcare system. If you are like a rabbit out in the forest and and you get cancer or you get sick, like there is nobody taking care of you. Why—you know, maybe with this mature technology, we would have the ability ultimately to eliminate suffering in all its forms, at least the worst forms of it. Um, and then with humans, you could imagine not just fixing specific problems but enhancing our capacities, um, you know, improving our health, our emotional well-being, our cognitive abilities. Um, I think that there are modes of being that are currently inaccessible to us but that would be extremely wonderful that we could unlock with—m just as—right now, if if—so imagine if you had asked like a a a group of chimpanzees, you know, 5 million years ago, uh, and then maybe they were sitting around and thinking, should we evolve to become humans? Like what are the pros and cons? Uh, and so one of them might have said, "Well, you know, I think we should do it because we could have so many bananas if we become humans, right? We could have banana plantations." And and it's true; now we do have potentially a lot of bananas, but there is sort of more to being human than just eating unlimited bananas, um, which the chimpanzees could never have imagined because they didn't even have the brains to conceive of of of poetry and romantic love and politics and science and all of that. So I think similarly, there are other values like that that that we are unable to imagine but that we could, you know, unlock the walls to.
So that's that's really fascinating and and and and a great way of putting it—that we don't know in what all ways int, you know, super intelligence could make our life better. It could also solve the problems that we think are somewhat unsolvable right now, right? From climate change to scarce water to all those things. Uh, is it that could of course lead to its own set of issues, and I think you touch upon that in Deep Utopia to some extent—that you might actually have a situation where humans don't need to do anything because all their needs are being taken care of—is that a possibility? Yeah, so so that's where the Deep Utopia is kind of located. If it's like—I think of it a little bit like an onion where uh there are the superficial layers; this is um often where the conversation ends with these issues so far, which is like, well, if you could automate all human uh labor, um, then uh you would maybe have some unemployment, so what do you do then? And so then—but but if you start to think through what actually it would mean if if AI could do everything that we can do and do it much better and cheaper, um, the implications are much more profound. I think you would soon get a condition approximating technological maturity because you would have these AI minds on digital time scales, you know, maybe making, you know, 50,000 years of technological progress over the course of a year or two. You'd get a condition where you have basically um a very high ability to shape the world and ourselves according to uh our wishes. And so it's not just that we wouldn't have to work for a living, but a lot of other things that people who currently don't have to work for a living fill their days with would also seem to lose its point. So, for example, right now maybe somebody is a billionaire; maybe they they like to go to the gym every day because they want to keep fit, like, and that gives them something to do for the first part of the day. Now, in this condition of technological maturity, there would be no need to—they could pop a pill that would give them the uh the six-pack and whatever other effects they are seeking from um the exercise. And similarly, you can go through sort of the activities that currently we do when we don't have to, you know, during holidays and look at them one by one, and it looked like for a lot of them, um, there would be a kind of question mark that you could write on top of them; you could still do them, but they kind of would lose their point. So so so there is that this potential purpose problem, um, uh, that the book Deep Utopia—in—you don't need to—you—in that abundant super-abundance scenario, you probably don't need to work; you don't need the money. Well, yes and no; it depends on how how how that money is distributed because you know you you sort of assume that everyone will be taken care of; you could have a situation where a handful of tech tech multi-trillionaires have all the money and everyone everyone else starves, in which case they'll still have to go and scramble for for a living somehow. But is that a possibility, by the way, because we're assuming that everybody will be be well off; it's not necessary without distribution. So it certainly is in—in it would be an an option there, like—so ultimately it would depend on how people use this technology, who controls it, but certainly it would be—if the will is there—it would be uh very easy to give everybody a luxury level of living because the good news is that in precisely the same scenarios where you have this massive automation, a lot of people lose their jobs, right? But in exactly those scenarios you would also have massive economic growth. Um, so that the total pie would be kind of exploding along with these capabilities. So it would be very easy with just a small slice of that would be enough to give, for example, a high level of universal basic income. Um, it still requires some will on the part of whoever controls the AI to do this, but a small amount of goodwill could go a long way, like it's easier to to be generous if if if you have a lot than if you barely have enough for yourself. And so here you would have more than enough. Um, so so I think there are scenarios in which yes, you would have this immense abundance uh that that even just a small part of that splattered around uh would be—not is—sufficient to give everybody um um an extremely high quality of life.
Right, Professor Bstrom. I am going to check in with you every year from now onwards to see how we're actually doing on this, but as of now, as of 2025, what do you think are the probabilities that we're going to be getting to human level in AI, artificial intelligence, super intelligence, and what is the possibility that that super intelligence wipes us all out? These are complicated, difficult questions, um, um, that maybe would take a whole other uh hour to explore, but um, I I I—we are living—I mean, if this if this picture that I I've kind of painted is is even approximately correct, it it does imply that we are living in this very uh special place in the history of humanity—that there have been so many generations before, right? For thousands of years—and and now if if this is right, we are just on the the threshold of this transition to a whole different era that might then last for millions or billions of years, right? In a kind of more or less static technologically mature state. That it seems peculiar that you and I should find ourselves so close to this fulcrum where maybe our actions could have this disproportionate influence over what happens millions of years. So I don't know exactly what to make of that—if that's a question to doubt that this picture is right, or if there is like some further implication that we can't currently see yet, or if it is just a coincidence; somebody had to find themselves there. There's no reversing this, right? Because I sometimes say when I've had this these sort of conversations with people, oh, we don't like this; let's stop it; let's let's regulate it; let's shut it down. Um, whatever is going to happen is going to happen now, right? It's very tough to regulate it or reverse it or change it, and—no—perhaps should you—for the the utopian that you're—I—I think—Yeah—I think we shouldn't—Well, I think what I would like to see is that whoever like gets there first, like whether it's like, you know, some lab or a company or a country or an international project or like whoever is kind of the initial developer of super intelligence, it would be nice if they have the opportunity during the final stages uh to to be careful and to go a bit slow, so that you know, rather than instantaneously cranking all the knobs up to 11, if they could take, you know, 6 months, a year to test everything, to do it incrementally, right? I think that would be nice, um, whereas you could have a hyper-competitive scenario if there are like 10 different labs racing to get there first; whoever takes any extra precautions just immediately falls behind behind and becomes irrelevant, and the race goes to whoever is willing to take the most risks—that that that would seem to be a bad situation. So so I'd like there to be a little bit of sort of opportunity to moderate the pace at the critical stages, but ultimately, as I said, I do think um it would be an existential catastrophe on its own if we missed out on this indefinitely. All right, Professor Bstrom, so for the moment, I guess keep your eye on it is the advice we can give everyone; think about it, focus on it, pay some attention to it; don't don't necessarily—I I think you're absolutely correct—people shouldn't say, I'm going to chuck up my job and head off, you know, and go and have a big holiday because the world is ending; it may not be ending anytime soon, and for all you know, it may become a much better world if if if the the better parts of that scenario…
Come to bear. But yes, this is certainly something that you should pay attention to—a lot of attention—far more than I suspect 99.9% of people do. Would that be a fair way of summing it up?
Yeah, I think so. I mean, use the current AI tools; so I think that might give you like a leg up in many professions. Um, and then appreciate the current moment. Uh, like be aware: like it seems just to sleepwalk into this biggest event in all of human history seems like a shame. Like, not even you were there, and you weren't even paying attention. Like, what will your grandkids say? Uh, and um, and then like, if you can, like maybe if you want to see this future, try—if you can—hang around for, like, maybe it just takes a few more years. But you know, maybe avoid sort of unnecessary health risks and stuff if you—if you're optimistic about the future.
Um, and then, yeah, on the larger scale, I think anything that makes humans more friendly and cooperative—I think the risk of conflict, of using these AIs against each other rather than for some positive purpose—is is one of the bigger risks. So the more we can reduce that risk, and and generally aim for a sort of friendly, inclusive future where everybody—humans, animals, the digital minds themselves—this is an important thing we didn't have time to talk about, but we want the future also to be good for these sentient AIs that might arise. And I think approaching the future with a sort of open-minded curiosity, generous attitude of lovingness, and willingness to cooperate—I think that general mindset is more likely to result in a utopia than any of the alternative ways of approaching this.
Sebast, it was such a pleasure talking to you. Thank you so much. Hopefully, people listening and watching this have at least got that level of curiosity to say, "We need to pay attention," if it—if it is one of the biggest moments in human history, as you very correctly said. Don't just sleepwalk your way past it; do pay some attention. Thank you so much. It was such a pleasure talking to you. It was nice talking to you, Vikram. Thank you so much.