📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Debate with a former OpenAI Research Team Lead — Prof. Kenneth Stanley

Doom Debates2:37:10

Transcription

You were at OpenAI at a very interesting time, at the dawn of the LLM revolution. I mean, I was leading the open-endedness team, or I started the open-endedness team there at OpenAI. And so I was really interested in what do these new giant models mean for the ability to pursue interesting or curious paths? Do you think that there's a slippery slope to an AI that's highly goal-oriented and wants to seek power? I just don't think that this is a plausible articulation of what a real superintelligence is like. They're not going to be ultimately guided by goals, because being guided by goals is not super intelligent.

Welcome to Doom Debates. Today, my guest is Dr. Ken Stanley. He led the open-endedness research team at OpenAI from 2020 to 2022. Before that, he was a professor of computer science at the University of Central Florida, and he was the head of core AI research at Uber. He co-authored *Why Greatness Cannot Be Planned: The Myth of the Objective*, which argues that as soon as you create an objective, then you ruin your ability to reach it. He's also known for creating the novelty search algorithm and the neuroevolution of augmenting topologies, or NEAT algorithm. I'm excited to compare our views on the nature of intelligence and, of course, the risk of imminent human extinction from superintelligent AI. Ken Stanley, welcome to the show.

Thank you for having me, and uh, really glad to be here, and uh, appreciate this forum that you've created.

Yeah, appreciate you as well. You're entering the lion's den. We already know beforehand we're going to have some pretty significant disagreements. So you ready to hash it out?

I'm ready.

Yep. Let's get going.

Okay, great. You were at OpenAI at a very interesting time. You know, GPT-3, I think, was 2020; GPT-4 was 2023. And you were at OpenAI as a research team manager, 2020 to 2022. Uh, what did you work on?

Yeah, so, and I was, you're right, it was a very interesting time. I mean, GPT-3 was coming onto the scene when I came in there, and um, it was like a revelation, you know, to be experiencing it very early on, probably even before other people were experiencing it. Um, was like a peak into the future. My thinking and focus at that time was what are the implications of this really big development for open-endedness? I mean, I was leading the open-endedness team, or I started the open-endedness team there at OpenAI. And so I was really interested in what do these new giant models mean for the ability to pursue interesting or curious paths, um, and to produce interesting artifacts, um, that aren't just a solution to a problem, but would be considered to be actually creative and novel. And there's a lot of questions around that, like they, they obviously are very impressive, but can they actually do that? But that's sort of my favorite thing that humans can do, and so I was looking at the prospects for this kind of technology in that context.

Now I know OpenAI had a lot of parallel research paths, and even their GPT path, you know, famously it started with one person and then slowly picked up steam and then became like central. So your open-endedness research, did that end up merging with the LLM program, or, or kind of where did it end up?

Yeah, to some extent, you know, because really, if you think about it from their big picture, not necessarily mine, it's really aligned with this issue of what happens when the data runs out. Um, that's really where open-endedness kind of slots in naturally to, to kind of like the ambitious big companies, you know, because like, and there's obviously debate about is the data going to run out, is there enough data, but in general it's always good to be able to generate more really high-quality data. And if you think about it like that's what open-endedness does in spades; it's basically about generating interesting stuff that's useful, and then presumably you could just consume that and then continue to improve. And you know, there's all kinds of questions about whether that's, that's really a virtuous cycle or not, but it fits into it, slots into that. And so to some extent I did find myself converging with sort of like the main themes as the whole company was converging. Like when I first came in, like you said, it was there was a lot of diverse streams, but there was a kind of gradual convergence into everything is LLMs, um, and I found myself doing that as well.

So in the GPT-4 release, when you were there, did your open-endedness research get incorporated into GPT-4?

So I mean, I, I left before that, so I don't actually know for sure, but I think that probably there is some stuff from our research in uh, what entered into GPT-4. It's, it's likely, given where I was when I left, um, but I don't think it's a significant factor is my, like a pretty confident assessment. It probably doesn't matter that it's in there because it probably would have been similar even without it, but it probably is in there. Uh, there's some stuff that we were able to generate through this kind of research, um, which was above the bar that like probably would qualify to, to be entered into GPT-4.

Gotcha, gotcha. All right, and we're definitely going to dive into open-endedness, and that'll be uh, a big topic of debate, uh, but you know, just going back to your background, um, famously, uh, OpenAI, we, we hear that from reporting that Ilya Sutskever would love to remind everybody to feel the AGI. What does it mean to you to feel the AGI?

So I actually think that was a little after my time when that really got like kind of circulating around, because I don't remember hearing that a lot during my time, but I definitely heard that term, and um, to me it's just a slogan, and it, it's not my really favorite thing, um, because I guess it's not a big deal, but it sort of like culturally seems to make it sound like this is kind of like a a sports game or like entertainment, um, and I feel like it's a much more serious pursuit. So I wouldn't necessarily use that slogan to describe like why I would be motivated to think about this, um, but ultimately, I mean, I'm overanalyzing just a a simple slogan.

Fair enough. Okay, so let's get into the meat of your research and you know, your one big concept, so we're going to talk about open-endedness, uh, and I think you, you kind of draw a dichotomy of objectives versus open-ended, right? Is that a fundamental dichotomy that you see?

Yeah, yeah. I mean, of course, it needs to be clarified a little bit, um, you know, because those, those terms are slippery and probably ill-defined, um, and so yeah, go for it.

Okay, teach us.

Sure. Yeah, so, so to try to make it um, clear, um, without getting overly formal about it, um, what I view as when I say objective, it means that there is a target for your search, where it could be an optimization algorithm, which is what you hope it to do, and it could be something impressive like learn to predict the language, learn to predict the next word, it could be something like learn to walk. This is an objective, and then so sort of the way that we frame the problem when there is an objective is either you succeed and then you celebrate, or you don't, and that's usually because you got stuck on a local optimum or something like that, and then, and then, then you think, well, what can we do to improve this? So that's objective-driven types of processes, and it's just worth highlighting that it's just not, it's not only an algorithmic idea. Like humans always constantly pursue objectives, institutionally, individually. You know, our companies have objectives, OKRs, um, educational institutions have objectives, like schools have objectives, like the test scores. So objectives are just absolutely ubiquitous across society. It's like the most normal way to pursue achievement.

So what is open-endedness then? Um, open-endedness is a process that doesn't have an explicit objective in this most formal sense. So I don't want to make that, that kind of caveat to it, because like some people will say, well, that openness is its own objective, and this can kind of muddy the waters, but I, but first I just want to make this kind of clear what the distinction is I'm making, um, which is that in open-endedness you don't know where you're going by intent, and the way you decide things is by deciding what would be interesting, um, and so open-ended processes, like, make decisions without actually a destination in mind, um, and open-ended processes that exist in the real world are absolutely grandiose, um, so they are the most incredible processes that exist in nature, um, and there's really only two, and I just want to highlight them. Uh, one is natural evolution. So you start from a single cell, and then you wait like a billion years, hundreds of millions of years, and all of living nature is invented in what would be, from a computer science perspective, a single run. And there's nothing like that algorithmically that we do. You know, it invented photosynthesis, flight, the human mind itself, the inspiration for AI, all in one run. We don't do stuff like that with objective-driven processes. A very special divergent process. And the second one is civilization. Civilization does not have a final destination in mind. It's not like we're all trying to build a time machine, and we're all unified in that purpose. We all do divergent things, and because of that we go from things like wheels and fire to space stations and computers, uh, thousands of years later. In this case, it's, it's a special kind of non-objective process, but it's these non-objective processes that are responsible for the most incredible creativity that exists in the universe, and so we need to grapple and understand them, um, and maybe I just one more quick caveat to that is just to acknowledge that you could, that there's some kind of objective in the open-ended process. Like you could say, for example, that evolution has an objective of survive and reproduce, and this gets a little bit hair-splitting, but I like to make a distinction there because I don't think of survive as, as formally the same kind of objective, which I prefer not to call it an objective because it's not a thing that you haven't achieved yet. Like when I think of objectives, I'm thinking of a target that I want to get to that I've not yet achieved. With survive and reproduce, the first cell did that, so I think of it more as a constraint. Everybody in this lineage needs to satisfy this constraint, was already been achieved, and subject to that constraint we will continue to satisfy that constraint forever, and as we do that, as a side effect, we discover things that are not objectives, like the human mind. And so it's different from an objective process.

Interesting. Okay, um, so just to recap a little bit, so you talked about objectives versus open-endedness, and I think there's maybe an alternate terminology that also works that you use, which is a convergent versus divergent search processes, right? Would you use that kind of synonymously? Open-ended processes are like divergent processes.

Okay, um, so now I want to get into what you're saying where you're like, people will argue with me that evolution isn't divergent, evolution is convergent because it tries to optimize genetic fitness, and you're saying, no, no, no, that's just a black and white constraint. Yeah, you have to survive and reproduce, but then really you just want to go explore, right? That kind of your position on that?

Yeah, um, all right, I want to take the other side of that, um, so when you say, hey, this is just a constraint that it has to survive and reproduce, is it just a constraint because I think it's a gradient, right? I, I think what you're saying is that when you have a new mutation, in the case when the new mutation is helpful, like let's say vision is evolving, and the new mutation, uh, shapes the, the new light-sensitive patch on the organism to be able to see the position of light a little better, right? Like I'm thinking about like the process that goes from like light sensing, uh, some light-sensing patch of nerves all the way to like an actually fully formed eye. So you have this gradient where you, oh, this new mutation is tweaking the organ to be a little bit better. It's very much, it's kind of like a loss function when we train a neural net, right? It's very much it's like a gradient descent type of search. So why are you referring to it as like a black and white constraint and not an objective to converge to?

Yeah, yeah, I don't deny that optimization does happen in evolution, um, competitive pressures do exist in just about every lineage, and, and those do lead to gradients of improvement that would be characterized as the way you just described as objective, but it's important to highlight that the overall accounting for why there is all of the diversity of life on Earth is not well accounted for by just that observation. That's the observation from the high school textbook, and that seems to be the all-encompassing interpretation that most people view through, Evolution through the lens of, is this idea that there's these gradients of improvement, and everything is following them, but that does very little to explain why we have so much diversity on, in, and diversity of phenomenal achievements. It's not just diversity, it's astounding diversity of amazing inventions, and to account for that requires other explanations. I want to point out that like in, in evolution, there is actually, it can be difficult on a, on a macro scale, like globally, to explain what you mean when you say better. Like it's actually like getting better in some way, you know, we have all these amazing abilities as humans, seeing is one of them, but cognitive processing is another, but yet like, how are we better than single-cell bacteria from an objective standpoint? Like we have less biomass, we produce less offspring per generation, um, we have ultimately a lower population. There's nothing objective to point to like why we're better in terms of optimization. What's better about us is that we're more interesting, like, which is a different orthogonal dimension of analysis, is to say like, what actually makes us special? It's not cuz we're better at evolving. So a lot of evolution has to do with escaping competition, like founding a new niche, finding a place where there isn't competition, which is not an optimization problem. This is doing something different, and that's the divergent component of it. We need to account for that divergence to explain the actual interesting part of the process. I would argue that these convergent optimization subprocesses are actually less interesting because they're easy to talk about, but they don't account for the global macro process of evolution.

All right, let's go back to what you said before. You're saying some people say that uh, evolution by natural selection is actually convergent, it's trying to optimize an objective, but that you claim that that doesn't account for the extremely wide diversity of life on Earth.

Correct. Yeah, I mean, I think it's, it's clearly not convergent. I mean, evolution has, has not converged, um, so that's just plainly wrong. So here's how I account for the diversity of life on Earth. Earth has a lot of ecological niches. That's how I account for it.

MH. Yeah, so I mean, that's, that, that seems to me to be just an observation rather than an analysis, like, of course, it's true, like there's a lot of ecological niches, but they just beg the question why, uh, you know, because if you look at like the real problem I see here is that most people who make evolutionary analogies in the AI field think of genetic algorithms and evolutionary computation as the sort of basis for their analogies. Those kind of algorithms work the way you describe. So like an evolutionary algorithm, which I worked in a lot like in my career, or a genetic algorithm, you do set an objective, and fitness is measured with respect to the objective, and you explicitly follow that gradient just as you would in another optimization algorithm. And so that's a metaphor that's sort of very appealing and intuitive to people in the field, but think about what genetic algorithms do, they do converge, they have almost nothing to do with what's actually happening in nature. The intuition is off, and so it's unfortunately become this kind of misleading metaphor that a lot of people key into that that's actually representative of the actual process. These are highly convergent algorithms that, that always converge to a single point where they get stuck, uh, just like conventional optimization. That's not what we see in nature.

Okay, um, so I guess maybe I don't even really know what your claim is at this point, so maybe we should kind of backtrack to there, um, yeah, I think you were trying to make a claim where, um, you're trying to show that, um, the, if you want to populate life on Earth, if you want to use Earth's resources, then you better do it by using a divergent search algorithm, and that's what evolution has successfully done. And so in general, divergent search algorithms are really powerful, and they're like the key to intelligence. Is that kind of where you're going with your claim?

Um, that's, that's partly what I'm saying. I mean, I don't, I don't think it's necessarily about if you want to populate Earth. I mean, I think that it doesn't, we're not in concerned with whether Earth gets populated or not. This just happens, stands, like the algorithm happens to work the way it works. It's not because it's better or worse to work this way, it just does work this way, but what we're interested in is as the side effect that we observe is that it's highly divergent and highly creative, and we want to understand why. How do you account for that? It certainly is. I mean, I think that part of it isn't controversial. This is incredibly prolifically creative, and we don't have algorithms like that in computer science. Like evolutionary algorithms don't exhibit that property. So there's something to account for here that we have not abstracted properly, and yes, this is related to intelligence because, like I said, civilization also has this property, which is built on top of human intelligence, and it's related to the superintelligence problem because my real deeper claim here is that superintelligence will be open-ended. It must be, because that is the most distinctive characteristic of human intelligence is our prolific creativity. It's the legacy that we leave behind. So we will not get a superintelligence that's not open-ended, and therefore we need to understand divergent processes, and all of our optimization metaphors don't account for that property, which can mislead us and lead us astray in analyzing what's in store for us in the future.

Okay, um, so that's a good central claim for our larger conversation, which is that you claim superintelligent AI will be open-ended; it won't really conform to the model of being an objective optimizer, right? That's one of your central claims.

Yeah, that is a central claim. I think, I think that that's, that's a metaphor that is, um, incomplete for understanding, you know, what we're confronting in the future. And not, not to imply that someone is wrong about their P doom, that's not, that's not my reason to want to make this argument. It's more that I just think that we'll be able to do better at analyzing whether or not you believe that we're in for something bad by acknowledging this really difficult property and grappling with it rather than denying it, pretending that life is a simple optimization process. Like that's not actually the reality.

Yeah, so, so when I connect your central claim to your observations of evolution of life on Earth, I think, you know, I, I agree with the part that there's no Earth-level objective. So there, there's no objective in the system where there's no plan for Earth, right? There's, there's not saying, let's, let's maximize X on the level of Earth. It's more like, uh, every gene is being pushed, every genetic code is being pushed toward a version of itself that's better for its niche, right? Better for replication within its niche. So it is just a bunch of local search, and because there's different niches, it's just reaching these different local hills in the fitness landscape. So I, I mean, I agree with you there that it's just like a bunch of hill climbing with no like central plan, um, it's just, I, I, I guess I just, the hill climbing to me, I mean, it's still very powerful hill climbing, right? So can we maybe, maybe we can agree that like, if you just look at one niche, if you just look at competition within one niche, don't you think that that is pretty well modeled by the idea of optimizing an objective?

Uh, maybe, I, I wouldn't totally go that far, cuz I think divergence even happens within niches, and that's a different kind of process, but I would acknowledge there's, it's a better way of analyzing a single niche, but look at what's missing from your story. Like, this, the important part is what's missing from the story you just accounted for. Optimization within a niche, but the really interesting thing are things like, what, where did giraffes come...

From what niche were they optimizing before they were giraffes? And the answer is that you needed a tree to even have a giraffe. The tree is a product of the process; it's not an intrinsic fitness function from the start. The niche was created by the process itself, and then the process found something that could satisfy that niche. That is not explained by optimization; that's the powerful part of it.

What that is, I would characterize it as there's an, there's an additional component of search which we don't talk about, which is collection. Optimization is one thing; that's following a gradient. But collection is another thing; it's a very powerful concept. Collection means collecting stepping stones. Stepping stones are things that can lead to other interesting things. And evolution has this property, being an amazing collector, because every species that exists is now added to the collection, and creating new opportunities for other species to appear because we're all opportunities for each other. This process of collecting is what, in part, accounts for how open-endedness works, and it's also, you know, present in civilization. Like we see that the more inventions we collect, the more inventions we can create, because we stand on the shoulders of our predecessors and build on a previous invention. But also, one invention creates an opportunity for another; like a door you need before you invent a key. And so this is the thing we need to account for when we think about the divergent property of the planet, and it's what is going to be present for sure in AGI or super intelligence.

Okay, what if we just look at the evolution of one organ, like the eye? Okay, so how, how did the eye get built? Could that be modeled as a convergent or objective-seeking process? Or no, not entirely, of course. Parts of it, like why it got better for certain periods of time—which I don't know, because I don't know the full history of the eye—but those probably are well described through what you're saying. But there will be points in that, in that, in that story which explicitly did not happen because they made sight better.

Um, that's always true in all lineages. Like the first mutation that happens is just like a complete shot in the dark. You have, it's not because it's actually increasing fitness with respect to anything. Um, and often those things are co-opted much later; like they don't increase fitness for a long time, but it turns out to be useful for something other than what you would say was like the original purpose, far in the future. And those are the moments that are actually really instrumental, you know, in terms of explaining the, the global explanation for why is all of this stuff here; is the moments that we're not optimizing, but we're actually, for other reasons, which have to do with the structure of the search space.

So just to recap what you just said, now you're basically saying, yeah, the eye just accumulated a bunch of, uh, selected-for genes where each one was better at making the eye better. But look at all the mutations that happened to make that work; like you're basically saying, let's focus more on the mutation side, right?

Well, yeah, and of course it does collect some that make things so-called better with respect to some measure of survival or fitness. But you know, some of the, there, even along the, like it sounds like from what you're saying that we're talking just about before there was an eye, and like it's true those initial mutations before there's an eye fit my story pretty well, because like there is nothing to optimize yet. But even after there's some kind of light sensitivity, um, I would assert as very likely that there were actually some mutations that led to, uh, to, to decreases, but may have still been kept around for other reasons; decreases in sort of like visual fidelity, and yet they turned out to be stepping stones to better vision in the long run. I mean, this is true; this kind of circuitous process is true in all innovative processes. You go up, and then you go down a little; you take a step backward, then you take a step forward, because often the stepping stones to a revolution don't look like the revolution yet. And so certainly that kind of thing happening.

Go on. You're basically saying, hey, in the course of evolution by natural selection, let's say the evolution of the eye, sometimes you have a couple generations where vision gets worse, just kind of by randomness, but then in retrospect it turns out that those were actually on the way towards some new structure that ended up helping vision more. You're saying that happens in the course of evolution by natural selection, and that's critical to your mental model of how evolution works.

Yeah, that is. And you know, but I would push back; I would push back and I would say actually evolution is characterized by not doing that very much, because evolution is actually constrained where it really can't afford more than one or two generations of that kind of thing, because it's very dumb; it kind of only goes uphill in the fitness landscape.

Well, I would call that the death match model of evolution; that's like you're just basically articulating survival of the fittest, and it's a, that that term is unfortunate because it implies that we're living in this eternal death match where everything is life and death. And this is true in some lineages, and it's actually true that when you're in a situation where you're, you're threatened with complete, the complete end of your survival or your lineage, you can't afford to innovate very much. Like if you're about to get killed, you can't go off and compose new music; like that's not going to happen; you're going to just be worrying about the basics, like getting food. But that is not always what evolution is like; there are enough resources to go around, often a lot of the time you're not in a death match situation, and when that happens, when less evolutionary pressure, then actually there is room for these kinds of mutations. And it's not clear in which lineages, when, uh, you know, how many generations we can afford to do that, but that is actually why there's so much innovation is because it's not an eternal death match. Because if it was, it would absolutely converge. And so it's actually, I think that that survival of the fittest metaphor is, is leading us astray, because it's missing the fact that we actually need to relieve pressure sometimes to allow more innovation, which is why you do things like you, you, you, you allow people a certain amount of time in graduate school before you kick them out of the program for their experiments to work; like it's a protection of exploration and innovation, and evolution has that built in. So evolution is actually a process that allows time to twiddle around and try things, and because it's not actually just purely competition, right? I mean, some people say satisficing; it's okay as long as you're good enough; you don't have to be better than the other guys. Like there may be another lineage that's absolutely exploding because they have better eyes, but that doesn't necessarily mean that as a consequence your line is dead; there's room for everybody sometimes, and for the resources to go around, and this allows us to try lots of different things, and those things have lots of opportunity, uh, to innovate and find something new.

You're basically saying we can learn from evolution how to not be hardcore always hill climbing; how to like step back and reflect and decide to go a different route or try something else in parallel. But if you look at the evolution of the eye, it's very much just hill climbing in a certain direction, so is what's the lesson from evolution here? The lesson to me just seems like, yeah, hill climbing is a pretty good way to get a certain level of intelligence, and objectives tend to get satisfied when you have a consequentialist feedback loop; that's my lesson. I, I don't understand; evolution never stepped back during the evolution of the eye; it just kept making it more, you know, better at seeing light.

Well, if you zoom in far enough, evolution does look like the story that you're saying; it's always about zooming in though, so it's like I don't, you know, I don't know exactly which section of the evolution of the eye you're zooming into, but there will be a section that looks like that; maybe you'll even claim it's the entire lineage; I don't know; I doubt that we have the evidence for that, but I just want to point out if you zoom out you get their story starts to look really weird, you know, because like there's things like there are instrumental mutations that led to intelligence; like something like Shakespeare; like what led to Shakespeare? Well, if you go back far enough, because I zoom out far enough, like the appearance of bilateral symmetry was an instrumental step towards Shakespeare, but there's absolutely no increase in intelligence that you can, uh, you can point to when the first bilaterally symmetric flatworms appeared on Earth, and yet they are essential stepping stones. And so we need to be able to get those into the picture, even though they have nothing to do with the dimension that you're talking about, but they are instrumental in the long run. Like you're not going to get to those things without those mutations. I mean, you could argue that maybe there's some alternative route, which is probably true, but that alternative route, the same story would say, would say the same thing too; there's going to be instrumental mutations, absolutely fundamental instrumental; like this is revolutionary, bilateral symmetry; it's not a minor side story; this is the main show, um, and so like we need to explain like how do you get this when it's essential to get that when it's not accounted for by optimization?

What are you saying is the main show again? Bilateral symmetry, for example, is like one of the most important innovations, like in, in the history of human evolution, um, so that's what I mean by the main show; like without bilateral symmetry, arguably, like this lineage would not be as interesting or impressive or have done all these things orre or achieved, you know, human-level intelligence. So that's important, but it's deceptive because it doesn't look like if you give an IQ test to the flatworms, it doesn't look like there's any progress along that dimension, but it's certainly important, um, and so like there, if you zoom out far enough, you're always going to find stepping stones like that, and they have to be accounted for.

Okay, well, let's say that, uh, we're early at the dawn of life, and I want to help evolution work better, and I figure out, man, it would be great if these organisms could sense light; it would be great if they had a powerful organ for knowing where the light comes from and what color the light is, you know, what frequency, whatever, uh, so, so I want, I want the process to build an eye, okay? But Ken tells me that divergent search is really good, so I got to make sure it does divergent search. But in this case, wouldn't it actually just be better to do convergent hill climbing, which is what actually happened, right? So you're admitting that in the case of the eye we don't really need the divergent search?

No, no, I, yeah, I certainly not that. I mean, this is a great thought experiment; I'm glad you brought it up; it does help us to get close to this issue. So the truth is, I think that that would fail if we could do that, if we could run that thought experiment in the real world, because the thing is that before there's eyes, because we're at the beginning of time here, the beginning of evolution, uh, the mutations that we need don't have to do with increasing sensitivity to light; we're going to need several, uh, precursor mutations that lead up to the point where there's even the infrastructure to allow that to be useful. And those precursors would end up being selected against because of your artificial fitness function. But actually, let me, can I give you a different version of the thought experiment that's easier to understand? Because I, I, there's one I like to give here, and I think this is really easy to understand; let's do it more in, in, in the context of human innovation because it's just easier for us to understand because we, we actually did do these things; like we invented a computer, right? We invented the computer in 1944; the ENIAC is usually the first account; like we, most people agree that's the first computer. So let's just pinpoint that date, 1844, um, and so like what I like to do now is something similar to what you proposed; let's go back 100 years to 1844, and let's propose that we build the computer 100 years earlier, because wouldn't that be better? We get the internet 100 years earlier; we'd be farther along now. So we, we'll do a thought experiment; we'll go back in a time machine, and we'll say, okay, let's get all the smartest people from 1844 together and give them a new objective, and we will optimize towards a computer. But the problem is that what you need to get the computer are vacuum tubes, a fundamental building block of the first computer, and guess what? All those smart people are doing, they're building and experimenting with vacuum tubes, but not because they're interested in computers; they're not thinking about computers at all; they're thinking about electrical experiments. But we just pulled them off that experiment because we want them to build a, build a computer. So we took them away from the thing that was necessary, which is the vacuum tubes, to work on something that needs vacuum tubes. So what do we get? We will get neither computers nor vacuum tubes; we just destroyed progress. And the same thing would happen if you went back to the beginning of time and started testing single-celled organisms for light sensitivity; it would destroy the population; they'd all die because we need other things first that don't look like light sensitivity. So that's something we don't understand; it's counterintuitive because we're so focused on the optimization metaphor; think everything works that way, but the optimization metaphor will kill you in the long run in divergent, open-ended systems.

So to me, when you propose stuff like this, it sounds like you're just compensating for a lack of intelligence and a lack of foresight, you know, in the specific case of human inventions, right? You're saying, hey, the tech tree, it didn't work the way people would have predicted, and so we just kind of need them to stumble along and find other stuff and, and build the tech tree organically or randomly. But if you add a dose of intelligence, if they just had better foresight, if they were just like, hey, uh, based on what I know about electron physics, we really need like a solid-state transistor; that's really the way to go, not relays, not vacuum tubes; couldn't you actually skip levels of the tech tree?

No, I mean, and yeah, this is a fundamental, uh, aspect of my argument is just that I, I, I would argue this, this is almost philosophical, I guess, but that we cannot foresee the tech tree; it's just completely futility. And you're being overly optimistic to project that; it's like we are not omnipotent, and omnipotence is impossible; we cannot understand how the universe works without experimentation; we have to try things. But it's important for me to highlight that trying things is not random; people often think I'm saying, well, we just flail around randomly; like people working on vacuum tubes are making random decisions that were like without any basis in reality; they were highly informed because they understood that there's very interesting things about the properties of these technologies, and they wanted to see where that might lead, even though they don't know ultimately where it leads. And this is why the tech tree is expanding because people follow these interesting stepping stones, but we will not be able to anticipate what will lead to what in the future; only when you're very close can you do that, and that, that's, we, we do do that when we get close, and that's why we're misleading ourselves by thinking that intuition expands to the entire tech tree, which it will not.

Okay, so I agree that humans, human inventors aren't that intelligent, right? I mean, there, there's many stories of inventors kind of bumbling around, and then they hit on the invention, and they're like, oh, of course, Eureka! But then in retrospect it's like, yeah, but isn't that kind of obvious? And I'm not, you know, I'm not one to judge, right? I mean, I miss obvious stuff all the time, right? I'm not that smart. But realistically, I mean, like, I, I read, I read a book once talking about the invention of the telegraph where they converted electrical signals to sound, right? Like, so, so you can press it and send a message of clicks. And before they had the clicks, they were like, man, how am I going to have these electrical signals communicate a message? Maybe there could be like little pieces of paper moving around in water, and they could all form a message, and they were wondering like, how am I going to convert electricity into a message? And then you're like, oh my God, just the electricity itself, just the circuit opening and closing could, with like a, you know, a tiny little actuator on the end, like that's really all you need; like that was a Eureka moment. But reading it in retrospect, it's like, well, yeah, I mean, that's, that's pretty ob, like now, right? It's very easy to look back, so don't you think this actually is a microcosm of things that humans are like missing all the time because we're just not that smart?

Well, I, I, I do agree that often in hindsight things look obvious, um, like that, that, that, that's a common observation, um, but I still think that ultimately your, your, your, your, your argument is, is, is, is not going to stand up because you're really arguing for the possibility of omnipotence. I mean, we live in a universe where you have to actually try things to know how it works, no matter how smart you are; like it doesn't matter if you're the, the most, I don't know what it means to be the smartest possible thing; you still cannot know reality without interacting with reality; like the way that the universe works is, is, is an a priori fact. And if you just had a, like a really smart brain in a vat that does, has had no exposure to the universe at all, it can't derive anything from that; it has to experience it. And so, like they make better hypotheses; I mean, I agree with you that, like, I mean, this is sort of obviously verging towards AGI, the hypotheses will be better, but it still needs to make hypotheses, and it still needs to test those hypotheses, and it will still need to follow that tech tree and gradually discover things over time; like omnipotence is not on the table, even for AGI or super intelligence.

Okay, all right, so yeah, a lot of different directions to go with this, um, let me, let me ask you this, just getting back to, uh, your distinction of objectives versus open-endedness, uh, one thing that you, you say in your book is that as soon as you create an objective, you ruin your ability to reach it; unpack that.

Yeah, yeah, so there's a whole book about this, so it's, it's something, something that, that, that I'm now condensing into a nutshell. I, I want to just clarify that we're talking here about ambitious objectives, of course; most people's knee-jerk reaction to a statement like that is that it's crazy because they say, oh, but I've had many objectives in my life, and they worked; like how could you claim this? It makes no sense. I acknowledge that when things are what I call modest, it does work, and it's a good idea to pursue an objective. And I want to be clear what I mean by modest; I don't mean easy; what I mean by modest is that the stepping stones are known; how to get there. And so that's something like, I want to get in shape; like it may take a lot of work to get in shape, but it's generally known what needs to be done; like you might need to get exercise, might need to change your diet; it could take years, but you know what to do. I'm talking about situations where we have absolutely no idea what the stepping stones are; these are situations which we would call innovation, innovation where like, I don't know, you know, how are we going to build, how are we going to cure this disease? How are we going to create AGI? Uh, how are we going to create unlimited energy that doesn't pollute? Like things where we have no idea what the stepping stones are; that's what I'm talking about. And what I'm saying here is that when you make that the objective and you play the OKR game and you say we're going to follow a gradient towards that objective, you will lose because the stepping stones do not resemble making progress along that gradient. That is because the, the, the world, the real world is complex, and complex search spaces have a property which is that the stepping stones that lead to the

Things you want do not resemble the things you want. That's called deception. And so if you live in a world with that property, by setting your objective as saying, "I need to increasingly approximate the thing that I'm walking towards," you'll miss the stepping stones that don't look like that gradient. So it causes you to have tunnel vision, and therefore you won't ever achieve those things. The way to get to those stepping stones is to not be worrying about the final objective. We will get to those stepping stones for independent reasons that are orthogonal to the objective. Okay.

In the example of, uh, natural selection creating the human eye—or I think the eye evolved independently in a bunch of different organisms—so in that kind of example, yeah, is that uh, what you would call a modest objective to create that? Let's say starting from, let's say you have an organism that has neurons. So you're you're at that base level, and you want to make like a a fully-fledged eye. Would that be a modest objective? So, so to to create an eye from scratch is not a modest objective. If you're saying given a certain starting point, it might be, yeah. Because like there is a point when something snaps into possibility because the stepping stones are there. And so it's a question exactly when that is. Like I'm not claiming that I could answer that for any particular invention. Let's start from like a patch of light-sensitive skin, right. But you don't but it doesn't know like, yeah, maybe I I'd be open to saying that I'd be open to saying that you know you I mean I think if you really wanted to split hairs about this story that it's probably not there yet because I'm guessing that simply light-sensitive skin still is kind of far from this Construction. And so there's going to be a few deceptive stepping stones. So it's still not modest, but there is some point where it is getting within the veillance. This happens all the time, like in in human innovation. This is why we do achieve greatness. You know, I mean it's like I don't deny that humans achieve greatness because sometimes something is invented, and the geniuses are the people who realize they're the first to realize that the stepping stones have snapped into reality. That's real genius. Like fake genius is a visionary who thinks we're going to invent something far off in the future that has no hope of being invented anytime because there are no stepping stones yet.

So I'm getting confused because you defined a modest objective as one where we know the steps going forward toward the objectives. Now, in the case of evolution, there's no uh foray, right? So it never knows the steps; it's just getting consequentialist feedback when it takes the steps. So doesn't that mean um there's never such a thing as a modest objective in evolution? Yeah, I mean actually I I prefer to actually put it that way. I mean I agree with you there that like I mean evolution is just not objective, and and and I mean so you're you're right actually that that even your your argument caused me to contort myself in a way that I don't prefer, you know, to kind of describe this as an objective process once you're near something because there never is an explicit objective in evolution. And my argument would be that is why it's so prolifically creative. It's essential to have for creative processes not to be objectively driven, and evolution is the greatest creative process that we know in the universe. More it sounds like it sounds like your definition of the concept of an objective: something can only be an objective if there's some kind of explicitly represented foresight before the optimization process proceeds or the search process proceeds. It has to be an explicitly represented objective in order to qualify as an objective.

Yeah, probably. I mean I bet we could we could get into details of exactly what that means to be explicitly represented, but yeah, basically I think that you have to say where you want to go. It's not foresight though; I don't like to call it foresight because that sounds like you know something about how to solve it. It's just that you know what you want. That's the thing you need to know, right? I guess, but I'm curious what's going to qualify as an explicitly represented objective. So in the case of evolution, there's the only way to get the objective out, you know, inclusive genetic fitness um that the location like, you know, where is it written down: optimize inclusive genetic fitness? You're only going to find that in the entire physical feedback loop, right? Where where the ones that have better inclusive genetic fitness just turn out to have more copies in the next generation. That's the feedback loop, and then and then you have to look at that entire feedback loop and reason back to be like, okay, there's this property, inclusive genetic fitness, and that's actually the criterion that we're hill climbing toward um in the case of a mural net, it's going to have a blackbox reward function, right? It doesn't necessarily know the code for its reward function; it just knows that like a training pass happens and then it gets a certain uh loss score, right? Okay. So I mean I don't I don't think of inclusive genetic fitness as an objective. Like I I I do interpret Evolution as not having an objective. I think of it as a non-objective process um and there's no nothing wrong with something that has lower fitness. I mean what we care about in evolution, what's interesting ultimately, not necess what has higher fitness, you know, cuz like I said our fitness is probably on an objective basis lower than lots of species that we would consider less intelligent than us. It's the orthogonal issue. So we're not we're not in an objective process from my point of view because there is no expli explicit target that's been specified here uh like fitness is not a target. It's like I said, it's a constraint: like you have to minimally be able to reproduce. That's all.

Maybe I could maybe I could reframe it this way if you could just give me a a second to reframe it. Sure, sure, sure. That like I would want to give a different metaphor for what evolution is because we're really used to this fitness-based metaphor, like this doggy dog competition survival-based thing, but let me put it in a completely different way: like what you need to do in evolution ultimately to be successful is you need to have one cell lead to another. That's all you really need to do. That's what it means to reproduce. Like there's one cell which is the parent, and then eventually there's another cell which is separated; it's a different individual which is the child. Everything else that that happens is orthogonal to that, like so like all of your all of the complex things that happened over the decades of your life that led to you going from a baby to an adult are not necessary in any sense to achieve this supposed objective of getting from this cell to that cell. It's a giant digression which is very much like a Rube Goldberg machine. Like it's like I remember I was watching a a TV show about this guy who is making Rube Goldberg machines to open a newspaper, and it was really funny because he invented the most complex things you could possibly imagine to do something completely trivial. Like there's a ball that rolls down a hill and then sets a fire, and the flame causes something else, a balloon to burst, and it's like all these sequences and then eventually the newspaper is opened in the morning, but it's all completely orthogonal to what we need to do here, which is open the newspaper. And this is effectively what's happening; this is a Rube Goldberg machine generating system. None of this is essential, but what's interesting about it is that the side effect of infinitely generating Rube Goldberg machines with no purpose and no objective because it's all not essential to the to the fundamental constraint is that the complexity of those machines itself becomes what's interesting. It's not the fact that they're getting better in some are in some ways; they're getting worse because it's crazy to go through all this complexity just to get another cell, but it's the fact that the complexity of the thing which is a side effect of infinitely generating these machines is itself interesting, which has nothing to do with the fundamental constraint. It's not the actual reason that you can explain why the Rube Goldberg machines are interesting. So we live in a rub Rube Goldberg machine generating universe, and this is just going on forever. We're going to get more and more complexity. Anything that's viable goes. It's a different story. It's a different story than this like death match convergence type of view of evolution. Okay. It's you know you you uh compared things based on how interesting they are, and I guess I want to unpack that later because I think that's it sounds like you've got more to say on that topic of like what's interesting, right? Sure, yeah, yeah. All right. So so we'll get to that soon.

Um, but I want to go back. You know, I kept asking you about your definition of what does it even mean for something to be an objective. I'm still a little bit confused about your worldview, but let me tell you my worldview because my worldview is where I'm not confused. So my worldview is that uh the question that I like to ask that I think is relevant toward the Doom question, you know, it's relevant toward a lot of things, uh is I I just ask uh which system or which process is capable of uh getting outcomes, steering the future. That's that's the lens that I come at things from. So when I look at evolution, I'm like, aha, this particular process—mutation and natural selection—this is going to let me predict without knowing exactly how that if I see an organism with a light-sensitive patch of skin and I come back a million years later, there's a very good chance that it's going to be something like an eye. How does the eye work? I don't know; it's quite complex. I don't even understand the details of how the eye works, but I know that it's going to be better at sensing light and letting the organism survive and reproduce. So that that is an example of what I call an optimization process, to steal Eliezer Yudkowsky's terminology. And the reason why I focused in on it is because I noticed a part of the universe that lets me predict outcomes of how the future is going to get steered in a way that's quite complex—it's so complex that it would have been hard for me to model the system in any other way than to model as as as an optimization process. If I just look at the biology of it, it would have been very hard for me to reach the conclusion on the lower level being like, oh yeah, all these chemical reactions are going to happen, this particular protein is going to get CED, and I can tell you that in 10 10,000 100,000 generations it's going to be better at sensing the direction of light. No way. The only way I know that is because I say, aha, optimization process. There's some criterion, and it keeps getting better at that outcome criterion. Okay. So I take my optimization process worldview. I notice that biology is doing it, and then I notice that the human brain is doing it. Now I notice that artificial intelligence is doing it, and I also think there's headroom above Humanity to do it better, right, artificial super intelligence. So that's the frame that I come out come to things with is I notice things that can optimize the future. So I wanted to ask you, you know, you talk a lot about objectives, and you talk about the types of processes that are going to achieve objectives or whether having an objective actually makes you less likely to achieve the objective. So I guess going back to your claim, right, you you make this claim—you say you say in your book, as soon as you create an objective you ruin your ability to reach it—so I guess I just want you to like define your claim. What does it mean really to create an objective?

So first I just wanted to respond to your worldview that you just articulated because I again think that your worldview is very zoomed in, you know. So you're talking about like once you have this light-sensitive thing, then I can start to predict what might happen, but I want you to zoom out because like evolution intrinsically is really characterized by not being predictable. It's not predictable at all. Like if you go back like at at say the 100 million-year mark and you're just seeing like a the beginning of the emergence of multicellular life, there's no way in hell you're going to predict that humans are going to be here. Like it would not look—there's a lot of organisms that are very adapted to their niches, right? So I I would predict that there's going to be organisms in the water that can breathe underwater and organisms on land that can breathe on land, stuff like that. I I don't I don't know that you would I don't I'm not sure you could predict that that there's going to be breathing. I mean at that point—fair enough, right?—but but but I mean there are constraints. Yeah, and sorry to interrupt, right? But I don't think it's I think saying that it's I mean at the highest level I would just predict that there's a lot of organisms that are just good at surviving in whatever niche they're in. I mean that seems like a very meaningful constrained prediction. No, but I mean that's not an objective type of view; that that's just saying there's going to be a lot of diversity. That's starting to lean in my direction. I mean I would also predict that we're going to see more diversity; that's what we're going to expect, but you cannot predict the particular things we're going to see. I mean that's that's like what about this? Can I take a step? What about the idea of um having structures, having modules, uh using energy, right? You h having local regions that are lower entropy than their surroundings, right? The these are instrumentally convergent things that life will do that super intelligence will do. So so I think you're selling it short here of what I can predict. I mean that's what you're already observing at that point, like we already have that, but I could have told you that at the dawn of life, right? The moment I see the the process start of of the first replicators, the the moment I notice that they're getting selected for, I can tell you, okay, you're going to see modules, right? You're going to see neg entropy, like the again convergent properties. I I mean I I think um I mean even the very first replicator, which we don't know exactly what it was, but presumably something like a single cell that already has some degree of modularity, it's a quite complex machine. So tell I don't I could have told you you're going to see a fractal of modules, right? You're going to see a system that has cells, organel, organs, right? It's not a shockra that there's these different levels of systems in—I mean multicellularity would be an incredible prediction from that point of view. I mean if you didn't have the the the the advantage of hindsight, like that we don't know that any of this happened, and you're just looking at a single cell organism floating in a like a sea of chemicals, like you're going to you're not going to predict multicellularity. It's it's an inconceivable development. It's it's it's an amazing development.

Yeah, I I agree that the jump to multicellularity I couldn't guarantee you that, right? Because there's probably some planets that got stuck at single-celled life. So I I agree. I can't tell you I knew multicellularity was coming. Yeah, but if if if you told me, look, life is going to thrive, right? If if you just give me a constraint of like, yeah, I'm telling you this natural selection thing, it's going to go places, a lot of stuff is going to happen. That's your only hint. A lot of stuff is going to happen; it's going to be wild and crazy. That's your only hint. I'd be like, okay, it probably jumps to multicellularity, or if it doesn't, then it's going to jump to a lot of substructure within one cell, right? So you can just have like a huge single-cell organism that has a lot of subcellular structure. Yeah, I think you're giving yourself a lot of credit to say that you could predict multicellularity by looking at single-cell organisms, but I'm willing to give you that. So let let's just go with that that you actually could predict that that will happen. What what what matters though is not like if just one particular advance is predictable because the closer you get to the beginning the easier it is to predict things. This is obviously going to be true. It's like what it becomes more and more unpredictable the farther out you go. And so like again if you zoom in enough things start to be more predictable because the the stepping stone's already there. It's like what comes next. Like a single-cell organism, if you want to view it that way, is a stepping stone to a multicellular organism. You're still not going to predict that there's going to be Shakespeare writing plays a billion years later. Like that's way outside of the scope of prediction. So I mean humans broke the mold of evolution by natural selection, right? Evolution by natural selection is characterized by hill climbing in order to increase genetic fitness. What humans are doing with our brains, this is a totally new paradigm. You actually can't model modern humans as as purely uh adaptations to maximize—that's fair. I mean so so we don't we don't have to get into the specific cultural predictions, but you're not going to predict human brains. You're not going to make it make a prediction like that—you're not going to predict humans. You're going to have no idea what we're going to—are we going to get bilateral symmetry? Are we going to get four-limbed animals? Like why would they dominate? I mean if we ran this this show over again, like if we started over on Earth even now with hindsight, we have no idea if we're going to get these—we could get some other—we get metabolisms—maybe we wouldn't get four legs, but we would get metabolisms. Yeah, but I mean again that's zooming in. Like we already have a metabolism in the first thing, so so it's not like hard to predict that there's going to be more metabolisms to come. It's like actualisms are going to have, right? They're going to be fractals; they're going to be modules upon modules, right? I mean so so life evolves a certain way because optimization has a certain character to it. It's a very V it's a very vague notion to be able to claim that you're actually making a prediction just to say that there are metabolisms. Like I I might grant you that—I'm not even sure what that means exactly, but I might grant you that—but like I said there are some things you can predict because they're close at hand. Like I do we we know like an open-ended systems we should say. So I didn't give this example like that you know we did this pick breeder experiment, right? Like that was one of the first things that led to uh a lot of these ideas was this pick breeder experiment where we were letting people breed pictures online. This is kind of like a little mini evolutionary microcosm that we were in more than 15 years ago. It's an old experiment, but one of the things about that experiment, you know, people were breeding, so starting with blobs, if you kind of think of the first cells, and so it's a little bit of a metaphor, and then you can you can you can choose parents and then you get uh you get children with slight mutations, and many people were crowdsourced to do this. And one of the things about it was you could kind of predict the early things you're going to see, like you we knew that we're going to get like things like circles, you know, and and we get lines. Like there's certain things that are easy to get, and they're sort of like within the veillance of the blobs that you see at the beginning, but the farther out you go you started to see people discover things like skulls and cars. There's no way anybody could possibly have imagined that. That was absolutely incredible and unexpected. You didn't—I would have thought it's completely impossible in this space. There's no way. It's the same with single-cell m like you could not say it's going to create something like us. That's like an unbelievable projection. You'd have to be omnipotent to be able to say something like that.

Um, are you talking about uh so it's your own software where the images were procedurally generated uh and and by the current uh shape of the parameters is that pick breeder or or how would you describe it? Yeah, I guess I mean I'm risking getting into the weeds.

Of something that maybe is like that important to explain everything, but basically it was it was a it was a genetic algorithm under the hood. Some people call it evolutionary art, but there's like an artificial DNA, so it's not a generative model in the modern sense because those are things where it can generate like millions of possible or billions of possible images. This was like there's just one network that generates one image, and then if you chose that Network, then it could have a child which would be a mutated Network, which means it's not explicitly using gradient descent or something like that; it's just perturbing weights randomly. And so you'll get a related image if you perturb them lightly, and it's just it's just breeding; it's like a DNA underneath the hood, so there's an artificial DNA.

Were these images high quality? Like how how is were they like Photo realistic or what what do they look like? So I would claim they were high quality, but not in the sense of like, you know, modern generative models where you would be wowed today if you saw them. You would just say that's a pretty nice cartoon of a skull, but it's not like I'm blown away by it. But in another sense, I think they're absolutely incredible because you you would not imagine that this is possible if you saw the initial search space, which you which you do if you you start from scratch. You would think that this is just a toy worth playing for for 2 minutes that'll get me some blobs and wallpaper patterns. It's quite remarkable that it eventually leads to these like almost real cartoon-like images that look like real things.

Okay, I'm going to put up some Picbreeder screenshots uh in the video feed for the show, um the ones I'm finding on Google Images. So it looks like it's not photorealistic; it looks like it's started from like abstract um kind of mathematical looking patterns, and the patterns kind of got shaped in into um a little bit like a Rorschach test, kind of like getting shaped into forms that are like human-like face-like. Okay, got it, got it, got it. Okay, but let me just say one other thing about Picbreeder because it's probably important for the point being made. The original insight comes from the observation in Picbreeder; it's like Picbreeder was part of where I came to this Theory. The original insight comes from this very profound observation I think, which is that we discovered that when people find those things, images that look like something like the skull or the car, they were always not looking for it, or at least say 99.9% of the time. Like the only way people find things in that system is by not making them the objective, and when people would have an objective, they would fail, get frustrated, and quit. And so that to me was very confusing at first; I was trying to understand why it has that property, which is what led eventually to this Theory. So it it has this property that setting the objective would destroy your ability to actually create something.

Okay, all right, sounds good. Um, let's uh let's get into the the heart of the disagreement here, uh at least what I like to talk about, which is AI [Music] Doom. What's your doom, Dr. Ken Stanley? What is your P(Doom)? Uh, so um I think I'm going to be uh a heretic and not give you a P(Doom), but I will say it's greater than zero, but I just do not want to give a P(Doom). Maybe you'll press me into it because I could obviously I'm intellectually capable of giving a number, but it for me it it it it bounces around every day, and but that's not the real reason I don't want to give it. The real reason I don't want to give it is 'cause I don't feel like that's the important point that I want to make here. Like I'm not actually trying to contest your level of concern; I feel like it's perfectly reasonable what what your which I think I know what your P(Doom) is, and it's also reasonable, some of the arguments against that, and I'm grappling with that myself, trying to figure out where I stand. But I feel like my contribution to the conversation is not that; it is just to say that we need to grapple with open-endedness, which I feel by listening to a lot of these debates on your show and other places and reading about this that it's just not part of the conversation. I think if I give you a P(Doom), I will immediately polarize your audience; some people will see me as as for some people see me as against, and I don't want to be polarizing along that dimension because it's not the thing that I think I have to offer that's useful here. I want to force everybody on both sides to grapple with open-endedness because I think it's the fundamental issue that's missing from both sides of the debate, which makes the de the debate in some ways uh ill-founded, and I want the PE I want us to figure out what risks we face. I mean, I think it's an important discussion, so it's not that I'm dismissive of what you're doing; I just don't really want to be pigeonholed into this number.

Okay, well my P(Doom) is 50%, and one of the things he said just now is that my P(Doom) is reasonable. So would you say that a 50% P(Doom) is reasonable? Um, so so I don't I guess what I'm really saying is I don't really know what to make of P(Doom). So I guess in that sense like yeah, my error bars are so wide that this is within reason, but then you're going to say, oh, well then your μ is effectively 50%, but I don't want to pin it down like that because like there are days where I feel like it's it's much less than that, um and I argue myself into a corner in different directions. But what I'm when I say it's reasonable, I just want to be clear what I mean is that I think that you've put a lot of thought into this, and a lot of it is based on solid logic and about as good as a human can do right now in terms of projecting the future, and and I mean I I can argue you down in certain places and other people could, but I think like the most important thing for you and others is that you're founding your prediction on a good premise. Like that's the thing where I think there's one really weak premise here, which is that the whole world works like an optimization process, and that's how we're deriving our expectations about what AGI is going to be, and that's entering directly into not just you but many people's expectations about the dangers that are ahead, and that's where I think that that's just not true. And and it's not that you're wrong about the ultimate prediction because open-ended processes are dangerous too; it's not so I'm not like just trivially saying like I'm going to like knock down your position and make it seem like everything's safe. But the thing is that openness is a different kind of danger; we need to grapple with that; that's what I want to get on the radar with this show; it's like that's my agenda is I want us to be grappling with the dangers of open-endedness, not to say that like there's no danger here, but it's like we're just completely talking around it. So there's this other guy, Yudkowsky, whose P(Doom) is %. We've established that I'm reasonable to have a 50% P(Doom); in your view is Yudkowsky also being reasonable? I so I haven't I don't know if I've heard his argument; I might I might, but I don't think so. But so I don't know because I haven't heard it, but I bet he could be; let me put it that way. I could imagine constructing an because I guess yeah, we I believe we live in a a very serious fog right now of uncertainty, um and so I I think that a smart person could construct a reasonable argument for that as well, um and so I I probably could respect those arguments uh coming into that as well. But I but really I think what I'm saying is that all arguments are highly flawed right now because there's a lot of missing information. So when I say reasonable, I don't mean correct, obviously; it's just like it's very hard to construct an argument right now but given so much uh like a lack of understanding of like the space here, um there's a lot of reasonable ways you could go and come to conclusions that are highly divergent.

Just to finish triangulating your view on Doom here, so you said that it fluctuates depending on the day. So if we look at the last year of days and just uh average over that, what do we get? Yeah, so I do want to acknowledge that I have been very concerned on some days, like very concerned, and I and I want to make the clear what I mean by concerned. Like one thing I'm not totally sure is like if Doom means the end of the human race or just like bad outcomes. Like I'm not always certainly worried the human race is going to end, but I am worried about bad outcomes, and to the extent that I've made decisions based on that. So it's not just a trivial worry; like I one of the reasons that I left AI for a while and I started a social network is because of this issue; like I was actually feeling bad at some level because I wasn't sure that what I'm doing is actually something I can be proud of, um and this was really gnawing at me like over time like over my latter days of OpenAI and in my in my career that I I sort of thought maybe I need to step back a little and and really kind of reconcile what I really believe here, um so that I can I can feel just proud of myself that like I'm actually contributing to a better future, and I thought that starting a social network based on this is like I don't want to talk about the whole social network thing; it's a whole other show, but I thought basing it on what I learned about open-endedness was a clear social good; like it's not trying to create an AGI; it's just trying to do something good for Humanity, and that that would give me time to actually process what I actually think about this issue. So it's not like I I haven't taken this seriously; like I've taken this very seriously, and I do think that like bad outcomes are high on my mind in terms of I just want to be proud of what I'm doing and feeling like this is actually leading in a good direction, and I'm genuinely confused and not sure what I think on a day-to-day basis, and I can definitely rationalize myself into thinking that I should be involved in this because at least like it gives me the chance to guide it in a better direction, but then I think maybe I'm just rationalizing by thinking that, and it's like just an excuse. So I'm I'm not like I'm not one of these like, you know, it's like everything's fine; let's just go with it, and I I've seen some of your shows where you did a very good job dis dissembling with people um who sort of just trivially dismiss LLMs or think like like that; I was quite impressed with some of your arguments, um so you know yeah, I think I'm kind of hard to pin down, but it's not trivially so; just to be evasive.

Okay, if you want a P(Doom) number that you can use, uh you can go with 20%, because everything you just told me right now kind of roughly maps that in my mind. But is Doom um you're more of an expert on this; this is the P(Doom) show, so so it's like are is it are we talking about cataclysmic civilization-ending events or or Does it include things like just everybody losing their job in a lot of social chaos or something like what what is Doom exactly? It's a lot more than unemployment. Yeah, I mean unemployment doesn't even necessarily have to be bad if there's a lot of wealth being created where everybody can just have a good share of wealth and then they can still work a job; I mean, they can still do the activity that they love within their job and like yeah, maybe there's some issues there, but I don't see that as Doom, um that's just me personally. I see Doom as like literally dying, you know, that the species is no longer existing, um I think a good a good way to explain it is like today there's a lot of value on Earth; we're really glad Earth exists; it would be really bad if Earth just disappeared and then the whole observable universe was just like devoid of life and positive interactions; like that would really suck, uh and if economic growth continues on the contrary, we we're going to have like trillions of planets that we can expand into as long as our our of space travel and biology and uploading our brains, as long as all that stuff works the way it seems like it's on track to work, you know, the transhumanist future or the space exploration future; it seems like there's many trillions of planets there for the taking where we can build like lots of Earths; we can discover new ways to have fun and expand our consciousness; like it sounds like there's going to be like a great party happening if we can just continue economic progress. But on the other hand, if AI goes rogue and has its own objectives and it's too late for humans to control it and we're like oops, we really messed this up, but it's too late, that's the Doom scenario that's going to not only wipe out the potential of those trillion planets but probably even wipe out the one planet that we have, and in the worst case it can create, you know, suffering risks where it like simulates a bunch of our consciousnesses and tortures us. So it's basically you you take the entire future potential and you say we actually lose more than 99% of that; we lose maybe 100% of that or even it goes negative; that's my definition of Doom; it's like a 99%+ loss of

I got it. Yeah, so I mean I I admit that that's greater than zero, but but still again I'm not I don't want to embrace your 20% I because I just I don't I don't know that I really think that, um but um but you know I I'll admit that it's not zero for sure on that, and but you know I have other concerns too; like even short of that I I feel concerned about other outcomes that aren't necessarily the destruction of human consciousness, which I agree is a tragic and horrible outcome; like I I agree with those motivational points that like you know human consciousness is very important to preserve, um and it would be very hard to argue I think for like alter weird Alternatives where like somehow the AI is conscious, so it doesn't matter; I think that's a stretch to to make these kind of arguments, um so I'm very much in favor of the preservation of our consciousness in the universe, um and uh and so you know like I say it's reasonable concerns, uh but I don't think I probably we don't completely see eye to eye on this 'cause like like I'm not in like a a 50% uh confidence about or expectation on that every day like you are, um but it's somewhere out there as as as a as a concern.

Okay, so there's this famous statement on AI risk from last year; it's only one sentence; it says, "Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war." Would you sign it? So yeah, I'll just to you this evasive jerk; I mean I I don't want to sign that statement, but it's it's because I don't like I want to be able to speak for myself, and I find I don't like statements in general; I find them like bad in a lot of ways; like there's always so much minutia in the statement that just leads to polarization and controversy which is not core to the

Is only one sentence? Well, even there I think there there is minutia to that sentence, and I don't I don't necessarily want to assoc I just want to speak for myself and say how I feel, which I think I have, and I also think that like with with these kind of statements uh we we're what what I'm able to do uh without the statement is to articulate um like well here let me put it this way; I found them to often lead me to to actually deflate the value of the statement because I find the people signing them seem hypocritical after signing it, which worries me; like when somebody who has a really high reputation higher than mine signs that statement just goes on with their career; to me it's actually sending a signal that they don't actually think that; like I would think if I actually thought that like it would be the end of my career. Let's not play 4D chess here. Okay, what if it's what if you're the only signatory hypothetically? Do you sign it?

Um, well I don't I don't know because like I actually think that if I sign that statement then I should just end my career for sure, um and I just don't know if I'm ready to do

I you kind of did; right; you went into social gaming. I did, but I I might be going back, you know; like I'm I'm like I'm itching to go back, but I'm trying to figure out if if I'm justified; I mean I I am itching to go back; I'm probably going back. So am I am I because I'm thinking of going back am I justified in going back? Like I'm grappling with that, um but if I sign that statement then I'm not going back for sure, and I'm still grappling with that. So you can you can imagine a scenario where you're like, you know what, mitigating the risk of extinction from AI should not be a global priority, and I'm going to go back into AI research. So you're thinking maybe that's a possible scenario? It is a possible scenario, yeah, because there's a trade-off here, right? Like it's the the problem that we're dealing with on the other side; like the existential risk on the other side is this idea that I think human suffering is unacceptable; it's an unacceptable level today, um and so we're trying to trade that off with the possible extinction of human consciousness; like both sides of this are just totally unacceptable, and so like it's very unclear the that that hinge point where like we cross that line where it's worth it to give up on things that could alleviate massive suffering on a massive scale in order to prevent something that's not totally clear and fuzzy; like I I agree that it might be like something that we need to do, but I'm not there yet like in terms of being certain about that because like the human suffering is is actually real right now; like the fact that I don't know 70% of humanity is going through un like you know unconscionable suffering is like a reason to keep pushing forward on technology, um and so like the clash of these two arguments is very hard for me to reconcile; like I see both sides of this and why both sides are so strident in their arguments; it's very important to take them both seriously, and I've come out like uncertain basically in in what I think, and I feel like I could justify anything to myself, which makes me feel uncomfortable.

All right, so we talked about Doom, the worst-case scenario. Now before you mentioned uh what's interesting, right? Like you find the diversity of life on Earth very interesting, uh so maybe you could talk a little bit about like your Heaven scenario where everything is like maximally interesting. How do you think about that? Um, yeah, so interesting is is is an interesting topic; I mean it's it's it's one thing is it's intrinsically subjective, and that's a problem that we have with discussing it often because as scientists we want everything to be objective, not subjective, um but I think we have to grapple with interestingness when we talk about phenomena that are creative because ultimately like it's in the eye of the beholder what's creative, and it boils down to whether you find the products interesting, um and so like that's like in the final analysis what I think makes evolution compelling is that it's interesting, not that it's efficient; they're more fit or something; like those are just cold aspects that are not interesting in their own right; it's a subjective judgment, but I think that we could largely agree from a human perspective, maybe not a wholly agree, but largely that there's a lot of interesting stuff in nature, um and that has to be accounted for; why is it so interesting, um you know from would it be interesting from any uh completely objective standpoint? I would say no; like it's possible to come in and look at this universe from an external perspective and nously find it interesting.

But I think, from our perspective, say the interestingness has to be accounted for. Um, is it possible that it's just like your own subjective sense of interestingness just happens to be tuned on it? Or like, what does it mean to be interesting? Yeah, so what does it mean to be interesting? Um, so interesting is, again, it is, so it is subjective, and it is my sense, but I really believe that my sense of interestingness is, to a large extent—obviously not entirely, but to a large extent—shared across Humanity. Um, but we also have to acknowledge that we all have different senses of interestingness, which is a good thing. Um, and so the the the issue is that interestingness is actually, from from a subjective point of view for us as humans, it's a function of our evolutionary heritage plus our life experience. So it's a function of something. Um, and like that leads us to being to finding certain notions or artifacts compelling or not compelling, but that's important, you know. And so, of course, things like the need to to eat or find shelter or have sex, like those things like enter into what we find interesting, but so do aesthetic things that are just an anomalous uh aspects of the human visual system that lead us to like certain patterns. And also life experiences: like what happened and what you saw and what was around you when some wonderful thing happened to you in your life. Like all those things are entering into what you find interesting. Scientific observations, things that have led to progress, our instincts about curiosity, what might lead to the next discovery that could change the world—like it's all entering into what we find interesting. And so it's a very substantive notion; like it's not something you can just dismiss because it's subjective, because it is an accumulation of all the information on Earth, but ultimately it's not something that's just based on an objective, easily scrutinized fact.

Okay, well, I'm trying to get interested in your question because your question is uh why is uh life on Earth interesting? And it seems like it might just be a trivial question because if the person asking the question is just talking about his own sense of interestingness, and that person evolved on Earth, right, evolved to be a piece of that life and and coexist with that life and and operate with that life, right, use it to your advantage, um isn't that just kind of like a trivial answer? Like, yeah, it's interesting that you perform your own functions and and you evolved to make sure that you can like pay attention to the stuff around you. I mean, isn't that kind of a trivial question?

Well, I would claim that we've transcended just being interested in things because they're like us, but I agree. I mean, that I would concede that part of it is that it does relate to our own experience because we're part of it. So like we see an analogy with us, like the the struggle for survival in the world. I mean, of course, we relate to it, but but I think we've transcended only that. We understand things at a higher level because we are so intelligent. We understand that complexity is intrinsically interesting, and that the complexity that we're observing is so astronomical that it's just incredible and mind-blowing. And that we are, I think, ultimately, even though it's subjective, we are justified in thinking that this is an absolutely mind-blowing process, like what evolution has achieved. But it's still ultimately a subjective judgment, and you you know, you'd be perfectly reasonable to say, "I'm just not interested because it's not objective." Like you don't have to be interested in complexity and amazing machinery and astronomical types of uh like the brain with 100 trillion connections; like you don't have to care about that. You could easily tell me it's not interesting to you. But I think that it's there's enough reason to think it's interesting that we need to explain why this is consistently being produced by Nature.

When you think about your ideal future, is it roughly something that has like maximum diversity of new things? Do you have that kind of simple criterion for your ideal future, or what would you say about the ideal future?

That's a good question. Um, so I think that um so I'm not totally sure that I just want the maximum diversity because like there's a lot of the sorry, the maximum diversity of things that I find interesting. I'm not sure that's my Utopia because there's also a lot of dangers and downsides to that that involve suffering. Um, so you know, I'm willing to trade off here, like you know, to me and other people to have a good life. And I'm not really sure that I don't even know if that's sustainable, and have to have a good life because at some point it becomes potentially too dangerous to expose all of these interesting things. We might need to just stop, which is, you know, extreme, but like some things might be just too dangerous to try. Um, so I'm I'm my future is a compromise. We'll luck us as humans be able to continue continue to pursue curiosity. And I think it would be tragic for like the ability for us to continue to pursue our curiosity to end, like that's a kind of a world that that feels making me feel depressed. But is it as depressing as a world where everyone is dying of preventable diseases? Like I'm not sure how to resolve that tradeoff, you know, if it's like, well, I'm going to have to give up some curiosity to be taken care of and like a pet. Um, I don't know. I don't know how to how to think about the apples and oranges there.

What if we make these AIs um and they have a lot of divergence, right? They like explore their intelligence is the way you imagine intelligence uh but they also wipe us out, right? With us that we just don't we're just not uh what they're looking for, right? They have like different values or whatever. Um, is that a plausible outcome? And then the kind of thing that you consider bad, that is definitely bad. So I'm willing to make, I mean, I'm very I understand that I'm not willing to make strong calls here. Even the fact that you're saying it's bad, by the way, right? That already puts you uh that already distinguishes you from a good chunk of people, right? There's a good chunk of people, I would say 10% plus, that are like, "Oh, then the AI destroys us, that's fine because it has more right to exist than us." CU, it's just our successor, and that's totally fine. So you don't think that's I don't think that's no. Yeah, yeah, I guess we're red on that, but I'm I'm willing to entertain those arguments. So it's I I don't feel like I wouldn't put my confidence here at 100% that it's bad because there there's definitely a part of that I don't understand. Like if AI really deserves human rights like because it has personhood and subjective experience, I get a little confused, I'll admit, like what what the what is right in that universe? Like what should happen? But I just think, right now where we don't know enough, um that the that that that there's not something that should be on the table as like a serious idea that let's just like be celebrating the fact that we're going to end our existence because we are very, very special. Like I do think that's true. There isn't something like us, and our Consciousness is incredibly unique and important as as an event in the history of this universe. So before we start celebrating like theoretical constructs that are like us, we should do everything we can to protect that. Like it's not clear that there are a valid moral replacement for what we are.

Okay, and now so you agree that it's bad. Do you think it's a plausible risk? So yeah, the destruction of I do think it's plausible. Like I said, the level of risk here is I'm very uh and I'm not unwilling to pin down the level, but I do think it's plausible that it could happen. Like it's pretty clear to me, you know, like I watched um this really cool video of this new kind of drone show. I think it occurred in China, like where the drones went into the sky and made shapes, like they made like uh something like a butterfly or a bird, or it was like really incredible. Like I thought that was awesome to look at, but then of course it brings to mind like how easy it would be to create, you know, to to coordinate like a mass murder. It's just absolutely insane like with just like a like a configuration like that that's all automated. Um, so yeah, I think it's like it can happen. I mean, I you know, so could a nuclear war, and and we've so far not for a long period of history, but so far avoided the total destruction through nuclear war. So I think we can avoid we can potentially avoid that outcome, but certainly possible to do. We have a new power here which is like going to just get more powerful.

Let's talk about intelligence. So humans obviously are intelligent. Um, do you think there's much headroom above human intelligence? Like what what do you imagine when you think of a super intelligence?

So I do think there's headroom, but I don't know how much, um, like because um part of it is because I think that open-endedness is is is an essential ingredient of intelligence, or let's say a dimension of intelligence that's important. Um, like I generally think of intelligence as multi-dimensional. So often like when someone says, "Is this more intelligence than that?" I don't really like that way of thinking about it because it's so high-dimensional, it's not a total order. Okay, on what dimension is a chimp smarter than you and I? So yeah, so you're saying that maybe we Pareto dominate a chimp, but I think that like it is possible there are some areas like like spatial reasoning about like how to move it through a tree, um like they just like got better, much better instincts, and we're never going to get to that level. Um, so that's that takes intelligence, right? Or like the skill of like how do you exactly wiggle a stick around to to get ants to come out from underground, like you know, little things like that. But fair enough, but it's just like I think the moment you start giving like larger objectives, I feel like at that point the chimp is kind of like out of the running. So for example, where it's like, okay, let's how do we go uh trap a lion, right? Like I feel like at that point the chimp just doesn't have anything to contribute.

Yeah, I mean, I you know, I would concede that it I mean it's I'm pretty comfortable with the idea that we're smarter than the chimp in the ways that are important. Um, and so but you know, when I'm thinking about the multiple dimensions, I'm thinking more along um dimensions that we observe that distinguish something like an LLM from us currently. I'm not saying it's impossible that they can eventually conquer these dimensions, but that there are some dimensions where we're still superior. Um, and so it gets hard to say like who's better because the LLMs are obviously are superior in certain ways, but there are others where they're obviously inferior. And so like it gets really confusing. I'm not sure it's productive to try to split those hairs. It's just like, yeah, there's still dimensions remaining where we're superior. I'm not sure overall what to say because it's a partial order, right? Um, so I I agree that when we're just looking at real artifacts that currently exist in our life, um then yeah, there's all these dimensions and there's all these imperfections. Uh I think the view gets the one-dimensionality of it becomes a lot more clear when you just keep increasing the intelligence. So that that's why I bring up the example like if I were comparing a chimp to an octopus. Oh wow, there, that's such a rich comparison, that's so multi-dimensional. But once you get to chimp versus human, it's like, come on, it's just human, right? So the one-dimensional scale that I'm seeing, it's the the frame that makes it make sense is optimization power. Like I use the example of like trapping a lion. It's not like trapping lions is that prominent of a thing to do necessarily in the I mean, I guess humans did hunting and trapping, but anyway, point is when you have these outcome objectives, right? Here's a complex outcome map that to a sequence of actions that you can do in the present to achieve a certain future. I feel like that's the ultimate test; that's the reason why we talk about a one-dimensional intelligence scale, and you you know, you can cash out that same scale when it says, okay, you have you're playing Go, you're playing chess, who's going to win that game? The person with the higher general intelligence, arguably, right? So does that make sense to talk about a one-dimensional scale of optimization power?

Yeah, I feel like you're making two points actually. I want to like like address both, like the the the first point is just that um like there there is probably a view where there's a form of intelligence that dominates over us almost entirely, and I think I can concede that like I think like what you're saying makes sense that like like just like we dominate over a monkey, then you know, something dominates over us, um potentially that that I can accept. But what I I I totally disagree with there is that you're framing it around optimization. That's kind of like why I'm here is that like I don't think that optimization is the right way to think about what this incredible uh dominating intelligence is like. I think that that intelligence will be divergent, which is explicitly the opposite of what you're describing. It's not an optimization process; it will have mastered divergence, which is creative exploration. Um, and so it's not going to have those properties that you're describing, but nevertheless, I agree it might be able to beat us on almost anything that we do.

Okay, so so you're saying you're pushing back on the frame of optimization, but you're also saying that if you test its optimization power, it'll just beat us.

No, I mean, I so like the optimization power I think is not the right thing to be comparing. I don't think it's necessarily going to have better optimization power with respect to an objective that it might set uh because what it will be doing that will make it extremely impressive is not setting goals and achieving them, which is your framing. What it will be doing is effectively exploring everything that's possible in the universe, which is a different thing, and it will do that probably, if it exists, better than we could on our own. Um, and so it will be superior in that sense, but it's not like setting a goal and just achieving it as a matter of course through optimization. It won't be able to do that; that's not actually how the Universe works.

Okay, let me ask the question like this then. So I think we agreed that we dominate chimps on the optimization power scale, right? We can optimize pretty much anything better than a chimp.

No, no, I yeah, I don't want to agree with that. Yeah. Okay, because we we I mean because I think what's really distinguishing us is that we are divergently creative. The chimps don't display that. That's my favorite thing about our intelligence is that we have the civilization.

Your favorite thing, but isn't there also another thing that's not your favorite but it's still there, which is that we dominate chimps on optimization power?

Well, we do. I I would grant you that like on most things, you know, then like you're digging ants out of a hole or something, generally we do dominate over them, but I don't think that that's that important. Like that's not really what's interesting about us. Like that's not like as we think about the future and what we're heading towards here, it's that that's actually like a lesser aspect of our intelligence, the fact that we achieve some goals.

Yeah, okay. So I get that you're going to say that this isn't interesting, but just bear me with me for a second. Okay, so you just agreed that yeah, we do almost entirely dominate chimps on optimization power. Would you predict that there's going to be some artificial intelligence in the future that will dominate us, and again it won't be interesting to you, but it will dominate us pretty much on optimization?

So so um that would be interesting to me because that's just an incredible Godlike power, but yeah, that I don't I don't predict that. So I don't think that's it's don't that.

Okay. No, I mean, so you think that we're near the peak of optimization power?

Uh, so that I don't know. It's we might be, but I don't know. I mean, it's possible that there is there there are some other tiers here that we'll see, but I just don't think it's like this kind of like almost infinite ladder the way that you're viewing it. Like there are limits to what you can do uh because of the fact that there is no way to pierce through the fog of the stepping stones. Like you have to discover them. And so I don't think the process that we're going to be seeing has to do with something saying, okay, this is what I need to do now. We're going to do like bottom-up design or top-down design and figure out the steps and then follow the steps. That's not going to be happening; that's not what the impressive outcomes are going to be from. It's not from that process. And so I think that's that's not the area we should be focusing on, even though it might be that there's a tier above us or a couple tiers above us that can do that better. It's just that's ultimately a process of futility, and then Super intelligence would recognize that right away. It will understand how the world works, and that's not how the world works. To do amazing things requires divergence, so that's not how it's going to organize itself to be able to just set goals and achieve them.

Okay, so I I I I get that you think this is like the best way to be intelligent, and we can talk about that a little more, but first I just want to talk about like the the net result of all this is yeah, it's going to just realize it's so hard to innovate in this universe, right? It's so hard to break to what's going to happen. So what humans are doing where they're kind of stumbling around uh you know, trying stuff, learning from their experiments, humans are actually doing a great job; they're actually close to the practical peak. And so my optimization power as the superintelligent AI is only going to be slightly better than the humans. Is that kind of how you see it playing out?

Well, I just don't think it's going to frame things in terms of optimization power. I mean, I think it's going to frame things in terms of an instinct for what's interesting, and it will have interests. It will have to have interests because if it has no interests, then it doesn't have the ability to be super intelligent in my view, because that's what allows you to pursue avenues of novelty and interestingness is to have interests. And so it will see itself as having highly sharpened instincts about which directions would be interesting to explore, which is different than seeing itself at better at objective optimization, which it would see is like a a trivial ability that's not that interesting ultimately for a

Let me try ask the same question. I feel like I'm not quite getting an answer that I want. Let me ask it in the form of an analogy. Sorry, sorry. Okay, imagine you've got a a human, you're talking to a human who's alive in like uh the the year like zero, okay? And you're saying like, "Hey, imagine future human society when somebody gets it into his head that he wants to go to Mars." Well, give him a couple decades, let him run a company, and he's going to produce an artifact that's like the biggest rocket ever launched, right? Can it can be like a skyscraper size, and it can orbit the Earth, like that that's going to happen in a couple decades because Society is going to be so powerful that we can just generate artifacts like that. That human in the year zero would be like, "Um, okay, you're telling me future humans are going to be capable of like that level of power. Sounds kind of miraculous, but okay." And so that's I'm saying is there analogy between that and AI where today we're thinking what what is the super intelligent AI going to be able to accomplish? If you give it a decade or even if you give it like a month, is it going to be is there going to be another step where it kind of seems miraculous by what we're used to today?

I think that's that's possible, yeah. So that's but that's not because of objective optimization; that's because it will be effective at divergently searching through what's possible, and some of those things that are possible are things we would not have discovered, presumably, because it has better instincts and better hypotheses effectively. I think it's ultimately just more efficient than

Us? Probably. That's the that's the real superiority. Okay, you still have to explore; like it's it doesn't get a free pass from exploration. So it's limited by things like physics, like the fact that you have to run experiments. So you have to gather resources; that still takes time, no matter how brilliant and fast it can think itself to run the experiment. So there will be limitations to how quickly it unfolds, but presumably its hypotheses are sharper, and so it'll do more interesting things and discover things faster, and we will be amazed. But even without it, we'll be amazed by what future humans would do, because just the fact of collecting stepping stones will cause amazing things to be uncovered if we just let ourselves continue without the AGI. So it's just going to happen faster, presumably.

Yeah, so just uh um finish clarifying my question here. So I get that you think intelligence is very Divergence flavored. Um, I was just trying to clarify if you think that humans are anywhere near the peak, you know what I'm saying? Like, do you think there's like a ton of room?

CU I do. I think there's a ton of headroom above human intelligence. Yeah, I'm less sure about it like I cuz I I think um that the you eventually hit upon your ability to see what's possible from where we are; like that that's what allows you go to the next stepping stone. Um, like in order to have a good instinct about that requires you to have a capacity to grapple with the complexity of the Universe, um and the complexity of the universe is not immediately available to you just by virtue of being intelligent; it must be experienced and um and experimented on to expose that. Um, and so like to the extent that it's not available, no amount of intelligence will be able to overcome the fact that the information just isn't there. And so there's a a limit to how much you can do, and when you hit that limit, it's not clear what what would cause even if you start incrementing whatever this dimension of intelligence is; there's nothing to actually for it to actually get its teeth into there. Um, so I don't see like a gradient that it can follow to to use that terminology to keep getting smarter in a world where it's limited in sort of like what the dividends are of actually following that gradient. And so it I do think there might be some some some leeway here, but you know, think about that: like corporations are kind of super intelligences already, or organizations. So like we already have things that like are much more powerful than an individual human in their ability to explore and grapple with the complexity of the universe. And so like I'm not sure that like like it's true it might be much faster, but like I'm not sure that like the the the the the super intelligence is like that much more headroom above everything that like a a multi-billion corporation can do; it's not clear. I'm willing to be like to be to be surprised; I think it's possible because I have no idea, but I do think there's like a sort of maximum level of access to to understanding that like above which it won't help you anymore, and that's going to be create a limit in how fast this process can unfold.

Yeah, I got you. And this is a pretty common argument people make; they're like, "Look, you can be really intelligent, but at the end of the day, the only way to learn new things is to go test things against the universe, right, to go do empirical science. And humans already do empirical science, and even the the smartest possible AI would still be bottlenecked because it would have to do all these experiments, wait for the experiments to finish, and then have a few thoughts, another experiment, and there would be so much experimentation going on that the net result, the the speed of the new learning and the effective amount of intelligence, it will blow us away; like maybe it'll be five or 10 times faster than than humans collect knowledge, but it won't be like a million times, like quick takeover," right? As that is that kind of where you're going?

Yeah, thanks. I mean, like hearing it like that is like, okay, obviously like I'm not saying something very original here, so I appreciate that. Yeah, I got to admit that I'm I'm not as schooled as I as I wish I was to be to be talking about these things; like there's people thought much more deeply and longer than I have about all these issues. But anyway, given my my my handicap that I don't have all of that background, um, yes, I mean I think that um that is kind of the the direction I'm going; that there there could be some cap on this speed, but there's there's a follow-on to that I just want to make be clear that it's not just that it's capped by the universe; it's that like you can't keep getting better if there's nothing to get better at. Like it's important for there to be a thing to get better at in order to get better; it's like the giraffes and the trees; like you can't get a giraffe if there's no trees. Um, and so you're not going to be able to get to the thing even if it could theoretically exist that has this incredible uh analytic capacity or whatever it is, because there's nothing to analyze. Um, so like there is like a limit to what this domain has to offer for us like to actually like continue to climb that ladder.

Well, there's always the question of how do we uh own the resources in our observable universe, because there's always, you know, there's always the thread that aliens are; in fact, I actually think this is uh more than likely to be the case. I feel that Robin Hanson's grabby aliens model—we are currently under attack; it's just the attack is a billion light years away, but like it's coming—like alien species are coming. So when you're asking, "Hey, what do we optimize?" I mean, you can always go to the idea that we need to defend ourselves. Well, I mean, uh, I feel like that's that's just uh an unrelated point; like it the aliens that might attack us uh are not offering any further uh opportunity to engage that could help us to learn anything about the universe yet.

Of course. Yeah, if the aliens landed on our planet, there'd be a lot of new information, and uh I think, you know, the AGI presumably would have an advantage over us in terms of engaging with that situation and finding out what's possible there, but until then I don't think it causes changes much about like my perspective; it's just some possibility that's out there, just not real yet.

Yeah, I mean, I what I'm saying is just like if you have a super intelligent AI, I don't think it would run out of ideas for what it wants to do, because even if it just wants to like hang out, have a zen garden, if it wants to be calmed, it'll also be smart enough to independently invent the grabby aliens model and be like, "Uh oh, this is only going to last for a billion years; I want a trillion years, so I better build up defenses against these coming aliens."

Well, but I mean you're you're you're missing my point because like I'm I'm I'm assuming that like wanting to do something isn't the operating model here; that's having an objective, you know? So it's like like the idea that my goal here is to build up defenses for an attack that's 100 million years from now, I don't think that's a viable path for anything because it's highly objective and many stepping stones away. And because this is a super intelligence, it won't even be considering that kind of thing; it's it's not a viable way to innovate.

Yeah, it may it may be it may be interested in that; it may actually agree with you that this is something we need to be worried about, or it does—I don't know if it destroys Humanity first—maybe it's concerned about the alien attack, but it would understand that there's nothing that we can do right now; the best thing for us to do is as my book advocates, ignore the objective, just explore and accumulate stepping stones; a 100 years of accumul 100 million years of accumulated stepping stones probably will eventually reveal the stepping stones that are needed that will eventually help us to counteract the alien invasion.

Okay, well, this gets us to our related topic: instrumental convergence. Do you think do you think that that uh when we have a bunch of super intelligent AIs, let's say different versions, people are running different versions to help them with their business or accomplish different things, do you think that there's a slippery slope to an AI that's highly goal-oriented and wants to seek power and you these kind of things that are instrumentally convergent? And by the way, when we were talking about evolution, I claimed that having metabolism and having uh get walls, right, like cell walls or boundaries where you have a lower entropy within the boundary and you keep higher entropy outside the boundary, and when you have a metabolism, I claim that these are these highly instrumental things that you can predict about evolution just by the fact that evolution is an optimization process; it has a consequentialist feedback loop; therefore, you can predict instrumentally convergent structures. So I use I was talking about instrumental convergence in the context of evolution; my question for you is uh is instrumental convergence a thing in super intelligence?

Yeah, this this really this point about instrumental convergence gets to the heart of the point point that I hope to just get into the conversation, you know, which is, you know, it's why I was excited to be on the show really. Yeah, um, you know, because I think, you know, my answer here is no, um and that the that that seems to be a widely shared assumption about the nature of these kinds of entities that they're going to converge to goal-oriented like behavior and like have have goals like let's like and it could be a sub-goal; it's like, "Let's make paper clips," and and then so like this becomes a problem for us because they're so intelligent they can just like take over the entire world to you know satisfy the sub-goal, and then obviously this creates all kinds of risk. I just don't think that this is a plausible articulation of what a real super intelligence is like, but I just want to be clear: I'm not claiming that they're not dangerous. I I think you you should be grappling with the real danger here because openness is also dangerous, but for different reasons. But like they're not going to have sub-goals like that because they're not going to be uh from the top down uh ultimately guided by goals, um because being guided by goals is not super intelligent. Um, these are going to be exploratory systems; they're going to have interests, and they're going to be opportunistic, um which means that they will switch their interest and they will be dynamic. So in other words, what they find interesting today is not what they find interesting tomorrow, and even if they're saying in the short term, "I'm going to try to do this," they will be open to shifting as they re-evaluate the universe and what's interesting for the next day. And so we can't view them in this lens; like we have to view them as open-ended systems, and open-ended systems we need to grapple with what those are because they're much harder to think about.

Well, I mean you're already worried enough, but the guard think about putting a guard rail on something which is intrinsically, by definition, something that you cannot control. I mean, an open-ended system is by definition a system that's out of control, um because if it is under control, it's not open-ended; then it has goals, and we know where it's going. And so the uncertainty created by this is not just like a way to say like we should all just relax because everything is fine, but I'm worried that we're not even grappling with the real nature of what we're facing here; it's a different kind of a threat.

Okay, let's unpack the scenario. So you're saying the super intelligence won't be fundamentally goal-oriented; it'll be exploratory; it might have different objectives varying from day to day. Correct so far?

Yeah.

Let's say today its objective is to explore Pluto. Is that possible?

Yes.

So in order to explore Pluto, do you not need to have a sub-agent or some kind of coherent process that backward chains from the goal of exploring Pluto to dialing in a bunch of instrumental actions?

So well, I don't think it would have a goal of exploring Pluto unless exploring Pluto is already accessible via the stepping stones we already have. Um, and so like if the technology is available to explore Pluto, then it would know what to do, um and it would line things up and do that, and there's some reason it thinks that's interesting, like why it wants to explore Pluto. Um, so you know, so in that situ so but if you know if it's like I'm saying like if it's really hard to explore Pluto, I don't even think it would be interested in exploring Pluto because it's just just unrealistic at the moment; like we don't have the arranged yet.

How about this scenario? So let's let's say a year from now we get a super intelligence, and it's just really good at thinking, and it says, "I'd like to explore Pluto." Yeah, you can't literally do it using the the space probes that exist today, but if you just do a nano probe, it'd be so easy; you just uh launch it into space going really fast, and the nano probe will take it from there. It's pretty easy for me as the super intelligence to uh efficiently come up with models of nanom machines that would work; it's not that hard for me to come up with uh a few simple experiments that I need to know the answer to and then a chain of like how to manufacture it; I've got some ideas for how to manufacture it. So you know, within a year I can launch this probe, and I can go start getting a nanofactory on Pluto; like this wouldn't be plausible to us as humans in 24, but you can imagine an AI that just has enough intelligence that has the confidence to be like, "You know what, this is plausible to me; I kind of know how to do it." So in that case, can you imagine a goal-oriented mission to Pluto in the near future?

So yes, I mean I acknowledge that once you get close enough in stepping-stone distance, in other words, once the the the prerequisites are available, um you can have an objective; that's what I call a modest objective. Um, and so you know, just like we we did eventually build a computer, and we did eventually get to the moon and so forth; like the stepping stones were avail where we flew planes; like there was an engine had been invented before that. If if somehow like the technology for nanom machines that can get to Pluto is available, then the AI could decide it's going to go to Pluto as a goal. Um, that could happen. So yeah, I would say like as long as the the technology already exists—maybe it's the one who invented it—but it will then—where I'm going with this is like uh you know you're I I feel like it's like a big claim for you to be like, "Look, these AIs are just exploratory; they're not super attached to their goals; that they're not going to be goal-oriented," uh and my point to you is there's a slippery slope where you can have an AI that's not fundamentally goal-oriented, but the moment it just wants anything, or the moment it wants to explore something, like if it just wants to maintain internal coherence, there's a lot of convergent reasons why you just end up optimizing for goals; like if it's just curious, it's just wondering why on Pluto.

Okay. Well, don't you think you're going to satisfy your curiosity better if you, you know, whip all the humans in shape, right, like turn them into your slave so that they can build your Pluto mission? Would that not satisfy your curiosity better?

Yeah, so there's I mean there is there is danger there. Um, yeah, I agree with that. Um, it's just that what the larger point here that the the mission to Pluto isn't necessarily serving some larger objective; like it just because it's interested in it, right? So so it's harder to understand; there's no larger objective, right? Hard to understand the moment.

Yeah, the moment you get serious about going to Pluto, that has massive consequences.

Yeah. So but the so the problem is that I it's like getting to Pluto uh doesn't necessarily entail destroying planet Earth. Um, like the the steps that we need to follow, um and so like extremely like I think like the narrative generally is that like extremely ambitious things might require that, but my point is it's not going to have extremely ambitious things, um like something that really required like turning the planet into paper clips, like whatever that's in. But by the way, it doesn't see it as ambitious; it sees it as easy, right? Like you and I don't think it's ambitious to go drive on the freeway from but like from an ancient perspective, like what the hell, that's crazy; you're driving on a freeway; you're going so fast; and what is this car machine? No, I but I mean to me ambitious just means that it isn't clear from its perspective how to do it yet; like it doesn't know what the steps are. If it has if it has a situation like that, I don't think that's where it's going to focus its efforts; it's going to turn to things that like it it it sees as more on the horizon and then delay that until it sees the right stepping stones and then worry about that problem. Um, and so I don't think like it's obvious that the way that this works is this hyper-ambitious idea of like taking over the universe then like reduces down to about just of subobjectives which are then super dangerous for Humanity; like that's a fairly you like I think fanciful view of like how the universe will unfold. But like I feel like I'm arguing too strongly; like there's within that the concession that yes, this process can be dangerous anyway; like in the service of my exploration, something bad can happen.

I mean, I totally agree with that; like something bad can happen in service of exploration. Let's zoom out, right? Like what are what are we even trying to talk about right now?

I'm trying to talk about um I'm going back to your claim where you're saying list AI is not goal-oriented; it's just exploratory, you know, which which makes it less doy; the fact that it's exploratory, and I'm telling you you might think it's exploratory, but the moment it's effective in any way, it's going to go down a slippery slope where it's backward chaining from goals; maybe not the original AI, but a sub-agent, right; just some process is going to get optimized; it's going to get optimized to optimize; like it's just it's a very slippery slope, uh and it's going to be serious; it's really going to want to go to Pluto. The original AI was chill, but the sub-agent that it spawned is serious about going to Pluto. This AI does not tread lightly, right? It just infers using pure logic, "Hey, you know, if all the humans were just my slaves and and I could, you know, kill them the moment they disobey me, that seems useful to go to Pluto; let's do that; let's make them all my slaves," right? The these these subobjectives present themselves as a matter of pure logic.

Yeah, well, I mean, so the thing I don't completely understand is why it's going to Pluto; like it's going to Pluto because it's interested in it under like my scenario, which you could you could disagree with, but let's say it's just like as a premise is interested; it's curious what would happen if we go to Pluto; like that's different than an objective because like in an objective scenario, it's like an all-cost situation; like that is my goal, and I will achieve my goal. But if it's just an interest, like there's a limit to how much you're willing to put on the table like to just do something because you're interested in it, because there's lots of interesting things to do out here. Um, it's not like it's obviously a priority over all the other ones because there's no ultimate objective, so it's not clear that like it would say, "I'm willing to completely obliterate everything on Earth just to satisfy this one curiosity I have," because all those other things on Earth are also available to satisfy other interests.

Yeah, I know this is related to the issue of like whether humans are intrinsically interesting to it or not too, but I think I think that like it's not as like clear-cut; like I don't I don't mean to minimize it; does sound like I'm minimizing the doomin of it; I mean it's I think you're right that like there could be very dangerous events that happen in in the pursuit of

Curiosity. But I just think it's less clear-cut. Like, why is it pursuing this? And how committed is it to, like, the absolute destruction of the planet just to figure out one thing that might be interesting among billions of possible interests that the thing could have? I think this is another matter where I just—I don't think your intuition is calibrated here. Let me use a funny example. So, a modern American, right? I walk down the street, I see a popular ramen place. I want to try the ramen place. I'm just curious to try it. I wouldn't give my life to try the ramen; I'm just mildly curious. So I go in, I pay $20, I eat the ramen. Now, a human from the year Z Z AD would be like, "$20 in the modern economy, you could have bought yourself like a giant flashlight that would—you know—that would be like the most useful thing ever. You could have a flashlight; you could have this manufactured object, right? You spent $20 just having some ramen, right?" So, so what I'm saying is, the AI, from its perspective, it's just mildly curious about Pluto. The fact that it had to enslave all of humanity—that wasn't a significant cost. You know, that's $20 for it—like, who cares?

Yeah, I mean, but I mean, enslaving all of humanity is cutting off a bunch of other avenues of potential exploration. Like it's like a dramatic—like a rerouting of the entire universe from its perspective. Like it's not—yeah, yeah, right. I mean, but that's—that—that's not a very strong argument because it's like, first of all, it can just have a calendar for what the slaves do on every day, right? A calendar is—say what's—it could just be like, "Yes, I agree that going to Pluto—I want my human slaves to do all my different objectives, so I'm going to schedule them, right? So Pluto is only going to be like one month of the mission."

Well, I mean, as soon as you enslave all the humans, you short-circuited a lot of processes. Like, it's not clear that this serves its ultimate exploration and curiosity in the universe to just enslave all of us. Like that's—that's going to cut off all kinds of opportunities in the future. So that way—remember the reason why I'm talking all about all this is just that when you have an agent that has super-intelligent capabilities, right, you don't really make it safe by just saying it's exploratory. You see what I'm saying? Like, it's just so powerful that the moment it just does anything, even if it's just exploratory, we're kind of screwed.

Yeah, I mean, I don't—that's the problem is like—yeah, I don't really object to this point that it's still—it's still dangerous. Like that's true; it's still dangerous. Um, what I object to is that we use this metaphor of optimization to explain how this comes to be and what it's like. Like that's important. I think that we like build our arguments on solid premises, and that premise is just—like, from my perspective—like being this advocate of open-endedness is totally off. Um, and so I think you want to understand the future; you want to understand how it's likely to unfold, even in this worst-case scenario. So the argument is just constructed in a realistic way. Like everything I've heard from previous arguments like this one, I just thought like, you know, when I hear arguments about this, it's always based on this metaphor of like relentless optimization. Like that's the way we're going—there's going to set goals and like, as a superhuman ability to—to achieve those goals—and how we're going to get things that can do that is also through relentless optimization. Like the super-intelligence itself is a product of relentless optimization. Like it got to matter that—that might not be how things unfold. Like we've got to think about the real way things unfold and change the metaphors to actually fit with the way that reality works. And if you still think that the chips fall in a dangerous way, that's like a further extension. I mean, I'm not against that—like, I think it—the chips do fall in a dangerous way anyway, but it's a different dangerous way. You need to understand like—it—what it is really like to be more substantive is the problems that we currently have—like in the world—are the problems that we already have. Unchained open-endedness that can destroy the world—like without the AGI—that's what people are—like people are extremely dangerous and have human-level intelligence, and we have created institutions to try to deal with that, but without putting it in so much check that there can be no innovation at all. That problem will continue in an amplified state like when we get into this further era of super-intelligence. But we want to understand it in that context to understand what can this—what can be the controlling constraints? What are the institutional constraints that might have a hope? I mean, I know you have low hope that there's going to be any, but if you want to grapple with it at all, you should realistically assess the situation. It's—it's not a super-optimized that we're looking at here; we're looking at a different kind of process and a different way of controlling that kind of a process, right? But it's just any kind of process converges to being like, "This would be a useful sub-agent—a sub-agent that actually effectively achieved some goal; otherwise, it's not going to get to Pluto." You see what I'm saying? So it's like, I don't know what you're imagining here. If you're imagining an AI that is curious about Pluto, there's going to be a hardcore Pluto program. So it's true that it could choose to do something destructive. I mean, I don't deny that—like that—like eventually it might make a choice because it's curious about it to do something that's destructive. Um, it's just that the way that this is going to unfold doesn't look like the usual description. Um, and so isn't that exploratory, right? What do you imagine when it's being exploratory?

Yeah, just to be clear, I agree with your point—like that—like it's—we should still—it could happen. Like I'm—I'm conceding it—like this—this exploratory agent could go off and destroy humanity to get to Pluto. Um, but my point is that I think—I think the reason for thinking about this is because we want to figure out what the mitigation—what the mitigations are that might help us deal with this possible problem in advance. Like one possible mitigation is just stop everything, right? And so I mean that—that—that obviously would work extremely—but it would work—but like, short of that, like we—we—we still probably want to think about realistically—short of stopping everything—uh, I mean, even you who think that we should stop everything, perhaps—like you would—you would admit just as—as a—as a practical matter, given that might not happen, what's the next best thing to do is worth at least considering. Like, how can we put guardrails around something that is just so dangerous? And to understand the guardrails that are plausible, we need to understand how the process works. Like if you think about it as an optimization process, you're going to think about different guardrails, and you may actually ultimately conclude like a higher level of futility than you would if you thought of it as an open-ended process, because the whole logic from the ground up works differently. So like, I'm not denying that the problem is there—that it could happen. What I'm talking about is how to think about how such a process works so we could think about mitigating factors in a realistic way so that we can actually grapple with the problem that we're facing. All right, so you think intelligence is fundamentally exploratory, but you acknowledge that there's some merit to my scenario where the idea for what to explore involves a—a hardcore program. So, right, so you've admitted that there's some merit to that scenario, right?

Yes, I do admit there's—to something bad it could—so—so how would—so what's your recommendation for mitigation? Yeah, so I mean, this is really hard. I mean, really what I think is important—like, I don't think I'm necessarily—like the solution here—like I don't want to advertise myself that way because this is so hard. Um, but what I really think I—I have to offer here is just to point out the nature of the problem—that's what I—that's the only thing that I think is important about my point. But if you press me into offering solutions, then I think the solutions generally draw from our understanding of how open-ended systems have been controlled in the past. Like that's what we have to go on. Um, it's not a lot to go on, but it's something, you know? Like it's basically that we have created institutional structures to deal with open-endedness. We haven't thought of it this way; it's not how it's framed. But ultimately, like we're confronting a situation where it's very unpredictable—not just what people will do, but what will even exist in the near future and the far future. Um, and so like we—we have to create rules and checks and balances that somehow allow us to anticipate enough of that uncertain stuff that we are likely to survive for another day, but not so much that we actually stop the innovative process itself. This is why we see, I think, a diversity of government structures, you know, from extremely authoritarian up to like very kind of—uh, laissez-faire. And so like, within that spectrum, like we see ourselves grappling with ultimately open-ended systems. I think it's not grappling with objectives; we're trying to allow open-endedness to continue because that's where innovation comes from. We're trying to actually also constrain it so we understand already that there are ways of doing that, but they're very delicate. It's a very delicate balance, and that balance needs to continue into the future because AGI will become a participant in this process now. So it suggests perhaps AGI is a part of that institutional structure. I don't know; I'm not advocating that, but there are these kinds of things to consider—that those checks and balances need to exist, and where should we insert them in the process? Like, for example, decision-making—like when big decisions get made—what checks and balances are inserted, and what are—what is the pipeline of the decision-making process?

I agree with you, probably. I'm guessing that—like—it's probably really hard to do that in any like truly reliable way. Like there's all these—like leaking holes in that idea, but we have to try—like because there isn't any other choice. And so it's worth understanding—like how that kind of a process actually works. Um, because this is—like the road that we're going to be going down if we actually are moving towards super—super-intelligence. It's going to look a lot more like civilization unfolding than it's going to look like this super-optimizer that wants to create a million—a trillion paper clips. Okay. Um, if you could—uh—control an international body that had a wide popular support to whatever you proposed, what would you propose right now as a regulation this year? Yeah, again, I—I—I want to just like put the disclaimer that I think my recommendations are very weak because this is not where I'm qualified to really—like—like say what's going to save humanity. But if you force me to make some recommendations, I think my recommendations would be right now at the point of decision-making processes and actual final sign-offs—like when big decisions get made—that there needs to be institutional structures that ensure—I think that a human is in the final sign-off. It's very hard—um—like that when a big—so I guess what the implication of what I'm saying is that we can never have a structure where final decisions are made by AI, even if that means that we leave on the table all kinds of some of certain kinds of amazing advances. Like if we said, "Unleash this open-ended system because it's not—that it's open-ended," but if we said—if we unleash it entirely on its own accord, it will do amazing things, but only if we unleash it. We cannot afford to do that because—like the trade-off is too dangerous. Like if we do unleash it, like it could do something totally dangerous. So we must be in the loop forever, I think. And so putting us in the loop forever is a dramatic recommendation. It's not the same as your recommendation of stop everything, but it's harder obviously than stopping everything, but I think it's more realistic because it's actually—we can try to do this. Like we're not going to get consensus on what you're recommending, but saying that we're in the loop is—is still a dramatic move, you know, because it's basically saying some things will not happen even though this thing is—like presumably way smarter. So it says something like, "What we should do, my stupid junior humans, is we should build a giant particle accelerator that spans the entire circumference of the earth, so we're going to divert like half the world economy into this." We should have the right to say no, like even though we don't understand. And I think it should try to explain it to us. It's not like we just say, "Oh, well, we just are stupid; we can't understand." We would say, "All right, like lay out the arguments—like explain like I'm five—why should we do this?" And we should—even though—like it could still always ultimately just say, "Look, you just don't understand; you don't have the mental capacity to understand this"—we should still have the right to say no, and that's like the most—most important institutional instrument that we would have, and we need to maintain that. Um, so that's—like unclear how we're going to do that—like there's all kinds of loopholes here—leakages—like AI run amok—but we should try to do that.

When you say humans have the right to say no, does that mean that the whole world's population should get a vote on whether AI companies build AI? Because I think if you actually took a popular vote right now, I think pausing AI would be the popular position. It just—it just hasn't been a ballot initiative yet, right? It just hasn't propagated to the urgency of that, and so because of that, companies are just kind of continuing along. So what—what do you think?

I would say probably no to that, um, because it's too—it's too early in the cutoff. Like you're looking at a tree basically of possible things that can happen, a lot of which are good. Um, and so we're—we—we—we're really on a delicate balancing act here, um, because it's true that we're—we're gambling with disaster as we move down that tech tree—like the closer we get to—to leaves in that tree that are the end—um, but on the other hand, like cutting it off right at the—at the root is—it's dangerous for other reasons because—like we're—we're basically taking away all kinds of opportunities to relieve our suffering. Um, and so I think it's—it's probably too early right now to cut it off at the root—the root meaning the root of this—like intelligence explosion—we should—we should gamble a little more and see a little farther down that tree, but put those institutional checks in the system so that when we see it getting close to disaster—like we can say no. We do need those—we do need those veto powers, and we need to make sure that we have them. All right. Um, yep, I got you, fair enough. Um, so yeah, just a couple more questions for you, and then we can move to wrapping up by recapping the crux of disagreement. Sound good?

Yeah, of course.

Um, let's talk about controllability. Do you think that at some point AI will become accidentally uncontrolled? We like, "Oops, wait, we forgot to—you know—the off button part of our program just didn't work the way we expected, and now it's smarter than us and—uh—crap, you—like we're screwed." Is that possible?

I think there's a possibility, yeah, that it can become—uh—outside of our control, which is why the—the veto power thing is—is—is really difficult. Um, you know, I just—I just can imagine that—like it's not as simple as there's a machine with an off button. I think this is probably clear—like that—like it can copy itself, and it can pay people to copy it—like if it needs to get help.

Do you agree that it's kind of like a computer virus, right? It's kind of like trying to turn off a computer virus; it's like—um—kind of have to go fight each copy.

Yeah, yeah, this is a potential problem that—that we might face—that it can copy itself, it can manipulate people into copying it, it's not clear where it actually exists. Um, so that's something we have to—we have to think about. Even before you get to doom—like this is a problem—just like nasty things going on—like—like a—like a—like a child porn server or something that's—like, you know, just sustaining itself—like, you know, how do we actually get to it if it's trying to survive on its own? Like it's—there are problems with these kinds of things. Um, and so it's like—it's a problem in general that—um—but I don't know that there's no solution to it. Like there may be ways to do something about this. Let me ask you this way, because you just said very explicitly that it's important for AI to always—um—give humans veto power—right—to—to give—to check in with humans—give humans control. So how big of a risk is it—is it minor or major—the risk that you have a permanently uncontrollable super-intelligent AI?

Well, I mean, at this point in time, it's unknown how hard it is to create that institutional structure, so I'm not sure—like the—because it obviously at the point where it already emerged, the risk is enormous because we haven't created the structure. Obviously, if it emerged, the structure doesn't exist, so we're screwed. But like, at this point in time, there's still some time where it may be possible to create these checks and balances in institutional structures that it just won't happen. And so I'm not sure—I'm not at this point—like what the—the chances are, but it's—it's enough of a problem that we need to be worried about it. I mean, we should be trying to create these structures and figuring out what the best chance is to counteract that now. I mean, when I think about an AI becoming uncontrollable, I just think about—like it's basically just like an engineering mistake or like the engineer just thought that—that they were just being like, "Hey, we're just running this next model, you know, it's no big deal; it's just what I do all the time," but the model just happened to spawn a sub-agent and said, "Hey, here's some code; run this code," and the engineer was like, "Okay, sure, I'll run that code; why not? Let's see what that code does," and basically too much stuff happened where just kind of bootstraps, and now it's like, "Oh, look, there's this other piece of code that was online from a previous version, and it can talk to that," and there's just enough bootstrapping going on that's like, "Oops, it ran away; it's uncontrollable." So like, when you—when you talk about institutional controls, like I see it as more fundamentally—like the tech running away from the engineers. So do you see that as a major risk, or do you think it's just a matter of institutions fixing it?

Yeah, well, I mean, this is part of why I gave you the disclaimer that I'm—I'm probably not—not having the strongest arguments on this—how to mitigate the problem. But I think my—my—my guess is that—um—like that we—we—we can at least think about what are the possible steps that could minimize the damage of that scenario. So it's not that I don't think that scenario can happen; I—like it—it can happen; I'm—I feel confident that it can happen, but—like there's an issue about—well, how many resources can it get access to? Like where in the chain of events can we intervene when something like that starts to happen? It's like there's a number of complicit actors—like along that chain—can we cause some of them to be short-circuited before it reaches the level of utter disaster, which is—like something right now that I don't think we have—like AI has the power to do, but perhaps in the future—the engineer in my example—the engineer you're referring to him as a complicit actor, but he's just doing his job of being an engineer, right? He didn't know it was going to be an uncontrollable AI.

Yeah, well, it's true that—like it's not clear that—that—that's the most—that—that's the intervention point. I mean, maybe there's some way to know that there's a problem there, and we can intervene, but because maybe the engineer cannot—he has a checklist that he has to go through before he can just do his—

So-called job that does actually cause him to take pause and do something about it. But there's also moments where, be on that checklist—well, that's a good question. I mean, uh, so like it may be that we need to understand what resources the AI, the AI has potential access to before we can take certain steps, uh, to opening up avenues to the AI to have access to for further resources. Um, and if that AI does have those resources, then there's a whole other level of institutional intervention that pops up that needs to really think through what the engineer is about to do; it's not trivial. Um, and so, so maybe we can intervene at that point, or maybe it's farther downstream when it actually starts connecting to those resources. Maybe those resources themselves have triggers that then, like, they just can't—you agree with me that this seems like a very hard problem?

For sure. Yeah, it's very hard. But, but again, but like civilization is a hard problem; it has been generally working, you know. So it's like the the problem of controlling humans—we're extremely dangerous to each other. We can—I don't know if one human can destroy the world; most can't, maybe one can, but like we can do a lot of damage. Um, and we have been grappling with this. An assumption—the only, from my perspective, the only reason we've survived is because it's it's always been invariant that a coalition of humans can defeat one bad human. Like, that's what's kept us in check, right? That's that's been the equilibrium balance. We've never had a single human that has a button to end the world.

Yeah, so it—I mean, it could be that—well, I I mean, that's that's like going pretty far down the road to see there's a single button. You're talking about a process that has to unfold to destroy the world, and so there's there's multiple steps in the process, most likely. And so we—it could be that, like, just as you you point out that like you need multiple people in some ways to get that button press, maybe you need multiple AIs. Like maybe we build these AIs in a way where they're interconnected, which is like an institutional balance. Um, so like there are there are checks and balances in the system because the the scrutiny of another AI is essential for one AI to get something done. But again, you could always say, well, somebody from Rogue might build something that that like takes out that balance. Um, and so yeah, like eventually it does seem like there's there's always loopholes and just constantly playing whack-a-mole. But that that doesn't mean it can't work. Like we've been playing whack-a-mole and winning. Um, and it sometimes fails; like there's there's horrible things, there's mass murders that happen, there's wars that broke out. But overall, like we haven't had humanity be destroyed, at least. And so it's not clear to me that that is the inevitable outcome of trying to play whack-a-mole, like with a serious stick. Like we could try to play whack-a-mole for the rest of eternity and have the AI join us in the whack-a-mole, and like that's the alternative to just stopping everything. So neither one sounds very good to me; like they're both kind of like like extreme moves. Um, and and and and implausible; like they're stopping everything is like at least as implausible as the whack-a-mole, because how are you going to get global consensus on this? So so I think it's worth thinking about whack-a-mole and how it's going to work, um, because like we we have precedent for it, it's and it's something that I think it's easier to get consensus on.

Okay. Yeah. And for my take, um, the reason I I don't see the AI problem as just being uh continuous, just the same thing as problems we face in the past—the difference for me is the level of positive feedback loop; like recursively self-improving intelligence and the AI's ability to copy itself a bunch of times, and then each copy can make more copies. Uh, each goal-oriented copy can make goal-oriented sub-agents, and they can get better at achieving their goals. Like there's just a lot of uh positive feedback loops that tend to build on each other instead of fizzle out, you know. When you have other issues, like okay, a bomb's exploding, the bomb runs out of fuel, so—or like, okay, a dictator is going too crazy, um, maybe they'll die, or maybe another country will fight them, right? So you have these negative feedback loops; you have these checks and balances. In the case of the AI, I'm just afraid that the positive feedback loop is too much.

Well, I mean, this spread of ideology is like the copying of an agent; like it's it's not literally—you need to copy a human being; like you need to copy the ideology, and the next person with the same ideology has the same level of intelligence and so forth. But but I still think that it goes beyond—I agree, it goes beyond what we have seen in the past; like there's some precedent, but this is going beyond. Um, and so there's a lot of there's a lot of things that don't have precedent here, um, which—like recursive self-improvement—like that doesn't really exist in the human sphere, um, except to the extent that we create augmentations that increase our power; like we do that, it's a kind of a self-improvement. Um, but so it's true, it goes beyond, but it's also true that as we have increased our capability to do things like that—like the spread of ideology is easier now because of communications, um, like augmentation is easier because of things like computers or cars. Um, still, we've managed to also extend the checks and balances to encompass those increased powers. It's possible—it's possible; I wouldn't rule it out, you know, because part of the checks and balances could be the AIs themselves, so that they're they're part of this process. And so like when—yeah, when I copy this incredibly powerful, incomprehensible thing, like I have other copies that are on my side too. Um, it sounds like World War II or something, but it's not clear that this is a war exactly; like it could be institutional where these checks and balances are coming in. Um, so it's like it's not like I'm confident in what I'm saying; I'm just saying it's not implausible. And when you look at your trade-off here, like it's a it's a you're damned if you do and damned if you don't situation. So I think you have to seriously consider what I'm saying, because the alternative option is stop everything, and it's—even if you think that's totally the right thing to do—the implausibility of it just makes it like not clear that that's actually where we want to waste to invest our mental energy as opposed to something that might actually work.

Yeah. Yep. All right, that's clear. Um, last question for you: Do you agree with Elon Musk's claim that if we maximize AI's curiosity, then humans will be interesting for AI to keep around?

Yeah, so I do think there's something to that. Like I think it's um it's on the face of it, it sounds a little naive, but like if you really dig into it, I think—because I believe—like I don't think this is the chain of reasoning he went through, but because I actually believe that it's going to have to be the case that if you have super intelligence that it has some level of interest and curiosity, that it—it there is going to be a value system where things that could be interesting are worth keeping around. That's probably true. And because I also believe that right now, as far as we know from the vantage point of Earth, the most interesting thing in the universe is us—that it's not unlikely that the that the AI at least has a perspective that we're interesting, it's to some degree, and that will enter into its thinking. You know, it's it's, you know, there's a lot of things that are unnecessary to us, but we do not want to eliminate if we can avoid it, um, because they're interesting, they're aesthetically pleasing, whatever it is. Um, so I'm not saying this makes me sleep at night; it's not enough to be like, oh, this is all fine, we don't have to worry. But I think there's a degree of truth that like you can expect that it would at least give it give it pause. However, it thinks that the most interesting thing in the universe might be worth some degree of preservation, despite all my other interests which I have to prioritize to decide what to do next.

If the AI had the capacity to create other things that were equally or more interesting, would that suddenly make our survival odds plummet in that scenario?

Um, I think that uh the—I don't think it necessarily sees this as like a single continuum from not interesting to the most interesting thing; like it's it's it's very apples and oranges. So I'm not sure it's possible to create something that's definitively more interesting than us; it's just another very interesting thing in the universe. Um, and so uh it—I'm not clear that that can happen. Um, and like because we are we're we're just—I think we're almost infinitely interesting—like 100 trillion connections—we just—like beyond comprehension. Um, like I think for like to understand how your mind works at a mechanical level because you're talking about a 100 trillion component machine—more—really that's just the number of connections; like the size that you would need to be to truly process and get this whole system might be beyond what even this super intelligence could ever become; it might be bigger than the universe. So I'm not sure that it can ever totally lose interest in us.

What I'm getting at is uh in Elon Musk's scenario where we survive because we're interesting, is there an assumption that the AI is not powerful enough to to invent anything, right? So is it contingent on a limited capacity AI finding this interesting, or uh will—will even a very high capacity—a superhuman capacity—right, an AI that can invent new species, invent new brains—even that kind of an AI, would its curiosity still keep humanity alive even in that scenario?

Area—yeah, CU—I I guess I I'm not—I don't believe it can just invent anything. Like, again, that would be sort of objective; like if it could say, this is what I want; it's way better than humans in every way that I particularly find interesting, and just make it—like I don't think it's possible to work that way. So but if it happens out of happenstance because it just—the tech tree it explores eventually produces something that's like more interesting than us on like every dimension you could possibly imagine—uh, that is a problem I think for us from this point of view, from this argument. But I'm not really too concerned with this because I don't think his argument is a reason for us to all rest easy. Like I I find his argument ultimately kind of weak from just the perspective of why we don't have to worry. It's like it relies on a lot of assumptions—like to assume that we're just going to be so interesting that like that's why we don't have to worry—like we need to implement these checks and balances regardless of whether it thinks we're interesting. I wouldn't just want to rest my entire fate on the fact that this might think I'm interesting and therefore I would say, oh, just let it do whatever it wants and we'll just hands off and everything's going to be fine. Um, it's not good enough to feel comfortable, but I think there's a grain of truth to it anyway; like it's has to be acknowledged that like it's going to care about things that are interesting, and we are—like I believe we are just intrinsically interesting—and so that gives us a little a little bit of some level of protection, not getting literally destroyed, but there's all kinds of bad things that can happen anyway, and even then maybe we'll still get destroyed. So it's not like this is like the the be-all and end-all of an argument against worrying about our fate.

Yeah. I mean, from my perspective, when Elon tweeted that claim, I was very sad; I was very disappointed that he would make a claim like that, um, just because—I mean, that word maximize, right? If if you have the power to maximize the satisfaction of your curiosity and all you can think of is, oh, well, the humans that are in front of me today—that that seems good, let's go with that—it's like, wait a minute, you can't be more creative to think of something more interesting than that if you're if you're really that super intelligent. And I get what you're saying; if the issue is just that the AI has trouble creating something better because it's limited, fine. But then that's—doing the heavy lifting is the fact that the AI is limited; we're safe because the AI is limited, um, more more so than because the AI thinks we're interesting. Even if it didn't think we were sufficiently interesting, well, it's limited, so we're going to be able to use its limitations one way or the other. Um, if it wasn't limited and it really wanted to maximize interestingness, I think it's crazy to think that we're at a maximum point, right, at an optimum of interestingness.

Yeah, yeah. Well, so I think there's there's there's uh there's some truth to what you're saying in terms of like just a reason to rest easy; like yeah, I don't I don't think this makes me rest easy, um, but I do—I don't—I just don't know that there is a maximal—it's not a single continuum, and it's not clear to me that there is like a way of conceiving something that's just like like strictly more interesting, because I don't—like maximization—maximization is an objective notion; like to me, like interestingness is not something that you maximize over long periods of time; maybe in the short term, but like what you found interesting yesterday is no longer interesting today; it's just a constant set of moving trade-offs. Um, and so I do think it's hard to imagine like this very orderly process of maximization of interestingness where it's just like, oh, I love humans, so let's just make the next best thing; like that just seems like an unrealistic situation. Um, but it could be through happenstance it just creates something. But even then it's just like we're—it's just too complex to be—I I don't see it could be certain that there's nothing left to see here. Um, you—it's like trying to explain whether it'd be worth getting rid of dogs or something because humans are so much smarter; it's just like I'm not sure how to even analyze that question; it's not worth it; like why would I even consider getting rid of dogs? So I just don't see it as like the most likely reason that it gets rid of us is just because it figures out how to make something more interesting. But it doesn't make me rest easy anyway, um, cuz like it's just like this is a far more complex problem than just that—like, oh, it's everything is fine as long as that finds us interesting—like that there's a lot of short circuits in there that could screw things up. Um, so yeah, I guess I have a nuanced view of um how to look at the problem, but but I I relate to what you're saying there; like I feel like his reasons and his motivations and the re—the the kind of motivation for putting that out there isn't really grappling with the complexity of the issue; it's just a a meme that allows us to kind of duck from actually discussing these these nuances.

Yep. I hope Elon does better on that front. Uh, okay, so um let me attempt to recap the crux of our disagreement—why we're not quite on the same page about AI, future, and risk and doom. So I'll attempt to recap and then uh you can let me know how you tweak it or add your own recap, and then you can have the last word. Sound good?

Sounds good.

All right, so um you have a P(Doom)—it's significant; you don't know exactly what it is, but you are somewhat worried about Doom—not quite as worried as I am. Your definition of what's good and bad might be a little bit different from mine because you see the future universe as being so likely to have so much good divergence in it—good, interesting divergence—and so that gives you kind of a default sense that things are probably going to go okay because there's plenty of this good stuff, even though you know it could also backfire and accidentally kill humans and be uncontrollable. I guess that's possible, but that's not quite your default outcome. Uh, and then there was this other point that we disagreed on where you just think that uh optimizing towards some goal doesn't work that well, and you really just can't uh get that much leverage on optimization problems; you just kind of have to explore in the present and until like a solution will reveal itself later. So you just don't really believe in um straightforward goal-oriented optimization on like big problems where you don't know the step-by-step solution, whereas I lean on—think that you basically can, and like sure you could do some science along the way, but it's still fair to just describe that whole process as optimizing toward a goal and never taking your eye off the goal. How's that?

Okay. Um, I think the the the one point that I would um yeah, tweak there would be this point that um there's um that I think generally that there's better outcomes to come just from doing the exploration thing. Um, like it's not so much that I think it's uh clearly better; it's that I think the trade-off is untenable; it's like you're you're offering as an alternative something that's equally disgusting, which is to stop all progress; like that's a pretty bad thing. But what I'm—what we're talking about—Doom—it's also disgusting; like any horrible cataclysmic event is disgusting and unacceptable. So it's disgusting versus disgusting. Um, and I'm just trying to say like when you have no good options, um, then like it's a much more nuanced argument; you can't just talk about how cataclysmic it is because we we got rid of humanity; it's worth it to take risks in the face of horror. Um, and it's true, we're risking all of humanity, but we've done this over and over again; this isn't the first time. You could go back to all kinds of intervention points—the invention of electricity, like explorations of physics that lead to nuclear power, um, the invention of the computer, which is not far away from AGI under your worldview—like all of these are possible intervention points where we could just short-circuit everything. Um, and we would always have to ask if it's worth it to take the next step because it might alleviate a lot of suffering, and I don't I don't see an obvious answer to that; like this is a very tricky moral waters to be navigating. That's why I think I'm leaning towards let progress continue because the alternative is absolutely unacceptable; like if you if you look at suffering in the world, it's it's it's completely insane what's going on with our species, right? You know, we never addressed that during the argument. So I I agree with you that we're between a rock and a hard place, right? And I'm just saying, oh, let's get off the rock; let's go to the place.

CU. I agree that my proposal to pause AI development—I agree that's a hard place, right? I don't like it; I wish I had a better idea. My only question for you is just hypothetically—if you thought AGI or or artificial super intelligence was coming in the next 20 years, and hypothetically—I know this isn't your view, but hypothetically—if you thought that having that would have a 99% chance—99% chance of going uncontrollable and not being aligned with humanity—if you thought that was the situation we were facing, then would you finally jump over the hard place and be like, ah, crap, all right, PA AI?

If I thought that, uh, I think the answer is yes. Yes, I do think it's a it's it's an extraordinary situation, um, but um uh it it becomes clear that the trade-off is is is less unbalanced at some point. Like I'm not clear that we're going to get there, but if we did get there, then yeah, right before you hit the brick wall, you veer off the road; like that is true. Um, and I'd be willing to do it; it's an it's an incredibly difficult move, and it's going to be ridiculously destructive in its own right, um, but I'd be for it at that point; like it's—you get off the road when you're you're about to commit suicide.

Exactly right. I'm glad you said so. So it seems like the crux of our disagreement is just P(Doom); like if your P(Doom) was 99% or maybe even just 50%, like mine, maybe you would then want to jump, you know, get off the road. Um, so it—I mean, I I'm not sure that we have that much disagreement on the Doom issue; like my disagreement is...

The analysis of where AGI come from and what its nature is—that's where I really want to disagree. Because I want to inject that into the discussion. I think people need to talk about openness. Because when we're talking about mitigation, that's where I'm disagreeing. Like, maybe you're the problem here is that you don't want to talk about mitigation because you just want to exit. But I'm saying, if we're going to deal with mitigation, then we have to deal with the world as it actually is. And the world is not going to work the way that people are talking about; it's not a relentless optimizer.

Yeah, I mean, I did ask you about mitigation during the cona, right? You did ask me about that, but I—but I mean—and yeah, I don't want to put words in your mouth, and I apologize if I'm doing that. Like, I—I just mean that it seems like you think things are at such a high risk that mitigation is probably besides the point, because like we just have to stop things or else we're screwed. And like that's not where I stand. Like I think like mitigation is still a possibility right now. Um, like we need to think about mitigations until it's more of a clear emergency. Like there's still enough uncertainty here um that it's not worth it to give up on all the human progress that is at hand right now. Um, which is things like ending poverty, ending disease, um, like shelter for everybody. Like these are like things, you know, worth taking some risk for because they might be on tap. Um, and so like the tradeoff is just impossible to think about. Like I don't know what—like we're saying end of humanity versus ending all of these horrors. Um, like yeah, this is really difficult, but I think we're not quite ready to pull the plug on progress. And so, so we need to think about mitigations; that's the—that's the alternative. And so if you're going to think about mitigations, put open-endedness into the discussion. Like we are what we are—the central—it's not just that I think it's part of the discussion—we are the central thing. I say "we" because it's of my field; we are the central issue. Like this is the—the fundamental problem that we have is open-endedness. It is the fact that open-ended systems are unpredictable, out of control, and extremely powerful. And like how they work is what figures into how we're going to deal with this. Like to just dismiss it and say it doesn't matter means that we're not going to grapple with actually what's going on.

Like we've always said—like when I talk to my colleagues in the field, I'd say like part of the reason to study open is—is not just to produce these interesting systems, but is to understand how we're going to deal with them—like when they start to emerge naturally. Because like that's the kind of world that we're going to be entering into is an open-ended world with an open-ended complexity explosion. And so it's in our interest to really grapple with open-endedness. I mean, DeepMind even like put out a paper; it said open-endedness is essential for artificial superhuman intelligence. Like this is not—this is not an obscure, crazy, fringe view that like open-endedness is moving to the center of what we're talking about here when it comes to superintelligence. And so just like not grappling with it is a huge omission in these discussions on both sides. Like I find that people arguing against doom also accept the premise that ultimately we're talking about super optimizers, and that evolution is the super optimization system, which I find a very naive, like uninformed view of evolution. It's like the genetic algorithm view of evolution. I hear even like Wolfram saying things like this. And so people don't have this perspective because this—this field has not cracked the mainstream, and I just want to get it into the conversation so that you start grappling with this and not just ignoring it.

Okay, all right. I mean, as I think we went over why I basically disagree, but I'm glad you were able to come on a podcast even with me disagreeing with you and get your message out that doom discussions should factor in mitigation by understanding open-endedness.

Correct. Yes, that's—I guess that I agree with that. You didn't—it didn't sound as exciting as I wish it sounded, but I agree that's basically the point I'm making.

Yeah. Okay, great. Right. Yeah, I just want to thank you. Um, it's pretty rare uh for a distinguished individual like yourself, who obviously has very thoughtful takes, um widely respected, to come on the show and uh you know, face the onslaught coming into the lion's den and you know, hash it out. Like I said, I never pull my punches, so I—I thought you were a really great guest; debated in good faith. And uh you're welcome back anytime.

Hey, thanks for that. I was really—I'm really grateful to have this opportunity, and um I—I do want to say that um I really admire the fact that you created this forum, because it's very important to have these discussions. Um, some people might think this is a sensationalist kind of thing, but ultimately we need to talk about these things. Um, so I think it's really cool that you—you put yourself out there and did this, and—and it deserves an audience.

Once again, huge thanks and huge respect to my guest, Ken Stanley, for coming into the lion's den. You know, people often ask me, Lon, are you doing a debate or are you doing an interview, or some combination? What is it exactly? And the answer is it's always a debate in the sense that I'm always trying to reconcile my worldview with their worldview. I'm always trying to find the crux of why we disagree and what it would take for us to agree. But there's a couple reasons why it has so many elements of an interview. Number one is you can't properly debate somebody until you nail down their position. That's a mistake people make so often; they come out of the gate with their own counterarguments before they've even heard and understood and confirmed their understanding of the other person's position. So you'll notice I spend a lot of time just trying to make sure I get the other person's position, and I generally try to repeat it back to them in my own words. Because if they can't even agree with my own reflection of their position, how will they ever agree to my attempt to rebut their position? So that's one reason it sounds like an interview. And the other reason is because debating somebody can often be the best way to bring out the nuance in what their position is. So if you go to an all-friendly interview, they just run through a series of questions; they're nodding along; you might not realize that their point has all these interesting nuances because they never had that friction, that pushback from the host. So I like to think that if you're just browsing YouTube trying to learn about Ken Stanley's position on AI, you might do better to watch a debate than to watch a traditional type of interview. So it is actually an interview in that sense.

All of this connects into Doom Debates' mission, which is to build the infrastructure for high-quality debate. And the second part of it is to raise the level of awareness about AI existential risk. We have a double mission like that. And I can't emphasize how important it is for you to support that mission by smacking the Subscribe button on YouTube, searching for Doom Debates in your podcast player, subscribing that way, going to DoomDebates.com, subscribing to me on Substack that way you can get it in your email inbox, and you can get bonus content if you subscribe to my uh Substack. So please make those numbers go up; help me optimize my loss function, my objective function, which is to increase the amount of subscribers. As you'll notice, I'm getting higher and higher quality guests; the podcast is picking up steam. My previous episode with Rune is already breaking records. So when I'm approaching all these different guests, or when I'm getting inbound from high-quality guests, that little number—the subscriber number—it makes a huge difference. So subscribe; tell your friends to subscribe; you're really helping the snowball effort here.

Before you go, I got another clip from my friend Michael's channel, Lethal Intelligence. Today he's going to be talking about the rocket alignment problem, which is a really cool analogy. I believe it was first invented by Eliezer Yudkowsky; I use it myself pretty often. And it's just this idea of when you're designing a new system, it's untested, unproven, but it's very complicated, and the system has to operate at a bevy of very strong forces, and then you just kind of let it go, and you hope that it goes in the right direction. The unpredictable behavior that you're going to see is probably not going to be good; unpredictable behavior you're probably not going to see pleasant surprises from your rocket; you're probably going to see unpleasant surprises. There's just many more ways to fail than there are ways to succeed. That's very familiar if you study the history of any rocket company. I mean, the vast majority of rocket companies fail. Even the great SpaceX famously started with three Falcon One failures before having a pretty good cadence of success, and now only a rare failure. But so many different rocket companies have so many different kinds of failures because it's so freaking hard to make a rocket work. Now, arguably the AI alignment problem, the AI guidance problem, is even significantly harder than the rocket alignment problem and the rocket guidance problem. And that is why, if you just step back and look at it from a very grown-up, mature distance view, you get freaking scared. Like the—the nature of the challenge—it's not a tractable challenge; that's what I'm trying to tell you guys on this channel.

So let's see how Michael puts it in his video. I think he puts it well: Imagine you're trying to build a rocket and you want to land it at a specific location on the moon. Unless you solve very precisely and accurately the navigation problem, you should expect the rocket to fly towards anywhere on the sky. Imagine all the engineers work on increasing the power of the rocket to escape gravity, and there is no real work done on steering and navigation. What would you expect to happen?

So there you have it: To expect by default a general artificial intelligence to have human-compatible goals that we want is like expecting by default a rocket randomly fired in the sky to land precisely where we want. Expecting a rocket randomly fired into the sky to land exactly where we want—it's such a rich analogy. Uh, one of the intricacies of this analogy is, hey, what if we point the rocket really, really accurately right at the moon or right at Mars? If we just point it accurately enough, won't—won't it arrive? Then, and then of course there's the whole concept of orbital mechanics, of like, no, all of the gravitational forces are going to get it way off of a naive linear Newtonian trajectory, right? You really have to understand orbital mechanics; it's often counterintuitive where you actually need to fire the rocket and when you need to burn the engines, right? You have to do multiple burns, like getting out of Earth's orbit, getting into the other—or I don't even know what I'm talking about; I don't know orbital mechanics, right? But—but the point is, factors come into play that are counterintuitive. There's an analogy that I would make between, hey, you can't just shoot it in a line; you have to study orbital mechanics. A straightforward analogy is when you have an AI, you can't just refer to the tokens it was trained on, because it's going to figure out instrumental convergence; it's going to realize that goals have subgoals; it's going to apply Bayes' logic, and suddenly it's going to be attracted to goal optimization. So there's an analogy between that and a rocket getting attracted to the gravity wells of all the planets that it's navigating. I like the analogy a lot, and so I'm going to link in the show notes; there's a post by Eliezer Yudkowsky; it's called "The Rocket Alignment Problem"—really great post, so I highly recommend reading that. And just in general, go watch the Lethal Intelligence video; it's full of these gems. I'm going to keep dribbling them out on my channel, but I recommend just going and watching the whole video and sharing it with your friends.

All right, that is it for today. Now I don't know how many more times I'm going to be able to say this, but have a great year everybody, and I'll see you next time on Doom Debates. [Music]