Transcription
In the industry, there is a concern that people don't understand this, and we're going to have to have some reasonably, I don't know how to say this in a not cruel way, a modest death event, something the equivalent of Chernobyl, which will scare everybody incredibly to understand that this stuff. We see the possibility. We, we're doing everything we can to hold this stuff and prevent it.
So, Google's former CEO, Dr. Eric Schmidt, just gave a mind-boggling interview at the SRRI forum. He gave some massive warnings about what AI is going to do with humans in the coming years. Watch this. We talk about AI, but what you should really be talking about is scale computing.
So, I was talking to Elon, of all people, um, a week ago, uh, about Grok 3, and I said to him, "How is it that Grok 2 went to Grok 3, which is so good?" It's well, it's within the leaderboard, right? And his answer was that he had the largest collection of computers doing one model training because all the other ones are piece-parted out. And that's again another data point that's from a week ago that the scaling laws of scale compute allow you to do these extraordinary things. And that's why again, you go back to the arc of, of the Alto and the Ethernet and so forth, and you just compound it. When do the scaling laws end? Dario wrote a very good paper, which you all should read, at Sario Omodai, um, on the, on essentially DeepSeek and scaling laws, and in it, he proposes that there are really three scaling laws going on.
The first scaling law is the one that you know about, deep learning. Deep learning is what ChatGPT works on, and that scaling law says that as you add more hardware and data, you get more emergent behavior, right? And we haven't found the limit all that. Everyone says there is a limit, but we haven't found it yet. I'm sure there will be a limit, but we haven't found it yet. So, I'm joining the club, and remember that the fact that there is a limit does not mean that it's soon.
The second one has to do with reinforcement learning. Reinforcement learning is essentially, uh, it's what AlphaGo did, and it's basically what it does is it scores different paths. It looks forward and back, and forward and back, and it gives a reward function for that. And then the third one is called test time compute, which is where, inside the system, while you're delivering answers, you're also updating your answers. He argues, and I think this is correct, that the latter two are just at the beginning of their scaling laws.
So, this means that the hardware folks in the room, build us more hardware. And those of you that are working on electricity and energy, we need a lot more energy. Yeah. The energy demand, I'll give you an example. I spent a lot of time in the Arab world recently and in India. These people are sufficiently insane that they're talking about one to five gigawatts, up to 10 gigawatts data centers. Now, let me remind you how big a gigawatt is. The largest dam is about two gigawatts. The big ones, right? The really, really big ones, the ones that are too far to drive across. Um, a nuclear power plant is around one and a half to two gigawatts. So, we're talking about two nuclear plants and one data center. That's what we're discussing. And by the way, one gigawatt is 45 plus or minus billion dollars of hardware inside of it. So, if you have 10, it's a half a trillion dollars.
And then with that kind of scaling, um, I know lots of people are working on algorithms and other approaches to reduce the energy demands of these clusters, particularly for computing with AI, but then as it becomes more cost-effective in terms of energy and also more efficient algorithms, people may use it more. So, which curve do you think will win in the end in terms of that aggregate, right? We're well, there's an old saying that, um, "Grok giveth and Gates taketh away." The old-timers will know I'm talking about. Intel makes the hardware, and Microsoft uses it all, and, and, and it just seems to arrive just in time, like if the thing doesn't work, the better hardware comes along, and boom, the software people are happy, and then they're complaining they need more hardware. I, that's how it works. We're certainly in the middle of that, but we're not at the end.
Um, so one scenario, uh, which I think is the most likely, is that we will end up with very large data centers. And yes, there are improvements. Another example is DeepSeek got all this attention, which we can talk about, um, that improvement in, um, so the cost per query and the improvement in cost per query delivery is actually faster on Gemini than on DeepSeek. This has been well established. So, and, and the improvement of these now essentially deep learning algorithms is going up at about a factor of 10 a year. So, at a factor of 10 a year, each year, you have lots of room for innovation. What DeepSeek did is they came up with a completely new algorithm to do something called SFT, and it's, it's an actual invention, but because they released it in open source, all of their competitors are just copying it. And furthermore, they based it, their algorithms, um, they've acknowledged, on Meta's open source model Llama 400. It's also the case that many people believe that they distilled illegally or inappropriately from OpenAI and so forth. The way you can tell this, if you ask DeepSeek who it is, it says, "I'm OpenAI." Right? It learned what it was told.
So, I was thinking back the, okay, so in this next part, the host asked Dr. Schmidt, "What happens when we run out of data to train AI?" And his take on the issue is quite interesting. Watch this. Much of the data you're training on is processed data. This, the industry actually believes that we've gotten all the data. But, but I was thinking also about the speed. You were talking about the hardware. And you know what's amazing with regard to what's happening with AI is not just faster chips, but it's computer architecture. It's the algorithms that's causing it to enabling it to accelerate much faster. I think that was one of the comments in your book when you were writing the afterward of the first book, that it's moving faster even than you anticipated in that first one. And then, so now I think a lot of the conversation is the extrapolation, right, beyond the, "What algorithm can it do this? How much power does it use? Where is it brittle? Where does it elucinate?" But fundamentally, how does it change people interacting with the computer? Because I was thinking about the fact that all this was possible because people developed algorithms, they developed chips, hordes of people had to label data to be trained on. Now, people are being hired to basically double-check answers, and is this the best answer in terms of reinforcement? But at some point, that relationship's going to change in terms of people with the tools. It's changing already. It may be in a whole, I mean, that's sort of what you were, but let me make pondering in terms of, let me make a, a stronger argument. Yes.
So, 45, no, 40, no, 50 years ago, in this building, what is known as the Windows, Icons, Menus, and Pull-down interface was invented. I know, because I used it. I didn't invent it. I used it. That WIMP interface is 50 years old, and it's what you use every day when you're on a personal computer, and there's derivatives of it on the phone. Um, doesn't it make sense that AI should be able to generate enough bits to have an on-demand interface that's tailored to you for your task, and that the whole notion of raster graphics and fixed bounding boxes and all that stuff that was invented here is completely replaced by storytelling? Right? That's, and, and you start thinking about that and thinking about it, excuse my previous joke, in, in the person's culture and language and so forth. It's a profoundly different experience. And I was thinking about that also in terms of how AI is different than analytics, because we think about big data and all of that, and in many of those cases, people were defining the algorithms and approaches that the computer supercharged. But then with, uh, all the things that Google had done and related to that, it was not training on algorithms people were providing, but teaching it how to learn. And that's how it won chess and Go and so many other things, and doing protein folding. Um, so how do you see that people interacting with the computer because of that might change beyond just a computer as a tool, as an extension helping do those tasks, but a collaboration of a different type we've never seen? And this again, I know is something that you've worked on at SRRI for a long time. Um, if you think about it for a while, at least for the next decade, okay.
So, in this part, Dr. Schmidt kind of gives us a timeline for how long humans will be able to control AI. And then he hints at ASI's arrival, a point where AI is suspected to practically rule over the human population. Kind of like a dystopian era. It'll be us plus computers working together. There are scenarios which we call artificial super intelligence, which I can talk about, where that might actually change. But at least for the next, for the foreseeable future, it's going to be us and computers. And both the first book and now this book are fundamentally about what does it mean for human identity to be working with an intelligence that is higher than ours in some areas and not as high in others. And we collectively believe that the industry is bringing the stuff forward faster than government, society, civil society can adapt it. You see this? That I can assure you that the people who invented social media, who were largely liberal people from San Francisco, did not anticipate the polarization globally that social media has allowed, because it allowed you to find your, your, your friends, right? And, and if you're in the narrowing of your society, um, and the algorithmic boosting that started roughly in 2015, um, in the various tech companies, um, is shown to be at least correlated, not necessarily causing, but correlated with a rise in mental illness among teenagers, especially teenage girls. There's lots of evidence or at least correlation of that. So again, when, when we were all inventing this stuff 20, 30 years ago, it never occurred to any of us. Now, Henry would say in his German voice, "This is because you didn't take any of those classes." And I said, "Yes, I was too busy programming." Right? So his position, which is what the book to some degree was about, was that the solutions to are not to let the tech bros make these decisions. Sorry for the stereotype, but it, it's true that this requires a concerted effort. So the reason we wrote the book was as a call out to say, we need interdisciplinary teams, people who understand history, right? Who've not had charmed lives, who understand how people are victimized, especially women, u, women and children. Um, and all of that as part of the conversation. What are the appropriate guardrails in those things? AI is both, the benefits of AI are so profound. Let me just list them. Curing every disease. I have one group that has decided that in two years, they can define all human druggable targets. That's a pretty big deal, right? And then the drug industry would then try to drug those targets. And we'll see if they can do it. But at least they're ambitious, right? That's like an, that's a Nobel Prize achievement. If they manage to do it, we'll see if they can, or multiple. Um, there are, there are many, many things in biology. Biology is to AI the way math is to physics, right? In the sense that that biology is so complicated, you're going to need AI to interpret it, and that's sort of what's going on. Another way to say it for biology is that sequence prediction, which is what again, ChatGPT is, deep learning, the way the way these algorithms evolve is they take a sentence, they take a word out, and then they score whether they put the correct word back in. Any system of physics, math, biology, and so forth, that is sequence-based is amenable to those techniques, and that's why, for example, AlphaFold and so forth did so well. It's a natural next use of essentially predict sequence prediction.
So, things like climate change, um, climate change is still occurring, guys, and, um, the development of new energy sources, materials, so forth that are carbon capture are all predated to that. I can just give you example after example what it can do. Um, and, and health is most, is the obvious one. What about an AI doctor for the world? Somebody who works with the nurse practitioner. Many, I was in Burma, uh, and I was talking to the guide lady, and I said, "I didn't see any hospitals." She said, "Oh, we don't have any." And I go, "What happens?" Well, somewhat like, if they're really sick, and she said, "Oh," pause, and then she says, "They walk into the woods and they die." Very sad. I mean, again, it would never occur to me, but that's probably true for a large portion of the population of the world. Why can't we improve that? And the answer is, think about an AI doctor in their language that serves them. Teachers, right? AI, there's lots of evidence of excellent use of AI in the high end, but what about getting the average person's education up again, in their language, and at their level, right? To help the teachers in rural schools? Why do we not have these? All of these things are very powerful. On the negative side, um, I've worked with the aforementioned, usual suspects, to work on this question about AI safety, something again, you've looked at quite a bit. And the core problem today is bio and cyber. Cyber, because it's pretty obvious, right? That if you, if you do enough reinforcement learning, you can eventually find a zero-day exploit, as they're called. So imagine if you have a super intelligence, it can find them all very quickly. We'll see if that happens. And in bio, it turns out it's relatively easy to take existing viruses and modify them so that the various solutions and detections don't work, and then they kill a lot of people. Now, thank goodness, this hasn't not happened yet.
In the industry, there is a concern that people don't understand this. Okay, so in the upcoming, Dr. Eric Schmidt compares the current AI situation to needing a modest death event like Chernobyl for people to be alerted to the risk. I guess it's a massive statement coming from a person who knows a thing or two about AI. Let's listen. And we're gonna have to have some reasonably, I don't know how to say this in a, a not cruel way. A modest death event, something the equivalent of Chernobyl, which will scare everybody incredibly to understand that this stuff. We see the possibility. We, we're doing everything we can to hold this stuff and prevent it. But does it take a tragedy? Hiroshima and Nagasaki were a tragedy. That's what, and Henry did, as you know, right? The early, and this is a Rand, Rand project where they invented essentially mutual assured destruction, right? In the 1950s, nuclear. So, so we're gonna have to go through some similar process, and I'd rather do it before the major event with terrible harm than than after it occurs. And because it, you sort of paint that there's both amazing goodness that can come and risk, and often the risk we talk about, and you talk about in your book, is how do we align these systems so that they reflect and protect human values. Uh, and also, I thought it was very compelling the discussion about how productivity will improve, tremendous growth in society overall, but it could have very disruptive effects with regard to equity and balance in terms of opportunity for people and things like that. So maybe one last comment, then we'll take some questions with regard to what are things that can be done along the lines of what you described in the book for the strategy for how do we begin to bound this? So, so that we get the benefit and minimize the risk.
So, Henry and I travel to China to meet with Xi, to warn them. This is well before DeepSeek, and as a result, the US and China are having what are called track two discussions about this. They're hilarious because the Americans are all on the Zoom, kind of normal Americans, kind of disheveled, and the Chinese are all lined up in a row with their little ties, very organized, very precise. Uh, I guess it's how each side shows love to the other. Um, and, um, I don't think I'm going to be a very dip, dip, very good diplomat. And my co-author Craig leads this and does it very well. Just as long as you have to eat shrimp with chopsticks that's not peeled. I had that once when I was over in Beijing. Um, I think that the real question, first place, I don't, maybe this is why again, I'm not a diplomat. To get the other side to give up something while you're in a race turns out to be really hard. I'll give you an example. One person said, "I have an idea." I said, "Okay, what's your idea?" "What we're going to do is we're going to have a treaty where each side has a bomb that is on the power supply to the data center, and it can remotely detonate the bomb on the, not on the data center, but on the power to the data center whenever it's really upset." And I said, "Good luck. Try that." So again, people are thinking about how do you, how do you create a situation? I published a piece with Dan Hendrickx, um, last week on super intelligence, which, um, I hope you all get a chance to read. And just to, just to finish the, this is actually important. Let me get it out. Um, the industry believes that in the next little while, we're going to get to the point where we're going to not just have human scientists, but we're going to have AI scientists. So, the industry believes that you'll have like a thousand people at OpenAI, and you'll have a hundred thousand AI scientists. Now, let's assume these AI scientists are as good as the humans, or even better. And this is the slope of innovation. And I already told you it's a factor of 10 per year, which is mind-boggling. What happens when you add a, a million AI scientists? Presumably the slope goes like this. Okay. Now, let's assume for purposes with that the US gets its act together, highly unlikely. And we're actually doing this. We have all the data centers and so forth, and we've just done this, and China is six months behind. Now, everyone here would say, "No problem. Six months is not very much." In network effect businesses, when the slope of growth is this, you never catch up. Okay? So, this means that when America gets to the point where something new that could completely destroy the country of China occurs, China would have a six-month latency or more, or vice versa. Or vice, and obviously it's vice versa. So, the first thing you conclude is that America should win the, the race for super intelligence. That's kind of obvious. But the real question is, how do you manage the global partnerships? And, um, one obvious thing to do is say, well, what would China, in the scenario where we're ahead? What, what is the first thing China would do? The first thing is that they would try to steal the intellectual property. The second thing that they would do is try to modify the weights. That's called an adversarial attack. Um, that might work. So let's say that they don't work, and we're nearing the point of total intellectual dominance, right? That we're building a thousand Einsteins and a thousand Leonardo da Vincis much ahead of them. What are China's options? Preparatory attack, a preliminary attack. So you see in this logic, this is inherently destabilizing to world order. And this is a problem that Henry did not have. He used to, when he negotiated with the Soviets, uh, what he would do is tell them how many missiles that they had in the beginning, um, we, our classified information about their classified information, if you will. So they knew what we knew. So they knew, and then, and the negotiations consisted of, well, you have 5,000 of those times this many kilotons, and that's too much, and we have less, but we have more kilotons, and so forth. Maybe we can reduce both numbers. You can't have that conversation in a network effect business. But he told me one day that he went into the meeting, and, um, they're sitting there, all the, all the Soviets are here, and the Americans are here, and he's leading it, and he starts, and immediately they all start screaming at each other on the other side, and one guy is carried out on the Russian side. He was not cleared to know what the Russians were doing that Henry was about to tell them, right? Just, you know, just imagine any of that happening today. So as we pivot to the questions, do you make a distinction between super intelligence and generalized intelligence?
Okay, so this next part is where Dr. Eric Schmidt claims we're three years away from AGI, and then he compares AGI to ASI, which is the most feared form of artificial intelligence. Yeah, the general, the term that's evolving, yeah, the, the definition is evolving in the industry. And I should say, by the way, that I call this the San Francisco school. The San Francisco school believes that we'll get to the equivalent of something close to AGI within three to four years, which, if so, is a huge event. It's like even more important than the founding of Xerox PARC, excuse me, but not by much, right? Just a little. Yeah. I mean, PARC, PARC is really important. And that's a really, the arrival of a super intelligence that is equal to the union of, of so intelligence is something greater than that. It occurs when you take the union of everyone's intelligence and you say, "We're even smarter than that." Okay.
So, in this part, Dr. Eric Schmidt makes a very interesting statement on AI versus the human mind. Watch this. A question about misinformation. What is truth? There's more. When, um, when new information scales beyond known human information, information, and the question is provocative, "Aren't we creating reality?" So, Henry was very interested in the move 37 and AlphaGo. The Go game had been around for 200 years, and AlphaGo invented a move that had not been seen or at least documented in human history that worked and caused them to win the, yeah, it's like game two, right? And, and he kept, he kept saying to me, "How does, is that an alien intelligence? I mean, is that what is that?" Right? So, so I answered it by saying, it's an algorithmic derivation of a particular path. That's not what he was looking for. He was looking for why did it occur there and not with, with humans. So, um, with respect to the question of truth, Google, at Google, I answered this all day by saying, "I have no idea, but we know how to rank people's attempts at truth." As a scientist, I will tell you the reason science works is because of essentially the ability to falsify. So, the moment anyone has a science result, everyone tries to see if they can falsify it, and if it's attacked enough times, then it's, uh, true. I was, I remember learning about relativity, Einstein's relativity theory, and I thought, "This is the stupidest thing ever." So I talked to my graduate student friend who did it in physics. "I use it every day. It works. It works all the time." So, in other words, Eric, you're wrong. So, it sure looks like the way we understand truth in science is from repeatability, falsifiability, and so forth. Um, with respect to speech, we're in a situation where people have lost track because they, the distinction between online and offline is going away because they live online and offline. It's quite seamless. They're losing perspective as to what is truth, and what the, what are the most important things in your life? Your health, your family, the people around you, your safety, you know, your immediate environment. We're sort of losing all of that. And it's made much worse by AI, AI algorithms, because AI algorithms maximize revenue by maximizing outrage.