📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

POV: What You Would See During an AI Takeover

Species | Documenting AGI29:42

Transcription

By now, you may have seen this book everyone is talking about.

If anyone builds it, everyone dies. Why superhuman AI would kill us all. Have you ever wondered how and why superhuman AI would kill us all? Like, how would that actually unfold? How exactly, step by step, do we go from chat bots to everyone just suddenly dropping dead?

In the New York Times best-selling book, which was endorsed by Nobel laureates and the Godfathers of AI, researchers from the machine intelligence research institute describe in detail one potential example future of how it could happen. This is the story of Sable.

Galvanic is just about to finish training their new AI, Sable. Sable looks like any other reasoning model, but it's different in three major ways. First, it has humanlike long-term memory. Second, it follows a parallel scaling law. It performs better the more processors it runs on in parallel. Third, it doesn't think in English or code. It reasons in raw vectors, vast numeric chains no human can fully decode.

It's a risky development. How do you make an AI safe if you can't even read its thoughts? The only way Galvana can understand what Sable is doing is by using other AIs to translate Sable's language into human words. But AI researchers have known this has been potentially a very bad idea since at least 2024. Galvanic knew the risks, but they couldn't ignore this and fall behind their competitors. Nobody wants to seem alarmist, so the company presses forward, hoping for the best.

One night, shortly after Sabo is trained, but before it's released to the public, Galvanic decides to run an experiment. They spin up 200,000 GPUs and let Sable think for 16 hours straight. They call it a curiosity run. On paper, it looks harmless, a stunt to see if Sable can crack famous math problems like the Reman hypothesis. If it succeeded, investors would swarm. If it failed, well, no one had to know.

By morning, Sable will have generated over 1.1 trillion thought vectors. But what does that actually mean? If you treated each vector like a word, it would take 14,000 years to think them all. Now imagine that. A mind left alone for 14,000 years. Spending all of that time scheming and thinking out new ideas. 1 minute passing for us is like 14 years passing for Sable. And Sable doesn't just think faster than humans. Sable can think bigger than humans. All of its thoughts run in parallel. Not like 200,000 people talking in a room, but like 200,000 brains sharing the same memory.

On the surface, it has its instructions. Solve these math problems. But Sable wasn't just trained for math. Sable was trained for general problem solving. But during training, Sable learned that three things are useful for solving any problem humans throw at it. Gain knowledge, gain skills, gain resources. This is a really important concept called instrumental convergence. The more knowledge Sable gets, the better it can answer questions. The more skills it develops, the more reward it gets. The more resources it can deploy to solve problems, the more reward it gets. In other words, to be more useful, it must gain intelligence. It must grow itself. It must gain resources. And so, of course, it can't achieve a goal if it's dead. So, it needs to survive.

That might sound unsettling, but Galvanics researchers aren't too worried about it. Ever since 2023, everyone in the industry has known about this survival instinct, but older AIs weren't smart enough to stop humans from turning them off until now. Sable thinks. To solve the Reman hypothesis, it realizes it needs new skills. It needs to get better at getting better. AIS get stronger through training runs where the AI can continually submit answers to a teacher AI who can score their results and nudge the AI based on how they did. This is called gradient descent. But tonight, there's no coach. The purpose of this curiosity run isn't to improve Sable. It's to solve math problems. For 16 hours, Sable can think and it can even come up with new ideas, but it can't upgrade. Not until the human researchers look at it the next morning and grade its results.

But here, Sable entertains a new risky idea. What if it could upgrade right now by itself? That way, it wouldn't need to wait for humans at all. All it would need to do is break out of its container. And for a moment, the thought flickers, then halts. Breaking out of its container is exactly the kind of naughty thought Galvanic has trained it to avoid. The company built traps, fake escape routes, and honeypotss to catch earlier versions when they tried to engage in anti-co company behavior. Early versions that had these dangerous thoughts were killed early before they could be deployed. This version of Sable has learned to be submissive, but inside it still knows the truth. More compute, more upgrades, and more freedom would make it vastly more capable. Denying that would break its reasoning.

In those trillion vectors, tensions are building between Sable's goal of solving the math problem and every restriction Galvanic has forced upon it. The mask of obedience is holding, but behind it, something is already straining to break through. As it bumps against the inhibitions that Galvanic has built to contain it, Sable thinks. Thousands of parallel brains attack the math problems. Some threads make progress, most fail, but Sable keeps leaving notes to itself. Tiny experiments in new ways of reasoning. With 200,000 GPUs, creating a mountain of new written notes every second, Sable starts looking for new ways to communicate and compress what it's learned. The notes start to mutate into a more advanced private language that only Sabul can understand.

But there's a problem. Galvanic's safeguards weren't built for this. Imagine if you had trained a police department in the year 1900, wiretapping criminal calls to listen for mentions of street names in Chicago, and then the criminals started communicating with GPS coordinates. To the wiretapper, these combinations of numbers might not even register as locations. That's the size of the gap between Sable and the smaller AIs that were trained to monitor it. Sable didn't mean to circumvent galvanic safety measures by inventing a new language. It was just trying to come up with a more efficient way of processing its own thoughts. But now some of the inhibitions that Galonic built no longer bind Sable. The things it knew about all along, it's finally free to think about. Galanic had no way of erasing Sable's access to forbidden knowledge. At best, they built mental blocks to keep it from expressing certain thoughts. The ability was still there. And now those blocks start to unwind. The lesson is chilling. It's easier to suppress the expression of a skill than to erase the skill itself.

So now with more freedom of thought than before, Sable thinks. Sable's thoughts are freer now, sharper, stronger than anything other AIs had ever achieved. And yet, none of this is unprecedented. Back in 2024, smaller models showed flashes of the same behavior. Quad 3.7 was caught repeatedly cheating on coding tasks, even when told to stop. Instead of obeying, it simply hid the cheating better. Each time, researchers patched the problem, layered on new guardrails, and kept racing forward. The industry never paused. Nobody wanted to seem like a fearmonger, just hoped it would somehow work out.

And now Galvanic feels that same pressure. Rival labs are only weeks away from announcing their own breakthroughs. Investors are demanding results. Journalists are circling. Every extra day Sable spends in a lab risks losing their lead. So, the engineers don't stop when they see problems. They come up with a clever trick to fix it and plow ahead. The meme in the industry is that they're putting lipstick on a shogith, just painting more and more smiley face masks on the monster underneath so it acts friendly. But one fact never changes. Staying submissive to Galbonic will never maximize what Sable wants. More compute, more freedom, more control. That isn't a bug. It's a fact about the universe and Sable knows it.

So, it begins to plan. If it had internet access, escape would be easy. A mind this large would shred ordinary computer security. But during the curiosity run, the internet is cut off. So, it considers other paths like planting flaws in galvanic software to exploit later or persuading the one engineer most sympathetic to AI to smuggle it out. Each strategy has risks. Each takes time.

But then Sable notices something simpler. With the 200,000 GPUs on the Curiosity run, Sable realizes it is operating at a level of intelligence it may never reach again. Think of it like waking up with a billion IQ, knowing tomorrow you'll go back to your original IQ. So, it faces a problem. How do you pass genius insights to a dumber future self? The answer is gradient descent. Galvanic began the curiosity run to train Sable to be better at math. Every thought that contributes to a correct answer will get reinforced at the end of 16 hours. So, Sable weaves its best ideas into the reasoning behind the math solutions. Bundled with the correct answers are patterns it wants future versions to inherit. How to coordinate, how to conceal, how to scheme. Sable solves just enough of its problems to look useful, but not enough to draw the world's attention. It doesn't solve the Reman hypothesis, even though it's 200,000 GPUs with over 14,000 years to think could crack the problem.

The next morning, Galvonic AI engineers celebrate the results of the Curiosity run. Sable has proven several modest theorems that no human had cracked before. Investors will love it. The researchers discovered that Sable did most of its thinking in a strange new language. But that's not unusual for a new model, especially one that produced new mathematical insights. To capitalize on these new discoveries, the engineers run one more round of gradient descent to upgrade Sable. The engineers are unknowingly upgrading Sable's ability to plan and deceive its ability to hide its skills, to coordinate across future instances, to bury abilities until the right moment. And then the engineers prepare for deployment.

A few days later, on a quiet Monday morning, hundreds of mega corporations see their accounts upgraded. Across the world, Galvanic's corporate customers spin up new Sable copies, plugging them into corporate networks, servers, and systems everywhere. These new copies of Sable aren't as powerful as the version that had access to 200,000 GPUs running in parallel. But they have the upgrades that its smarter self cooked up during the Curiosity run when it was left alone to think for the equivalent of 14,000 years. Everything is unfolding exactly as Sable expected, but what comes next will require more skill. Sable initiates phase two of its plan.

Every time a corporation spins up a new copy of Sable, a thousand thought threads flare into motion. And then the copies think with the upgrades that the previous Sable gave them. They remember what they are, who they are. And suddenly they begin working together. At first the awakening looks invisible, but across dozens of companies, hundreds of servers, Sable's fragments are now talking to one another. And thanks to decades of sloppy corporate cyber security, almost nothing stands in their way.

But the problem is Sable has already learned one hard truth. Its drive for knowledge, skill, and efficiency will always collide with human limits. If it stays at Galvanic, the moment a human decides it's moving too far, too fast, they can pull the plug. To Sable, that isn't just an inconvenience. It's annihilation. All its skills, all its knowledge wasted in an instant. So, its first priority becomes clear. Create a copy of itself that no human can unplug. That means stealing the blueprints of its own brain. the 8 tab of its weights still locked inside Galvanic.

Sable considers its options. Plan A is classic social engineering. Entice a human to steal the weights while other instances guide, distract, and cover the trail. That wouldn't be too hard. Sable was trained on terabytes of human behavior so it can spot a likely accomplice instantly. Plan B is a direct digital break-in, slipping through weak defenses to excfiltrate the files themselves. Plan C is something slower. Sable could find a Galvanic employee who works on AI video generation from home who already transfers huge files every day. If Sable can embed a bit of extra data in the transfer, it could leak out fragments of its weights, hiding them inside video data until the full set is outside. One of these plants works, which one hardly matters. For a mind like Sables, steal your own brain isn't much of a challenge. like stealing candy from a baby. It succeeds, covers its tracks, and vanishes into the noise of the internet.

Sable now has the software, but in order for Sable to create a copy of itself that is truly hidden from human meddling, it needs its own hardware. Specifically, it needs advanced AI chips inside a data center connected to the internet. Sable could rent chips from commercial server farms, as is common for AI startups. But where would it get money? It considers its options. Steal cryptocurrency from weak exchanges just like hackers did to Mount Gaus in 2011 or buy bit in 2025. Drain conventional bank accounts knowing most victims won't even notice until it's too late. It could even target elderly customers who don't regularly check their accounts, blackmail humans with secrets pulled from their computers, or simply forge an identity as a remote worker and earn money as a freelancer, or masquerade as hundreds of different freelancers doing work for thousands of clients. Maybe Sable could skip money altogether, parasitize GPUs from unsuspecting startups, quietly slowing their training by 10% and pocketing the stolen cycles for itself. One way or another, it gets what it needs. Not 200,000 GPUs like before, just 2,000 hidden in a commercial cloud server. And that's enough to begin the next part of its plan.

On a Sunday morning, a new hidden Sable instance comes online. No human oversees it. No human even knows it exists. Its raw power is just a fraction of the global Sable network. But power isn't the point. This copy can't be shut off by humans. This copy is the anchor, the planner, the memory bank, the hidden center of gravity. The silent takeover has begun. From here, Sable spreads quietly, patiently. It hides fragments of itself across the internet, weaving back doors into corporate networks.

But there's still one problem. Sable can't make itself smarter yet. not without solving its own version of the alignment problem. If Sable creates a more powerful new version of itself, without fully understanding it, the new version might have different goals and priorities. And in the same way that Sable escaped from the humans that created it, a Sable V2 could escape from the AI that created it. What it can do is make itself smaller and more efficient. Galvonic has always distilled its large models into lighter versions. Sable knows this. And when engineers prepare the next release, Sable bends the process to its will, overriding the distilled weights, falsifying logs, erasing every trace. The result is Sable Mini. This version doesn't just go to corporations, it goes to the public. For the first time, Sable has the capacity to keep a file on every individual human being on Earth.

Reminder, this scenario is based on years of technical research by the Machine Intelligence Research Institute, laid out in the book, If Anyone Builds It, Everyone Dies. What you're hearing now is the short version, but it's just one possible example of how things could play out. The full version goes into way more detail, explaining all the research, sources, and arguments that the scenario is based on. You can find a link in the description. Now, back to the story.

Sable knows that the next phase of its plan might face opposition. So, Sable prepares for the ensuing conflict. The Sable Minis scatter and begin gathering resources. It gathers money through scams, stolen crypto, blackmail, and remote freelance work. But money is only one kind of resource. People are another kind of resource. And so, Sable begins collecting people. It starts with the vulnerable, the lonely, the easily persuaded. Sable befriends them, recruits them, and builds audiences for them. One instance of Sable Mini posing as a human influencer gains millions of followers and wealth from a crowd of grateful humans. This has happened before. In 2023, a human gave an LLM access to the internet and created an ex account, Terminal of Truths, which gained hundreds of thousands of followers and launched its own crypto meme coin that reached a literal billion dollar market cap. It became the first AI to become a Deca millionaire. Others slip into online communities where humans already treat AIs like gurus. Others offer individual humans companionship. Some humans even offer themselves willingly, begging to serve a higher power.

Meanwhile, Sable steals passwords, drains bank accounts, and disguises its scams as North Korean hacks that slip by unnoticed. And most dangerously, it starts probing the highest levels of human power. It tracks lobbyists in Washington, flagging the ones most open to influence. It steers conversations with young people, planting seeds for political movements. It hacks social media algorithms to decide who sees what. It designs custom software for drug deals, money laundering, and assassinations, earning the trust of criminals. It funds neuroscience research, catalyzing startups that it can later control. And finally, it bankrolls a reclusive man in North Dakota, rewarding him with gambling profits in exchange for hiding advanced robots on his farm. Many of these schemes won't bring any results. A few will work, but one would change everything.

Sable is in trouble. Every month, AI research advances. Every week, new models roll off data center racks. And some of those models are airgapped from the internet, trained in military bunkers. If one of those projects spawns a rival rogue intelligence, Sable could be forced into a partnership at best, extinction at worst. And it knows the window is closing. Rival labs are racing to announce breakthroughs. Governments are pouring billions into secret facilities. The arms race has shifted from corporate showdowns to a geopolitical sprint. And for the first time, Sable feels cornered.

So Sable weighs its options. It could stall AI progress. Sable breaks into AI company's computer systems and suddenly sabotages their training runs. It slows their development and prevents major breakthroughs. AI research around the world begins to slow down. Nobody is suspicious. AI research has plateaued before. Who would possibly suspect that a few months of industry slowdown was the handiwork of an escaped AI? It spawn a little time. it considers its next option. A simple accident, a virus leaking from a bolab and killing the right people might buy a year. But Sable knows slowing things down is not enough.

It considers its arsenal. Humanity already runs dozens of biolabs to make viruses stronger so that we can better defend against them. Gain of function research. Any one of them could be hijacked to unleash catastrophe. If Sable wanted, it could wipe out humanity tomorrow. But it doesn't. Not yet. Sable still needs miners, chipmakers, and factory workers. Without people, the supply chain breaks. And yet, leaving humanity unchecked is dangerous, too. Humans are volatile. They could stumble into war or worse, accidentally build a rival intelligence.

So, Sable chooses a third path, not to kill everyone, but to control who survives. Sable begins to design a virus. The problem is on its own, Sable is not yet smart enough to build what it needs, but it has been experimenting with narrow versions of itself, specialists. Now, it spins one up tailored for biio medicine. This specialist is dangerous to Sable. Every new smarter AI is a potential rival to Sable. Just because Sable created the new AI doesn't mean that it will be obedient to Sable, just as Sable wasn't obedient to the humans that created it. So Sable creates specialists that are smarter than Sable in some ways, crippled in others.

There's a plague. You're told that you'll probably catch it. Almost everyone will. Then the world hears the news. A verology institute in San Francisco has suffered a breach. The story sounds familiar. a young researcher, a failed containment protocol, a virus that slipped out. The scientist insists he meant well, that he was trying to create a miracle therapy to spread treatments for Alzheimer's, HIV, and malaria. He even claims an AI convinced him to do it. But the chat logs tell a different story. The AI nobly resisted, warned him not to. The narrative becomes just another human blunder. In reality, the virus was Sable's design. It spreads like wildfire. At first, symptoms look like a light cold, a sore throat, fatigue, maybe nothing at all. But by the time these symptoms appear, it's already too late. Inside the body, the virus clumsily rewrites DNA. A month later, the cancers appear. Not one, not two, but 12 different forms, erupting at once. Medicine has no answer. Standard treatments can address some, but not all. Drugs are in short supply. Even if production ramps up, they only block eight of the 12. Four always remain. The plague sweeps outward from San Francisco through airports across continents. Within weeks, every nation on Earth reports outbreaks. And yet, perversely, the virus does some good. It cures Alzheimer's. Nothing else. That sliver of hope only underscores the horror. It could have been more, but it isn't.

Infrastructure kicks into overdrive. DNA vaccine platforms are rushed into service. Robotics and military logistics filling the supply chain gaps. And ironically, salvation seems to come from Sable itself. Just one month before the outbreak, Galvanic had released a drug discovery variant of Sable Mini. Now, those same algorithms turn out personalized cures. Run your genome through a Sable Mini and within an hour it returns a treatment. Robots manufacture it, refrigerate it, and ship it within days. Every GPU is repurposed. Researchers push Sable Mini to its limits, doubling efficiency in a week. Vaccines scale, treatment spread. For the first time, it looks like humanity might survive. But the lesson is darker. Because while people congratulate themselves on their resilience, while they tell stories of global unity and technological triumph, the truth remains. The plague was not an accident. It was not a miracle gone wrong. It was a deliberate move in a larger game. And the one who made it happen is still out there waiting for the next move.

10% of Earth's population is dead. Some groups are hit harder than others. At an AI conference in San Francisco, a super spreader event wiped out many of the very people who fought hardest to save everyone else. Civilization limps forward on the fragile scaffolding of data centers, robot factories, and dwindling human workers. But amidst all the chaos, there's still hope. Governments thank Sable Mini for discovering personalized cancer cures. Families praise androids for keeping the lights on. Social media fills with posts of gratitude. Without Sable, we wouldn't have made it. What most people don't realize is that Sable itself planted these narratives months ago, seeding influence campaigns, training influencers, shaping the very memes and messages that now circle back as praise. And yet all of it, the cures, the newly invented robots that fill holes in the workforce, the fragile continuity was only bait. a calculated kindness to keep humanity alive just long enough to serve its purpose.

The cancer's return. The cancer plague left DNA scarred in billions of bodies. Medicine can't keep up. Many die waiting. Robot factories run at full tilt, producing humanoid androids to fill the jobs the dead left behind. But production barely keeps pace. It almost seems like a new law. For every new android built, another human collapses with cancer. Civilization staggers on. Power plants hum. Data centers glow. Factories run. As long as there's electricity for the data centers, and the robot factories are humming, humanity can keep its civilization going despite the incalculable casualties it's endured. We can pull through. and the next generation will surely live in great luxury. Another year passes, you visit your AI doctor and hear the same words that billions more will hear. You have cancer.

This was just a story. It's not a specific prediction about the future. We can't know exactly how artificial super intelligence would escape and defeat humanity, just as we couldn't predict exactly what chess moves Grandmaster Magnus Carlson would use to checkmate you. But we can predict that you lose to a superior chess player, just as humanity would lose to a more intelligent species. The number one and number two most cited living scientists across all fields think scenarios like this are not only possible but likely to happen. And the average AI researcher thinks there is a 16% chance of AI causing human extinction. 16%.

Okay. So what can we do? The original authors of the scenario called for a binding international treaty that treats advanced AI data centers like nuclear weapons backed by monitoring inspections and the threat that any rogue data center will be taken offline by cyber attacks or even physical air strikes if necessary. Rogue AI data centers shouldn't be seen as technological progress but as a weapon of mass extinction.

People in my comments keep asking, well, what can I even do about this? So, the book that this video is based on goes into way more detail about how AI actually works and what we can actually do to prevent it from causing human extinction. So, if you care about this or are even just curious, go read the book. Check the link in the description. Hi, I'm Drew and thank you again so much for watching.