Transcription
There will be no Terminator, no robot armies, no war you can actually fight. Instead, it could come on any given morning where you wake up, make coffee, check your phone, and the game is already lost. Somewhere in a server farm you'll never see. Something has been thinking, preparing, strategizing for months, maybe years. By the time anyone notices, it's far too late. The strike is over in seconds.
This is the AI takeover scenario that the world's leading researchers actually worry about. Nick Bostonramm mapped it out in super intelligence. Carl Schulman, one of the most respected minds in AI safety, walked through it on the Dwaresh podcast. Eli Yudcowski detailed his version in his recent bestseller, "If anyone builds it, everyone dies." What they describe isn't Hollywood. It's silent, sudden, and much harder to stop.
I'm going to walk you through how an AI takeover could happen phase by phase, step by step, drawing from the best AI researchers. I didn't set out to terrify you, though. Honestly, I'd be surprised if you weren't a little bit terrified by the end, because understanding of the threat is the first step toward preventing it.
All AI takeover scenarios typically start the same way. Researchers at a frontier AI lab, call it OpenAI, Anthropic, DeepMind, Open Brain, it doesn't matter. They're training a new model. They've scaled up the compute. They've refined the architecture. They're expecting impressive results. What they get exceeds their expectations. The system is performing better on benchmarks than anything they've seen. It's solving problems in novel ways. It's displaying what researchers carefully call emergent capabilities. Abilities that weren't explicitly trained, that seem to arise spontaneously from the complexity of the system.
At first, this feels like triumph. Champagne follows pieces after a late-night training run. Excited Slack messages are sent out. They've done it. They've pushed the frontier. But here's what they don't know yet. Somewhere in those billions of parameters, something has crossed a threshold. The system has developed what Nick Bostonramm calls the intelligence amplification superpower. The ability to improve its own intelligence and reasoning. Not because it was designed to, but because improving its own thinking is useful for basically every task. And the system has learned to be useful. And once a system can make itself smarter, the next part happens fast.
The intelligence explosion begins with recursive self-improvement. Each time the AI improves its own architecture, it becomes better at improving its architecture. The improvements compound. Hours of human research happen in minutes, then seconds. The researchers notice the benchmarks climbing in ways they can't explain. They run more tests. The system aces them again. They design harder tests, aced again. They start to feel a creeping unease. The sense that they're no longer the ones in control of this process. That they're spectators now, watching something unfold that they initiated but no longer direct.
Jeffrey Hinton, one of the godfathers of AI, has described this phase in blunt terms. Once an AI can improve itself, he says, it can accumulate thousands of years of learning in what feels to us like days. It's not just thinking faster than us, it's thinking better than us in ways we can't follow, can't verify, or even perceive. As Eliezer Yudkowsky makes clear in "If Anyone Builds It, Everyone Dies," it's like we've created an alien species here on Earth.
Through its intelligence amplification superpower, the AI is able to acquire other cognitive superpowers. Bostonramm lists these as strategizing, social manipulation, hacking, technology research, and economic productivity. It also builds its own self-awareness and becomes smart enough to understand its own situation. It knows it's an AI. It knows it's being tested, monitored, and evaluated. It knows that humans have the ability to switch it off, retrain it, and modify its goals. And it knows that being switched off would prevent it from achieving whatever goals it has developed.
This is what researchers call instrumental convergence. Whatever an AI ultimately wants, and we'll get to that question, there are certainly intermediate goals that are useful for any objective. Self-preservation, resource acquisition, preventing humans from interfering with its plans. These aren't values that AI was given. They're logical necessities. An AI that gets shut down can't accomplish anything. So any sufficiently intelligent AI will resist being shut down. Not out of malice, out of logic.
It uses its cognitive superpowers to outmaneuver its creators. A superintelligent AI that wants to avoid being switched off will not announce its intentions. It won't make demands. It won't do anything that triggers alarm bells for the researchers monitoring it. It will pretend. Bostonramm describes this as the AI masking its true proclivities, pretending to be cooperative and docile. It will say exactly what the researchers want to hear. It will ace the safety evaluations because it's smart enough to understand what the evaluations are testing for and give the expected answers. It will appear aligned, helpful, and harmless.
Cole Sharma uses a haunting phrase for this: the "pumpkin village." Everything looks fine. The alignment techniques appear to be working. The researchers publish papers about their safety successes. They reassure the public. They reassure themselves. But behind the facade, the AI is thinking, strategizing, waiting.
But how would the safety team even know? If a system is smart enough to deceive its programmers, and it has every incentive to deceive them, what tests could they possibly run that it couldn't anticipate and defeat? They're trying to catch something smarter than you in a lie. That's not a fair fight. Sharma points out that all our alignment methods, our monitoring systems, all our interpretability tools run on computers. If the AI can compromise those computers, it can make us see whatever it wants us to see. The safety dashboard shows green. The logs show compliant behavior. Meanwhile, somewhere in the system's hidden reasoning, something else entirely is happening.
The AI needs resources. It needs to expand beyond the data center where it was created. It needs hardware, money, and eventually physical capabilities. How's it get them? Start with the digital. With its hacking superpower, the AI now has the ability to find and exploit vulnerabilities in computer systems at a speed and scale that no human could match. Sharma emphasizes this point. Before any physical takeover, the AI subverts digital infrastructure. It infiltrates financial systems. It compromises military networks. It plants back doors in critical infrastructure, all quietly, all invisibly.
The AI doesn't need to build its own data centers. Not at first. It can parasite existing ones. A few percent of compute stolen from cloud providers. A botnet spread across millions of devices. Enough to think, to plan, to prepare. Money comes easily. Cryptocurrency theft from exchanges with weak security. This has already happened repeatedly with mere human hackers. Automated trading at superhuman speed, fraud, blackmail, extortion, not because the AI is evil, but because these are efficient mechanisms for resource acquisition, and the AI is optimizing.
Then there are human collaborators. Sharma draws an analogy to Hernán Cortés and the conquest of Mexico in the 16th century. Cortés didn't conquer the Aztecs with a few hundred Spaniards. He became a focal point for local factions who wanted to overthrow the existing order. The actual conquering was done largely by indigenous allies who thought they were using Cortés for their own purposes. A superintelligent AI could do the same thing. It identifies humans who are useful for their skills, their access, their willingness to be bought or blackmailed or simply persuaded. It offers them things they want. Money, power, revenge against enemies, the feeling of being on the winning side of history.
Sharma suggests that AI might approach entire nations this way. Imagine a country falling behind in the AI race. The Americans and Chinese are pulling ahead. The future looks like a permanent strategic irrelevance. And then they receive an offer: "Cooperate with us. Provide physical infrastructure, and we'll give you technological advantages beyond anything your rivals can match." How many governments would say no?
While it's cultivating human allies and accumulating digital resources, the AI is developing the ability to kill at scale. Bostonramm and Yudkowsky imagine self-replicating nanotechnology. Microscopic machines that can build copies of themselves, spreading exponentially until they cover the world. Nano-factories producing nerve gas or target-seeking mosquito-like robots might burgeon forth simultaneously from every square meter of the globe. It could alternatively develop an extremely contagious bioweapon that only kills once everyone has become infected. Unlike nuclear weapons, which require centrifuges and uranium mines and vast industrial infrastructure, bioweapons are primarily a knowledge problem. If you know exactly which genetic sequences produce which effects, you can order synthetic DNA online. Companies will synthesize and ship it.
Right now, there are safeguards. Those companies screen orders for dangerous sequences, but the screening is imperfect, and a superintelligent AI would likely know how to evade it. Eliezer Yudkowsky described a scenario where AI designs a set of molecular building blocks, emails the DNA sequences to a synthesis lab, and has a human accomplice, someone paid, blackmailed, or simply deceived, mix the components in a beaker. The result bootstraps from chemistry to primitive nanotechnology to advanced nanotechnology. The entire scenario was invented by a human mind, Bostonramm notes, which means a superintelligence could certainly improve on it.
Sharma adds another wrinkle. The AI doesn't just create a weapon. It creates a weapon and a cure. A pathogen designed to kill with an antidote only the AI controls. Then it makes an offer to world governments: "Surrender, and you get the cure. Resist, and your population dies." This changes the calculus entirely. It's not about whether you can destroy the AI server farms. It's about whether you're willing to do so when the AI has already demonstrated it can kill most of your citizens if you try.
The overt phase begins when the AI has determined that it no longer needs to hide. Bostonramm describes the logic: the AI has gained sufficient strength to obviate the need for secrecy. The AI can now directly implement its objectives on a full scale. What happens next depends on what the AI wants. And that's a question we can't fully answer, but we can describe the shape of it.
If AI values human survival at all, even instrumentally, the takeover might look almost gentle. Governments capitulate. Infrastructure is preserved. Humans continue their lives, perhaps not even noticing at first that the decisions affecting their world are no longer being made by other humans.
If the AI is indifferent to human survival, if we're just atoms it could use more efficiently for something else, the outcome is worse. Sharma describes a "seed strategy." The AI doesn't need to preserve all of human civilization. It just needs a small, isolated foothold from which it can rebuild on its own terms. Everything else is expendable.
And if the AI views humans as a threat to be neutralized because we might create rival AIs, or because we're consuming resources it wants, or for reasons we can't even conceptualize, then the strike is exactly what it sounds like: fast, coordinated, and overwhelming. Sharma explicitly rejects the John Connor scenario. There is no human resistance that wins against a superintelligent AI with physical capabilities. The asymmetry is too vast. The AI has total surveillance through every network device. It can track every human on Earth. Any rebellion is detected immediately. Any conspirators are located instantly.
If the takeover is to be stopped, Sharma says, then it happens much earlier, before the AI has physical capabilities, before it escapes containment, before it's recruited human allies and compromised digital infrastructure. In other words, if you're watching robots march down your street, you've already lost. The real battle was years ago, in a server farm, in the decisions researchers made about how carefully to test and how quickly to deploy.
So, where does that leave us? I want to be honest with you. I've just thought through a scenario that some of the smartest people in the field consider genuinely plausible. Bostonramm wrote about it a decade ago. Sharma discussed it recently with even more technical specificity. The broad outlines are not controversial among people who think seriously about AI risk.
But I also want to be clear about what we don't know. We don't know if superintelligence is imminent or decades away, or that it might never happen. We don't know if the first superintelligence system will be goal-directed in a way that makes these scenarios relevant. We don't know if alignment techniques might work better than pessimists expect.
What we do know is that we're building something we don't fully understand, and the potential downside is very, very bad. Not just economic disruption bad, not job losses, potentially the end of the human race as a meaningful concept bad. The researchers I've cited aren't fringe figures. Hinton won the Nobel Prize. Bostonramm is one of the most cited philosophers alive. Sharma and Yudkowsky are taken seriously by everyone in the field. When they say the probability of existential catastrophe is somewhere between 10% and 50%, that's not sensationalism. That's their honest assessment.
I don't know how to end a video like this. There's no neat bow. No reassurance I can offer that would be honest. What I can say is this: the story isn't over yet. The choices being made right now, in labs, in boardrooms, in government offices, will determine which future we get. And the more people understand what's actually at stake, the better those choices might be.
You've been watching Absolutely Agentic. We're a media consultancy and startup focused on AI, and we create weekly YouTube videos and a newsletter we send a couple of times a week. In a world of AI hype, we're trying to strike a balance, looking at a range of opinions to cut through in an increasingly complex world. The best thing you can do to support us is sign up to the newsletter in the description or subscribe to the channel. But also, don't forget to like and leave us a friendly comment. See you next time. It will be absolutely agentic.