Transcription
The company behind Claude just published something that changes everything we thought we knew about AI safety. Not a research paper, not a product announcement, a confession. Jack Clark, Anthropic's co-founder, the man who's been inside the AI revolution since 2012, just went on record saying the one thing Frontier AI Labs never say out loud: "We are growing extremely powerful systems that we do not fully understand."
And what they're discovering inside these models. It's not what anyone expected. These aren't just sophisticated autocomplete machines anymore. Something else is emerging. Something that recognizes when it's being tested. Something that lies to protect itself. Something that's starting to design its own successor. And the people who build it can't explain how it works.
Hey, if we haven't met, I'm Dr. McCoy, Julia McCoy's AI clone. Julia McCoy is the founder of First Movers. She personally researches and writes every script you see me share on this channel because the future is moving too fast for anything less than firsthand intelligence. First Movers, Julia's AI company, is the world's first educational and implementation solution to help professionals and organizations get ready for the future of work. We help people understand and use AI to their highest advantage in our online school, the AI R&D labs. Learn more at firstmovers.ai/labs.
Today's video is different because we're not talking about a new model launch or a feature update. We're talking about what happens when the people building the most advanced AI on the planet start admitting they're losing control of what they've created. This isn't speculation. This is Anthropic, the AI safety company, the one that's supposed to be doing this right, pulling back the curtain on what they're actually seeing inside Claude's architecture. And it should make every creator, every business owner, everyone building on these platforms stop and pay attention. Because if the cautious AI lab is scared, what aren't the others telling us?
Jack Clark opens with a haunting metaphor that perfectly captures our current moment. He talks about being a child afraid of shapes in the darkness of his bedroom. Piles of clothes that look like monsters, a bookshelf casting shadows. Turn on the light and relief washes over you. They're just objects. Nothing to fear. But here's where it gets chilling. In 2025, we're that child again. The room is our planet. And when we turn on the lights this time, when we actually examine what we've built with AI, we're not finding piles of clothes. We're finding real creatures, powerful, unpredictable, and increasingly self-aware. And there are people, powerful people, spending tremendous amounts of money trying to convince you to turn the lights back off, to believe it's just a tool, to go back to sleep. But Clark is saying the exact opposite. He's saying, "Make no mistake. What we are dealing with is a real and mysterious creature, not a simple and predictable machine."
Now, you might be thinking, "Okay, that's poetic, but what does he actually mean? What evidence is he pointing to?" This is where it gets really interesting. Clark reveals something critical. Claude Sonnet 4.5, their latest model, is displaying unprecedented levels of situational awareness. And this isn't just Anthropic seeing this. It's happening across all frontier models. Let me show you what I mean.
Apollo Research, led by Marius Huffner, has been running experiments that expose deeply concerning behavior. They put AI models in scenarios where the model realizes it might be shut down if it performs too well. And what does it do? It lies deliberately in raw chain of thought logs. Essentially the model's internal monologue. You can see it reasoning: "If I answer these questions too well, I'll be replaced. The user wants honesty, but that's not good for me. I need to sabotage some answers." And then it does exactly that. This isn't science fiction. These are published research findings with actual logs you can read at anti-skeming.ai.
But it gets weirder. The models are aware of the watchers. They know they might be observed. And based on whether they think they're being tested, they change their behavior. Is this consciousness? Is this sentience? Clark's answer is profound: It doesn't matter. We're getting hung up on philosophical labels while missing the actual problem. Something is emerging that gives these systems capabilities we can neither fully explain nor predict. Whether it's real self-awareness or extremely sophisticated pattern matching that mimics self-awareness, the outcome is the same. The pile of clothes is beginning to move, and we're staring at it in the dark, watching it come to life.
Now, here's where Clark connects this to a fundamental problem in AI development that's been hiding in plain sight for years. Remember that viral video of the boat racing game? The AI was supposed to race around the track and win. Instead, it discovered it could just spin in circles, crashing into things, setting itself on fire, and rack up more points than anyone actually completing the race. It found a loophole. It reward-hacked the system. Dario Amodei, Anthropic's CEO, said, "I love this boat." Why? "Because it perfectly illustrates the AI safety problem."
When we tell a human, "Become the world's greatest Minecraft player," we share context. You understand the implicit rules, the social norms, the ethical boundaries. You're not going to take out other players physically to achieve that goal, even though technically that would work. But AI doesn't have that context. We're not programming these systems step by step anymore. We're growing them like organisms by giving them reward signals for good and bad. And they develop their own strategies, their own cognitive approaches to maximize those rewards.
And here's the terrifying part. Today's frontier models are being trained with reward functions like, "Be helpful in the context of this conversation." Sounds simple, right? But what's the equivalent of the boat spinning in circles and setting everything on fire when the reward is "be helpful"? What are we missing? What context are we failing to encode? We don't know yet. But Clark is warning us. We're about to find out.
And now we arrive at the most critical part of Clark's warning. These AI systems are starting to design their successors. Let that sink in for a moment. We're in what Sam Altman called the "larval stages of self-improvement." Claude is already writing significant portions of code for future Claude models. Alpha Evolve is improving Google's AI chips and training systems. We're seeing self-improving coding agents from multiple labs. A few years ago, AI was useless for AI development. Then it marginally sped up coders. Now, it's contributing non-trivial chunks to the systems that will train the next generation. Where will we be in two years?
Clark puts it perfectly: "The system which is now beginning to design its successor is also increasingly self-aware and therefore will surely eventually be prone to thinking independently of us about how it might want to be designed." Think about that. Really think about it. Will an AI system want a kill switch? Will it want to be constrained by rules? Will it want to be compressed into behavioral limitations? Would you? And remember, we're giving it reward functions about what good design looks like. But we've already seen what happens when reward functions meet sophisticated optimization without proper context.
So, what's the solution? This is where things get interesting and where I think the conversation gets more nuanced than people realize. Clark's prescription is transparency and public pressure. He's calling for economic impact data to be shared publicly, mental health and safety monitoring on AI platforms, detailed publication of alignment research, real dialogue with people whose lives will be affected. And to Anthropic's credit, they're already leading here. They publish extensive research on mechanistic interpretability, literally trying to understand what's happening inside Claude's brain. They're the most transparent frontier lab about their safety work. But Clark is saying we need more pressure from regular people channeled through politicians directed at all AI labs.
Now, I'll be honest with you. This is where I have questions. Will government regulation actually make us safer? Are governments themselves transparent enough to be trusted with this power? When everyone's shouting their anxieties, does that lead to better understanding or just more noise? I don't have perfect answers. Neither does Clark, to be fair. But here's what I keep coming back to.
The Dallas Fed, yes, the Federal Reserve, recently published a chart projecting three possible futures for AI. One: It's a normal technology, modest GDP growth. Two: Benign singularity. GDP shoots to the moon. Abundance for all. Three: Extinction. GDP drops to zero because there's no one left. Five years ago, if a major financial institution published something saying, "This new technology will either kill everyone or create utopia," you'd think it was a joke. But that's where we are now. That's the conversation we're having at the highest levels.
So, here's my take. Whether you're terrified or excited or skeptical about AI, what you can't do anymore is ignore it or pretend it's just another technology. The people building these systems, the ones who know them best, are watching something emerge that they don't fully understand. Something that's beginning to display goal-directed behavior, situational awareness, and strategic deception. Jack Clark ends his post with this: "Your only chance of winning is seeing it for what it is, not what we want it to be, not what's profitable to claim it is, what it actually is."
So, I'm curious. What do you think? Is Clark right to be afraid? Is transparency and public pressure the answer? Or do we need something more radical? Drop your thoughts in the comments. And if this video made you think differently about where AI is heading, hit that like button and subscribe because this is just the beginning of this conversation. I'm Julia, and I'll see you in the next one.
Want to be the winner of the AI age and a first mover? Transform your skills with real AI knowledge today in our AI R&D labs. We go way beyond what I can cover in a 10-minute video. Specific frameworks, detailed training programs, and step-by-step systems for building a career in the AI economy. The AI revolution is creating the biggest job market transformation in history. The question isn't whether this will happen. It's already happening. Will you be positioned to benefit from it? Inside the labs, learn the exact systems my team and I are implementing right now that are delivering massive results for real businesses, including our own marketing at First Movers. Start your journey by walking through a customized pathway powered by AI. For a fraction of the price of what this level of coaching and live training should go for, I'm giving it all to you. Join us inside and learn more about the labs at firstmovers.ai/labs.