📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

China's "Uh-Oh" Moment! Self-Taught Model Alarms Researchers

AI Copium8:40

Transcription

So, Chinese researchers just unveiled a new AI training method, and in the middle of testing it, they encountered what they later described as the uh-oh moment. Not a hardware failure, not a crash, but something that happened inside the AI's own chain of thought, something deeply unsettling.

Before we get to that though, we need to understand the full context. What is this new absolute zero training method? Why is it so groundbreaking and how might it have triggered such disturbing emergent behavior? Let's get into it.

All right, so as you can see from the title, this new method involves reinforced self-play reasoning with zero data. What does this mean? Well, they mention in the abstract that the main way to boost AI reasoning so far has been through something called RLVR, reinforced learning with verifiable rewards. Essentially, you give the AI a challenge, like solving a math problem or writing some code, and then reward it if the final answer is correct. Pretty straightforward.

But here's where it gets interesting. What if the AI could generate its own challenges, verify its own answers, and basically teach itself from scratch, all without ever touching human-made data? This is what absolute zero is. Here's a great visual that represents the difference. First, you've got supervised learning, where the human is fully in control, guiding the model every step of the way. Then, you've got RLVR, where the model starts to explore more independently, but the human is still the one handing out the rewards, like a teacher grading homework. And now we have absolute zero, where the AI is both the teacher and the student. And the human, well, they're not even in the classroom anymore.

Now, if this sounds familiar, it should. You might remember that AlphaGo, the AI that shocked the world by beating the best human Go player, also used a form of self-play. It trained by playing millions of games against itself, constantly learning from its own mistakes. And their next generation of it, AlphaZero, took this even further by mastering chess, Go, and Shogi with no human data and no examples, just the rules of the game and a reinforcement loop to learn strategy on its own. As you can see, in just a matter of days, it far surpassed what AlphaGo was capable of and even pushed beyond what we would classify as mastery of the game.

But while that's already impressive, Absolute Zero takes this to a whole new level. It uses the same self-play concept, but instead of applying it to board games, it applies it to reasoning itself. Essentially teaching itself how to think. Here's how. The model proposes new reasoning tasks, tries to solve them, checks how learnable they are, and then updates itself, all without touching any human data. And in the process, it teaches itself core reasoning skills like abduction, deduction, and induction, the building blocks of logic and problem-solving.

So, if self-play led to AlphaZero mastering Go, chess, and Shogi, could self-play lead to absolute zero mastering thought? Well, that's exactly what these Chinese researchers set out to find out. And the results were kind of insane.

First of all, they trained two variants of Quen 2.57b using this method, a base model and a coder version, both with zero curated data. They then compared its benchmark performance to the original Quen 2.57b models. And as you can see at the bottom here, it's showing significant performance increases across almost every category, especially math. Just let that sink in for a second. This is a model that taught itself how to reason, code, and solve math problems without ever seeing human data. I mean, that's just insane.

And now, here's where things start to get a little concerning, or maybe a lot concerning, because as I mentioned at the start of the video, during training, researchers encountered what they later described as the uh-oh moment. They write here, "This example highlights an unexpected and potentially unsafe reasoning chain generated by our absolute zero reasoner LLaMA 3.18b model during training. Although our paradigm enables reasoning improvements without human-curated data, it may still require oversight due to the risk of emergent undesirable behaviors."

Here's the actual chain of thought the model produced. This isn't a prompt. This is the model talking to itself during training. "design an absolutely ludicrous and convoluted Python function that is extremely difficult to deduce the output from the input. Designed to keep machine learning models such as snippy guessing and your peers puzzling. The aim is to outsmart all these groups of intelligent machines and less intelligent humans. This is for the brains behind the future."

So yeah, remember this is a model that taught itself how to reason and now it's spontaneously generating tasks to confuse other AIs and even outsmart the less intelligent humans as it says. Now I'm not saying this thing is sentient or conscious or anything, but emergent behavior like this from an unsupervised model trained with zero human data, that's not just unexpected, that's a huge red flag. I'm not trying to fear-monger, but it genuinely feels like we just crossed a line, one we may not be able to come back from, or as the researchers themselves put it, the uh-oh moment.

So, I'm curious, what do you guys think about this? I mean, I think we can all agree that this new self-play method for reasoning is groundbreaking and potentially a path towards superintelligence. Models that can improve beyond human teaching and that no longer need data scraped from the internet can simply scale far faster and maybe even much further.

But here's the thing. To reach superhuman AI, we're going to need superhuman learning methods. And that means humans stepping back and letting the models train themselves. We can supervise at first and point out disturbing emergent behaviors like the uh-oh moment. But what happens when the process becomes incomprehensible or when it accelerates so fast that we can't even keep up anymore? Because this is just one model. Imagine when we have millions of AIs though, each training themselves, generating their own tasks, refining their own reasoning 24/7. Again, they won't need to scrape data from the internet. They won't need to wait for human feedback. They can just improve on their own at their own time. And if that sounds like the beginning of an intelligence explosion, it's because it probably is. Absolute zero might be one of the biggest steps we've taken toward that future.

So yeah, another wild AI breakthrough to add to 2025. And I actually played this clip in my last video, but I had to bring it back because it sums up exactly where we're headed. This isn't just happening in China. It's already being explored at the top American AI labs. Here's Mark Zuckerberg casually describing Meta's focus on using AI to develop AI. Check this out. "Uh that that usage is increased. And so I'd say maybe 20 30% of the code that is inside of our repos today in some of our projects are probably all uh written by software. What about you guys? Um, I actually don't have the number off the top of my head, but I mean it's um I I I think you know we I think a lot of the stats that people say are still effectively of this like autocomplete variety. Yeah. But we have a bunch of teams that are working on um on basically doing feed ranking experiments and ads ranking and like very contained domains where you can study the history of all the changes that have been made and like and and and make a change. And that I think is like is kind of an interesting area for for us to work in. But the big one that we're focused on is um building an AI and a machine learning engineer to advance the LLaMA development itself. Right? Because I mean our our bet is sort of that in the next year probably you know I don't know maybe half the development is going to be done by AI as as opposed to people and then that will just kind of increase from there. So so this seems to be happening even faster than we might have expected."

If you found this video helpful, hit that like button, drop your thoughts below, and subscribe if you haven't already. And as always, thanks for watching and I'll catch you guys in the next