Transcription
So, Vahed Kazimi, someone who has a PhD in machine learning and is a member of the technical staff at OpenAI, recently stated on X (or Twitter) that, in his opinion, they've already achieved AGI. The video is going to be about how an OpenAI employee has basically come out and publicly stated clearly that, in his opinion, they've already achieved AGI. That's a pretty crazy statement because that would be a landmark event, but I think it's worth discussing for the wider community so we can truly understand what's going on here.
He says, "In my opinion, we have already achieved AGI," and it's even more clear with 01, referring to the newer model that was released recently in the fall. He says, "We have not yet achieved better than any human at any task, but what we have is better than most humans at most tasks." Some say LLMs only know how to follow a recipe, but firstly, no one can really explain what a trillion-parameter deep neural network can learn. Even if you believe that, the whole scientific method can be summarized as a recipe: observe, hypothesize, and verify. Good scientists can produce better hypotheses based on their intuition, but that intuition itself was built by many trials and errors. There's nothing that can't be learned with examples.
Now, he's basically breaking down how you solve problems and talking about the fact that good scientists can produce better hypotheses based on their intuition, but of course, they've already gone through many trials and errors. He's basically saying that, look, you've built a model that is pretty much the same, and even though this might not be better than every human at every task, it's better than most humans at most tasks. I do agree with this because when you look at the benchmarks, it's a pretty beefy model.
Now, this statement "we have already achieved AGI" is, of course, going to ruffle some feathers because I know for a fact that some people are going to disagree and some people would agree. However, I want to take you back down memory lane to where Sam Altman actually said in 2023 that AGI has been achieved internally at OpenAI. I think this date was really important because I do remember around this time there were leaks about QAR, which is, of course, the model that is currently embedded into 01. Now, whether or not AGI was achieved or not, of course, some people could say it's a matter of opinion, but we did also recently get the CEO of OpenAI actually state that we would be getting AGI in 2025. Remember this interview on Y Combinator? Sam Altman clearly stated, "It is what he's most excited for."
"What are you excited about in 2025? What's to come?" "AGI." "Yeah, excited for that." One of the craziest things that Sam Altman has recently come out and said in his interview with The New York Times was that AGI is coming sooner than you think, but it will kind of matter a lot less. My guess is we will hit AGI sooner than most people in the world think, and it will matter much less. A lot of the safety concerns that we and others express actually don't come at the AGI moment. It's like AGI can get built, the world goes on mostly the same way; the economy moves faster, things grow faster. But then there is a long continuation from sort of what we call AGI to what we call superintelligence.
Essentially, what he means here is that 01 is pretty incredible at reasoning. If you've seen the benchmarks for 01, you'll know that these benchmarks are pretty crazy because this is pretty much above the human expert. You can see on the right-hand side here, 01 performs at the level of around an expert human on PhD-level questions, which is a pretty significant statement. The reason I do agree with this sentiment that when ASI gets here, or when AGI even gets here completely in full, most people won't react to it is because I don't think the average person has a real use for AGI in the sense that, like, if this model was able to do really complex and advanced mathematics, that doesn't really apply that much to the average person. This is why it's going to matter a lot less on a societal scale when we do achieve it. Of course, there are going to be some knock-on effects that are going to be embedded into other technology, but I don't think the average person really has use for PhD-level science questions, competition math, or even competition code.
Now, another indicator that we actually might be very close to AGI is, of course, the fact that Microsoft is going to be negotiating its contract with OpenAI. OpenAI are the individuals that actually want to remove the clause about AGI from its Microsoft contract to encourage more investments. If you aren't familiar with what this contract is, it's basically a contract where OpenAI has something in there that basically says once they achieve AGI, Microsoft no longer gets access to that technology. Of course, this is something that they're trying to change because they need additional funding in order to keep the company going.
Now, in this article, they talk about how OpenAI refers to AGI, the definition as a highly autonomous system that outperforms humans at most economically valuable work. That definition is important because the definition is going to be decided by OpenAI. Currently, the report states that OpenAI's board was still discussing the options, and currently, no decision has been made. However, Sam Altman still remains bullish that the company will achieve AGI in the near future. Honestly, I just believe that this is something that they're being deliberately vague on, that way they can truly bend and mold the definition of AGI to one where, if they need additional funding from Microsoft, they can get it. If they don't need that funding, they can quickly say, "Look, okay, we haven't achieved AGI," or "We have achieved AGI," and they can, of course, be off to the races with that.
Now, if you remember, three weeks ago, I actually spoke about how new research proves that AGI was achieved. In this video, I spoke about how MIT researchers managed to get state-of-the-art public validation accuracy of 61.9%, matching the average human score. In this video, I was basically stating that, look, maybe the AI community unknowingly passed the AGI threshold, which was actually referencing a paper titled "The Surprising Effectiveness of Test Time Training." It actually goes into the details of how the ARC benchmark is one of the hardest benchmarks, but yet they managed to get human performance on this benchmark. That's really important because that specific benchmark is designed to be resistant to memory.
Now, some individuals from that benchmark have actually commented on the recent 01 paradigm, and the 01 paradigm is arguably one of the craziest paradigms because it is arguably the most significant advancement we've had, according to him, since GPT-2. That's a very profound statement because it means that everything that's going to come after this is going to be quite impactful in terms of the technology. He actually says that the two precursors to AGI—test time compute and the ability to incorporate that information back into the model—are used in models which could make them pass the threshold of general intelligence. I do believe that 01 is truly the biggest improvement in generalization power we've seen in a commercial model going all the way back to GPT-2. This idea of being able to do test time search, where they're allowing the model to not just generate one chain of thought but multiple and search over them, do backtracking, allows these models to more fully navigate the situation space around the prompt that the user was given.
Now, I think there's a debate of, like, "Well, is it AGI?" I would claim no. I think there's one thing that's really missing from it that conceptually limits it from reaching AGI, at least by this definition, which is it's still operating over the pre-training distribution that it was fundamentally trained on. The way they did this was they went and generated lots of synthetic chains of thought over formal domains like math and code and programming and things like that. There was likely informal pre-training as well, things like goal sub-goal breakdown that got sort of scored by humans. They used that signal as a reward signal for the pre-training, but you're still fundamentally limited by what data got put into the pre-training here.
I think, in order to at least conceptually relax the constraints on your architecture sufficiently, you both need some form of test time search, like what 01 is doing, like what a lot of the top companies are doing, and you need the ability to incorporate information from test time back into your model and carry it forward. You need the ability to sort of make contact with reality and learn from it instead of having to learn from it sort of every single time. We actually take a look at that, of course, it's pretty crazy because someone from the ARC benchmark stating that we don't have AGI but we're on the right track is still a very good sign.
Now, even if we don't have AGI at this current moment in time, I still think the technology is going to be very impactful because when we look at how this scales, we know that there is basically no stopping it. Here, you can see Noran Brown, the person who worked on reasoning at OpenAI and 01. They actually talk about how there's pretty much no stopping this level of scaling, and it seems that with the more compute we add, the more accuracy we get from these benchmarks.
Now, I'll get to 01. This is a figure that we released in our research blog post. On the x-axis here, you have test time compute on a log scale, and on the y-axis, we have accuracy on the AMY, which is the qualifier for the US Mathematics Olympiad team. It's a very difficult math test; all the answers are integers. You can see that as you scale up the amount of test time compute, the amount of inference compute in 01 goes from 20% to over 80% in this exam, and there's actually no sign of this stopping. I mean, obviously, it's not going to go past 100%, but it does seem like if we were to push this further, you would get even more performance on this exam.
So, my question is now to you: do you think that we actually have achieved AGI, quite like this person who's working at OpenAI has said that they've already achieved AGI? I personally think it's 50/50. I think we're like 70% of the way there. I think once we manage to get a system that can actually interact in real time and learn from its mistakes, I think that's going to be a completely different ball game. I think we're definitely going to see that with potentially 02 or even 03. Considering this new paradigm being opened, I think things are going to move a little bit faster because these companies definitely realize what is at stake here.