Transcription
This is not just more of the same that we've seen in the past. We have now an existence proof that computers are able to do something that they've never been able to do before in in in history of humanity. Arc version 2 has just been released, and even the Frontier Foundation models are failing spectacularly.
Today we are releasing RGI 2, the next version of The Benchmark RGI. And now RGI 2 is pretty much the only unsaturated Benchmark that is feasible by Regular People, and so it's a very yardstick to measure how much fluid intelligence uh these models have, how close we are to true AGI.
And alongside that, we're really excited to be welcoming everyone to Arc prize uh 2025. The contest kicks off officially now; it's going to run all the way through the end of 2025. Um, the structure of the contest is very similar to last year. We're going to have the Kaggle leaderboard running. We're going to be testing uh this year on the semi-private data set and wrapping up and testing the final leaderboard on the private data set. Still got the big prize; it's unclaimed. Uh, in order to get the big prize, you have to open source your solution, have high high degree of efficiency running on Kaggle, um, and and yeah, we're really really excited to see all the new ideas. I think there was a lot that came out last year in 2024 that really pushed the frontier.
The next version of The Benchmark, it's more challenging; it's extremely unsaturated. Uh, all Frontier models are scoring effectively uh within single-digit percentage, um, and it's the first time where we've uh calibrated uh the human-facing difficulty of the tasks. So we actually hired uh roughly 400 people; we tested every single task, and every single task has been solved by at least two people. So we know it's very visible for humans; it's extremely out of reach for any system today. This is the frontier.
The Arc Benchmark forces us to confront an uncomfortable truth about our pursuit of artificial general intelligence, that the field has been overlooking: intelligence is not just about capabilities; it's also about the efficiency with which you acquire and deploy these capabilities. Intelligence is about finding that program in very few hops, using actually very little compute. Like, look at the the amount of energy that a human expends to solve one AR task over, you know, two, three, four minutes, uh, it's it's it's almost zero, right? Um, and compare that to a model like, uh, or on high compute settings for instance, which is going to use like over 3,000 bucks of compute, um, so it's never just an economic problem; efficiency is actually the question we're asking, right? Efficiency is the problem statement; it's not compute. The goal post is is AGI, uh, that's like what we're here to do. That was the whole point of launching Arc prize in the first place, was to raise awareness that there was this really important Benchmark that I thought shared showed um something important that like the the sort of uh research community and it was missing about the sort of nature of of artificial intelligence.
You know, this is like one of the things I think makes Arc special and and very unique and important, I would argue. You know, there's uh a lot of uh benchmarks in the world today, and to to my understanding and knowledge, pretty much every other Benchmark, you know, all the frontier benchmarks basically are trying to test for these like superhuman capabilities, right? These like PhD Plus+ type skills that you need to have in order to succeed at The Benchmark. It's not just compute; it's not just scale; you have to be scaling the right thing; you have to be scaling the right ideas, and maybe you have them. I personally just keep getting like surprised and impressed by the AR prize Community, um, how much folks are pushing the frontier, and I think it's really exciting too because it means that individual people and individual teams can actually make a difference. Um, if we're kind of in an innovation-constrained world, an idea-constrained world, which I think Arc AGI shows, um, that means you out there could actually make a significant contribution to the frontier of AGI. So if you're going to enter the contest, uh, go to ar.org, uh, and uh, good luck, good luck, see you on the leaderboard. [Music]
The these sort of like test-time optimization techniques or test-time search techniques, that's the current Frontier of AGI, right? And and there are many ways to approach it, of course. You can do you can do just uh test-time training or or in this case, you know, test-time, um, uh, you you can you can do search; you can do search of a symbolic space; you can do search over a uh Chain of Thought space, over a token space, or you can do search in latent space as well, right? So you have you have many different ways to do it, but really the frontier is: how do you adapt uh uh to novelty at this time by recombining what you know into some novel structure.
MLST is sponsored by Two for AI Labs. Now, they are the Deep Seek based in Switzerland. They have an amazing team; you've seen many of the folks on the team. They acquired Min's AI, of course; they did a lot of great work on Arc; they're now working on 01 style models and reasoning and thinking and test-time computation. The reason you want to work for them is you get loads of autonomy; you get visibility; you can publish your research; and also they are hiring as well as ML engineers; they're hiring a chief scientist; they really really want to find the best possible person for this role, and they're prepared to pay top dollar as as a joining bonus. So if you're interested in working for them as an ML engineer or their Chief scientist, get in touch with Benjamin Cruser, go to tabs.ai, and uh, see what happens.
Well, uh, Mike, it's uh it's amazing to have you on MLST. Welcome.
Yeah, thank you so much. We're very excited for to be here today. Mike, I I hear that you guys have got some very exciting news today. Tell me about it.
Yeah, super excited. To today we're back; we're really excited to be launching both Arc AGI 2 alongside an updated Arc AGI Arc prize 2025 contest. Both are going to be launching today, and you can go to Arc pri.org to to learn more and any of the contest.
Okay. And in a nutshell, what what is V2 and how is it different from V1? The way I to sort of think about it is: RCGI 1 was a benchmark that was designed to sort of challenge deep learning, and RCGI 2 in contrast is really a benchmark that's designed to challenge these new AI reasoning systems that we're starting to see from pretty much all of the frontier labs. And one of the really cool things about RGI 2 is uh we're basically seeing, you know, uh models, AI systems that are purely based on pre-training effectively scoring 0%, and some of the frontier um AI reasoning systems uh were in progress of testing them right now, and we're sort of expecting single-digit performance, um, so a really big update uh in terms of um over RG1 from 2024. So the the original version of of Arc was very much based at these kind of foundation models that didn't do reason; version two is tuned for the reasoning models. But what would you say to the charge that you're moving the goal post? I mean, how is ARV2 sort of like meaningfully um an evolution of The Benchmark?
Yeah, I mean, I think the way that I think about it is that the goalpost is is AGI, uh, that's like what we're here to do. That was the whole point of launching Arc prize in the first place was to raise awareness that there was this really important Benchmark that I thought shared showed um something important that like the the sort of uh research community and I was missing about the sort of nature of of artificial intelligence. Um, so that's kind of our goal post, and you know the the definition that I use for AI and the one that the foundation adopts is uh assessing this capability gap between humans uh and computers, and Arc prizes Foundation is really to drive that Gap to zero. I think it would be hard to argue that we don't have AGI uh if you look around and you can't find any more tasks that are very straightforward, simple and easy for humans uh that computers can do as well, um, and uh the fact is that we were able to find still lots of those tasks; in in fact, all of the tasks in RGI2 data set sort of fit into this category of things that are relatively easy and simple and straightforward for humans and comparably um very very difficult and hard for uh for AI today.
Okay, cool. So I know you guys have done loads of human calibration, and we'll talk about that in a minute, but the fundamental philosophy of of the of the arc challenge is focusing on human gaps, but at the same time AI models are becoming superhuman in so many respects. So is the big story the human gaps or is the big story The expansion of capabilities that are superhuman?
You know, this is like one of the things I think makes Arc special and and very unique and important, I would argue. Um, you know, there's uh a lot of uh benchmarks in the world today, and to to my understanding and knowledge, pretty much every other Benchmark, you know, all the frontier benchmarks basically are trying to test for these like superhuman capabilities, right? These like PhD Plus+ type skills that you need to have in order to succeed at The Benchmark. Humans like can't solve the problem that are in these benchmarks; you have to be very very uh uh like have a lot of experience, a lot of Education, a lot of training in order to be able to sort of even uh get close to sort of selling the benchmarks as a as a human. And I think those are important; those are useful, um, but I think it's actually more illustrating about something that like we're missing from the nature in the sense of like uh artificial intelligence by looking at the gaps, the remaining gaps between what's simple and easy for humans and what's hard for AI. I think that's much more of an inspiring story, um, I think it's uh one where um it's actually necessary uh to Target this in order to actually get AGI that is capable of um Innovation. You know, I think this is one of the main reasons I got into AI and AGI in the first place was um being really inspired and excited about trying to um build these systems that would be capable of like compressing science timelines, and if all we have is AI that looks like we had at the beginning of 2024, right? Based on pre-training, based on a memorization regime, um, you're never going to get to that because these are systems that are feely going to reflect back the experience and the knowledge that Humanity has sort of gained over the last, you know, 10,000 Generations as opposed to being ones that are capable of like producing new knowledge, new technology, adding to sort of Humanity's like Colossus of of sort of knowledge and Technology. If we want systems that can actually do that, we need AGI, and uh this definition that we've sort of used for this Foundation of easy for humans, not hard for AI, um, I think is uh if we can close that Gap, we'll actually get technology that's capable of doing that.
I wonder whether you think we are just about five discoveries away from AGI because there's going to be version three of the arc challenge; they'll presumably there'll be version four. Intelligence is multi-dimensional, and I can see this both ways, right? Because you know many critics of AI, they are almost gaslighting us; they're they're saying that this amazing technology that you're using it doesn't work, and I'm like, well, yeah, it does work. And will it be the case that the criticisms will become more and more kind of philosophical and they'll say, oh, you know, because because it's not biological or whatever, it's not the same thing, or do you think meaningfully we're about five steps away from AI?
I think think this is why benchmarks are important, um, and I had a similar question actually, you know, when I was starting to get back into AI in 2022, um, and trying to understand the world: like, are we on track for AGI or not? How far off are we? And I find that it's really really hard to get a sense of understanding of the capabilities of all these systems purely by using them. You can certainly get a sense by just interacting with them, um, but if you really want to understand what are they capable and not capable of, you really need a benchmark to discern this fact. This is uh, you know, one of the interesting things that I picked up from building AI products um at at Zapier as well: it's very different building AI with AI than it is classic software, right? One of the big differences is: when you're building classic software, you know, you can build and test with five users and know okay, hey, this is going to be like, you know, this product can scale to Millions; it's going to work the exact same way, and that's fundamentally just not the case with AI technology. You really have to like deploy it to a large scale in order to assess how it works; you need a benchmark alongside that scaling in order to tell you, hey, is this system working or not?
What were the main lessons that you learned from version one that you moved into version two? I think so. RGI2 has been in the works actually for several years; France started working on it, uh, crowdsourcing some tasks for it years and years ago. Um, there was a bunch of uh sort of inherent flaws we ran into with it that uh we learned as we sort of started popularizing the The Benchmark over the last year or so. You know, one of the things we learned was that a lot of the tasks were very susceptible to Brute Force search, so if that's something that has zero intelligence at all, and we wanted to minimize the sort of incident rate of tasks that were sort of susceptible to that, um, we hadn't human-calibrated it; we anecdotally we relied on some anecdotes to say that hey, RGI 1 is easy for humans; we had um, you know, a couple uh sort of STEM to STEM folks who had taken the whole data set, including the private set, and were able to solve, you know, 98, 99%, but you were relying on anecdote; we didn't have that calibrated across the sort of three different data sets that we had, and then we had all these AI Frontier reasoning systems come out over the last, you know, three, four months, and we've got a chance to study these um and learn what are the sort of qualities of Arc tasks that are remain very very challenging for these AI reasoning systems, which we can get into if you're curious, um, and so those are the main sort of insights and learnings that we took from RGI1 to try and produce an RGI2 Benchmark that I think will be a useful sort of signal for um development this year in artificial intelligence.
Can we quickly touch on the OpenAI situation? So in December they they La well they didn't launch but they gave you access to 03, and it got incredible performance on Arc V1, uh, human-level performance, something that we just didn't think really would be possible so quickly.
Yeah, surprise me; it came out of nowhere. I mean, can you just tell me the story behind that?
Yeah, um, yeah, I'll tell this is, you know, one of the reasons why I uh uh I'm always hesitant to make predictions uh in AI about timelines, um, you know, I think it's very easy to sort of make predictions along smooth scaling curves, um, but the nature of innovation is uh it's a step function, right? And step functions are really really hard to predict when when they're going to come out. Um, I think the the the best thing that I can say having spent sometime with 03 and looking at how it performs at ARC is that systems like 03 uh demand serious study. Uh, this is not just um, you know, more of the same that we've seen in the past; uh, this is uh, you know, we we that we have now an existence proof that computers are able to do something that they've never been able to do before in in in history of humanity, which is I think really really exciting. Um, I think there's still a long way to go to get to AGI, uh, but I do think that these things are important to understand, um, and sort of even discern how they work from a capability standpoint in order to make sure that future AI systems that were developing, building look more like this and not like the sort of pre-training uh pure scaling regime that we've had in the past. Um, so I I still remember the uh to sort of give you the anecdote; I still remember the two-week period, the Sprint we had on testing 03; it was right at the end of the contest; we had wrapped up our 20 2024 in I think early November last year, and we had a three- we four-week period where we were really really busy on judging all the final submissions, the papers, getting together the technical report, and uh we were dropping it on uh all sort of result on a Friday, and you know I was I was really hoping and anticipating that, you know, I was going to have a nice relaxing, you know, uh uh, you know, holiday period in December, and the day that we dropped the technical report we had a reach out from uh one of the folks at OpenAI who said, hey, uh, we'd really love you to test this new thing that we're working on; we think we've got some impressive results on on on on Arc AGI1. And so that kicked off a very very like hectic, fast, frantic two-week uh period to try and understand, okay, what what is this system like? Does it reproduce the claims uh that OpenAI had on testing it, um, and what does this mean for The Benchmark? What does this mean for for AGI? And um, you know, I think we're able to show the final result was that um 03 on its sort of high-efficiency setting, which fit within the sort of uh um budget constraints that we'd set out for our public leaderboard, um, got about 75% or so, and then they had a high-compute uh version which used I think like maybe 200x more compute than the low-compute setting which was able to score even 85%, um, and these are this impressive, and I think this shows the system like 03 has this, you know, uh sort of binary switch; it it we've gone from a regime where you know these AI models have no ability to adapt to novelty to something like 03 where uh in it's existence proof of now an AI system that can adapt to novelty in a uh in a small way.
Breaking this down a little bit, there were there were some interesting caveats that you just alluded to. So first of all, they did some kind of fine-tuning, and people at the time joked that, you know, isn't it scandalous that they were training on the training set?
So, yeah, this is this is like a very bad disc; this is a very poor critique; I think it misses the point of The Benchmark. I I think the folks who feel this way it's just because they're so used to thinking about benchmarks and AI from the pre-training scaling regime where like, hey, if if I trained on the data, um, you know, that's cheating them to test on the data, right? And that's true in the pre-training regime, but Arc is a very very special different Benchmark where it explicitly makes a training set available with the intention to train on it; this is very explicit; this is like what the Benchmark expects you to do; we expect AI researchers to use use the training set in order to teach their AI systems about the domain of Arc, and then what's special is we've got a private data set that you know very few humans have ever seen; the private data set it does not look like the training set; um, it requires you to generalize and Abstract the core knowledge Concepts that you learn through the training set at test time; um, fundamentally you cannot solve like the Arc AGI one or two uh private uh data sets um purely by memorizing what's in the pre-training set; this would be like, you know, maybe a a crude analogy would be, you know, if I was going to teach an AI system on grade school math and then test it on calculus; this is very similar to the type of thing that we do with with Arc, where you know the training set is much simpler, easy curriculum to learn on, uh, and then the test is a much more difficult one where you actually have to express true uh intellig; you have to have an actual capability of adapting to novelty at test time to solve it.
Okay, okay, all all of that is fine, but that you know there's a couple of things, right? They $25,000 per task or or more; that means they were probably doing samplings; they were doing a ridiculous amount of um completions; they were doing solution space prediction, which is very interesting, but but the main thing, Mike, just deep in your bones, deep in your bones, do you think that they were training on API data or or surely they were training on a whole bunch of data to do that? Well, and the extension of the question is: when they release the vanilla version of it, what performance would it get compared to their their tweaked version? We will test that as soon as it comes out, and I would love to report the results on that; they told us all they did was training on the training set, and I believe that's what they did.
Okay, very interesting. And just comment on the solution space prediction. I mean, I I was amazed that just predicting the output space directly they could do so well. I mean, doesn't doesn't that almost take away from the idea that we need to have discrete code, DSL type approaches if you can just predict the solution space so well?
I mean, effectively what 03 is doing is um it's able to use its pre-trained experience and recombine it on the fly in face of a novel task; it does this through like a regime called Chain of Thought; this is all informed speculation by the way; we don't have confirmed details; this is just, you know, my uh sort of personal assessment of how these systems work, um, particularly things like 01 Pro and and 03. Um, if you compare them with uh systems like 0 R1 or 01, these are systems that basically spat out a single chain of thought and then use that Chain of Thought in order to ground a final answer; this is distinct from how systems like 01 Pro and 03 work where they actually have the ability to do multi-sampling and recomposition at test time of that Chain of Thought; this allows them to build novel Coots that don't show up anywhere in the pre-training, um, not in the existing experience, and allows these systems to reach more sort of effect over more situation space effectively um based on what was in the original pre-training. Fundamentally these systems are a combination of like a deep learning model and a synthesis engine that is put on top, and I think the right way to think of them is: these are really AI systems, not single models anymore.
Yeah, and I agree with you; it's it's really funny though how you see the critique in the community because you know Gary Marcus is now saying, oh, it can't draw pictures of bicycles and labels and label the parts, whereas we see 01 Pro and 03, and it really does seem like a dramatic move towards um, you know, intelligence systems. But um, Mike, can we just quickly talk about the testing methodology? So the the real the real work that you guys did was you got a whole bunch of human subjects, and you you had the I think you had was it 400 test subjects, and they all needed like at least two people needed to solve every single task, and you had to do this experiment design, and you had to balance complexity of the tasks and so on. How did you do all of that?
Yeah, this was uh so this is one of the biggest things we wanted to fix with RGI 1; we didn't we never had a formal human calibration study on how do humans actually do on these things; you know, we relied on anecdote. So we had set up a testing center down in San Diego; we recruited tons of just folks from the local community, all the way from, you know, Uber drivers to single moms to UCSD students, um, and brought these folks out to go take Arc puzzles; a really cool like we to share some of the photos like these like testing shots where you have like, you know, dozens and hundreds of people like taking AR tasks on on on laptops. Um, our goal with originally with the data set was to ensure that every single task that we put in AR A2 is solvable by at least one human, and what we actually found was something even a higher standard I think which was that we found that every single task in the new V2 data set is solvable by at least two humans, um, under two attempts, and these are the same rules that we give to AI systems uh on the Benchmark, both on the contest as well as the the public leaderboards. I think this is a pretty good um sort of assertion of the sort of, you know, a straightforward comparison we can actually use now between: hey, are these tasks easy and straightforward for humans? Yes. Are they hard for AI? Yes. Like I said before, Frontier systems generally are getting close to zero or single-digit percentages on on these tasks now.
Okay, but the the idea though is there's more of X Paradox, right? Which is that, you know, um, basically while while we can select problems that are easy for humans and hard for AI, we haven't got AGI yet, but I was looking through some of your challenges, and and I I felt that some of them were very difficult; like, it would have taken me five or six minutes of deep thought to to get it. Like, are you are you finding that it's still easy to you know to find these things that are easy for humans and hard for AIs, or are you kind of scraping the barrel a little bit?
So I think easy for humans, hard for AI is a relative statement. Um, the fact is these ARV2 tasks uh were solvable by human on a $5 per task solve rate budget; they were solvable in five minutes or so, uh, and AI cannot solve these at all today. And so yes, I do think if you look at tasks, you have to think about them; there's like, you know, some thought you need to put in to sort of ascertain the rule, um, but the sort of data I think speaks for itself that you know we've got every single task now in the V2 data set from the public training set, I'm sorry, the public eval set to the semi-private set to the private data uh eval set, um, every single one is solvable by at least two humans under under two attempts, um, and yeah, uh these Frontier systems can't solve these things at all, or if they can, with a really really expensive budget to your earlier Point, thousands of dollars per task.
So you guys have been cooking; you're you're already working on version three of Arc. What can you tell us about that?
So the way I kind of think about the multiversion here: RGI1 again was designed to challenge deep learning as a paradigm; RGI2 is designed to challenge these AI reasoning systems; I don't expect that RGI2 is going to be as durable, right? RGI1 lasted for five years; I don't expect RGI2 is going to be quite as durable as that; um, I hope that will continue to be a very useful signal for researchers over the next, you know, year or two, um, but yeah, we've been working on RGI3, and uh I think the way to talk about RGI3 is: it's going to um challenge AGI systems that don't even exist in the world yet today.
Can you tell me about the the foundation that you're setting up?
Yeah, so this is one of the big cool things I think from Arc prize 2024 when we launched it; it was very much an experiment, um, you know, our Ambitions were were not quite what they were; I I think when we went into uh 2024, our main goals were just to raise awareness of the fact that
This benchmarks exists, and I think what we found was what I personally found: just kept getting surprised by the community around Ark. Um, you know, I, I remember this really specific moment when O1 preview came out, and there were thousands of people on, like, Twitter, like demanding that we test this, like, new model on Arc. And that was not my mental model of, like, what this Benchmark was or what the community was, and that was so cool. And that moment happened again when we ended the contest; that moment again happened when we launched O, uh, the results on O3. Um, and this kind of showed, I think, hey, there's a real demand for what Arc is providing; there's a real demand for benchmarks that look like this, that ascertain these, like, capability gaps between humans and computers. Um, and so we set up this foundation in order to, uh, basically be a Northstar for AGI and continue to produce, uh, useful, interesting, durable benchmarks in the sort of, um, Spirit of trying to, like, what are the things that are simple, straightforward, and easy for humans and still remain impossible or very, very difficult, uh, for AI. And we're going to carry that torch all the way till we get to AGI.
As you can see now, all of the large AI Labs, they're focusing on reasoning, and I'd like to think that Arc was at least a small part of that. And you folks are very focused on open source as well. So Mark Jen said specifically on the open AI podcast that there was that they'd been thinking about Ark V1 for years. There you go.
Well, yeah, exactly. But um, just, just tell me a little bit about that. So there's the industry, um, impact, but but you guys are really focused on open source as well. So how do you see those two things? So my, my sort of overriding philosophy at this point is: AGI is the most important technology that humanity is going to develop. And if it is true that we are in an idea-constrained environment, we still need new ideas to get to AI, which I think our, uh, GI to shows is true. If that's true about the world, then I think we should be designing the, like, most innovative sort of ecosystem and environment across the world that we possibly can. This is one of the reasons why we launched Arc Prize originally internationally, to reach solo researchers, um, to inspire researchers again to go try and work on these new ideas, to get past this pre-training regime, try something out. We knew that it would needed to be something beyond this and even beyond what we, what we have today. And I think if you look at, like, a really healthy, strong Innovation ecosystem, um, you're going to look at one that is very open, and there's a lot of sharing, and there's a lot of diversity of approach. Um, and this is in contrast to, you know, an ecosystem that would be very closed, very secretive, very dogmatic, very monocultural. And so those values of openness, those values of sharing are what the Arc Prize foundation stands for, in order to sort of increase the chance that we can get AI soon.
So, um, talking about the, uh, the version two of, of the Arc challenge, can, can you just give us an elevator picture of that?
Sure. So Arc 2 is basically a new version of Arc that keeps the same format but tries to address the main flaws that we saw in Arc one. So, for instance, in Arc one, we knew that, um, there was a little bit of redundancy across tasks. So we, we saw that actually very early on, as early as the 2020 K competition, and also Arc one was way too brute-forceable. So back in 2020, what we did after the K competitions is that we try to, uh, look at all tasks that were solved at least once by one entry in the competition, and we found that half of the private data set, uh, could be solved in fact, just, just via the sort of basic, uh, brute force program search methods that were, that were deployed during the first competition. So that means half the data set actually doesn't give you very good signal about AGI at all. So the other half was actually good enough, required enough generalization that The Benchmark of all was still useful and, you know, still lasted quite a few years after that, but it, it told you, you know, from the start that there were some pretty significant flaws, which is expected by the way, like, you know, when, when I started creating Arc back in, uh, 2018, 2019, I was flying blind. You know, I was, I was trying to capture my own thoughts, my own intuition about what does it mean to generalize, what, what is abstraction of this reasoning, and that turned into this Benchmark, but I could not anticipate what kind of AI techniques would be used against it. Um, and so yeah, as it turns out, a lot of it could be brute-forced. So Arc two, uh, completely addresses that. You cannot score, you know, higher than one, one or two percent at most, using brute force techniques on Arc two. So that's good news. Um, and other than that, we generally try to make it a little bit harder. So what we saw with Arc one is that, uh, it was very easy to saturate for humans. Like, if, if you're, if you're, you know, like a STEM grad, for instance, uh, you could very easily get 100% or within, within noise range of 100%, like something like 97, 98. And so that means that, uh, you were not getting a lot of useful bandwidth to compare, uh, AI capabilities with, uh, the capabilities of, of smart humans. And, um, if you just make it a little bit harder, then you get more range where, you know, if you're not, if you're not very intelligent, you will score lower; if you're very intelligent, you will score higher, and, uh, you're not super likely to completely saturate it until you're at the very, the very top end of the distribution. So that's, that's what Arc 2 is: same format, uh, same basic, uh, rules. So we're only using, uh, core knowledge, uh, you, you have these input-output pairs of grids that are at most 30 by 30, uh, but the content is very different. You're not going to find tasks where you only have to apply one basic rule that could be enumerated in advance, like, you know, some kind like gravity, uh, things falling task or symmetry task. All the tasks are very compositional, so you have multiple rules, you have more objects, the grids are generally bigger, the rules can, uh, be chained together, can be interacting together, and that makes it completely out of reach for brute force methods. And as it turns out, it also makes it out of reach, uh, for the, the, the, the base LLM training paradigm.
You're saying that you've made them more compositional and iterative and harder for humans. That's right. So could you give me a little bit more detail on that? I mean, if you think about it, there are, there are different dimensions of things that AI models can do, and there are different dimensions of things that humans can do. Have you sort of quite diversely explored that? Or, I mean, could you just give me a bit of a breakdown of the task characteristics?
So in Arc one, you had many tasks that were very, very basic where you just had one rule. Let's say, for instance, you have a few objects, uh, and you have to flip them, right? So this is an example of a task that's easy to brute-force because flipping is something that you can, you can acquire via pre-training as a concept or that you could just hard-code in a, in a brute force program search system. So if that's the only rule you have to apply and just apply it once, uh, that's not compositional; that's actually pretty easy to anticipate; that's easy to brute-force. So a compositional task is going to be a task where you have more than one concept, and, uh, typically they're going to be interacting together. Like an example of a very simple compositional task is, let's say you have, uh, object flipping, but also, uh, the objects are falling, right? So you have two rules to apply, uh, to each object at once, uh, but that again is a kind of task that could still be, uh, found by brute force program search if you have, as key elements in your DSL, gravity and flipping, for instance. And so you want to create tasks that are, um, that are where the rules are chained to a, to a sufficient level of depth that there's no way you could, you could find that chain by just trying every possible chain, where, where it would become too expensive. Of course, humans can still do it because humans are not just trying every possible combination of things they know on every problem that they see; they just have this, um, very efficient, very intuitive way of searching for, for a theory that explains what they see. You, you know, my co-host Keith, he had this idea of doing a recursive version of, of Arc, but the thought occurred to me is that even though we do this, um, systematic compositional reasoning, we still have some kind of cognitive limit. So if we nested, let's say four levels of, of Arc challenges within the same problem, wouldn't you find very quickly that humans just can't solve it, right? So if you just, like, concatenate two Arc tasks, for, for instance, uh, you get something that's much less brute-forceable; that's much harder because there are more rules going on. It's not quite what I would call compositional, though, because even though you have two rules at once, they're not interacting with each other, right? They, they, you can solve them separately, and they concatenate the solutions. And I think it's not, it's not a bad idea at all, like it will work as a way to make Arc, uh, more difficult, um, with again this caveat that you're not actually, uh, testing for, uh, depth of compositionality. Uh, one issue, though, is that it would only really work once because as soon as the person developing the AI system notices that the task can actually be decomposed into subtasks, then it's game over, you know, um, so I think it's actually more interesting to, to have multiple rules at once, but they're actually being chained together or they're interacting together in some way, where, for instance, one rule might be writing, uh, some information on the grid that needs to be read by the second rule, right?
What performance do the frontier models get on Arc V2?
So what we saw was a big gap between models that don't do any kind of test-time adaptation, like any kind of test-time search or test-time training, and models that do. And, uh, the base LLMs, even models like GPT 4.5, they're basically scoring zero. I think one of them, I think it was R1, maybe scored slightly above zero, uh, something like 1%, but, you know, it's, it's within, within noise range of zero. So any model that cannot, not do at test-time adaptation, that is, that does not possess fluid intelligence, starts at effectively zero. So in that sense, Arc 2 is actually a very strong sign that you have fluid intelligence, better than Arc 1. I think Arc 1 could already tell you that, but less perfectly. So on Arc 1, if you do not do test adaptation, you can still do up to roughly 10%, right? So on Arc 2, that's actually zero, so it's a better test. Now, when it comes to models that do test adaptation, so we tried, for instance, some of the, the entries from the K competition last year, you know, models that were doing test-time training in particular, or some, some kind of program search, and, uh, the best model, uh, the model that actually won the K competition, can do, I believe, 3%, uh, on, on Arc 2. And if you take an ensemble of the, the top entries from the competition, you get to 4%, right? Uh, so that's not very high. We also estimate that, so, so O3 would be the, the current, the art in terms of an AI model that does exhibit fluid intelligence. And so we haven't been able to, uh, test O3 on low compute settings on all of the tasks we wanted to test it on, but we, we've tested it on, on a subset, and so we can extrapolate which it would score on the, on the, on the entire set, and, uh, it sounds like it's going to be about 4%, right? Uh, so not, not, not super high. There's a lot, there's room to go higher than that, and we haven't been able to test O3 on high compute settings at all. So, you know, the model that was scoring, uh, 88% on, on Arc 1, so I can, I can make a guess, I guess, uh, based on what we saw from O3 L and, and, and other models, um, I think you might get up to like 15, maybe even 20%, if you were already maxing out, uh, the compute setting and spending like, you know, 10K, uh, per task, for instance, uh, but we see it be like far below, uh, average human performance, which should be more like 60%.
So that 4% that O3 gets on Arc V2, do you think of that as fluid intelligence, or do you think of that as a potential gap in, I mean, presumably you could have designed Arc V2, if you, if you selected the correct set of human calibrator challenges, you could have found a set which was still 0 for O3?
Yeah, absolutely. No, you, you, you could have, you know, adversarially selected against O3, and then O3 would do, would do zero. It would be very, very easy to go from 4% to 0%, right? It's just a few, few, a few tweaks that you need to change, uh, so we, we not, we not actually try to do that. Yes, I do believe that 4% does show that you have nonzero fluid intelligence, which is also something that you could, you could get as a signal from Arc one, um, and I think the, the, the, the sign that, uh, you see fluid intelligence in these models is the performance gap between the huge, uh, pre-trained-only models that don't do test-time adaptation that score effectively zero, maybe one, and you could say that that 1% is in fact a fluke, that, that sure, it should in practice be zero, uh, and the models that do test adaptation and do nonzero, that know 3%, 4%, 5%, right? And, uh, that means that there's something like 95% of the data set that will actually give you this useful bandwidth for measuring how much fluid intelligence the model has, and that's something you were not getting with Arc 1. Arc 1 was more binary, where if you don't have fluid intelligence, you're going to do very, very low, like below 10% roughly; if you do, you're going to score significantly higher, and getting above 50% would be very easy, uh, but because the, the, the, the measure would saturate very quickly as, as soon as you start having nonzero fluid intelligence, you did not get that, that useful bandwidth that you're getting with Arc 2. So I think Arc 2 should allow, uh, for, uh, you know, answering the question, is this model, uh, actually as fluidly intelligent as the average human, which is something you could not get in Arc 1.
I guess it's just an economics thing at this point. So if you spent, let's say, a billion dollars or half a billion dollars, you could saturate Arc V2. I'm not sure if you would agree with that, but if, if that isn't the case, I mean, what do you think are the specific things that are missing from O3 that, that are stopping it from doing it, doing better?
So it's never just an economics question because intelligence is not just about capabilities; it's also about the efficiency with which you acquire and deploy these capabilities. And sure, if you spend billions and billions of dollars, maybe you can saturate Arc 2, but that would already have been true back in 2020 using like extremely crude brute force program search. If you have a DSL that's actually true and complete, then you know that for every Arc task there exists a program that may not in fact be all that long, uh, that will solve the task, and all you need to do to find it is iterate over all possible programs in order of length, and then the first one you find is the one that's going to, that's going to generalize, right? Because it's, it's the, it's the shortest; it's the most, the most parsimonious. So if you spend, uh, unlimited resources, you already have, uh, AGI in that sense, just, just in the pure skill sense. You can always just try every possible program until you find one that works, but that's not what intelligence is. Intelligence is about finding that program in very few hops, using actually very little compute. Like look at the, the amount of energy that a, a human expends to solve one Arc task over, you know, two, three, four minutes, uh, it's, it's, it's almost zero, right? Um, and compare that to a model like, uh, O3 on high compute settings, for instance, which is going to use like over 3,000 bucks of computes, um, so it's never just an economics problem; efficiency is actually the question we're asking, right? Efficiency is the problem statement; it's not capability. So intelligence is knowledge acquisition efficiency. O3 did very well on Arc V1, and now that it does so badly on Arc V2, the whole point of, of your definition of intelligence is that given some basis knowledge, you efficiently recombine, you produce new skill programs. You're saying that in the absence of the base knowledge in V2, there is no intelligence; therefore, is O3 not actually as intelligent as we thought it was?
I think O3 is one of the first models, uh, perhaps the first model that does show fluid intelligence. So now what the results on Arc 2 are telling you is that it's not human-level fluid intelligence, right? But still, I would, I would consider O3 as a kind of proto-AGI with, with two, two big flaws to be curated. So one is, of course, efficiency. Efficiency is part of the problem statement, in fact, the central point. So as long as you're not as efficient in terms of, for instance, you know, uh, uh, data efficiency, compute efficiency, energy efficiency, then it's only a temporary solution; we will find a better solution in the future. Um, and also it's not, it's not quite human level. If it were human level, you'd expect it to score, you know, something like O3, O3 should score like over 60% on Arc 2, and we don't know what the exact number is going to be, but, you know, probably, probably like 4, 5%, right?
Do you think that general intelligence is a category or a spectrum?
So general fluid intelligence, it's, I would say it's both because there's a huge difference between just having memorized a bunch of skill programs that are static and, and, and knowledge factors versus being able to adapt to novelty to a nonzero extent. So that is a binary distinction: either you have fluid intelligence or you don't, right? And, and Arc one, uh, could answer that question for, in a system, uh, but once you have nonzero fluid intelligence, then the question is how much do you actually have, and how, how, how it compares to, to humans. Um, and that's related to the notion of, uh, recombination of the skill programs that you have, the knowledge that you have, uh, and, uh, depths of recombination. So if you do no recombination at all, you don't have fluid intelligence; if you do some recombination, you do, but then the question is how deeply, uh, can you recombine? Like, for instance, um, if you, using a, a program synthesis analogy, um, the, the question is how, how big of a program can you write on the fly to adapt to a new problem, right? Um, and, and of course, as well, how, how efficiently, how fast and how efficiently you can write it, right? So it is a binary, but it's also a spectrum. And Arc one was trying to ask the binary question, does the system have any fluid intelligence at all? And Arc two is, is more on the side of trying to measure how much fluid intelligence you actually have compared to humans.
How long do you think it will take for V2 to be saturated, and do you think it will survive until V3 comes up?
So that's, that's a question where you have to take into account resource efficiency. So if you're asking how long it would take before we have a system, uh, that can score, you know, higher than, let's say, 80% on, on Arc 2, uh, using, using less than, uh, uh, $10,000 of compute, for instance, um, I think probably around a couple of years. So it's very difficult to make predictions here. I think if you're just looking at current techniques and, and scaling up current techniques, I think it could take a while. I think the Arc 2 is actually way out of reach of current techniques, but of course, we're not limited to current techniques, you know, uh, in 2025, we're probably going to see new breakthroughs in the same way we saw, we saw new breakthroughs last year, and these breakthroughs are actually very difficult to predict. I was personally very surprised with the performance that O3 could get on Arc one last year; that, that came as a surprise. So maybe we'll have new surprises, uh, this year, um, but I would be extremely surprised if we see an efficient solution, uh, that's human level on Arc 2 by the end of 2025. I, I, I would basically rule that out. Uh, by the end of 2026, maybe, right? Which is why we have, we have a V3 coming, of course.
So on analysis of failure modes, I'm sure you saw the blog post that I read where it went through all of the different failure modes of, of O3, and of course, it was solution space prediction, which made it more surprising to me. My, my take on it was, I was really impressed that even when it failed, it was because the solution space got too big, or it was just getting minor mistakes, but broadly it, it got the direction of, of many of the problems quite well, uh, you know, similarly, tell me about the, the failure modes on V2, right?
So we were not able to test O3 as much, uh, on V2, uh, but I can tell you about failure modes based on what we saw on, on V1, and well, um, there are many, but generally this is a model where, uh, uh, reasoning abilities can, can decrease exponentially with, with problem size. If you have more objects, uh, in the scene, if you have more, more rules, more concepts, uh, interacting, you, you see this exponential decrease in, in capabilities. Um, it's also, you know, because it's a model that needs to, it works by writing a kind of natural language program that describes what it's seeing, that describes the problem, and the, the sequence of steps to solve it. So in that sense, it's 100% a natural language program, and that means that in order to solve a problem, it has to talk about it, uh, using words. And as a result, if you have a task where the rule is very simple to grasp for a human but in a nonverbal way, but it's very difficult to put it into words, it has no, uh, verbal analogy, uh, that's actually much harder to solve for, for, for this Chain of Thought model. Uh, other than that, we saw that, um, just one of the big challenges is, you know, compositionality: having multiple rules interact. There's also, it seems, there's a bit of a locality bias going on as well, where if you have to combine together bits of information that are spatially collocated together on the grid, that's easier for the model than if you have to do the exact same thing, but the two bits of information you have to synthesize are pretty distant. Um, so, uh, having to, to, yeah, combine together bits of information that are, that are separate, having to, so it seems as well that the, the model has a, a trouble, uh, simulating the execution of a rule and then reading the results. Like, for instance, if you're solving an Arc task and you grasp a certain rule, and then you start applying it, let's say it's like your, you're continuing, you're doing line continuation or something, and then, uh, you have to take another rule and use that rule to read a bit of information that you have written in the process of executing the first rule, that sort of thing is completely out of reach for the Chain of Thought models.
How multi-dimensional do you think intelligence is? You know, um, one school of thought, and, and I think you, you might subscribe to this, is that the, the universe is kind of almost made up of platonic rules that, that are disconnected from the, the world that, that we live in, and then there's this kaleidoscope idea you talk about, and they get combined together, and, and that's what we see. But another school of thought is that there will always be another dimension of intelligence; we'll always need Arc V4, V5, V6, and there'll always be something missing. Each step of generality that you cross, you gain a nonlinear amount of capabilities, right? And so after a few steps, you are so overwhelmingly superhuman across every possible dimension that, yeah, yeah, you can, you can say without a doubt that you have, you have AI; in fact, in fact, you have superintelligence. Uh, but yeah, intelligence is, in a sense, multi-dimensional, uh, and what Arc is trying to capture is really just this, uh, fluid intelligence aspect, this, to recombine, uh, core knowledge building blocks. So in my definition of intelligence, intelligence is about efficiently, uh, acquiring, uh, uh, skills and, and knowledge and, uh, recombining them, uh, to, well, again, efficiently, efficiently recombining them to adapt to novel tasks, to novel situations that you cannot prepare for explicitly. Um, and purely the ability to take a bunch of building blocks and, and recombining them, doing program synthesis, that's one aspect of that, that's probably the most central aspect, which is what, you know, this is why we, we're focusing on with Arc, but it's not the only aspect because this is assuming that you already have, uh, this, this pile of knowledge available. So it's not, it's overlooking, uh, the acquisition of, of these abstractions; it's also overlooking the acquisition of information about the task. In Arc, you provided all the information about the task at once, but in the real world, you have to collect that information; you have to take actions, set goals, uh, to discover what, what your environment is even about, what you can do within it, and you have to do these things efficiently, of course, and, and that efficiency aspect is very important because, you know, intelligence was developed, uh, by, by evolution; it's in evolution, adaptation. And when you're exploring the world, uh, you are taking on some risk, you know, you might get killed by, by a predator, for instance, and so you want to be, being, you want to gain the maximum amount of information and thereby power over your environment by taking on, uh, a minimum amount of risk and expanding minimum amount of energy. That's not something you can measure; that's not something we can capture with Arc, uh, V1 or V2 alone.
Can you just expand on the significance of the solution space prediction with O3, because that rather suggests to me that it's almost this rich-get-richer idea where it's nearly a blank slate, and it's very empiricist, and we just take the data in, and the neural network does all of the things. Um, I always imagine that we would need to have some kind of structured approach which took, you know, the core knowledge into account. Do, do you think that, you know, it, it's actually simpler than, than we thought, trying to directly predict the output versus trying to, uh, write down the steps to get the output?
They're not entirely separate things because, of course, once you've written down the steps, you can do what looks like transduction. And O3 is not actually a real transduction model because it's a, it's, it's much closer to a program synthesis model where the, the, it's searching for the right Chain of Thought to describe the task and, and list the sequence of steps to solve it. And once you have the Chain of Thought, you can just use the model to execute it, and that gives you the output. So from the outside, if you treat the entire system as a black box, it looks like transduction, but the same would be true of any program search system. What it's actually doing, and the reason why it's able to adapt to novelty so well, uh, it's, it's because it's synthesizing, uh, this, this Chain of Thought, which, which serves as a, a recombination artifact for, for the, the knowledge and, and, and the skills that the model has; that recombination artifact is adapted to the particular task at hand. So it's much closer to Bron's model. This is something that the community…
Found very confusing, because in the last interview you you were describing, I think it was 01 Pro, as being a kind of explicit search process. What seems to be the case is that it's, you know, there is some kind of reinforcement learning thing in in the pre-training, and then it maybe does some sampling at inference time. So you know we're doing a whole bunch of completions, and are you saying it's as if it's doing a program search, or are you saying it's somehow explicitly doing a a Chain of Thought, you know, program search? It's it's uh searching over the space of possible chain of thoughts and finding the one that seems most appropriate. So in that case, it's entirely analogous to a program search system where the program you're synthesizing is a natural language program, right, a program written in English. Okay. It just seems a bit strange, doing Auto regression on a language model, how that could be characterized as a search process. It's so uh a model like 01 Pro for instance, or 03 uh is not just autoaggressive; it has actually this test time search step, which is why it can adapt to novelty much much better than the base models are purely autoaggressive. That's why again you see on on so in general Arc, even Arc one has completely resisted the uh pre-training, the purely autoaggressive pre-training scaling Paradigm like from 2019 to 2025. We scaled up these models by like 50,000 x, like from from GPT-2 to GPT 4.5, and even on Arc one you went from 0% to something like 10%, and on Arc two you you're going from 0% to 0%, right? And meanwhile, if you have any system that's actually capable of doing test time adaptation like test time search like 01 Pro uh or or 03, then you're getting uh uh you're getting much much better performance there, this huge performance Gap. So generally you can tell the difference between a model that does not do test time adaptation and the model that does by looking at this performance Gap, this generalization gap on Arc, also by looking at latency and by looking at cost. So of course the model does time search is going to give you your answer uh um it's going to take much longer, like if you look at 01 Pro for instance, it's taking 10 minutes uh to answer your queries, and it's going to cost you much more as well because of all this work he's doing.
So I could download the Deep C car1 model and I'm running it on my machine, and as far as my machine is concerned it's just a normal LLM; it's doing greedy sampling, autoregressive, that's right, which is why which is way does not adapt to novelty and it scores basically zero on Arc, or like 1% maybe. Oh, so so so you're saying that is something different about 03 is qualitatively different.
That's that's correct; it is qualitatively different from all the other models that came before it. It is actually a model that has fluid intelligence; it has nonzero amount of fluid intelligence, and and and R1 for instance is not. Okay, so categorically it's doing some kind of active search process at inference; that's what it looks like. So of course I don't actually know how it works, but that's that's what I would speculate it looks like. Yes, and you see it in the latency, in the cost, and of course the Arc performance. Would you be shocked and surprised if if it came to light that it was just doing like autoaggressive greedy sampling? Honestly, I think it's very very likely because it's completely incompatible with the the characteristics of the system that we know of, that that we were exposed to when we tested 01. Awesome. And do you think that there will always be human gaps? Probably not always. Uh today they are very clear, very significant gaps, right? Like we're not we're not actually that close to AI right now, but eventually, you know, as we get closer and closer there will be fewer and fewer gaps, and at some point uh you know we're going to have AI system that are just overwhelmingly superhum along every possible axis you choose to look at. So you know I don't think that there will be gaps forever, right? Tim, thank you so much for doing this; appreciate it; looking forward to seeing it in a couple weeks.