📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Grok 4 is really smart... Like REALLY SMART

Matthew Berman19:41

Transcription

Grok 4 just dropped, and yes, Elon was right. It is the smartest model in the world, at least currently, and it is a pretty significant leap from other Frontier models. So, first, let me walk you through the progression of the Grok series of models. This was a slide from last night's live stream. We can see Grok 2, which, by the way, was only like 2 years ago, and we have it right here. It was just next token prediction. Here's the amount of compute. And with Grok 3, they 10xed their pre-training compute, and it was a really good model. Then they had Grok 3 reasoning, which took the pre-training compute, and in yellow right here, we see the reinforcement learning compute. But then the massive jump to Grok 4 reasoning. This is what Grok 4 is all about: reinforcement learning. And we've talked a lot about that on this channel, so hopefully, this is no surprise to you.

Here we have pre-training, and here we have post-training. They put a ton of compute behind reinforcement learning, and this is really the power of reinforcement learning with verifiable rewards. And the verifiable rewards is the important part. These are problems that have known solutions. The most basic example: 2 plus 2 is the problem. Four is the solution. If we're using that problem and solution to train a model, we can tell it, "Okay, try to figure out what 2 plus 2 is." And when you get the answer of four, we're going to give you a reward for it. And now, do that a bunch of times over a bunch of very hard problems, and the models are getting really good. This is also what elicits the thinking behavior from these models: this reinforcement learning with verifiable rewards paradigm. And so, if we thought there was a wall with RLVW, Grok blew right through it. In fact, reinforcement learning with verifiable rewards was so critical to their workflow, they started running out of problems. They were actually struggling to find enough problems that we've written down with rewards that we know in the world.

And that's when Elon Musk started talking about reality being the ultimate test. These models are great; they can pass these benchmarks really well, but we are limited by how many problem and answer sets we can give them because there just are a limited amount in the world. But when you put these models in the real world, and usually that's going to come in the form of a humanoid robot or something that can interact with physics, that's when we have essentially unlimited verifiable rewards.

All right, so now let's get into the benchmarks. The first benchmark they talk about is Humanity's Last Exam, and this is a very hard benchmark. These are frontier knowledge questions that, if you can imagine, only an expert or team of experts would ever be able to get right in a single domain in this exam. However, this is an exam that spans math, physics, biology, social science, computer science, engineering, chemistry, and other. So, picture this: the smartest PhD postdoc in the world and a team of them working over hours, days, weeks would be maybe a few of the questions in a single domain, right? Grok 4, on the other hand, well, let me just show you. And before we get into that, if you want to learn how to get the most out of Grok 4, definitely download the similarly named Humanity's Last Prompt Engineering Guide, created by myself and my team. It is completely free. Download it today. Link in the description below.

And the way that they revealed the scores for Humanity's Last Exam for Grok 4 was pretty awesome. They showed the progression of different features and different abilities given to Grok 4 and then what it was able to accomplish. And let's walk through that. So, here are Humanity's Last Exam's top scores based on current Frontier models: Gemini 2.5 Pro coming in at number one at 21.6%, Claude 3 Opus at 20%, Claude 3 Sonnet at 18%. All good scores, right around the same score, though. Now, switching back to Grok 4 with no tool usage. Grok 4, 26.9%, already substantially ahead of the other frontier models. But it doesn't end there. Then they gave Grok 4 tool usage. This is things like web browsing, more sophisticated memory, an environment in which it could write and execute code. And with that, it was able to achieve a 41%. That is a massive improvement from 26.9% and double what the next inline best model can achieve. But that's not it. Then when they scaled up test time compute, they reached 50.7%. So, with tool usage and scaling up test time compute, 50.7%, breaking the 50% barrier and really just demolishing all the other models that have been tested against this benchmark.

But what does scaling up test time compute actually mean? Previously, my association with test time compute was just giving it more time to think, having it output a bunch of chain of thought, and then from that, coming up with the best possible answer. But Grok 4 seems to have taken it a slightly different direction. What they do, and this is specific to something called the Grok 4 Heavy version, is they spawn multiple agents. Each one of those agents goes out, tries to solve the problem, they actually work together, they share notes. When one of them figures out something that works, they share it with the other ones, and each one of them gets better. And then at the end, they select whichever answer, whichever solution is best. And with all of that, they got the 50.7% number. And by the way, if you want to test out Grok 4 easily, check out our sponsor, Abacus. If you're like me, you probably have subscriptions to a bunch of different AI services and you jump between them all the time, and it's kind of frustrating and not only that, pretty expensive. And that is where Chat LLM by Abacus AI comes in. It is an all-in-one AI platform that includes all of the latest and best models from the leading model providers. And they also have something called Route LLM, which automatically picks the best model to send your prompt to, dependent on the actual prompt. So, it is routing your prompt to the right LLM. And of course, you can also chat with PDFs. So, download any documents that you want and easily ask questions, extract insights, gather data, whatever you need from your existing documents. And not only that, they also have text-to-image and text-to-video models. So you can generate awesome images, awesome videos easily. They also recently introduced Deep Agent, which is an incredibly powerful AI agent that can basically do anything. So, building websites, building apps, creating presentations, research reports, chatbots, or even building games. And all of this for just $10 per month. So, check it out, chatlm.abacus.ai, or click the link in the description. Let them know I sent you. Much appreciated. Thank you again to Abacus AI. Now, back to the video.

And by the way, I already paid for Grok 4 Heavy. Let me show it to you really quickly. All right, so here's Grok 4 Heavy, and I'm actually going to give it one of the math problems from Humanity's Last Exam. And let me just tell you up front, I don't even understand what this question is asking. The point right now is I just want to show you these multiple agents being spawned and coming back with an answer. I'm going to be doing full tests in another video. So, here is "Compute the reduced 12-dimensional spin board of the classifying space of the..." I can't even read this. Okay, so here's the question. Let's kick it off. And now, here we go. We can see four agents have spun up. They're initializing, and each of the four agents are now off and running their own solutions. So, this could take a while. I just wanted to show it to you really quick, just to show you the interface. I actually think the UI is really cool looking. But yeah, that's what Grok 4 looks like. It is spinning up multiple agents, allowing them to go out. Each of the agents is sharing their knowledge, and then they come back with the best answer. So, as you're thinking about the naming conventions, you think about Grok 4 is the single agent version, and Grok 4 Heavy is the multi-agent version. And it's not cheap. I'll tell you about the pricing later.

And they also showed off a few really cool demos during the live stream. Let me show you some brief clips from that. First, during the demo live, they had Grok 4 predict who was going to win the World Series, and it gave it all the tools and all the compute it needed. Take a look at that.

"Everyone knows Poly Market. Um, it's extremely interesting. It's the, you know, seeker of truth. It aligns with what reality is most of the time. And with Grok, what we're actually looking at is being able to see how we can try to take these markets and see if we can predict the, the future as well. So, as we're letting this run, we'll see how uh Grok 4 Heavy goes about uh predicting the, you know, the World Series odds for like the current teams in the MLB. And we can see uh here we can see all the tools and the process it used to actually uh go through and find the right answer. So, it browsed a lot of odd sites. It calculated its own odds comparing to the market, the market to find its own alpha and edge. It walks you through the entire process here, and it calculates the odds of the winner being like the the Dodgers uh and it gives them a 21.6% chance of uh winning uh this year. So, and it took approximately 4 and a half minutes to compute."

Next, they had Grok 4 make a visualization of what two black holes look like when they collide.

"We asked it to generate a visualization of two black holes colliding. Um, and of course, you know, it took some, there are some liberties. It's in my case actually pretty clear in its thinking trace about what these liberties are. Uh, for example, in order for it to actually be visible, you need to really exaggerate the the scale of the, you know, the uh the uh the waves. And yeah, so here's like, you know, this kind of inaction. Um, it exaggerates the scale in like multiple ways. uh it drops off a bit less in terms of amplitude um over distance and um but yeah we can kind of see uh the basic effects that, you know, are actually like, you know, correct. It starts with the inspiral, it merges, uh, and then you have um the ring down, and like this is basically um largely correct um yeah uh modulo some of the simplifications that need to do um, you know, it it's actually quite explicit about this, you know, it uses like post-post Newtonian approximations instead of actually like computing the general relativistic effects at like, uh, near the center of the black hole, which is, you know, incorrect, um, and, you know, will lead to, you know, some incorrect results, but the overall, you know, visualization is, uh, yeah, is basically there."

And of course, Grok, what it's really known for, or at least what I really love it for, is real-time information. So, here is Grok 4 going out and getting all of the announcements and the timeline of announcements of model scores released for Humanity's Last Exam. Take a look.

"Let's uh create a timeline based on expost uh detailing the, you know, changes in the scores over time, and we can see, you know, all the conversation that was taking place at that time as well. So, we can see who were the, you know, announcing scores and like what was the reactions at those times as well. And we can see like, you know, it defines the date that like Dan Hendricks had initially announced it. We can go through, we can see, you know, OpenAI announcing their score back in, uh, February. And we can see, you know, as progress happens with like Gemini, we can see like Kimmy, uh, and we can also even see, you know, the leaked benchmarks of, uh, what people are saying is, you know, if it's right, it's going to be pretty impressive. So, pretty cool."

Now, some more benchmarks. Take a look at this. Here's GPQA. Here's Grok 4 with no tools, 87%, and Grok 4 Heavy, I assume with tools, 88.9%, as compared to the next best model, 86%. Not a huge jump. Claude 3 Opus 2025 with Grok 4 Heavy scored a perfect 100%. This is insane. These are some of the hardest math questions in the world. A perfect 100 score. Claude 3 Opus actually did quite well, 98.4%. Here's Live CodeBench at 79.4%. So, a really good coder. Gemini 2.5 Pro at 74%, which is, in my opinion, the best coder, but I haven't tested Grok 4 yet. So, we'll see. Here's Math Arena, 96.7%, and USA MO, which is the Math Olympiad test. You can just see Grok 4 Heavy is demolishing the other models.

All right, switching back really quick. I just wanted to show you the progress. We're 5 minutes 48 seconds into Grok 4 Heavy trying to solve this problem. We're about halfway through, if this progression bar is accurate, and we can just see it running and running and running. Now, unfortunately, I can't see chain of thought. You can only see the progress of each agent.

All right, next, ARC AGI. This test is built to be easy for humans to solve but really difficult for AI to solve. It's essentially looking for patterns, learning many skills from those patterns, and then applying them to new tests. So, as we can see here, you kind of look at these different visualizations, learn how they're changing, and then try to figure out how this one would change based on the patterns you saw here. And Grok 4 absolutely crushed this test. So, here's the ARC AGI V1 coming in at 66.6% as compared to Claude 3 Opus, 60.8%, and ARC AGI V2, 15.9%. Double. Claude 3 Opus, 4 second place. So, you can see right here, just in a league of its own on this benchmark, and it was independently tested. So, Greg Cameron, the president of the ARC Prize, said, "We got a call from XAI 24 hours ago. Let's test it." They walked them through their testing policy: no data retention, model checkpoints must be intended for public use, and temporary increase in rate limits. And then let's look at his take on it. "Grok 4 is now the top-performing publicly available model on Arc AGI. This even outperforms purpose-built solutions submitted on Kaggle. The top previous score, 8% by Claude 3 Opus. Below 10% is noisy. Getting 15.9 breaks through that noise barrier. Grok 4 is showing non-zero levels of fluid intelligence. Absolutely crazy. This is true generalization."

But again, all of these are kind of ethereal benchmarks. They're not real. They're not in the real world. So, they tested it against this new benchmark called Vending Bench. And these models are essentially put in charge of managing a vending machine in the real world. And so they're given a budget, they're given inventory, everything. And here are the results. So, Claude 3 Opus has a net worth at the end of the test of about $1,800. Gemini 2.5 Pro has a net worth of about $789. A human comes in at $844. Claude Opus 4, which was a big jump, about $2,000. But Grok 4 coming in at $4,700. So, this again is a real-world exam. This is how is it going to interact and how is it going to actually perform on a real-world test. So, very impressive.

And over the last few months, the XAI team has talked a lot about AI creating video games. Elon Musk has said, "We're going to be creating AAA video games in the near future." Whatever you think about his timelines. But they gave a Vibe Coder access to Grok 4 and said, "What can you make in just a few hours?" Here's what that looks like.

"So, Denny is actually a video game designer on X. So, uh, you know, we mentioned, hey, who wants to try out some uh uh Grok 4 uh preview APIs uh to make games, and Danny answered the call. Uh, so this was actually just made a first-person shooting game in a span of four hours. Uh, so, uh, some of the actually the unappreciated hardest problem of making video games is not necessarily encoding the core logic of the game, but actually go out, source all the assets, all the textures of files, and and, you know, to create a visually appealing game. So, one of the core aspects Grok does really well with all the tools out there is actually able to automate these like asset sourcing capabilities. So, the developers, you can just focus on the core development itself rather than like, you know, so now you can run a, you know, entire game studios with a game of one with like one person, and then you can have Grok 4 to go out and source all those slot assets, do all the maintaining tasks for you. So, it's a pretty cool game, a shooter, lots of cool graphics, lots of different rules and logic, and it just looks really cool. So, very good."

Now, Elon Musk said, "I would expect the first really good AI video game to be next year." I don't really believe that. I think these games are fun, but they're definitely like one-off games. We're not going to see an Assassin's Creed. We're not going to see the next Halo being created by AI. Not yet. And certainly not by the end of next year. And specifically, Elon talked about, so it has to have very good video understanding so it can play the games and interact with the games and actually assess what whether a game is fun and and actually have good judgment for whether a game is fun or not. And that gets into the realm of taste. And in my opinion, at least for the foreseeable future, taste is the realm of humans. Humans are the best at curating experiences for themselves and other humans. That's why I actually do think humans are going to be in the loop for a good bit of time longer.

So, if you want to test out Grok 4, it's available today, and it's also available via the API. So, hopefully, they're going to be plugging into all of the agentic coding applications out there. That would be awesome. It has a 256k context window, multimodal reasoning, real-time data search, and enterprise-grade security, which I don't exactly know what that means, but fine. But it is not cheap. For Super Grok, it's $30 a month. So, more expensive than ChatGPT Plus, a Claude subscription. And Super Grok Heavy is $300 a month or $3,000 a year. And with that, you get everything in Super Grok. You get Grok 4 Heavy, higher rate limits, and early access to new features.

So, coming back one more time, it's still running. We're at almost 15 minutes, and it still looks like three of the four agents are not close. So, this is really long horizon thinking, but remember to subscribe because I'm going to be testing Grok 4 thoroughly.

All right, so last, what comes next? What can we expect? Well, Elon said Grok 4 is currently based on their foundation model version 6 and their in-progress training version 7 right now, which should be done by the end of the month, and that will improve multimodal reasoning and understanding. So, we have the Grok 4 release, which just happened. They're going to have a coding-specific model coming out in August, a multimodal agent in September, and a video generation model in October. We'll see if these timelines hold. I'm not going to hold my breath, but I'm really excited. I have more videos coming on this, so stay tuned. If you enjoyed this video, please consider giving a like and subscribe.