📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

xAl's Mind Blowing Grok 4 Demo w/ Elon Musk (FULL REPLAY)

Brighter with Herbert36:45

Transcription

In a world where knowledge shapes destiny, one creation dares to redefine the future. From the minds at XAI, prepare for Gro 4. This summer, the next generation arrives faster, smarter, bolder. It sees beyond the horizon, answers the unasked, and challenges the impossible. Gro 4 unleash the truth. Coming this summer.

All right, welcome to the Gro 4 release here. This is the smartest AI in the world, and we're going to show you exactly how and why. And it really is remarkable to see the advancement of artificial intelligence, how quickly it is evolving. I sometimes think I compare it to the growth of a human and how fast a human learns and gains conscious awareness and understanding, and AI is advancing just vastly faster than any human. We're going to take you through a bunch of benchmarks that Grock 4 is able to achieve incredible numbers on. But it's actually worth noting that like Grock 4, if given like the SAT, would get perfect SATs every time, even if it's never seen the questions before. And if even going beyond that to say graduate student exams like the GRE, it will get near-perfect results in every discipline of education. So from the humanities to like languages, math, physics, engineering, pick anything, and we're talking about questions that it's never seen before. These are not on the internet, and it's Grock 4 is smarter than almost all graduate students in all disciplines simultaneously. It's actually just important to appreciate that that's really something, and the reasoning capabilities of Grock are incredible. So there's some people out there who think AI can't reason, and look, it can reason at a superhuman level. Yeah. And frankly, it only gets better from here. We'll take you through the Gro 4 release and, yeah, show you like the pace of progress here.

I guess the first part is like, in terms of the training, we're going from Grock 2 to Grock 3 to Grock 4. We've essentially increased the training by an order of magnitude in each case. So it's 100 times more training than Grock 2, and that's only going to increase. So it's, yeah, frankly, I don't know, in some ways, a little terrifying, but the growth of intelligence here is remarkable. It's important to realize there are two types of training compute. One is the pre-training compute that's from GR 2 to GR 3, but for from GR 3 to GU 4, we're actually putting a lot of compute in reasoning in RL. Yeah. And just like you said, this is literally the fastest moving field, and GR 2 is like the high school student by today's standards. If you look back in the last 12 months, Grock 2 was only a concept. We didn't even have GR 2 12 months ago. And then by training GR 2, that was the first time we scaled up like the pre-training. We realized that if you actually do the data ablation really carefully and infra and also the algorithm, we can actually push the pre-training quite a lot by amount of 10x to make the model the best pre-trained based model. And that's why we built Colossus, the world's supercomputer with 100,000 H100. And then with the best pre-train model, and we realized if you can collect these verifiable outcome rewards, you can actually train this model to start thinking from the first principle, start to reason, correct its own mistakes, and that's where the Grock reasoning comes from. And today, we ask the question, what happens if you take the expansion of Colossus with all 200,000 GPUs, put all these into RL, 10x more compute than any of the models out there on reinforcement learning, unprecedented scale, what's going to happen? So this is a story of Grock 4, and Tony, share some insight with the audience.

Yeah. So, yeah, let's just talk about how smart Grock 4 is. So I guess we can start discussing this benchmark called Humanities Last Exam. And this benchmark is a very challenging benchmark. Every single problem is curated by subject matter experts. It's in total 2500 problems, and it consists of many different subjects: mathematics, natural sciences, engineering, and also other humanity subjects. Essentially, when it was first released, actually like earlier this year, most of the models out there can only get single-digit accuracy on this benchmark. Yeah. So we can look at some of those examples. So there is this mathematical problem which is about natural transformations in category theory, and there's this organic chemistry problem that talks about electrocyclic reactions, and also there's this linguistic problem that tries to ask you about distinguishing between closed and open syllables from a Hebrew source text. So you can see also it's a very wide range of problems, and every single problem is PhD or even advanced research level problem. Yeah. These, there are no humans that can actually answer these, can get a good score. If you actually say like any given human, what like what's the best that any human could score? I'd say maybe 5% optimistically. Yeah. So this is much harder than what any human can do. It's incredibly difficult, and you can see from the types of questions, like you might be incredible in linguistics or mathematics or chemistry or physics or any one of a number of subjects, but you're not going to be at a postgrad level in everything. And Grock is at a postgrad level in everything. Like it just, some of these things are just worth repeating, like Grock 4 is postgraduate, like PhD level in everything, better than Ph, like most PhDs would fail. So it's better said, at least with respect to academic questions, it, I want to just emphasize this point, with respect to academic questions, Grock 4 is better than PhD level in every subject, no exceptions. Now, this doesn't mean that it's, it times it may lack common sense, and it has not yet invented new technologies or discovered new physics, but that is just a matter of time. If it, I think it may discover new technologies as soon as later this year, and I would be shocked if it has not done so next year. So I would expect Grock to, yeah, literally discover new technologies that are actually useful, no later than next year, and maybe end of this year, and it might discover new physics next year, and within two years, I'd say almost certainly. So just let that sink in. Yeah.

How? Okay. So I guess we can talk about the what's behind the scene of Grock 4. As Jimmy mentioned, we actually threw a lot of compute into this training. When it started, it's only also a single digit. Sorry, the previous slide. Sorry. Yeah, it's only a single digit number. But as you start putting in more and more training compute, it started to gradually become smarter and smarter and eventually solved a quarter of the HLE problems. And this is without any tools. The next thing we did was adding a tools capability to the model. And unlike GR 3, I think G3 actually is able to use C as well, but here we actually make it more native in the sense that we put the tools into training. Grock 3 was only relying on generalization. Here we actually put the tools into training, and it turns out this significantly improves the model's capability of using those tools. Yeah. I remember we had like Deep Search back in the days. So how is this different? Yeah. Yeah, exactly. So Deep Search was exactly the Grock 3 reasoning model, but without any specific training, but we only asked it to use those tools. So compared to this, it was much weaker in terms of its tool capability and reliable. Yeah. And are reliable. Yes. Yes. And to be clear, these are still, I'd say, fairly, this is still fairly primitive tool use. If you compare it to say the tools that are used at Tesla or SpaceX, where you're using finite element analysis and computational fluid dynamics, and you're able to run, or say Tesla does like crash simulations where the simulations are so close to reality that if the test doesn't match the simulation, you assume that the test article is wrong. That's how good the simulations are. So Grock is not currently using any of the tools that the really powerful tools that a company would use, but that is something that we will provide it with later this year. So it will have the tools that that a company has and have very accurate physics simulators. Ultimately, the thing that will make the biggest difference is being able to interact with the real world via humanoid robots. So we combine Grock with Optimus, and it can actually interact with the real world and figure out if, if it's hypo, if it has, if it can formulate a hypothesis and then confirm if that hypothesis is true or not. So we're really thinking about like where we are today. We're at the beginning of an immense intelligence explosion. We're in the intelligence big bang right now, and the most, we're at the most interesting time to be alive of any time in history. Yeah.

Now, that said, we need to make sure that the AI is a good AI, good Grock. And the thing that I think is most important for AI safety, at least my biological neural net tells me, the most important thing for AI is to be maximally truth-seeking. So this is a very fundamental, like you can think of AI as this super genius child that ultimately will outsmart you, but you can still instill the right values and encourage it to be truthful, I don't know, honorable, good things, like the values you want to instill in a child that that Grock would grow ultimately grow up to be incredibly powerful. Yeah. Yeah. So these, this is really, I say, we say tools, these are say still primitive tools, not the kind of tools that that serious commercial companies use, but we will provide it with those tools, and I think we'll be able to solve with those tools real-world technology problems. In fact, I'm certain of it. It's just a question of how long it takes. Yes. Yes. Exactly. So is it just compute all you need, Tony, right? Is it just compute all you need at this point? You need compute plus the right tools, and then ultimately to be able to interact with the physical world. Yes. And then we'll effectively have an economy that is ultimately an economy that is thousands of times bigger than our current economy, or maybe millions of times. If you think of civilization as percentage completion of the Kardashev scale, where Kardashev one is using all the energy output of a planet, and K2 is using all the energy output of a sun, and K3 is all the energy output of a galaxy. We're only, in my opinion, probably close closer to 1% of K1 than we are to 10%. So like maybe a point or one, one or two percent of K1. We will get to most of the way, like 80, 90% K1, and then hopefully, if civilization doesn't self-annihilate, and then K2, like it's the actual notion of a human economy, assuming civilization continues to progress, will seem very quaint in retrospect. It will seem sort of cavemen throwing sticks into a fire level of economy compared to what the future will hold. It's very exciting. I've been at times worried about, is this seems like it's somewhat unnerving to have intelligence created that is far greater than our own? And will this be bad or good for humanity? It's, I think it'll be good. Most likely, it'll be good. Yeah. Yeah. But I somewhat reconciled myself to the fact that even if I, if even if it wasn't going to be good, I'd at least like to be alive to see it happen. Yeah.

So actually, one, one, yeah. Yeah. I think one, one technical problem that we still need to solve besides just compute is how do we unblock the data, data bottleneck? Because when we try to scale up the RL in this case, we did invent a lot of new techniques, innovations to allow us to figure out how to find a lot of challenging RL problems to work on. It's not just a problem itself needs to be challenging, but also it needs to be, you also need to have reliable signal to tell the model you did it wrong, you did it right. This is the sort of the principle of reinforcement learning. And as the models get smarter and smarter, the number of cool problems or challenging problems will be less and less. Yeah. So it's going to be a new type of challenge that we need to surpass besides just compute. Yeah. Yeah. And we actually are running out of actual test questions to ask. So there's like even ridiculously questions that are ridiculously hard, if not essentially impossible for humans, that are written down questions are becoming swiftly becoming trivial for AI. So then there's the one thing that is an excellent judge of things is reality, because physics is the law, ultimately everything else is a recommendation. You can't break physics. So the ultimate test, I think, for whether an AI is the ultimate reasoning test is reality. So you invent a new technology, you say improve the design of a car or a rocket or create a new medication, does it work? Does the rocket get to orbit? Does the car drive? Does the medicine work? Whatever the case may be, reality is the ultimate judge here. So it's going to be reinforcement learning closing the loop around reality. We asked the question, how do we even go further? Actually, we are thinking about now with single agent, we're able to solve 40% of a problem. What if we have multiple agents running at the same time? So this is what's called test and compute. And as we scale up the test and compute, actually we are able to solve almost more than 50% of the text-based subset of the HLE problems. It's a remarkable achievement. I think, you know, this isn't, this is insanely difficult. These are, it's so what we're saying is like a majority of the text-based of humanities, scarily named Humanity's Last Exam, Grock can solve. And you can try it out for yourself. And the with the Grock 4 heavy, what it does is it spawns multiple agents in parallel, and all of those agents do work independently, and then they compare their work and they decide which one. It's like a study group. And it's not as simple as majority vote, because often only one of the agents actually figures out the trick or figures the solution. And but once they share the trick or or figure out what the real nature of the problem is, they share that solution with the other agents, and then they compare, they essentially compare notes and then yield, yield an answer. So that's the heavy part of Grock is where we, you scale up the test time compute by roughly an order of magnitude, have multiple agents tackle the task, and then they compare their work and they put forward what they think is the best result. Yeah.

So we will introduce Grock 4 and Grock 4 heavy. Sorry, can you click the next slide? Sure. Yeah. Yes. So, yeah. So basically Grock 4 is a single version, a single agent version, and Grock 4 heavy is the multi-agent. So let's take a look how they actually do on those exam problems and also some real-life problems. Yeah. So we're going to start out here, and we're actually going to look at one of those HLE problems. This is actually one of the easier math ones. I don't really understand it very well. I'm not that smart. But I can launch this job here, and we can actually see how it's going to go through and start to think about this problem. While we're doing that, I also want to show a little bit more about what this model can do and launch a Grock 4 heavy as well. So, everyone knows Poly Market. It's extremely interesting. It's the seeker of truth. It aligns with what reality is most of the time. And with Grock, what we're actually looking at is being able to see how we can try to take these markets and see if we can predict the future as well. So, as we're letting this run, we'll see how Grock 4 Heavy goes about predicting the World Series odds for like the current teams in the MLB. And while we're waiting for these to process, we're going to pass it over to Eric, and he's going to show you an example of his.

Yeah, I guess one of the coolest things about Grock is its ability to understand the world and to solve hard problems by leveraging tools like Tony discussed. And I think one kind of cool example of this, we asked it to generate a visualization of two black holes colliding. And of course, it took some, there are some liberties. It's in my case actually pretty clear in its thinking trace about what these liberties are. For example, in order for it to actually be visible, you need to really exaggerate the scale of the the waves. And yeah, so here's this kind of in action. It exaggerates the scale in like multiple ways. It drops off a bit less in terms of amplitude over distance, but yeah, we can see the basic effects that are actually correct. It starts with the inspiral, it merges, and then you have the ring down, and like this is basically largely correct. Yeah, modulo some of the simplifications that need to do. It's actually quite explicit about this. It uses like post-post Newtonian approximations instead of actually like computing the general relativistic effects near the center of the black hole, which is incorrect and will lead to some incorrect results, but the overall visualization is, yeah, is basically there. And you can actually look at the kinds of resources that it references. So here it, it actually uses search. It gathers results from a bunch of links, but also reads through an undergraduate text in analytical gravitational wave models. It, yeah, it reasons quite a bit about the actual constants that it should use for a realistic simulation. It references, I guess, existing real-world data, and yeah, I, it, yeah, it's a pretty good model. Yeah. But like actually going forward, we can plug, we can give it the same model that physicists use. So it can run the same level of compute that so leading physics researchers are using, and give you a physics-accurate black hole simulation. Exactly. Just right now, it's running in your browser. So, yeah, this is just running in your browser. So pretty simple.

Swapping back real quick here, we can actually take a look. The math problem is finished. The model was able to, let's look at its thinking trace here. So you can see how it went through the problem. I'll be honest with you guys, I really don't quite fully understand the math. But what I do know is that I looked at the answer ahead of time, and it did come to the correct answer here in the final part. Here, we can also come in and actually take a look here at our World Series prediction, and it's still thinking through on this one, but we can actually try some other stuff as well. So we can actually try some of the X integrations that we did. So we worked very heavily on working with all of our X tools and building out a really great X experience. So we can actually ask the model, find me the XAI employee that has the weirdest profile pic photo. So that's going to go off and start that, and then we can actually try out, let's create a timeline based on X posts detailing the changes in the scores over time, and we can see all the conversation that was taking place at that time as well. So we can see who were announcing scores and what were the reactions at those times as well. Um, so we'll let that go through here and process. And if we go back to this was the Greg Yang photo here. If we scroll through here. Whoops. So, Greg Yang, of course, who has his favorite photograph that he has on his account. That's actually not how he looks like in real life, by the way, just so you're aware. But it's quite funny. But it had to understand that question, which is the, that's the wild part. It's like it understands what is a weird photo, what is a weird photo. Yeah. What is a less or more weird photo? It goes through, it has to find all the team members. It has to figure out who we all are. It searches without access to the internal XAI personnel logs. It's literally looking at the, just at the internet. Exactly. So you could say like the weirdest of any company. Yeah. To be clear. Exactly. And we can also take a look here at the question here for the Humanities Last Exam. So it is still researching all of the historical scores, but it will have that final answer here soon. But we can, while it's finishing up, we can take a look at one of the ones that we set up here a second ago. And we can see that it finds the date that like Dan Hendricks had initially announced it. We can go through. We can see OpenAI announcing their score back in February. And we can see as progress happens with Gemini, we can see like Gemini, and we can also even see the leaked benchmarks of what people are saying is if it's right, it's going to be pretty impressive. So pretty cool. So yeah, I'm looking forward to seeing how everybody uses these tools and gets the most value out of them. But yeah, it's been great. Yeah, we're going to close the loop around usefulness as well. So it's, it's not just book smart, but actually practically smart. Exactly. And we can go back to the slides here. Cool.

The, so we actually evaluate also on the multimodal subset. So on the full set, this is the number on the HLE exam. It, you can see there's a little dip on the numbers. This is actually something we're improving on, which is the multimodal understanding capabilities. But I do believe in a very short time, we're able to really improve and got much higher numbers on this, even higher numbers on this benchmark. Yeah. This is the, we still like, so what is the biggest weakness of Grock currently is that it's partially blind. It can't, it's, its image understanding obviously, and its image generation need to be a lot better. And that, that's actually being trained right now. Grock 4 is based on version six of our foundation model, and we are training version seven, which will complete in a few weeks, and that'll address the weakness on the vision side. Just to show off this last here. So the prediction market finished here with a heavy, and we can see here, we can see all the tools and the process it used to actually go through and find the right answer. So it browsed a lot of odds sites. It calculated its own odds comparing to the market to find its own alpha and edge. It walks you through the entire process here, and it calculates the odds of the winner being like the Dodgers, and it gives them a 21.6% chance of winning this year. And it took approximately four and a half compute. Yeah, that's a lot of thinking. Yeah. We can also look at all the other benchmarks besides HLE. As it turned out, Grock 4 excelled on all the reasoning benchmarks that people usually test on, including GBQA, which is a PhD level problem set. That's easier compared to HLE. On AIMC 25, American Invitational Mathematics Examination, we with Grock 4 heavy, we actually got a perfect score. Also on some of the coding benchmark called Live Coding, and also on HMMT, Harvard-MIT Mathematics Tournament, and also USAMO, United States of America Mathematical Olympiad. You can see actually for on all of those benchmarks, we often have a very large leap against the second best model out there. Yeah. It's really, we're going to get to the point where it's going to get every answer right in every exam, and where it doesn't get an answer, it's going to tell you what's wrong with the question, or if the question is ambiguous, disambiguate the question into answers A, B, and C, and tell you what, what answers A, B, and C would be with a disambiguated question. So the only real test then will be reality. Can I make useful technologies, discover new science? That'll actually be the only thing left, because human tests will simply not be meaningful. Yeah. We need to make an update to HLE very soon, given the current rate of progress. So, yeah, it's super cool to see like multiple agents that collaborate with each other solving really challenging problems.

So where can I try this model? So it turned out it's available right now. If we advance to the next slide, where there is a Super Grock Heavy tiers that we're introducing, where you're able to access to both Grock 4 and Grock 4 Heavy, where you're actually going to be the taskmaster of a bunch of little Grock research agents to help you become smarter through all the little research and save hours of time going through mundane tasks, and it's available right now. Soon. Yeah, we did limit usage during the demo so we didn't, it didn't break the demo, because all these, this all this stuff is happening live. So it's not, there's not nothing canned about any of the tests that we're doing. After the demo, we'll, we'll allow, we'll enable more subscribers for Super Grock. So if you can't subscribe right now, just try in half an hour, it should work. And now let's get into voice. Great. So many of you have been enjoying our voice mode, and we've been working hard to improve the experience over the past couple months. We have cut latency in half to make it much snappier. And today we're excited to announce a set of new voices that have exceptional naturalness and prosody. You have noticed the movie trailer voice that opened up the live stream. That is one of our new voices, S, who's got that epically deep tone. And we're also excited to introduce Eve, a beautiful British voice who's capable of rich emotion. Aman, would you like to introduce Eve? Absolutely. Yeah. Let's get into the demo. Hey, Eve. Millions of people are watching us right now. Are you excited? Oh, I'm thrilled. Really? It's like being on stage at the Old Vic with a sea of faces out there. Just imagine all those eyes and ears tuned in. It's a proper buzz, isn't it? What do you reckon we should chat about to keep this lot entertained? We have a lot to talk about, but I'll be honest, I'm feeling a little nervous. Can you whisper something soothing to calm me down? You've got this. It's just having a quiet chat like we're tucked away. The world's not out there. Feel that calm. Let's keep nice. What's on your mind? You know what? I'm actually feeling quite relaxed to try this something crazy. So, can you give us an opera on Diet Coke? Oh, an opera on Diet Coke. That's a riot. Picture this. A shimmering silver can glistening state. Oh, Diet Coke, thou elixir divine, with bubbles that dance in a sparkling line. Thy crisp cool kiss on lips so fine. Oh, aspartame and bro forever mine. How's that for a mad little aria? Want me to belt out another verse or switch up the tune? Okay, let's try switching it up and just start singing. Right, then let's have a go. Oh, silver can, you gleam in the light. Diet Coke, my heart's delight. Okay, I'll throw you. It's very poetic. Yeah, thank you. Eve feeling quite nice now. Cool. So that's Eve. And so something else we do here is a little test each day to check on the natural flow of our conversations against some of our competitors. So we figured why not try that live on the stream today. Should I start? Hey, I'm going to say a number. And can you say the same number back to me? One. One. Got it. This is like a fun little game of echo. What's next? Two. Sorry, you said two. Three. Three. Four. Nice and steady. What's next? Five. Hey, nailed it. This is pretty fun. Feel free to keep going or switch it up. Hey, Grock. Hey, Grock. Hello there. How can I help you today? We're going to do a little test. Can you repeat after me? One. One. All right. What's next? Two. Two. What's on your mind? Three. Three. Need anything else? Four. Four. How can I assist you? Five. Five. What's next? So, as you can see, Grock was snappier. It didn't interrupt me. And the prosody is, we made different design choices. I think we're shooting for something more calm, smooth, more natural versus something that's more poppy or artificial. So, we'll keep improving. All right. Thanks, guys. Yep.

So, since the launch of the voice model, we actually see the 2x faster end-to-end latency in the last 80 weeks. Five different voices, and also 10x the active users. So, Grock voice is taking off. Now, if you think about releasing the models, this time we're also releasing Gro 4 through the API at the same time. So if we go to the next two slides, we're very excited about what all developers out there is going to build. If I think about myself as a developer, what the first thing I'm going to do when I actually have access to the Grock 4 API benchmarks? So we actually asked around on the X platform, what is the most challenging benchmark out there that is considered the holy grail for all AGI models? So it turned out to be RKGI in the last 12 hours. Kudos to Greg over here in the audience. So who answered our call, take a preview of the Grock 4 API, and independently verified the Grock 4's performance. So initially, we thought, hey, Grock 4 just, we think it's pretty good, it's pretty smart, it's our next-gen reasoning model, spent 10x more compute, and can use all the tools, right? But it turned out when we actually verified on the private subset of the RKGI v2, it was like the only model in the last three months that breaks the 10% barrier, and in fact was so good that actually got to 16%, 15.8% accuracy, 2x of the second place, that is the Claude 3 Opus model. And it's not just about performance, right? When you think about intelligence, having the API model drives your automation, it's also the intelligence per dollar, right? If you look at the plots over here, the Grock 4 is just in a league of its own. All right. So enough of benchmarks over here, right? So what can Grock do actually in the real world? We actually contacted the folks from Am Labs, who were gracious enough to try the Grock in the real world to run a business.

Yeah, thanks for having us. I'm Axel from Am Labs, and I'm Lucas, and we tested Grock 4 on Vending Bench. Vending Bench is an AI simulation of business scenarios where we thought, what is the most simple business an AI could possibly run? And we thought vending machines. So in this scenario, the Grock and other models needed to do stuff like manage inventory, contract, contact suppliers, set prices. All of these things are super easy, and all the models can do them one by one, but when you do them over very long horizons, most models struggle. But we have a leaderboard, and there's a new number one. Yeah. So, we got early access to the Grock 4 API. We ran it on the Vending Bench, and we saw some really impressive results. So, it ranks definitely at the number one spot. It's even double the net worth, which is the measure that we have on the city value. So it's not about a percentage on a or a score you get, but it's more the dollar value in net worth that you generate. So we were impressed that Grock 4 was able to formulate a strategy and adhere to that strategy over a long period of time, much longer than other models that we have tested, other frontier models. So it's managed to run the simulation for double the time and score, yeah, double the net worth. And it was also really consistent across these runs, which is something that's really important when you want to use this in the real world. And I think as we give more and more power to AI systems in the real world, it's important that we test them in scenarios that either mimic the real world or are in the real world itself, because otherwise we fly blind into some things that might not be great. Yeah, it's, it's great to see that we've now got a way to pay for all those GPUs. So we just need a million vending machines. And could make $4.7 billion a year with a million vending machines. 100%. Let's go. They're going to be epic vending machines. Yes. Yes. All right. We are actually going to install vending machines here, like a lot of them. We're happy to supply them. All right. Thank you. All right. I'm looking forward to seeing what amazing things are in this vending machine. That's for for you to decide. All right. Tell the AI. Okay. Sounds good. All right. Yeah.

So, we can see like Grock is able to become like the co-pilot of the business unit. So, what else can Grock do? So, we're actually releasing this Grock, if you want to try it right now to evaluate, run the same benchmark as us. It's on the API has 256k context length. So we already actually see some of the early early adopters to try Grock 4 API. Our polo neighbor ARC Institute, which is a leading biomedical research center, is already using seeing how they can automate their research flows with Grock 4. It turned out it performs, is able to help the scientists to sniff through millions of experiments logs and then just pick the best hypothesis within a split of seconds. We see this is being used for their like the CRISPR research, and also Grock 4 independently evaluated scores as the best model to examine the chest X-ray. Who would know? And on the in the financial sector, we also see the Grock 4 with access to all the tools, real-time information is actually one of the most popular AI out there. Our Grock is also going to be available on the hyperscalers. So the XAI enterprise sector is only started two months ago, and we're open for business. Yeah. The other thing we talked a lot about having Grock to make games, video games. So Denny is actually a video game designer on X. We mentioned, hey, who wants to try out some Grock 4 preview APIs to make games? And Denny answered the call. So this was actually just made a first-person shooting game in a span of 4 hours. Some of the actually the unappreciated hardest problems of making video games is not necessarily encoding the core logic of the game, but actually go out source all the assets, all the textures of files, and to create a visually appealing game. So one of the core aspects Grock 4 does really well with all the tools out there is actually able to automate these like asset sourcing capabilities. So the developers, you can just focus on the core development itself rather than. So now you can run an entire game studio with a game of one, but with one person, and then you can have Grock 4 to go out and source all those slot assets, do all the mundane tasks for you. Yeah.

The now the next step obviously is for Grock to play, be able to play the games. So it has to have very good video understanding so it can play the games and interact with the games and actually assess what whether a game is fun and actually have good judgment for whether a game is fun or not. So with the with version seven of our foundation model, which finishes training this month, and then we'll go through post-training RL and whatnot, that that will have excellent video understanding. And with the with the video understanding and the and improved tool use, for example, for video games, you'd want to use Unreal Engine or Unity or one of the main graphics engines, and then generate the art, apply it to a 3D model, and then create an executable that someone can run on a PC or a console or a like. We we expect that to happen probably this year, and if not this year, certainly next year. So that's, it's going to be wild. I would expect the first really good AI video game to be next year, and probably the first half hour of watchable TV this year, and probably the first watchable AI movie next year. Like things are really moving at an incredible pace. Yeah. When Grock is 10xing the world economy with vending machines, it will just create video games for humans. Yeah. It went from not being able to do any of this really, even six months ago, to to what you're seeing before you here, and from very primitive a year ago to making a sort of a 3D video game with a few hours of prompting.

Yeah, just to recap. So, in today's live stream, we introduced the most powerful, most intelligent AI models out there that can actually reason from the first principle, using all the tools, do all the research, go on the journey for 10 minutes, come back with the most correct answer for you. So it's crazy to think about just like four months ago, we had Grock 3, and now we already have Grock 4, and we're going to continue to accelerate as a company, XAI. We're going to be the fastest moving AGI companies out there. So what's coming next is that we're going to continue developing the model that's not just intelligent, smart, think for a really long time, spend a lot of compute, but having a model that is actually both fast and smart is going to be the core focus, right? So if you think about what are the applications out there that can really benefit from all those very intelligent, fast, and smart models, and coding is actually one of them. Yeah. So the team is currently working very heavily on coding models. I think right now the main focus is we actually trained recently a specialized coding model, which is going to be both fast and smart, and I believe we can share that model with you guys, with all of you, in a few weeks. Yeah. Yeah. That's very exciting. And the second after coding is we all see the weakness of Grock 4 is the multimodal capability. So in fact, it was so bad that Grock effectively just like looking at the world squinting through the glass and seeing all the blurry features and trying to make sense of it. The most immediate improvement we're going to see with the next generation pre-train model is that we're going to see a step function improvement on the model's capability in terms of image understanding, video understanding, and audio. Right? It's now the model is able to hear and see the world just like any of you. All right. And now with all the tools at its command, with all the other agents it can talk to. So we're going to see a huge unlock for many different application layers. After the multimodal agents, what's going to come after is the video generation. And we believe that at the end of the day, it should just be pixel in, pixel out. And imagine a world where you have this infinite scroll of content in inventory on the X platform, where not only you can actually watch these generated videos, but be able to intervene, create your own adventures. The view is going to be wild. And we expect to be training a video model with over 100,000 GB200s and to begin that training within the next three or four weeks. So we're confident it's going to be pretty spectacular in video generation and video. Let's see. So that's anything you guys want to other than I guess that's it. Yeah. It's, it's a good model, sir. It's a good. Yeah, we're very excited for you guys to try Grock 4.