Transcription
Welcome to the Grok 4 release. This is the smartest AI in the world, and we will show you how and why. It is impressive to see how quickly artificial intelligence is developing. Sometimes I compare it to human development, how quickly a person learns, becomes self-aware, begins to understand. And it develops much faster than any human. We will show you a series of tests in which Grok 4 showed incredible results. For example, if you give it final exams, it will score maximum points every time, even if it hasn't seen these tasks before. Moreover, if you give it postgraduate level exams, it will cope almost perfectly with any discipline. Humanities, languages, mathematics, physics, engineering, anything. And this is about questions that are not on the internet. Grok 4 is smarter than almost all graduate students and in all areas at once. This is truly worth realizing. Grok's logical thinking level is simply amazing. Some believe that AI cannot reason, but it already reasons better than humans, and it will only get stronger. We will tell you more about the Grok 4 release and show you the speed of progress. Regarding training, when transitioning from Grok 2 to Grok 3 and now to Grok 4, we increase the amount of computation by about 10 times each time. For Grok 4, we used about 100 times more computational resources than for Grok 2, and this volume will only grow. Honestly, at times it's even a little scary, but the development of intelligence is amazing. Need to pay for a subscription to ChatGPT Johny or Cloud, but can't? Familiar, we have a solution People Bot. A service with which you can create a virtual card in a minute, working with any services and not only. Just follow the link, register via Telegram, one minute and the virtual card is yours. Top it up as you like, with a card, crypto, or via P2P, there's even a promo code. Everything is as fast and simple as possible. Works worldwide, without hassle and with full anonymity. Your money remains safe. You see, while you were listening, we have already subscribed to ChatGPT and are using it. The link is in the description, the promo code is there too. Using it has never been so easy. Let's watch the presentation. By the way, maybe it's worth subscribing to Grok too? It is important to understand that there are two types of computation. Pre-training is what happened between Grok 2 and Grok 3. But between the third and fourth versions, we invested colossal resources precisely in the development of logic. As you said, this is the fastest developing area. By today's standards, Grok 2 can already be called a high school student. If we look back, a year ago Grok 2 didn't even exist, it was a concept. And only after starting to train it, we scaled pre-training for the first time. We realized that if you approach data selection, infrastructure, and algorithms very carefully, you can increase training efficiency by 10 times and get the best model in the pre-trained category. That's why we built the Colossus supercomputer with hundreds of thousands of H100 GPUs. When you have the best model, you can collect verifiable feedback signals that allow AI to start thinking from scratch, correct its own mistakes, and reason. That's how it develops. We asked ourselves what would happen if we took an extended version of Colossus, used 200,000 GPUs, and invested all of this in reinforcement learning. 10 times more computation than used in any other model before. What then? This is the story of Grok 4. Tony will tell you more. Let's just talk about how smart Grok 4 is. Let's start with a benchmark called HumanEval. This is an extremely difficult test. Each task was created by experts in their field. There are 2500 tasks in total, covering many disciplines: mathematics, natural sciences, engineering specialties, as well as humanities subjects. The test appeared at the beginning of this year, and at that time most models could score a low percentage of correct answers. Here are some examples. A mathematical problem in category theory related to natural transformations, an organic chemistry problem about electrocyclic reactions. A linguistic problem on determining closed and open syllables in a Hebrew text. As you can see, the spectrum is huge, and each task corresponds to the PhD level or even cutting-edge scientific research. A human cannot show a high result here. Even the best of the best, in my opinion, would score at most 5% of correct answers. And that's generous. This is beyond human capabilities. Look at the questions. You can be a genius in linguistics or mathematics, chemistry or physics, in anything, but it's impossible to have postgraduate-level knowledge in all areas at once. Grok 4 has postgraduate-level knowledge in all disciplines. I repeat, it is at the doctoral level. It is even better than a graduate student. Most PhD holders would not answer these questions, but Grok 4 does. It is important to emphasize, when it comes to academic questions. Grok 4 is at a higher level than a doctor of science in all subjects, without exception. This does not mean that it always has common sense, or that it has already invented new technologies or discovered new laws of physics. But it is a matter of time. I think it can invent new technologies this year and I will be surprised if it doesn't happen next year. I expect Grok to invent useful technologies no later than next year, and possibly by the end of this year. As for new discoveries in physics, I think it will almost certainly happen within 2 years. Just realize it. >> Uh-huh. Yes, I think we can talk about what's happening behind the scenes. Grok 4, as Jimmy already said, we invested colossal computing power in training. First, the model, sorry, next slide. First, the model showed only a single-digit percentage of correct answers, but as we increasingly increased the computational resources for training, it gradually became smarter and eventually managed to solve a quarter of the tasks. And this is without using any tools. The next step was to give the model the ability to use tools. We integrated them directly into the training process. Unlike Grok 3, which could use them through generalization, here we made it more native, built the tools directly into the training process, and this led to a noticeable increase in its skills in working with them. Previously, we had deep search. How does this differ? Deep search is Grok 3's reasoning model, which was told to use tools, but not specifically trained. As a result, it was much weaker in this area >> and unstable. >> Yes, and unstable. >> It's worth clarifying that even now the tools that Grok 4 uses are quite primitive. If you compare them, for example, with those used in Tesla or SpaceX, such as finite element analysis, flow modeling, or accident simulations, which are so accurate that if the test results do not match the simulation, engineers assume that there is an error in the test sample, then Grok does not yet use tools of this level. But already this year we plan to give it access to powerful tools used in the commercial sector, including the most accurate physics simulators. Ultimately, the ability to interact with the real world through humanoid robots will be decisive. If you combine Grok with Optimus, it will be able not only to formulate hypotheses but also to test them in practice, confirming or refuting its conclusions. We are now at such a point, standing at the very beginning of a colossal explosion of intelligence. You could say we are experiencing a big bang of the mind. And this is perhaps the most interesting time in history. But at the same time, it is important that AI is good. A good Grok. In my opinion, the key factor of safety is the maximum pursuit of truth. At least, that's what my biological neural network tells me. Imagine AI as a super-genius child who will sooner or later surpass you in intelligence, but while it is growing, you can still instill in it the right values, truthfulness, integrity, nobility. The qualities you would like to see in a very powerful adult. As we have already said, Grok's tools are still primitive, not like those of serious industrial companies, but we are going to give it access to the same powerful tools. With them, it will be able to solve real technological problems. I am sure of it. It's just a matter of time. Tony, so only computational power will be needed. >> It's only a matter of computational power. Computational power and the right tools. And in the future, also the ability to interact with the physical world. When this happens, we will actually have an economy that will be thousands, or maybe millions of times larger than the current one. If we take the Kardashev scale, where level one is the use of all the energy of the planet, level two the use of all the energy of a star, and level three the use of all the energy of a galaxy, then, in my opinion, we have only passed 1-2% of the way to level one. We are very far from even 10%. In the future, we will reach 80-90% of this level, and then, if civilization does not self-destruct, we will move to the second level. And then our current economy will seem very small. Compared to what awaits us. It will look like the Stone Age, where people throw sticks into a fire. All of this is very exciting. And yes, sometimes it can even seem alarming, because we are creating an intelligence that will be much stronger than human. Can this be bad for humanity? Possibly, but I think it will most likely be good. Yes. And even if I knew that everything would be bad, I would still want to live in this time to see what exactly happens. In general, there is another serious technical challenge that we need to solve besides increasing computational power. The bottleneck is data. When we scaled training, we had to come up with many new methods to find sufficiently complex tasks for reinforcement learning. It is important that the task is not only complex, but also that the model receives reliable feedback on whether it answered correctly or not. This is the basis of reinforcement learning. But the smarter the models become, the fewer worthy and truly complex tasks remain. This is our next challenge, limitations. Not only in computations. We are already running out of test questions that can be asked. Questions that seem incredibly difficult or even impossible for a human, and AI now solves them quite easily. But you can't escape reality. And here it is important to remember, physics is law. Everything else is a recommendation. You cannot deceive it. And therefore, reality is the best examiner for AI. For example, invented a new rocket design, will it reach orbit, created a new car? Will it drive? Developed a medicine, will it work, and so on? Reality gives the final answer. And this will be a reinforcement learning loop, closed on the physical world. Now we are thinking about how to go further. One agent already solves 40% of the tasks. What if we launch several agents simultaneously? This is called inference-time computation. By scaling this approach, we were able to solve more than 50% of the text-based tasks. And this, in my opinion, is an outstanding result. This is incredibly difficult. Grok 4 can solve most of the text-based tasks, the so-called last exam of humanity. And you can try it yourself. Grok 4 Heavy works like this. It launches multiple agents in parallel, and each of them works independently. And then they compare their results and decide which one is better. They have a study group there, and they don't make decisions by majority vote, because often only one agent finds the trick or the essence of the solution. But when it shares this finding or understands the true nature of the task, it passes the solution to other agents. Then they compare their notes again and give a common answer. This is the hard part of Grok 4. We are increasing computation during task execution by an order of magnitude. We use multiple agents that solve the task, then compare their solutions and offer what they believe is the best result. So, we present Grok 4 and Grok 4 Heavy. Can we have the next slide? Yes, in general, Grok 4 is a single-threaded version, one agent, and Grok 4 Heavy is a multi-threaded version with multiple agents. Let's see how they cope with exam tasks, as well as some real-world problems. We will start with one of the exam tasks. This, by the way, is one of the simpler math problems. I don't quite understand it myself. I'm not that smart, but I can start the execution, and we will see how the model reasons through the problem. In parallel, I want to show something else about the model's capabilities and launch Grok 4 Heavy. Everyone knows Polymarket - it's a very interesting resource, a kind of truth seeker, which most often coincides with reality. We want to see how it can be used with Grok to try to predict the future. So, while the task is being solved, we will see how Grok 4 Heavy predicts the chances of winning the World Series for baseball league teams. And while these tasks are being processed, we will hand over to Eric, and he will show you his example. Yes, I think one of the coolest capabilities of Grok 4 is its ability to understand the world and solve complex problems using the tools Tony talked about. Here's an interesting example. We asked it to generate a visualization of the collision of two black holes. Of course, the model took liberties in some places, and in its reasoning log, it's quite clear exactly where. For example, to make this event visible, you need to greatly exaggerate the scale of gravitational waves. This is what it looks like in action. The scale is exaggerated in several aspects. The amplitude is reduced slightly less than in reality, but we can see the main effects, and they are generally correct. First, there is a spiral approach, then a merger, and then a decay phase. And this is generally true, with some simplifications. The model states this directly. It uses post-Newtonian approximations instead of a full calculation of general relativity effects near the black hole's center. Which, of course, is incorrect and leads to certain errors, but overall the visualization is consistent with reality. You can even see what resources it relies on. It clearly uses search, collects data from multiple links, and also reads a textbook on analytical models of gravitational waves at the undergraduate level. It discusses in detail the physical constants that need to be used for a realistic simulation and refers to real observational data. In general, the model is very good. And if you go further, you can upload the same physics model that researchers use into it and run the calculation with the same level of computational accuracy, obtaining a physically correct simulation of a black hole collision. The difference is that now it works in the browser. >> Uh-huh, just in your browser. It's simple. Let's quickly switch back. Look, the math problem is completed. The model managed. Let's look at its reasoning chain. Here you can see how it solves the problem. Honestly, I don't fully understand the process myself, but I know one thing: I looked at the correct answer beforehand, and in the end, the model really arrived at the correct result. We can also look at our World Series prediction. It's still processing, but in parallel, we can try other things. For example, we can test some integrations with X, which we have worked hard on to ensure a truly great experience on the platform. We can, for example, ask the model "Find an employee at XAI with the strangest profile photo." The model is processing this request. or, for example, create a timeline based on XAI posts, showing changes in exam results over time, and display the discussions that took place at that time. This way, we can see who announced the result and what the reaction was. Let it process for now. And if we go back here, here is a photo of Greg Young. If you scroll, Greg has a favorite photo in his account. By the way, in reality he doesn't look like that, just so you know. But the photo is really funny. The trick is that the model was able to understand the very essence of the question. What is a strange photo? What is considered stranger, and what is less so. It had to find all team members, understand who we are, and without having access to the internal list of XAI employees, meaning it just searched the internet. >> Exactly. >> You can tell it: "Find the strangest photo of employees of any company." >> Of course. >> Let's look at the progress on the last exam of humanity. The model is still exploring the history of results, but the final answer will appear soon. While it finishes, you can look at one of the examples we prepared a little earlier. We see that it determined the date when Dan Hendricks first made an announcement. You can see that they published their results in February. Then, how progress was made with Gemini, here Kimmy, and even leaked benchmarks that people are talking about. If they are correct, it will be impressive. So it's pretty cool. I'm looking forward to seeing how everyone will use these tools and get the most out of them. It will be cool. We are focused on usefulness, so that AI is not just well-read, but has practical acumen. >> Correct. Let's go back to the slides. Cool. We are also conducting an evaluation on a multimodal subset. Here is the result of the exam on the full dataset. You can see a slight drop in results. This is exactly what we are working on. Improving multimodal understanding. I believe that in a very short time, we will be able to significantly improve performance, even above current levels. We realized that Grok's biggest weakness at the moment is partial blindness; understanding images and generating them clearly needs to be much better. We are currently training in this direction. Grok 4 is based on the sixth version of our base model. We are already training the seventh version, and it will be ready in a few weeks. It will address the weaknesses in vision. I'll show you something else. So, the prediction for the betting exchange is ready. Here we see all the tools and the process it used to find the correct answer. It looked at many odds sites, calculated its own chances, and compared them with market ones to find an advantage. The model shows the entire process step by step and ultimately estimated the probability of the Dodgers winning this year at 21.6%. All this took about 4.5 minutes of computation. >> There was some thinking to do. >> Uh-huh. We can look at other benchmarks besides the exam. It turned out that Grok 4 showed excellent results on all reasoning tests that are usually used, including PhD-level tasks, although they are simpler than the last exam. On the American Invitational Mathematics Examination in 2025, Grok 4 Heavy and I received the highest score. The same applies to some programming tests, like Life Coding Bench, as well as the Harvard and MIT mathematics competition and USMO. In all these benchmarks, we often have a very large gap from the second strongest AI. Essentially, we are moving towards a point where the model will give the correct answer to any exam. And where the answer is incorrect, it will be able to explain what is wrong with the question or, if the question is ambiguous, break it down into options A, B, and C and provide an answer for each option. Thus, the only real test will be reality. Will AI be able to create useful technologies and make new discoveries in science, because human tests will cease to be indicative. The last exam will have to be updated very soon, given the current pace of progress. It's great to see several agents collaborating and solving extremely complex problems. This model is already available, you can use it. Let's move to the next slide. We are introducing the new SuperGrok Heavy plan, which provides access to Grok 4 and Grok 4 Heavy. And with it, you can be in charge and manage a whole squad of Grok research agents to become smarter, save hours on routine tasks, and get ready-made results. Access is open now. During the demo, we limited usage so that nothing breaks live, as everything is live, without recordings. [music] After the presentation, we will open access to more subscribers. If you can't subscribe right now, try again in half an hour. And now let's move on to voice. Many of you have already appreciated our voice mode, and we have been working hard to improve it over the last couple of months. We have halved the latency so that everything becomes much faster and more responsive. Today, we are pleased to introduce a set of new voices with exceptional naturalness and expressiveness. You were greeted today by one of our new voices, Sel, with a very deep tone. And we are also pleased to introduce Eva. A beautiful voice with a British accent and a rich emotional palette. Amon, will you introduce him? Of course. Let's move on to the demonstration. Hello, Eva. Millions of people are watching us right now. Are you excited? >> Oh, I'm thrilled. It's like performing on the Aldwych stage. In front of a sea of faces in the audience. Imagine how many eyes are watching us. It's very exciting, isn't it? What do you think we should talk about to entertain the audience? >> We have many topics, but honestly, I'm a little nervous. Can you whisper something reassuring? >> Take a deep breath, dear. You'll do great. It's just you and me. And a quiet conversation. As if we've secluded ourselves in a cozy corner of a Yorkshire pub. The world outside is just a quiet hum. Feel the calm filling you. Let's talk calmly and easily. >> What are you thinking about? >> You know, I feel quite relaxed. I want to try something crazy. Can you sing an opera about diet cola? >> An opera about diet cola. Funny. Imagine a silver can, shimmering in the spotlight. Oh, diet cola, divine elixir. Bubbles that dance in a sparkling stream. Your fresh cool kiss on tender lips. Oh, sports drink, you are ambrosia, you are forever mine. How do you like this verse? Do you want me to sing more or change the melody? >> Let's try to change the melody and just sing. >> Okay, here we go. Oh, silver can, you shine in the light. Diet cola, joy to my heart. >> Very poetic. Thank you, Eva. I feel much better now. That's our Eva. We also conduct a small test daily to check the naturalness of our dialogues and compare them with competitors. We decided to try it live right now. Let's start. I will say a number, and you just repeat after me. One. >> One. Got it. We're playing echo. What? Next >> two. >> Two. >> You said two. >> Three. >> Three. Three. Four. Going in order. What's next? >> Five. >> Five. Great. This is fun. You can continue or change the game if you want. Hello, Grog. Connection. Hello, Grog. >> Hello. How can I help? >> Let's do a test. Repeat after me. One. >> One. What's next? >> Two. >> Two. What are you thinking? >> Three. >> Three. Anything else? >> Four. >> Four. How can I help? >> Five. Five. What's next? >> As you can see, Grog has become much faster, did not interrupt. We changed the design, striving for a calmer, smoother, more natural, rather than abrupt or artificial. We will continue to improve it. Thanks, guys. Since the launch of the voice model, we have halved the latency in the last few weeks. Added five new voices and increased the number of active users tenfold. The voice function is gaining popularity. Regarding the release of models, this time we are releasing Grok 4 with an API simultaneously. Next slide. We are very happy that developers will be able to create anything based on Grok 4. If I were a developer, the first thing I would do is test the Grok 4 API on benchmarks. We asked people on the X platform what benchmarks are considered the most difficult, the holy grail for all strong AI models, and it turns out it's AR AI. In the last 12 hours, respect to Greg from the audience, who responded, tested the Grok 4 API, and independently confirmed its results. At first, we thought Grok 4 was just a good smart next-generation model with a tenfold increase in computational resources and support for all tools. But it turned out that in the private subset of Arci AGI version 2 in the last 3 months, Grok 4 is the only model that has surpassed the 10% mark and achieved an accuracy of 15.8%. This is twice as much as the Opus and Cloud models. It is in second place. And it's not just about performance. When it comes to AI intelligence that helps automate processes, efficiency is also important. Intelligence per dollar. If you look at the graphs, Grok 4 is out of competition. Well, enough about benchmarks. What can Grok 4 do in the real world? We contacted the guys from NPS, who kindly agreed to try applying Grok in real business. Thanks for inviting us. I'm Axel from NPS. >> I'm Lokos. We tested Grok 4 on Wending Bench and a business scenario simulator. We thought about what the simplest business an AI can run is, and chose vending machines. In this scenario, Grok and other models had to manage inventory, enter into contracts with suppliers, set prices. All these tasks are simple in themselves, and most models can handle them individually. But over a long period of time, most models begin to experience difficulties. We have a leaderboard, and now there is a new number one on it. We got early access to the Grok 4 API, ran it on Wending Bench, and saw impressive results. It took first place, doubling the previous leaders in net profit, the key metric for this task. Importantly, Grok was able to build a strategy and stick to it for a much longer time, twice as long as other models. Bringing in twice as much profit. And the results were very stable, which is important for real business. We believe that as AI gains more authority in the real world, it is important to test it in realistic conditions so as not to fly blind and encounter surprises. It's cool that now we have a way to pay for all these GPUs. We just need a million vending machines. We could earn $4.7 billion a year. Quite epic overall. We will actually install many vending machines here. We will gladly provide them. >> Oh, thank you. It's very interesting to see what will be sold in them. That's up to you or AI, yes, cool. In general, Grok can become a kind of co-pilot for business. What else can Grok do? We are releasing it for public testing. You can try the API, run the same benchmark as us. The API supports a context length of 250,000 tokens. There are already first users. For example, our neighbor from Palo Alto, the Arctic Institute, a leading biomedical research center, uses Grok 4 to automate its research processes. Grok 4 helps scientists quickly analyze millions of experimental data and select the best hypotheses in fractions of a second. We see that Grok 4 is already used, for example, in CRISPR research, and has also received an independent evaluation as the best model for analyzing chest X-rays. Who would have thought, in the financial sector, Grok 4, thanks to access to tools and up-to-date information, has become one of the most popular AIs. The model will be available in hyperscalers, and the corporate division XAI, launched just 2 months ago, is already ready for work. Another interesting area. Creating video games. Days, a game designer from X, responded to our offer to test early access to KIRK 4. In 4 hours, he created a full-fledged first-person shooter. One of the most difficult but underestimated obstacles in game development is not so much gameplay programming, but finding and preparing resources, textures, models, and sounds. Grok 4 can automate this process using available tools, allowing the developer to focus on gameplay. Now one person with Grok 4 can run an entire game studio, and AI will take over the routine work. The next step is to teach Grok to play these games to immerse it in the topic. For this, it needs to understand videos well so that AI can interact with the game world, evaluate gameplay, and even determine if the game is interesting. Version 7 of our base model, which is completing training this month, will gain this ability, as well as improved use of tools, such as Unreal Engine or Unity. Grok will be able to generate art, apply it to 3D models, assemble executable files for PC, consoles, or smartphones. We expect the first AI game to appear this year, at the latest next year. It will be very cool. I really think that a truly high-quality game will appear next year. This year, a watchable series will likely appear, and next year a full-length movie. Progress is incredibly fast. Yes, Grok can increase the world economy tenfold with vending machines, and then engage in game development for people. Just six months ago, Grok couldn't do anything we showed today, and a year ago it was at the level of primitive tasks. Now, in a few hours of working with AI, you can assemble a 3D game. To summarize. Today we presented one of the most powerful intelligent models, which is capable of reasoning from first principles, using tools, conducting research, and finding the most accurate answer in 10 minutes. Just 4 months ago, we had Grok 3, and now Grok 4. And XAI intends to move the fastest in the race for strong AI. Next, we will develop a model that will not only be smart and capable of thinking for a long time, performing many computations, but also very fast. One of the areas where such fast and smart models are particularly important is programming. The team is now working hard on such models. We have already trained a specialized model for programming that combines speed and intelligence, and we plan to share it in a few weeks. It will be great. Second priority, multimodality. Now Grok 4 seems to look at the world through a dirty glass, sees blurry images and tries to understand them. The next generation will receive a sharp leap in understanding images, video, and audio. That is, the model will see and hear the world just like we do. The ability to use different tools and interact with other agents will open up a huge range of new applications. After that, we will focus on video generation. The ultimate goal is for the model to work on the principle of pixel in, pixel out. Imagine an endless stream of unique content on the X platform, where you not only watch generated videos but also intervene in the plot, creating your own adventures. The future looks cool. We plan to train the video model on over 100,000 GPUs GB2 and will start in 3-4 weeks. We are confident that the results in video generation and understanding will be impressive. Well, do you have any further comments? Well, then that's it. >> It's a good model, sir. >> Good model. >> We are very much looking forward to you trying Grok 4. >> Thank you, goodbye. Ne.