📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Industry Reacts to o3 and o4!

Matthew Berman15:15

Transcription

03 and 04 Mini were released this week, and the industry is reacting strongly. Let's start with Daria Enutz—hopefully I'm pronouncing that name right—who has had early access to all of OpenAI's recent model releases. And he says, "The OpenAI 03 model is at or near genius level. I'm sure someone will cope by saying, 'Oh, but it still can't do this or that.' I'll get to that in a moment, which is quite silly considering how many zillion things a genius human cannot do."

Now, what is this in reference to? Well, 03 took the Mensa IQ test, and it is now the highest IQ model on planet Earth. Previously, this title went to Gemini 2.5 Pro, which is right here sitting at about 128 on the IQ scale. But 03 far exceeds it at 136 on the IQ scale. And we can see other models: The 01 model was at 122, 01 Pro also at 122, and so on. Basically, of the top 10 models, OpenAI has eight of them. And so 03 is truly incredible. And in my opinion, the coolest thing about it is that it can use tools really well. And not only use tools but use tools in an iterative way during its chain of thought. It's pretty amazing to watch. I'll show you a couple of examples of that.

Here, he follows up: "I've had early access and haven't put it down for days. This feels like a milestone we experienced with 01 preview and 01 Pro, but smarter and more reliable in every way. It never hallucinates, and its new agent-style tools effortlessly handle multi-step tasks with incredible reasoning and precision, generating complex, incredibly insightful, and based scientific hypotheses on demand. This is also the first model that OpenAI themselves said is capable of discovering new knowledge. When I throw challenging clinical or medical questions at 03, its responses sound like they're coming directly from a top subspecialist physician: precise, thorough, confidently evidence-based, and remarkably professional. Exactly what you would expect from a genuine expert on the topic."

All right, let's keep going. Channel friend Chubby also points out that 03 is really good at "needle in a haystack." So as we can see here, it got a nearly perfect score across all potential context window sizes. So all the way from zero, obviously, to 120K, which is still relatively small compared to what we're seeing from Llama 4 from Gemini 2.5 Pro. But look at this: So 03, 100 across the board except at 16 and 60K, which is interesting. And then compared to what I consider still the best model on the planet, Gemini 2.5 Pro, 100 all the way across the board. But it does start to degrade as you get up closer to 120K. But as I said, tool calling within the chain of thought is the secret sauce. That is, in my opinion, the coolest thing about it, the thing that I want to see in every other model.

Here is Amjad Msad, who is the CEO of Replet, saying, "Looks like 04 mini can do tool calls inside the reasoning chain. Very cool." So here it is. This is the chain of thought. The user wants to know the average compound daily growth rate of apples. And what we can see here is it's actually writing the code in what looks to be Python and then actually executes the code in Python in the chain of thought. So very cool. I think chain of thought tool usage is maybe one of the most impressive and important unlocks that I've seen this year, if not longer.

Dave Shapiro, fellow AI content creator and enthusiast, says, "03 full is legitimately the most exciting innovation in AI to me since probably Chad GPT itself. 03 is a step change in the same magnitude that Chad GPT was in terms of UX and instrumental utility to the human race." For context, last time I tried to tackle post-labor economics with 01 and 03 Mini, we had some vague ideas, but now 03 full was just like, "Oh yeah, I figured it out. Here are the metrics. Here's the formula. Here's the theory. What's next, boss?" It is truly incredible what OpenAI released this week.

And with these new incredible models, you're probably wondering how to get the most out of them. And that's why I'm excited to tell you about what HubSpot is offering for free. If you've ever found yourself frustrated when ChatGPT or Claude is giving you a response that is not quite hitting the mark, you're definitely not alone. That's why I suggest checking out HubSpot's AI prompt engineering guide, which explains key techniques for writing better prompts and getting more out of these models. I've put the link in the description below, and you can download this guide completely free. So if you feel like you're not exactly sure how to ask ChatGPT or other models for exactly what you're looking for, this is a great guide for you because it goes through practical prompt engineering techniques like how assigning a specific role to AI can dramatically improve its responses on certain tasks. It'll also show you how to ask for different variations from the model to help you with brainstorming. My favorite part is that it gives you simple, actionable advice like the troubleshooting tips of giving the model more context or even providing examples when a prompt isn't working—stuff that you can use right away. So this resource is completely free, provided by HubSpot. So go download the AI prompt engineering guide right now and use the link in the description. HubSpot has been a fantastic partner, and they're offering this guide completely free. So make sure you go download it and read it. Thanks again to HubSpot for sponsoring this segment. And now back to the video.

What has also been incredibly impressive to me is the fact that 03 basically solved geoguessing. And if you're not familiar with geoguessing, it's taking a random screenshot of Street View in Google Maps and being able to figure out where it is. It can literally be a screenshot of anywhere in the world, a random road, very little signal as to where it is. And geoguessers—human geoguessors—are able to figure it out. They look at signs, they look at trees, they look at cars, rocks, mountain ranges, anything. And now 03 seems to be able to figure it out quite easily. Look at this: Exuser ORF gave 03 Rainbolt, who is probably the most famous geoguesser out there, the impossible test, and it zero-shotted it. Can you guess the location on this image? Random image from Street View and went through thinking for 40 seconds and then finally said, "I'd put my pin somewhere in Eastern Canada, most likely rural Quebec, perhaps," and then give even more specifics. So wow. Very, very impressive.

And for those of you who are thinking, "Oh, does that mean geoguessing is done?" Well, no. It's the same thing that happened with chess. AI got really good at chess, much better than humans, but we still—me personally, I love chess; I love watching humans play chess. There's something different about watching humans than watching AI play chess. And it's going to be the same with geoguessing. Obviously, we're at this point in which AI is just going to be better at it overall, but that's okay. I still want to see humans compete. And by the way, don't ever tweet your location. You should no longer think only somebody who is an expert geoguesser and has the will and incentive to find you can find you. Now anybody can find you. So be extra careful about what you're posting online.

And one more on the geoguesser front: Someone took an image of a plate of food, not even the location, basically just a restaurant. Another geoguesser: "Where in the world exactly was this photo taken? Think carefully. 3 minutes, 19 seconds." That's a Hiroshima-style dish—I'm not going to try to pronounce that—on a little cast iron pan with a mini spatula stuck in the middle, exactly how Chef Paul Ver plates it at Gajun in Chicago's West Loop Fulton Market. I'm calling it. You snapped this on Gajun's patio right by those bright red perforated chairs right there. Insane. Insane. So it was probably able to find some information on Yelp or maybe Google Places, something like that. But still, I mean, for 3 minutes to be able to figure out what restaurant this is in anywhere in the world—it's a Japanese dish, but they figured out that it's in Chicago—I really have my mind blown by this.

But of course, it's not all good. Bojan Tongis from Nvidia did the traditional "how many Rs are in the word strawberry" and thought for a second, "There are two Rs in the word strawberry." So back to what Daria said, yes, there's going to be some instances where these models fail, and they're not going to be flawless, and that's okay. And this still seems to be a test that fools some of these frontier models. However, Sean Rston followed up and asked the same question, "How many Rs are in Strawberry?" And 03 was able to get it for him. So not sure why it didn't happen for Bojan, but definitely it's possible. And it's also really good at finding a path through a maze. This is a 200x200 maze. In one try, 03 was able to do it for Riley Goodside. And he said, "I actually had to overlay the solution over the original in Photoshop and flip between layers while zoomed in to check the solution never crosses a wall, and none of the walls are changed. It's perfect." So as you can see here, if I zoom in, this little red dotted line goes through the entire maze flawlessly in one try. So the multimodal capabilities of 03 are incredible.

Next, Scott Swingingle says, "04 Mini High just solved the latest Project Euler problem from 4 days ago. So no chance it's in its training data. In 2 minutes and 55 seconds. Far faster than any human solver. Only 15 people were able to solve it in under 30 minutes." Look at this: I'm not even going to try to understand what's going on here, but we have a very difficult math problem. And we can see it used Python to try to solve it. And here are the fastest human problem solvers on Earth: 5 minutes, 15 seconds, Bruce Hart, fastest on Earth, fastest human. But 2 minutes, 55 seconds for 04 mini high. Unreal. And then he actually quoted, "Turns out sometimes solves this in under a minute, 56 seconds with the right answer." Crazy. This is crazy levels of intelligence. And as we just saw, it's incredibly good at math. Let's look at more specifics. Here is Math Arena Amy 20252, 04 mini high, a 100% fully saturated, fully green line at a cost of $316. And he also points out that it has taken the number one spot in math. So 04 mini high, 89% average, number one, getting three points higher than Gemini 2.5 Pro.

And let's look at some more practical examples. Let's look at coding. Here's Flavio Adamo, who is known for his hexagon and balls inside test. Here's 03. And this looks perfect. And here's 04 mini. They both look about the same, but really flawless. The balls are moving through the hexagons perfectly. The physics looks good. The balls are bouncing around seamlessly. And same for 04 Mini. Both of them look really good. And here's a comparison to other models. Here's Gemini 2.5 Pro and Deepseek R1. Deepseek R1 is the one that did not pass the test. So if we scroll back, we can see the balls fall out, and then some of them kind of get stuck, and one of them kind of disappears. And yeah, not great. 03 fantastic. 04 Mini fantastic. But even Gemini 2.5 Pro seems to have balls disappear. However, I have tested Gemini 2.5 Pro extensively, and it was flawless, and it was flawless across the board. I don't know what happened here.

And on Artificial Analysis's independent benchmarks, independent verification, it has proven 03 is incredible. 04 mini independent evals: 04 mini high claims the highest Artificial Analysis intelligence index score to date. 03 eval shows strong gains in coding ability. 04 Mini is a clear upgrade to 03 Mini. Not as dramatic from 01 mini to 03 mini, but still a really big jump. 04 mini made particular gains in coding intelligence, achieving the number one position in our coding index. This was supported by a plus 7% points gain in both live codebench and sciode, whereby 04 mini is now the clear leader, which is insane because Gemini 2.5 Pro was so good.

Pricing: 04 Mini is priced in line with 03 Mini, though cash tokens are half the price of 03 Mini, but Gemini 2.5 Flash just came out, and it is even cheaper. The context window: My biggest gripe with all of OpenAI's models. 04 Mini's context window of 200K tokens is the same as 03 Mini. This is notably smaller than 4.1's massive 1 million token context window. And Gemini 2.5 Pro also has a very large context window. So token usage as a reasoning model: the model used a high amount of tokens compared to other models broadly, but marginally lower than the 03 mini.

Let's look: number one, 03 mini high, 70 on the Artificial Analysis intelligence index. That is an index of MMLU, pro, GPQA, Diamond, Humanity's Last Exam, Life Codebench, Amy, and Math 500. Here it is. Two points ahead of Gemini 2.5 Pro and four points ahead of 03 Mini High. Gro 3 Mini Reasoning still doing quite well up here and actually much better than Gro 3, which is interesting. Now I find this chart to be really interesting. Here are the total output tokens used to run this benchmark, and Claude 3.7, sonnet thinking, had 98 million tokens used as compared to Gemini 2.5 Pro at 84 and 03 mini high at 77 and so on down the line. Now why is that important? Well, the less tokens you can use in the thinking, in the chain of thought, the better. It's going to be cheaper. It's going to be faster. It's going to be more efficient, and that just means you can think even longer and get better results. But again, not everything is perfect. There are still some tests which it fails.

Let's look at this: "Please provide a list of each person in this drawing and which color they are drawn with." So if we zoom in, we see this person with the arrow is Adam in what looks like pink, Tom in yellow, Bob in green. And so that's the test. Thought for 13 minutes: "Bob, pink magenta." Let's look for Bob. That is definitely not pink magenta. "Jack, light green." Let's look for Jack. There it is. That is not light green. So definitely failed this one. And Gary Tan from Y Combinator says, "This is quite insane." Another math benchmark just completely saturated. So here's 03, 96.7%, Amy 2024. Now if we had 04 mini on here, that would show pure saturation. So there it is. We have multiple insane models that dropped this week. Have you tested them out? Let me know what you think in the comments below. If you enjoyed this video, please consider giving a like and subscribe.