📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

AI News: Gemini 2.5 Flash, o3 and o4, Claude Research, Kling 2.0, and More!

Matthew Berman15:46

Transcription

This week was all about new model releases; Frontier model releases. The first one we're going to talk about today is Gemini 2.5 Flash. This is the smaller, more efficient version of, in my opinion, the best model on the market right now, Gemini 2.5 Pro. You know the one that could build the Rubik's Cube in one go. And now we have a much cheaper version of it.

This video is sponsored by Replit, the easiest way to vibe code. I'm going to tell you about their new launch with Agent V2 a little bit later. Gemini 2.5 Flash is our first fully hybrid reasoning model, giving developers the ability to turn thinking on or off, which is phenomenal. So, as a developer, you have the option to just get responses for the more simple queries; or, for the more complex logic, reasoning, math, coding, you can turn on thinking. They also give you the ability to set a thinking budget, so a fixed amount of tokens to use within the thinking window.

Let's look at some of the scores now. One thing I like that they did was, in the benchmark comparisons, they included the models from OpenAI that were just released, 03 and 04 Mini. And I'm going to talk about those in a moment. It was literally released the day before, and a lot of model providers might have ignored them, but they included them, even though those models did better than Flash in many of the benchmarks. But let's talk about it. So we have Gemini 2.5 Flash. First, let's talk about the pricing, because that is really its standout attribute. It is $15 per million input tokens. That is as compared to 04 Mini at $1, Claude 3.7 Sonnet at $3, Gro 3 beta $3, and DeepSeek R1 at $55. So Gemini 2.5 Flash is even cheaper than open source. Then for output, we have 60 cents non-reasoning and $3.50 for reasoning. And once again, across the board, except for DeepSeek R1, it is much cheaper.

All right, let's look at the actual quality. Humanity's Last Exam: 12.1%. That is compared to 04 Mini at 14.3%, so 04 Mini is better. Claude 3.7: 8.9%, so it beats that. DeepSeek R1: 8.6%, so it beats that. The only model right now that is better is 04 Mini. GPQA Diamond, which is a science benchmark: 78.3%; OpenAI: 81.4%; so basically on par with all the other models. Amy 2025, Amy 2024 still doing very, very well. But as you can see, Mini, across the board, is the most powerful model, but significantly more expensive.

Okay, so look at this graph. On the y-axis, we have the arena score; and then on the x-axis, we have the price per million tokens. On the left side of the x-axis, it's more expensive; on the right side, it is less expensive. So hugging the outside of this quadrant right here is very good. So we have Gemini 2.5 Pro all the way at the top, the best model around. There's ChatGPT 4.0 latest, there's Gro 3 Preview. And the thing is, it's still decently expensive, but less expensive than a lot of these other models. Here is where Gemini 2.5 Flash preview comes in, right here. So still basically on par with their competition, definitely below Gemini 2.5 Pro, extremely inexpensive. I'm going to be making a full testing video for Gemini 2.5 Flash, and we will see how good it is compared to Pro.

Next, OpenAI this week dropped three different models, and two of them I'm going to talk about first: 03 and 04 Mini. 03 has the best tool use I have ever seen; in fact, it is able to use tools within its chain of thought process, which I have not seen any other model do. And 04 Mini is a different model: smaller, more efficient, less expensive, but both of them are very good. And yes, I am also going to be testing these two models thoroughly. Now I already made a video about these two models, so I'm not going to get too deep into them, but I want to show you one thing which absolutely blew my mind. So check this out. I was on vacation this last week, and here I am recording a video on vacation, different location. I am simply going to take a screenshot, just to ensure there is no location metadata in the image. I'm going to come back to GPT-03 right here. I'm going to drag the image, and I'm going to say, "Tell me where this person is exactly." Let's see if it's able to do it. All right. And actually, it took a lot shorter time this time than it did the last time I tried this. So here's the thinking: "The user requested a precise location, likely near Princeville, Kauai in Hawaii, based on the photo's details. I'd say it's probably Princeville facing the Hanalei Valley, with the views of the Namolokama Mountain. The surrounding vegetation and home structures match the Princeville area. There's a chance it could be Maui or Oahu." Okay, so that's all the thinking, and the final answer: Princeville, Kauai, on a lanai facing the Hanalei Valley and Namolokama Mountain. And that's true; that is exactly where I was. So essentially, geo-tagging has been solved. This is an absolutely incredible feature, and also kind of a scary one. Now the first time I tried this, it did something really cool; it would zoom in on the background of the photo, zoom back out, zoom in in other places, and essentially really identify each portion of the image and where it could be. This time it was just done a lot easier. And just to make sure it wasn't using any memories about previous chats, I deleted any mention of Kauai, Hawaii, Princeville, anything like that, and confirmed it was deleted.

Next, Replit has launched Agents V2, and they're also the sponsor of today's video. Replit is a phenomenal, fully cloud-based IDE, and they've been working hard on their vibe coding tools, specifically Agent V2. And I've used Replit a ton; it just makes coding so easy, especially the deployment process. You don't have to worry about setting up a database; you don't have to worry about deployment after you've coded something locally; it all just works because it's in the cloud to begin with. And with Agent V2, you have a substantially improved autonomous agent working on your behalf, as compared to V1. You will be five times more likely to successfully create what you need with Replit V2. And the best part is, it is fully cloud-based, so it doesn't matter where in the world you are; it doesn't matter which computer you're at, as long as you can log in via a browser, you have access to your entire repository, your entire codebase easily from anywhere. And you can even download the Replit apps on your Apple device or Android. So check it out: replit.com/refer/matthewberman. I'll drop the links down below. Use code Matthew to get 10% off your first month of Replit. Make sure not only to click the link, but to enter the code Matthew. Replit has been a great partner, so please check them out.

All right, next, let's continue talking about OpenAI. GPT-4.1 was released earlier this week. This is the successor to GPT-4.0: better, faster, cheaper, more efficient. And it almost was forgotten as quickly as it was announced because so many other models came out this week. But as we can see here, we got a family of three models: GPT-4.1 Nano, Mini, and the full version. And that is as compared to 4.0 Mini and 4.0 full. This is a chart with multilingual understanding on the y-axis and latency on the x-axis. Now, unfortunately, it's unlabeled, which is pretty much a crime against humanity, but it's another fantastic model in what was an OpenAI model release week.

Next, Anthropic, not to be left out of the conversation, this week released a few new features. One, they launched Research, which is essentially Deep Research, but they just named it Research. So here's what that looks like: "Hey, I'm planning to take a 3-month sabbatical to hike the Appalachian Trail. Turn on Research beta and go." It probably looks exactly like Deep Research on Grok, Deep Research on Google, Deep Research on OpenAI. But one thing that really stands out about it is they have an integration into the Google Workspace suite of products: that is Gmail, Calendar, Docs. And that is incredibly powerful. I cannot overstate how powerful that is. And I've been waiting for this; I've been waiting for an AI tool that can draft email responses for me. And I've already started testing it with Claude. Grok has something similar now; Gemini has something similar now; and they all really came out in the last week or so. So now you can use AI to search, to create, all through your Google Workspace. And I use many Google products, so I'm very excited about this.

Next, Grok GQ has released Compound Beta. Disclosure: I am a very small investor in Grok. It takes the open-source models, which they already power with insane inference speeds, and adds tool use as part of the API call. So the first two tools that these models will get: web search and code execution, really the two most important tools as of now. Compound Beta uses iterative server-side tool execution to answer complex queries. It can autonomously decide when and how to use tools such as web search and code execution and run them multiple times before returning a response. So this is already what a lot of the frontier closed-source models can do, but now we get it with Grok open-source and insane inference speeds. Compound Beta is powered by multiple openly available models already supported on Grok Cloud, including the latest Llama 4 models. It uses Llama 4 Scout for core reasoning, with Llama 3.370B assisting with routing and tool selection. So really cool; check them out.

Next, Cling, the text-to-video model company, has released Phase 2. First, Cling 2.0 Master here for video generation: even better prompt adherence than the 1.6 model, greatly enhanced dynamics, and improved aesthetics. So it takes an image: "The man first laughs happily, then suddenly becomes angry, pounding the table and standing up." This is the old version; this is 1.6. So pretty good; the hands look a little bit weird; don't look exactly natural. Okay, decent. Now let's look at the new version, 2.0 Master. Yeah, everything looks better; it's much more dynamic, much more fluid; the physics look better, the lighting, the smoke; everything looks better. Let's look at one more example. Here is Cling 1.6. We have a girl in a park; everybody walking by; everything looks okay, although if you look, the people walking by kind of look awkward; they're walking at an unnatural gait. And now let's look at 2.0. Now everything looks fast-forwarded; it's blurrier, and it looks like everything around the girl is moving quickly, and the girl is moving slowly. So this looks so much better, so much more natural. So they also say there's a significant improvement in dynamics: more range of motion on the character subject with fluid movements and natural speed; natural look with details even during the most complex movements for an immersive experience; also improved visual aesthetics; more dramatic expressions for professional-level acting. And so check out Cling; they make a fantastic AI video product.

All right, back to OpenAI. It is reported that they are in talks to buy Windsurf for $3 billion. Now I have some mixed feelings about this. Anytime that you're acquired by a company that provides the underlying infrastructure of your product, your product is going to obviously work a lot better with that infrastructure, with OpenAI models. And so if I want to use Claude, if I want to use Gemini, maybe in the future they don't focus so much on those models anymore. But I do have hope; I'm optimistic that they will continue to support all the models. Now, right now this is just being reported; it is not confirmed in any way. But for OpenAI to make this acquisition, I think it actually does make a lot of sense. Vibe coding is going to enable an explosion of builders, people who can build software who previously couldn't, either because the learning curve was too high or the cost was too high. Now you can basically build anything you want just with natural language. And you know, I'm all about vibe coding. And of course, OpenAI wants to extend beyond just the intelligence layer. Intelligence, as I've said, is becoming commoditized quickly, so they need to build applications on top of it. And Windsurf is a great application on top of the intelligence layer. So we'll see what happens with this story.

Next, OpenAI apparently is working on an X-like social network, which makes this tweet from Sam Altman make all the more sense. So back on February 27th: "Meta plans to release standalone Meta AI app in an effort to compete with OpenAI, ChatGPT. Sam Altman: Okay, fine, maybe we'll do a social app. Turns out that might have actually been true. And then lol if Facebook tries to come at us and we just Uno reverse them, it would be so funny." So Sam Altman frequently does this; he'll just say things that they're building. And I again think this is a great call. Building a social network is nearly impossible, especially today; the network effect is so difficult to build. But ChatGPT already has hundreds of millions of users, so being able to get that initial traction should be fairly easy for them. Now why does this make a lot of sense? Well, it's all in the data. The reason X is so valuable, the reason Meta's platforms are so valuable, is because they have so much data, and continuously new data to train their models on. OpenAI does not have that, so they have to buy data; they have to create synthetic data, but they don't have an organic data source like a lot of these other model providers do. So if they were able to build a social network successfully, they would have this self-generating data system. And so that sounds cool, and I hope they do that. And if it's AI-native, all the better.

Next, Microsoft just announced computer use in Microsoft Copilot Studio for UI automation. Computer use is the next frontier of agentic behavior. This new capability allows your Copilot Studio agents to treat websites and desktop applications as tools. With computer use, agents can now interact with any system that has a graphical user interface. Some of the examples that they show: automated data entry, market research, invoice processing. And here's the key: reimagining robotic process automation, RPA. This is a many-billion-dollar industry that they are essentially saying they are going to flip on its head. And yeah, I believe it. Browser use, computer use; this is going to change the RPA industry completely.

Next, also not to be left out of the conversation: Grok Jr. Okay, Grok adds memory. Grok now remembers your conversations. When you ask for recommendations or advice, you'll get personalized responses. This is an incredibly important feature for any personal AI. Maybe some of you don't want AI to remember things about you, but I personally do. I want to develop a shorthand with my AI assistant; I don't want to have to tell it everything about me every time I prompt it, and I want it to reference past conversations. Of course, you can always opt out of these memories; you can always delete the memories, as I showed earlier. But personally, I'm very excited about it. So memories are transparent; you can see exactly what Grok knows and choose what to forget. To forget memories, tap the little book icon beneath the message. Coming soon to Android; it's in beta, available on grock.com and the iOS and Android apps, excluding the EU and UK, for obvious reasons.

So that's all of our stories for this week. What an incredible week it was; all about the models this week. And thanks again to Replit for sponsoring this video. Check it out; I'll drop all the links in the description below. Make sure not only to click the link, but to enter the code Matthew. If you enjoyed this video, please consider giving a like and subscribe. And I'll see you in the next.