Transcription
So, at this point, you've probably seen that Gemini 2.5 Pro is topping the leaderboards. But what we're seeing here is just the tip of the iceberg. There's a number of big things that are happening just underneath the surface. Number one is that there are rumors of DeepSeek coming out with a brand new model soon, including R2, their next iteration of the reasoning model, supposedly really good at coding. Sam Alman also hinted at the fact that one of the models that's a stealth model—we don't know who made it, we don't know who is releasing it—but it seems like Sam Alman is hinting that it's OpenAI.
So here Sam is saying, "Quazars are very bright things." There's a stealth model in the LM Marina called Quazar that is very good. So that might be like an '03 or an '04 Mini, or an O4 Mini, or an '04 Mini High. We don't have all the details yet. Again, a lot of this is just rumors and speculations. These models are still in stealth mode. We don't know what they're going to be called, but we do have some hints about who they're from.
And here's the thing: there's yet another sort of series of models that are still in stealth that are being cooked up behind the scenes that are most likely Google. And specifically, the one that I want to mention to you today is Dragon Tale. So Google has an unreleased model named Dragon Tale that's outperforming everyone, even Gemini 2.5 Pro, on the web dev arena.
Now, first of all, take everything you hear with a grain of salt. These are rumors; these are anecdotes. But there's a number of people online that are all saying kind of the same thing: this thing seems to be very good at coding, especially high praise is given to it for its front-end design, web design. People are praising it for its ability to spit out landing pages very quickly that are very good-looking and very functional. So again, we're not like 100% sure about everything yet. Keep that in mind. But with that said, this looks like Google is cooking up something that's at least as good as the Gemini 2.5 Pro and potentially a lot better, by some people's estimations.
Here's what I think is the important insight into where Google is going with this. First of all, here are some of the models that are floating around that I think we believe might be Google's: Night Whisper, Dream Tides, Moon Howler, Dragon Tail (the one we just mentioned), Stargazer, Shade Brook, River Hollow. So thanks to AI for Success for providing some of these examples, and I apologize about the grainy image here, but a lot of people are saying that some of these models potentially could be better than 2.5 Pro. Some people are saying Night Whisper is better; some people are saying that Dragon Tail is better.
A recent interview from the CEO of Google Cloud might give us a hint as to what's happening here. We'll get to that in just a second. But you know, with this, the Night Whisperer is on the left, Gemini 2.5 Pro is on the right. And assuming you don't absolutely hate the color scheme here, it does look just a little bit more, I don't know, refined, somehow more modern. This feels a little bit outdated. Again, a lot of this is sort of aesthetic, personal preference, but overall, I like this one better, although the color scheme is a little bit throwing me off a little bit.
So here we have Night Whisper on the right, Gemini 2.5 Pro on the left. I think that's 2.5 Pro, although I feel like this says 2.0. It's hard to see, but certainly for web design, this looks a lot better. This looks like an actual web page: very clean, easy to see, easily organized. Here is a music visualizer. Again, they're comparing it to Gemini 2.0, but let's forget that for a second. The point is they're showcasing the abilities of the Night Whisper model specifically for web development. A lot of this is for the Arena web development, and I got to say, I mean, it looks good. It's very clean. I would certainly say this is good web design, you know, at least at first glance.
Here we have: "Generate a physics-based water simulation with balls and cups." So both get working physics; both had draggable balls. Night Whisper very clearly came ahead on markups and styling. So Night Whisper is the one on the left. Again, looks very clean, very sharp here. Night Whisper beats out Cloud 3.7 creating a three-dimensional calendar with today's date highlighted. Looks like it missed the date, but everything else looks really good. Cloud 3.7 wasn't able to render it.
So if you have—so if you haven't tried Chadbot Arena—basically, you have this side-by-side battle where you put in one prompt, and two different models kind of try to nail that prompt, trying to answer that prompt. You don't know which one's which, and you're supposed to look at the output and then see which one's better. Do you prefer Model A, Model B? It's a tie, or both are bad. I asked for a website that plays a minesweeper with fancy graphics, and Model A gave me this. I'm really liking it. Yeah, I know. This is great. Like the feedback, the little—this is great. Loving it. So Model A is really good. There's some like text here in the background that like describes it a little bit weird, but the graphics—it nailed it. Very impressed with the graphics. And then Model B created this. So all right, let's see. It does not work. So in this case, we of course say A is better. Let's see. DeepSeek R1 was the better model. Very interesting.
Now you can also do direct chat with a lot of the models. I don't know if all the um, secret models are available always in that, but here they will randomly pop up when those models are getting tested. So very often times, you know, Google and Anthropic and Grok and OpenAI will kind of throw their model into rotation to be tested out, that allows them to get some early feedback from people and also test it out against the other models, potentially to get fixed or maybe fine-tuned somehow. So could it be possible that, for example, Dragon Tale is the Google's model that's more oriented for coding, general coding? Maybe the other one was called Night Whisper, I believe. Maybe that's the one that's more fine-tuned, more oriented towards front-end development, website development, etc. And here Model B failed, but Model A produced this. So right off the bat, it's really good-looking. Definitely much better than I would have thought that it would be from right from the get-go. You know, everything's looking good. Great color scheme. And if you're wondering, yes, I know how to play this; I'm not actually trying; I'm just kind of just testing it out. So please don't yell at me in the comments that I don't understand how to play this game. I'm not sitting here playing it; I'm just testing stuff out, making sure that everything works. You don't get to put down the little question mark, just the red flags. But all in all, it's looking very good. This is the best one so far. I'm going to say A is better.
And as you can see here, A, we managed to get—so again, you're seeing all my takes here. This was the second one I've tried. This is River Hollow. River Hollow is one of the models. So again, it's likely—it seems like all of them are in rotation. Now, if I understand correctly, there's no way of selecting these hidden models in direct chat or direct side-by-side comparison, right? So you have to sort of randomly accidentally get it in the actual arena battle where they're assigned to you kind of without your knowledge. But you know, at least in this one small example, River Hollow is excellent.
By the way, it would make sense that the developers that are putting these models out there, like if they're testing it, they want to get early feedback. They don't want that stuff to be sort of corrupted or gamed by people if they can figure out how the model works, like if they can spot some telltale signs about it and then maybe skew the results one way or another. So yeah, I think that, if I understand correctly, for those new models that are being tested, you're not able to access them directly in any way.
Now let's really fast look at why this might be happening, to where you have so many different models appearing seemingly all at once in this chatbot arena. One thing to note is one of the reasons why coding is now getting such a bigger focus—I would say why so many more people are strictly interested in how well these models are coding—is simply that back in the days when we were starting, they weren't very good coders. If you gave them a simple coding problem, like a little script to write, it would be kind of a coin toss. It might get it right; it might not. And so we largely stuck to doing various word problems or reasoning problems where we try to give it some complex thing it had to think through. But now, as the coding aspect of getting better and these models are getting better and better at writing, it's almost a little bit more difficult to test them with hard problems that are purely text-based, like a mystery, who's who done it sort of thing. They might get it correctly if you use certain names, but then you switch the names of the characters, they might get them incorrectly. And that might simply be because it—that that specific problem wasn't in its training data.
And in fact, I think for myself, I've been leaning more and more away from simple word problems to really test these models, asking you to do complex coding tasks, and you know, complex—I say in quotes—but having it, for example, oneshot a complicated game. And then maybe what I'd like to do is create little little reinforcement learning and training pipelines to make the little characters inside the game learn how to play that game. Like if it's able to do stuff like that, that's fairly impressive. So it shows their ability to understand what you want, translate that into sort of how to build that project in code. And sort of on my side, I am able to easily see if it worked or not. Did the program run or did it not? And if it does run, then did it nail the thing that I wanted? Did it satisfy the requirements of my prompt? Also, if a lot of the text prompts, the whole thing kind of uh, just became, you know, Wes's story hour, where I sat here and read large outputs by these uh, large language models, which just visually wasn't the most exciting thing. With code, I feel like it's a little bit more exciting; it's a little bit more interesting. There's just more stuff happening on screen. So on all fronts, this is just a more interesting problem. Not to mention that automating coding, just from kind of an economic sort of perspective, could be incredibly lucrative. And I think that's why a lot of these companies—OpenAI and Anthropic and Google—why they're really targeting sort of these coding tools. They're targeting that as their number one goal to shoot for.
So this is Alberto Romero. So he is—so he runs the Algorithmic Bridge, and this is a recent post saying, "Google is winning on every AI front." I think the first few paragraphs here like really capture how I felt about this whole situation and also I think how the market as a whole felt a lot about a lot of this. In the beginning, Google was the favorite to win, sort of like the AI game. They had Demis Hassabis; they had AlphaGo, AlphaZero. Learning about Move 37 was kind of fascinating and kind of opened up a whole new paradigm, if you will, thinking about AI, the idea of it being able to come up with these novel moves and strategies that we as humans did not come up with or could not come up with. It was this new alien sort of intellect. And so as he says here, he's been low-key saddened by Google's constant fumbling. They had the tech, the talent, the money, the infrastructure. And the reason why—and this was kind of some speculation, but I think it's sort of maybe assumed or this is just kind of what most people believe—is AI could potentially present a big challenge to Google's main source of revenue, which was search ads.
And certainly we're seeing that now with Perplexity and Deep Research, and everyone has their own version of Deep Research. A lot of it is sort of aiming for the core thing that Google does. Why would I search something in Google and then have to find the web page and then find what I'm looking for on that web page while being bombarded with ads and popups and all sorts of nonsense? Like a lot of these pages are hard to navigate now; they're not very user-friendly. They're there to like shove as many ads in front of you as possible. And AI allows people to kind of get around that, and Deep Research—I don't even interact with a website other than, you know, if I'm on ChatGPT, I'm on that website or that app. I request some information; it comes back to me, you know, once the Deep Research process is complete; it does all the search and then notifies me when it's done. Obviously, that completely bypasses Google. So that's kind of, I think, the most reasonable explanation for why Google completely lost their massive lead that they had in terms of developing AI. They were worried that would kill a big part of their revenue model.
So as Alberto says, they didn't shoot themselves in the foot; they didn't shoot at all. And so now, you know, two and a half years after the ChatGPT moment, Google DeepMind is winning, and they're winning pretty hard. As Alberto is saying here, chances are that maybe everyone else—they don't even have a chance to win anymore. Sam Altman would of course love to poke Google whenever they were doing any new releases and try to frontrun them or release something at the same time or just before them to kind of take the wind out of their sails. But recently it seems like Google went all-in on AI, and they've been building up and shipping fast, and kind of this snowball is growing.
So where are we standing now? Number one: Gemini 2.5 Pro experimental is the best model in the world. I think most people would agree with that. I know some people still prefer Claude 3.5 or 3.7 for specific coding tasks. So of course, some of this stuff is subjective or it's relative to your specific use cases, but for the most part, in general, I think like 2.5 Pro is the current reigning king. Again, if we look at the LM arena, right? So in language, it's number one. In fact, here's like overall showing everything versus and every model. So as you can see, Gemini 2.5 Pro has number one sort of across the board. There's other models with the same sort of ranking, but as you can see, no one else is quite as dominant across the board as Gemini 2.5 Pro. It's also fast and cheap. They're giving away free access. It has a window of 1 million tokens.
Now we did have a Meta Llama 4 model, the smaller one, that had a context window length of 10 million tokens, making it bigger. There were some issues with it and questions as to how good it—it is, and it's not really like a direct competitor to the 2.5 Pro. Gemini 2.5 Flash is extremely fast and extremely cheap, more so than even DeepSeek, which was—that was its whole thing, like being the cheapest, fastest model, the cheapest to train, cheapest on inference. Gemini 2.5 Flash beats it out. This would be used in various edge applications and edge devices: phones and cars and thermostats perhaps. But just if you're thinking about like phone integration, Google has Android. Tons of Android users across the world, myself included. I'm not a huge Apple fan, or at least I should say I—I prefer the Android to the iPhone sort of ecosystem and phone and everything else. To each their own, but the Android sort of ecosystem is pretty big. Google also interestingly is the only company—I forgot exactly how they phrase it—but they sort of like have a model in every category, right? So they have their text models, the music—right, the LIA model produces music. It's not where Soundful is in terms of AI music, but chances are they're probably restricted; they're a little bit more careful with copyrights and how that whole thing is perceived. And again, eventually, they very well might catch up and do better, then especially if they're able to sort of work out some deals with the—with the music industry. They have Imagen 3, which everybody says Imagen—apparently Google says Imagine, which—Imagen 3, I guess. Let's call it that. We also have VQ, and we have Chirp for voice and speech. And while they're not winning in every single category, they are in the important ones. They're near the top in video; most people agree they're better than Sora in terms of AI video generation. Imagen is solid, maybe not—not number one, but it's—it's solid. And of course, the most important category: large language models. They're sitting at number one. The Deep Research mode is considered twice as good as OpenAI's Deep Research. And as we talked about on this channel last year when they first announced it, there's also Project Astra and Project Mariner. Project Mariner is computer interaction, so something like Operator or Anthropic's computer use. Project Astra is that assistant. You can actually go to AI to Google's AI studio and have that little on the left side stream have it use your web camera and interact with it. You can also use it on your phone, I believe. Give it access to your camera and just talk back live back and forth with the AI assistant, and it's—it's pretty good. It's very good. So that's likely going to be incorporated with the various devices like phones and various other Android things: Chromebooks, etc.
And as we covered I think just a few days ago, they're now—just now announced kind of their big thing for agents. They've created the agent-to-agent protocol, similar to MCP from Anthropic, like an open protocol for how different agents will interact with each other. They're also launching Agent Space, which is going to allow all the people around the world or the different companies that are building agents to kind of have a—I almost see it like as—as almost a marketplace. It's like what Google was for websites. You type in what you want, and then it searches for all the websites and it gives it to you. It almost seems like it's that but for agents. I'm sure different people will describe it differently, but they're basically almost building—it seems like—a Google 2.0 that's on the back of this new AI wave. They're publishing tons of papers. Demis and team received the Nobel Prize for their work with AlphaFold, and they're doing tons of sort of AI safety research and publishing things about what to expect, how to prepare for it, etc. And of course, Google is also a hardware company with their TPUs, the Tensor Processing Unit, recently announcing a pretty big breakthrough in that arena.
And this is I think a good time to play the clip from the CEO of Google Cloud talking about what it's like to work with Demis Hassabis and all the models that they're producing with Google's cloud infrastructure, how the TPUs connect to that. So here, for example, is a YouTube channel, Alex Kontraitis, where he interviews the Google Cloud CEO, Thomas Kurian, on AI competition, agents, and tariffs. About 11, 12 minutes in, they're talking about Google's in-house advantage. Because keep in mind, Google, or Alphabet as the company's called, contains all those things underneath that Alphabet umbrella. They have Google and YouTube and the hardware company and the Android and the cloud and also Google DeepMind. So what effect does that have on their ability to do this stuff? What advantage does it give them? This is a great interview, but let's take a listen to just a few minutes of this.
"What does DeepMind give you uh, that might be an advantage there, because it is in-house? We work extraordinarily closely with Demis and his team. When I say extraordinarily closely, our people sit in the same buildings. We work extraordinarily closely. My team builds the infrastructure on which the models train and inference. We get models from Demis and team uh, every day. In fact, we're staging models out to the developer ecosystem within a matter of a few hours after they have finally built. Uh, and then we take also feedback from users and move it upstream into pre-training to optimize the models. And one benefit we have at Google is all our services, whether that's search or us or YouTube, the inferencing of the same stack and same model series. So the model learns very quickly from all that reinforcement learning feedback and gets better and better. So there's a lot of close collaboration. Many times, if I can be frank, when we enter a new domain, like I'll give you an example: We built a solution for cyber intelligence using Gemini. So there's a lot of threats happening in the world. You want to collect all that threat feed. We do that using a team we have called Mandant uh, and also from other intelligence signals we're getting on what are the threats emerging. You then want to compare it to your environment to see if you've been—you know, you're at risk. And most importantly, you want to compare it to what parts of my configuration will somebody use to try and get in. And so we used our Gemini system to help prioritize and also help people hunt faster. We call it threat hunting faster. Now in that environment, the model has to learn how to find patterns in a large number of log files that people are ingesting, and that required specific tuning of the model to do that."
So at Google, you have the cloud people sitting next door to the Demis Hassabis, the DeepMind, the AI people. They have their own chips; they have massive amounts of data. And now it seems like they're fully focused on being the number one in AI. Again, it's likely not that they couldn't do that before; it's that they were hesitant to because of how it would affect their existing business. Now it seems like they're saying, "All right, let's get serious about it and let's sort of jump to the front of the line and secure our lead."
They've also recently released Firebase Studio. I've tested it. This thing is very exciting. If they keep developing and adding to it and iron out some of the issues, this has the potential to be huge. Currently, it's a little bit hard to work with; it's still a preview. So if you're going to test it out, lower your expectations a little bit. But it does seem to take the idea of something like Cursor. So it's an IDE, a way to develop code with the assistance of AI. So everything's built in; it's built on top of VS Code, an open-source developer environment. And you're able to, with just a click or two, very very quickly host those applications online. So if you have an idea for an app, you can very quickly build kind of a basic prototype and very easily host it online, have access to all of the analytics, all of that data, see how your users are interacting with it. So we went from something that could potentially cost tens of thousands of dollars or, if you're a developer, you know, many many hours of working on it yourself, to something that a child could potentially do and do it quite fast, maybe within a few hours from, you know, having the idea to developing the prototype to having it hosted online; your first user could be just a few hours after you kind of fleshed out your idea. It's still sort of—think of it as a beta—but if they keep improving it and, you know, Google can build stuff, and they really put their mind to it, this thing could be massive. Cursor, which is a similar idea, was one of the fastest-growing apps of all time. I think they went from 10 million annual revenue to 100 million annual revenue faster than any other app out there.
So all that's to say that Google is back in town. It's back on top, and they are now the ones to beat. So as you watch the space and we see the next generation of OpenAI models of Anthropic and DeepSeek, etc., kind of keep this idea in mind that Google probably has more resources and just a more broader reach into all the aspects that they need, including their own hardware to run inference to train these models. Not only that, but also as long as the search ads and online ads that Google controls keep working, they're not as dependent on having to produce revenues in these other areas. So let me know what you think. Do you feel like Google at this point is kind of—has enough of a lead to where it's unbeatable? Do you?
I think that a Google model will now be number one on the LM arena. You know, most of the time, the vast majority of the time. Maybe there'll be a few blips here and there when another model takes over.
Or do you think that we're going to still see big things from OpenAI and Anthropic and Deepseek and, of course, Grock? Right, Elon Musk and team building out a massive AI data center and trying to catch up as well?
Let me know what you think and stay tuned, because these models—all of the ones we're talking about, all the stealth ones—they're going to start dropping any week now, and they're going to start dropping fast. What a time to be alive, as they.