📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Googles AI Innovations, OpenAI Agents & More AI Use Cases

The AI Advantage24:16

Transcription

All right, so the common story amongst all the AI releases this week really has been things merging. On this show, I regularly point out that all the little features and GitHub repos or previews that we might look at will eventually be implemented into the software that you're already using, or the big player is just bundling it into one product. And we've seen a bunch of that this week, most of it happening by Google, in both the Google search and image generation arena. But beyond that, we have more interesting releases; some of them really hyped, like Manus AI, that claims to be the Chinese operator that actually gets things done. And then we have a brand new agentic framework coming out of OpenAI, again bundling multiple things together. So we'll talk about all of that and more in this week's episode of AI news we can use; the show that looks at all the generative AI releases from this week, and we have a first look or test of all the ones that, according to the Advantage team, matter.

All right, so I want to start this week's video with the Google AI Studio and the image generation capabilities that were added here. And even if you might not be interested in image generation, I think this advancement is really a big one because, as I pointed out, it's starting to bring things that we've seen in places isolated by themselves into one interface. Concretely, they've been stepping up the multimodal capabilities of their Gemini 2.0 flash experimental models; the main model that you also find in Gemini Advance, their ChatGPT competitor. The point here is this: up until now, if you wanted to generate an image, well, you tell it to generate an image, and you get a result. If you wanted to edit that image, there were other applications that did just that. And if you wanted to stylize the image, you could also do that with specific applications like Midjourney, but you have to use specific commands, and it just took a little bit of knowledge to get to the point, whereas here, all of it just happens through chat, at the highest quality level, which is probably the most user-friendly interface to do this.

So let me just show you. So I just switch this output format to image and text, and I say, "Generate an image of a cat with a hat." Voila, we have that there; pretty standard stuff. But this editing capability is the new thing here. So if I say now, "Make it hyperrealistic," let's see... 4.7 seconds. Oops, maybe the word was "photorealistic," so I'll just follow up the prompt by saying that. Okay, so there you go. That didn't work. How about, "Remove the hat?" Nice. "Turn the cat into a tiger." And as you can see, you can really iterate over this with text, just like you could over chat messages in something like ChatGPT. There you go, a tiger. What about "An angry tiger," and "Add the hat back in." There you go. Not very angry. Okay, it's not doing the angry thing, but I think you see the point here. I want to try one more thing, which is uploading one of my own images. So let's see, I just have one of these thumbnails laying around on my desktop. "Change the person's expression to be more excited." I'm really curious if it can do something like this. Oh my God, it did it! That's terrifying! Here's Johnny! But also kind of powerful. Okay, what if I say, "More excited," one more time? Oh my God, it's going crazy! This is starting to look like JD fans. Okay, it also changed the image, though. What about text editing? "Change the text 'ChatGPT' to 'Gemini,' for example." ChatGPT Gemini. Sort of, sort of. I'll try one more prompt to fix the text. Okay, now I removed everything else, but it sort of did it. So as you can see, it's not perfect, but honestly, this first one is really damn impressive. I haven't seen another tool executing this that well. I mean, their imaging model is one of the best, so this is pretty damn impressive, and I would expect an interface like this eventually making its way to all chatbots, because why not? It's just the best implementation. Let's round this out by saying, "Now turn this image into a cinematic movie scene," and seeing what it does on something a little more creative like that. Okay, not bad. Honestly, that background editing, the color shifting—I mean, that looks really good. If you kind of cover up the part of the screen that shows the mutated me, you could legitimately use this for graphics work if you need to get things done quickly. And the expression change is impressive. And this is me just using a random Google account here in AI Studio, and as you can see, I didn't even link an API key; a certain amount of tokens is free of charge here, so you can just go and try this for yourself and see what you make of this innovation.

Okay, followed by that, there's one more improvement here in the AI Studio which I really like, and that's the way to integrate YouTube videos. Let me show you quickly. If I go over to YouTube and let's just say I want to take this video from Google DeepMind, I can all of a sudden do this: "Summarize this," paste the link, and run this prompt like so, and it just needs to take in the video, and this just seems like such an intuitive thing to happen. But if you think about it, none of the other chatbots really do this. I mean, if I put this into ChatGPT, it hallucinates up something from a different video. Look at that; that's not even the title of this video. This one is Gemini Robotics and not AlphaGo, the movie. And it just immediately pulls up the video and summarizes it for you. Now Google products have always been best with YouTube videos, but I just wanted to highlight this because I know a lot of people do work with videos and online videos for either content creation or their research, and to me, this just does seem like the best interface to do that right now. And I reckon over time we'll just see a feature like this in every LLM platform out there.

Just quickly wanted to show you that this is available today for Google's AI Studio. By the way, just a little side note: I currently am a little bit under the weather. It was actually my birthday this past week, so I had a little birthday weekend. I had the best time, but I don't know, I suppose being 31 years old and doing a lot for a few days straight kind of gets to you. So I'm feeling better already, but yeah, not really at 100% yet. I hope you still enjoy the content, though. Let's get back to the next story.

Okay, next up we have another Google story, and this one is them essentially expanding their AI search features. And essentially what they're doing here—and this is just rolling out; some people have access to this; for other regions, it's coming—but they're adding a new tab to Google searches, and it just says "AI mode," and it's essentially Perplexity that's built into Google Search now. I myself don't have access to this, but team member Daniel actually got access to this already, and he tried it out for us, and right here in the screen recorder you can see how this performs in practice. So basically, the journey starts with google.com, as it usually would: "Chick-fil-A news 2025." I mean, what else would you be looking for in your free time? And then here at the top, next to the normal Google stories, they have a new tab called "AI mode," which gives you functionality just like Perplexity or ChatGPT search would, too. Now this is different from their initial approach where they kind of forced this AI summary into the top of their Google Search; now they're putting it into a separate tab. And considering how much traffic Google already has, from the looks of it, this is probably bad news for the competitors like Perplexity. Either way, they're just playing around with this, and considering how convenient AI searches can be and how much adoption they've seen already, I would expect this as a fixed new tab within Google searches sometime soon. It just makes sense. If you want different links, you still have Google, but if you just want the info, well, the AI mode will give it to you. And this is sort of part of a bigger trend; right, many of these services that we kind of point out week by week as they release, they come out, people try them, and if they're really sticky, these bigger players rebuild them, integrate them into their applications, as they already have the traffic and therefore the distribution.

Okay, now let's move on to the next one. Okay, and there's another big AI release coming out of Google this week, and that's the release of their Gemini free model that, as they claim, is the best single model that you can run on a GPU or TPU. They show this ELO score in Chatbot Arena, where people rank outputs of different models and how well this performs; better than Llama 2 7B, slightly worse than DeepSeek R1. But this model is really small; I mean, look at that: 27 billion parameters, and DeepSeek R1 has 671. And a 27B size you can usually run on a lot of MacBooks. But they do say that the performance is optimized for NVIDIA GPUs, and I got to say all the benchmark scores look really solid for a model this small. And if I read this technical report correctly, 64 GB of RAM should be enough to run this 27B model on your machine locally. So in practice, with 64 gigs, the full context size could not be used; this would require too much memory. But if you're okay with, let's say, 8,000 tokens of context, then yeah, this could probably be ran on 64 GB of RAM. But there's always so many factors involved with this that it will really depend on your machine. But, but the fact is, there's a new model like almost every week, and only time will show if people really like using this. But the trend of smaller and smaller models being more and more capable continues here.

There seems to be some ChatGPT update every single week, so I really do want to cover that because I do happen to know that most viewers are actually using that as their daily driver in terms of LLM platforms. And what they implemented this week was yet again something developer-focused: it's the ability for ChatGPT to write code directly in your IDE or code editor. For that, you do need the desktop app, and as you can see in this little demo, it will just connect to it. And with something like this, you don't need a Copilot in your IDE anymore. Sure, there's some functions that it still performs better, and no, I'm not a full-time developer that could appreciate every single little feature that a little Copilot might bring inside of an IDE, but I think for most people having a workflow like this where you have a free code editor like VS Code and you have it write code right inside of your code editor and you don't have to copy-paste things is an absolute blessing. And unless you need one of the specialty features that some of these Copilots bring, this is a fantastic feature that replaces the core functionality, which is seeing your code base and adding to it. And in the context of this trend of vibe coding that we'll also be talking about today, you don't even really need to know exactly what's going on in these files; you just need to be able to talk to ChatGPT and you need to have a vision for what outcome you want. But we'll talk a little bit about that in the section of vibe coding, as that is a trend that is increasingly becoming popular, and I also like to cover those on the show.

Yet another ChatGPT update is a tweet from Sam Altman. And this is not a release that came out this week, but I wanted to point it out, and it's essentially him saying that they trained a specific model for creative writing. The models we see right now are not focused on that; all the reasoning models, they're focused on math, science, and coding. You have to realize that writing is not the main focus here. And if you haven't seen this yet, I do recommend you pause this video for a second and read these three paragraphs. I won't be reading all of this out to you, but as an AI writing connoisseur myself, I can tell you this is like nothing we've seen before. GPT-4.5 right now, in my opinion, is state-of-the-art when it comes to writing, yet that model is still not optimized for it, and a model that is specifically trained to write well is just a whole different beast. I mean, this is high-quality, human-like writing. And one thing that I'm actually curious about, and I haven't done before: what happens if I take this entire text here and throw it into something like GPTZero, which is an AI checker? And there you go; it says, "We are highly confident this text is entirely human." And this is AI-written. So for anybody who has been doubting AI writing before, just think of the fact that these models are general purpose and they're meant to do everything; just wait for the ones that are specifically made for creative writing, or use GPT-4.5, and you might be positively surprised. But let me tell you, formulations like this—I've never seen an AI model do this before; this is really good. So whenever you find yourself criticizing outputs of LLMs, or somebody else is doing that, just ask the question, "Hey, are they using the best model for that? Or has that model even been created yet? Or is this a problem yet to be solved in the coming months?" And in the case of writing, hopefully we'll get to live test a model like this soon. I personally am looking forward to that. I thought this was important to cover, but now let's move on to some releases that are actually available today.

Another ChatGPT update this week is that they made Operator available to all regions. Now you don't need a VPN anymore; anywhere in the world. So if you have a pro account, which is a $200 account, you can use Operator now. But honestly, after testing it thoroughly, many of the things that would really be time savers and productivity unlocks just don't work yet. And that's where we enter another story of this week, which was Manus AI. So this little application out of China has caught so much attention upon release on X that people are freaking out that, "Hey, this next big Chinese release!" But then I think it's fair to say that the hype kind of died down quite quickly. And essentially what Manus is is a better Operator; that's how they present themselves, and that's also how it works in practice. They've actually released this, and although this is not publicly available to everybody, you can apply for an invitation code—that first they were giving out quickly, then slower. I did manage to get my hands on this and had about 3 hours of playtime in total because they weren't my own access codes and accounts; one of them was a community session where, one of my office hours, we kind of just played with the thing for one and a half hours and threw different prompts at it. And here's kind of my summary: especially when contrasting it to something like Operator or Anthropic's Claude; with both of those, I've spent a substantial amount of time and tested—I don't know—got to be over 100 use cases for both of those with different approaches. This thing, Manus, is better in certain ways and way, way worse in others. The main way in which it's worse is simply stated: it cannot use any accounts. And as most of the internet is behind an account, that's a massively limiting factor. I mean, if you think about it: every social media platform, your email, your Uber Eats account, heck, even something like Google Sheets or Google Docs—all of that is hidden behind your simple login—but it cannot use those because Manus is built upon a combination of: one, the Sonnet 3.7 on top of a Mopic model that's really good at multi-step thinking and coding, planning, logic; it's a really good model for something that's agentic like this. And then, two, they have their own version of the Chinese Qwen model, which they optimized to work here. And because they're using Mopic under the hood, Mopic is super, super limited in terms of stuff like using your accounts or saving passwords. And by limited, I mean there's a red line that you just cannot cross, and it doesn't let you do most things. Just a quick reminder that Mopic is kind of the most conservative company in the space when it comes to AI safety; I suppose Grok from xAI would be the other side of that spectrum. So even if you look at the very first example on their website—"Planning a trip to Japan," one of the classic examples that these companies love to do—it is super impressive because it uses a combination of Sonnet 3.7 and a web browser, just like Operator, and then it runs codes and to do certain calculations and it does research for you and puts it all together. And that's a good use case, but at the end of the day, you could be using something like Notion AI Pro to put together an itinerary like this, or you could be running a deep research. And as a matter of fact, let me tell you, it actually does a better job. Like look at that; it's way more detailed; it has the opening times; it makes more tailored recommendations; surprised me personally more often, but that depends on how you prompted. And then in the end, it created this massive table of all the different places and the opening times in the places for me. And I'm actually planning a Japan trip now for April, and what I'll do is I'll just simply print this table, bring that with me, and I have a more detailed itinerary than this put together. And I could keep speaking here for another 20 minutes and talk about different examples and prompts and how differently they perform, but essentially it all comes down to this: due to the fact that it cannot use accounts online, a lot of the things that actually work inside of Operator, like ordering groceries or booking trips for you, don't work inside of this. Fair enough, but you might say, "Those are not the most useful use cases." Well, true. And if you look at these, none of these are really things that you would want to do with Operator; these are things that you would more likely be doing with deep research. So what I kind of found by myself, without even referencing any other creators or anybody else's reviews here, is that this is more of a deep research competitor than it is an Operator competitor, because nobody really knows what goes on in the background of Deep Research; maybe it writes and executes its own code, and obviously that has a web browser and a thinking model plugged into it, too. It's just that the thinking model of Deep Research is LLaMa 2, and this is Sonnet 3.7, and LLaMa 2 is still ahead of the curve of everyone else; it just is. These responses just don't cease to surprise me, and I still run Deep Research most days of the week. So if you're looking for something like B2B suppliers or a list of YC companies or analyzing a stock, well, these are things you could be doing within Deep Research, but I think the thing that you get here is a more interactive experience; it shows you a to-do list; it uses different tools to complete it, and it works really well in many of those cases, unless you need to use some account, which is also a limitation of Deep Research. But that's why I kind of concluded that this is more of a Deep Research competitor, and it's just another slightly weaker Deep Research, in my opinion. And it's fine; it's just not the new open-eye killer or the next Chinese DeepSeek or whatever some people on X make this up to be. And again, this is just based off a few hours of usage, and I might change my opinion over time when I try more use cases or see specific things that go way beyond anything that I've seen so far. But so far, good product, but definitely overhyped, in my opinion. Just my two cents here; feel free to discuss in the comment section; maybe I'm missing something. And if you don't have an early access code and want to get your hands on something similar today, there's already an open-source repo called Owl. I didn't really play with this, but it's a workflow with multi-agent assistance that does all the things that Manus does, and you can kind of customize it. And as with any of the other open-source alternatives, I would probably expect this to perform a little worse. Again, this one I didn't actually try; I just wanted to point out that it just popped up on my radar, and I wanted to show you that people are rebuilding this already. And on this Gaia benchmark that Manus kind of boasts here, they achieved a score at 58, which I suppose puts it behind both Deep Research and Manus, but it's open source.

Okay, on to the next one, which is the OpenAI Assistants API and their new agentic framework. Okay, so this one is obviously a very developer-focused release from OpenAI, and they basically launched a brand new API which includes things that you really didn't get access to before, and it's all under the hood of one thing that they call Assistants API. And really simply explained, I would say that before you had an API that responded with text, and if you wanted to search the web and get responses from that, you needed to use a different API. And if you wanted to upload images, you needed to use a different API. And if you wanted to build a functionality into your app where you could upload documents and you would get an AI response, well, you guessed it: you needed a different API. Concretely, that was the Assistants API from OpenAI before. Now they kind of combine all of that into one thing, which is just called the Assistants API. So you just send it a request, and this Assistants API just responds; it includes internet search, file upload, and even Operator access, which allows the API to kind of use a computer and browse the web and do things. Pretty cool! I'm excited to see what people build with that. As I said, Operator is not extremely capable yet, but for little tasks across the web, it can be useful. I personally sort of stopped using it since it ordered like two or three dozens of bananas to the wrong address in Lisbon, and then I made another order, and it ordered another basket full of groceries to yet again the wrong address. But hey, all that is behind the API now, and you can kind of call it programmatically, which is amazing. And then they came out with this second thing called the Agents SDK. Rather than explaining that, I want to show you an example from Stripe that I really, really liked, and that is an example of how Stripe implemented this with their payment services. And it's an example of a freelancer having sort of a list of all the people he served that month, and in some form he expressed that, "Hey, these people have paid, and these people have not paid." This is a perfect example for an AI tool. I mean, look at this little table right here; it's a little check mark or a cross; maybe this was created in Notion, or heck, this could even be a handwritten note that you just provide to the agent. But as you can see, it lists all the clients, and then together with—have two little agents: one responsible for searching the files and one responsible for actually making the Stripe invoices from the Stripe side—the agents could kind of look at this little table and go for everything and automatically issue and send out emails with all the invoices that were still missing. And this way, you can weave together different functionalities, different apps, and create these little programs that run with agents in the background and get everyday tasks done. Now look at this, or the second example that Stripe also showed, which I thought was amazing, too, which is basically having an email agent that scans for emails like, "Hey, I need an invoice for this purchase," and then the agent automatically fetches that invoice and sends it back to the customer. I want to be clear: these aren't things that were not possible before; they were just a bit more convoluted; you had to use a specific platform and multiple API endpoints to kind of make everything work together. And now with OpenAI's own SDK, this just became so much simpler. And simpler also means it's easy to troubleshoot, and it's more reliable. So yeah, all these agentic workflows are really happening step by step, week by week; we keep progressing; they keep getting better, and I'll report back if I see any big unlocks that are relevant for a wider consumer base. But there you go. I think especially with that Assistants API, OpenAI really delivered for all devs that use AI in their applications.

Okay, next up is kind of a fun one, and we like to throw in live demos. And this is not affiliated, but we found this new application that does something new. We've seen similar things with ElevenLabs and Synthesia and other players doing AI video, but this company, Tavos.io, is really connecting a lot of these tools, and their claim is that, "Hey, all of these chatbots shouldn't actually be interfaced like ChatGPT; they should be a person sitting there talking to you, just like I am right now." And that's what their product does, and they have a little demo which I always like doing during this video, so let's give this a shot. "Start a video chat with Charlie" over here. Just get out my headphones so I can actually hear Charlie, and let's start this video chat. "Yo, what up, Charlie? How you doing today?" "I'm doing great, thanks for asking; it's always awesome to meet..." So, ah, listen, man, you're a bit robotic over there, and also I don't know, it doesn't feel so fluent; what you got to say for yourself? "I see what you mean; I can come across a bit robotic at times, but hey, I'm working it, man; I'm a large language model, so sometimes I might sound a bit stilted, but I'm always trying to improve and be more like a human, you know? Like a cyborg, but without the cool robotic limbs, brother." I love your... Okay, that's a great... Okay, okay, fair enough. Uh, please, first of all, don't interrupt me; secondly, can you tell me a joke? "Sorry about interrupting you earlier; I'll make sure to let you finish your thoughts from now on. Okay, here's a joke for you: Why did the neural network go to therapy? Because it was struggling to process its emotions. Hope that one computed correctly and brought a smile to your face." Um, actually, no, Charlie didn't, but there you go. I wanted to do this to show what a mix of these different tools might look like. I think it might be super early, and obviously this product is not at the stage where it's going to be changing the world yet, but I think we're getting closer and closer. And last week we looked at some of these voice generation models that sound so convincing; here they tried to plug multiple things together with a result that might not be that convincing yet, but I wanted to do this little segment to point out that, hey, we're getting closer every week, and this is something you can keep in the back of your mind as we review some of the releases; eventually they will all be combined into one thing.

All right, on to the next one, which is an overview of the top 50 generative consumer web products as published by a16z about a week ago. I just quickly wanted to spotlight this if you ever were curious what people are actually using. So as you might know, at the AI Advantage, we have our own rankings, but this is just sort of a big overview based on traffic, and also keep in mind they are a venture fund that specializes in technological investments, so obviously these will be tinted by their economical interest, you know? So take a list like this with a grain of salt. Nevertheless, it's a good list, and I'm kind of proud to say that I tried and tested most of these. There's a few curious ones which I haven't heard of before: What is Joyland, or CrushOn AI? Oh, and yeah, I've also haven't tried SpicyChat, but that's not really the domain of apps that I try. I mean, I haven't visited it, but I can imagine what SpicyChat would do. Naughty, naughty. But if you wanted to explore some new apps, well, this is one list to do that. And also, this is a great point to yet again remind you that we do a ranking of our own that we update every single month, though. But this is purely based off my team's and our community's personal preferences; these are the apps we see members actually using day-to-day, with reasoning why they're up here. And if you have a gripe with this list, you can leave a comment below, and we can discuss this down there. But we update this every single month; we do this for LLM platforms, image generation tools, and video generation tools on a monthly basis. Kind of considering a vibe coding ranking right now; that might come soon. Let's see if that makes sense. Either way, if you're looking to get oriented within the space, this is probably a great starting point. And also, I want to add one more comment here, which is: I know a lot of people come to this channel as complete newcomers to AI, and they want to just start from the beginning. And arguably this video might not be the very best place for that; we reference information and releases from past weeks, and this is really about staying at the cutting edge and finding out about the new things that keep coming out and not about building foundational skills. But I realized that many people want to do that, so if you're just getting into this, I have two recommendations for you that we built: one of them free and one of them paid. The free one is our weekly newsletter; we release it for free, and we put together our onboarding sequence, and you get a prompt template right in the beginning with some of our favorite prompts and techniques to get you kickstarted right away. Everything newsletter-related is completely free, and we don't even sell sponsored slots because we want to keep the experience as pure as possible there. But in the process of receiving some of those materials, we will recommend our community, which is really our premium experience and the ultimate answer to, "How do I get into AI? How do I actually learn some of these skills? Or where can I get my questions answered if I have any?" The community features a structured onboarding and multiple courses that will teach you foundational skills; you can pick the ones that matter to you, and then we run around 20 events a month and release weekly guides and resources to keep you up to date, even if you don't engage with them on a weekly basis. The structured workflow guides and the video courses—for example, in March we're releasing a brand new one on fine-tuning models, the very best technique in the world to make LLMs sound exactly like you—all of that comes with a community membership. So if you're interested in actually learning this AI thing from the ground up, then both the newsletter and the community are the very best ways to do that.

That, and that's really all I got for this week; a lot of developments, a lot of things merging, as I mentioned in the beginning. And again, I realized this stuff is not easy to keep up with, so if you're just getting started, sign up to our newsletter. And other than that, I will see you very soon on another episode of AI news you can use. But this is all I have for today; see you soon.