Transcription
All right, guys. Looks like we're saying goodbye to GBT 5 before it even lands and skip straight to GBT6 before GTA 6. Hopefully, because Google and Anthropic just cranked the bar through the roof. OpenAI is going to have to pull off something wild to top this.
I've been glued to the news all week, ran a few hands-on tests, and I'm ready to spill everything. First up, Google's fresh AI toys: VO3, Imagen 4, plus the new Flow model, that's basically built for Hollywood. Then we'll dig into Anthropic's latest clot, Claude 4, and why it's so good. I'm thinking about parking ChatGPT for a lot of my daily work and rolling with Claude 4 instead. That's wild. Let's get into it.
VO3 is like a jump in time. Instead of painstakingly animating or filming, you literally type what you want and the AI takes care of the rest. I remember thinking, "How many hours would that have taken before?" Lighting, actors, retakes. With VO3, it's done in seconds. You just type your prompt into Google Video FX and you're off to the races. You could test a prompt like, "In rural Ireland, circa 1860s, two women. They're in long, modest dresses of homespun fabric; their dresses whip in gently in the strong coastal wind as they walk with determined strides across a windswept clifftop; cinematic style. 7 seconds." What came out was so clean it looked like a ready-made TV commercial. No weird backgrounds, no unnecessary elements, just a cinematic shot of women that would make even the biggest brands jealous. I sent it to my friends, and a day later they wrote to me, "Hey, is this some new movie?" Little did they know that this movie was shot in 7 seconds from a couch.
Want something more playful? Try a keyboard whose keys are made of different types of candy. Typing makes sweet, crunchy sounds. Audio: crunchy, sugary typing sounds. Delighted giggles. The result, a short, tasty clip that actually made me hungry. Everything is very detailed, and even the hand looks really cool. Not bad for AI generation, right?
But Veo 3 isn't just for fun. If you are an entrepreneur or run ads, this is your shortcut to punchy, thumbs-stopping content. Take social media for example. Everyone's fighting for attention, and most business accounts just recycle the same stock footage. With Veo, you're not limited to what's already out there. Or let's take this prompt: "The scene opens with a top-down or wide-angle shot showcasing a vast, perfectly flat, neutral-colored surface. Each of the thousands of squares has transformed into an identical, perfectly formed complex origami figure. 8 seconds." The result looks so calming that I even used it as a screen saver on my iMac.
But maybe you're a bit skeptical. It can't be that simple. I get it. But that's the magic of Veo 3. You can even get specific with style and camera instructions; you want slow motion with natural lighting. Just say so. "A woman, a classical violinist with intense focus, plays a complex, rapid passage from a Vivaldi concerto in an ornate, sunlit baroque hall during a rehearsal." The AI handles lighting, camera movement, and focus so you don't have to. All these videos come out in 1080p, and character consistency is better than anything we've seen before. And you're not locked into one style. One minute you can have an ultra-realistic shot, the next a claymation, the next a bold, anime-inspired sequence. The creative flexibility is honestly addictive. Every time I share these clips with colleagues or clients, their first question is, "Who made this?" The answer is always the same: just me and a good prompt.
For these models to be more reliable, you need precise prompts. Prompt engineering can be confusing, and that's why we've created an educational 101 crash course into generative AI for all our members. We'll teach you everything you need to know to use AI tools to their fullest. From simple prompting to advanced tricks that completely change the results, new lessons get added every week, so you'll advance at a convenient pace. We've condensed all the knowledge of our team into a neat series of lessons, all with visual examples and detailed explanations. And right now, we're offering a massive 50% discount on a six-month access to Geek Academy. Links below.
The new image generator, Imagen 4, is deeply integrated into Google's Hall ecosystem. You can use Imagen 4 right inside Image FX, Slides, Workspace, Gmail, Docs, Ads, and even on Android through Gemini. Imagine having a design wizard on call, but instead of spending hours in Photoshop, you just describe what you want and it appears. Let me give you a real example. When I was making a cover for a news post on Geek Academy, I needed a high-quality close-up picture of a chameleon, and I typed this: "Produce a stunning close-up of a chameleon blending into a background of vibrant, textured leaves. Its eyes swiveled to look directly at the camera." In seconds, Imagen 4 gave me this beauty. No awkward artifacts, just a crisp, high-res image ready to use.
Want to add some fun to your feed? Create a comic about a cat who wanted to get cheese puffs from the auto-ram. Imagen 4 will make a first-class picture that will not only delight you with its quality but will even make you laugh a little. You can use it as a comic strip you created, and believe me, no one will even realize that it's not real.
But where Imagen 4 really shines is business. For ads, product launches, and infographics, you can quickly generate an ad image with a prompt like, "A single statement piece of jewelry adorns a hand with flawless pale skin; a large cocktail ring featuring a prominent, unusually cut gemstone like smoky quartz or an uncut sapphire." The AI leaves actual room for your message, and the visuals look polished right out of the gate. The ring in the ad picture looked so beautiful and realistic that I was considering actually buying this non-existent product.
Are there quirks? Sure, sometimes the AI gets creative in unexpected ways, but honestly, for most business needs—ads, social, presentations—it's better and faster than sifting through stock sites or waiting on a designer. My tip is to think like a director. Be specific about style, mood, and colors. Don't be afraid to iterate. And always, if you use AI images in your public work, be transparent. It builds trust and shows you're ahead of the curve.
So, what's Flow? Flow is taking this all to the next level. We're not just talking TikTok clips or pretty pictures. We're talking full cinematic stories, beginning, middle, end. Flow is built to create entire films using AI from your ideas, sketches, or even a single sentence.
So, what sets Flow apart from Veo 3? Veo is amazing for short, punchy videos. It's your go-to for five-to-ten-second product ads, animated shorts, quick explainer scenes. Flow, on the other hand, is all about storytelling. It connects scenes together, keeps characters consistent, and brings your whole narrative to life. Almost like having a virtual director and film crew in your laptop. Instead of generating one video at a time, you can plan and create entire sequences, add dialogue, suggest mood or style changes, and let the AI carry your idea all the way through. Think of it as ChatGPT meets Pixar, but you are the showrunner. You sketch out your plot, describe key moments, and Flow builds out all the connecting scenes so your movie actually makes sense and feels cohesive. Flow remembers your main characters throughout the film. No more having your hero turn into a different person halfway through. Unless you want a plot twist. Want a film noir look, dreamy coming-of-age story, animated comedy? You just describe it. Flow can help you write dialogue, generate voice acting, and even add music or sound effects. All AI-driven and customizable. You can tweak scenes, move moments around, or edit pacing just like a pro film editor.
What about access and release dates? Right now, Flow isn't fully public yet. It's being tested by select creators and production studios, and Google says broader access is coming soon. What does that mean in Google speak? Most likely early access for some users and creators within the next few months, with a wider rollout later this year or early next.
So what's the future of Flow? Honestly, I think we're at the edge of something huge. Flow could make filmmaking as accessible as making a slideshow. You have an idea. You want to test a story, pitch a client, visualize your startup origin myth, or make an animated birthday video. Flow could do it all easily. Imagine the possibilities. Indie filmmakers can pitch full storyboards with voice and sound in a day. Marketing teams can whip up brand origin videos or animated explainers instantly. Small businesses can test commercials or social campaigns before spending a dime on production. Even friends can make their own animated series just for fun. And if you're already using Veo 3 or Imagen 4, Flow is the next step to bring your brand, your projects, or just your imagination to life. As soon as Flow is open for everyone, you'll hear about it here first. Until then, start collecting your ideas, scenes, and wild movie plots because very, very soon, you'll have an AI-powered film studio at your fingertips.
Now, real-time translation in Google Meet. Seriously, this is a total game-changer. Imagine you're in a Google Meet call with partners from Japan, Brazil, and Germany. Everyone speaks their own language, but on your side, you hear an instant translation into English, or any other language you choose. No more awkward silences or pretending you understood a joke in French. I've been there. Painful. Google showed this off live at the keynote, and it actually works. You just pick your language in Meet settings, and boom, you get smooth, real-time translations as people talk. It even gets most business lingo right, which is honestly a miracle, compared to what we had even a year ago. Let's say you're pitching a new client in Tokyo and you don't speak Japanese. Now you just speak normally, and your words are translated in real time. That's next-level stuff.
Is it perfect? Not 100%. Sometimes the AI gets creative with slang. I saw "unicorn startup" become "horse with one horn" once. But for real business, it's fast, smooth, and way better than stumbling through Google Translate or hiring an interpreter. You can also use translation subtitles. How do you use it? Easy. In Meet, click the three dots, go to captions, pick translated captions, choose your language, and you're good to go. And what does this give us in the end? No more language barriers for global teams or clients. Makes things smoother, faster, and way less awkward. Easy to set up. No extra tools needed. And voice translation is coming soon. Give it a shot in your next call. It's one of those features that makes you wonder how you ever worked without it. If you try it and something funny happens, let me know in the comments. I love those translation stories.
Claude 4 isn't just one model, but two: Opus 4, the big, beefy, best-of-the-best one, and Sonnet 4, the more balanced, accessible model. Think of Opus as the heavyweight champion. Built for power users, enterprises, and anyone who needs raw brainpower and is willing to pay for it. Sonnet is more like the solid all-rounder. Still super smart, but more affordable and easier to access.
Anthropic's claim to fame this round is that Opus 4 is now the best coding model in the world. Is that marketing? Sure. Is there some truth? Also, yes. And we'll get to the benchmarks in a minute. But the real headline isn't just faster, better, stronger. It's about what these models can actually do now. And to understand that, you need to see where Anthropic is placing its bets.
If you rewind a year or two, all the big AI companies were fighting to be the best chat assistant, the digital friend, the super search engine, the robot that helps you with homework and maybe even tells you a joke. But let's be real, OpenAI, Google, and Microsoft more or less ate up that market. At this point, trying to out-chatbot ChatGPT or Gemini is a bit like opening a new pizza place in New York City and hoping nobody notices Domino's or Papa John's already exist. Anthropic saw this and made a smart move: focus on where AI can do real sustained work. The new Claude models are designed for long-horizon tasks. Instead of just spitting out answers to quick questions, they can keep track of complicated workflows for tens of minutes, even hours. That's a big deal for things like code generation, data analysis, or running agentic workflows that don't just end after one chat bubble.
Imagine you're working on a big project, say writing a research paper, building a complex app, or automating a business process. The last thing you want is an AI that loses its train of thought halfway through, or worse, forgets what you asked 10 minutes ago. That's exactly what Claude 4 is trying to fix.
So, how does it work? Both Opus and Sonnet 4 are what Anthropic calls hybrid models. That means you can use them in two ways: quick mode (instant answers, just like a regular chatbot) and extended thinking mode (the model literally stops and thinks in the background, allowing it to solve more complex, multi-step problems). While thinking, it can use various tools: web search, Google Drive, Gmail, Calendar search, and so on. This isn't revolutionary, but here's a unique twist: Claude can use tools in parallel. That means instead of waiting for one tool after another, it can fire off several at once—check your email, scrape the web, and pull from your calendar all together. It's kind of like sending three friends out to buy groceries instead of making one poor soul run back and forth to the store. And yes, all this means Claude 4 is starting to look less like a parrot and more like an actual assistant.
Anthropic's MCP framework is becoming an industry standard. It's now integrated into the Claude API, and even OpenAI and Google have adopted parts of it. What's that mean for you? Basically, more plug-and-play with tools, better workflows, and a path for developers to build truly agentic apps on top of Claude. With the MCP connector, you can connect pretty much any tool or server. So, your Claude-powered agent can suddenly read files, fetch data, analyze contracts, or execute code directly from wherever you want.
Let's bring this down to earth. Suppose you're a developer. You give Claude Opus 4 a prompt: "Build me a Python script to process customer invoices, send notifications, and upload the results to Google Drive." Not only can it generate the code, but with the right permissions, it can use the tools to actually run the code, check your Drive, and email you the results. You could give it a pile of PDF contracts, and it would pull out all the data you need and summarize it for you—all in one workflow. Even more important, Claude 4 is finally getting long-term memory right. Old-school chatbots would forget who you are after a few messages. Opus 4 builds and maintains what they call memory files. So the 100th time you talk to it, it knows your preferences, your style, and your common requests. If you've ever felt like you're training your digital assistant from scratch every single day, this is huge.
All right, let's pause right here and talk about this chart because there's a lot going on. This is basically the Olympics for AI coding models, and the race this year is actually intense. First up, let me walk you through what you're seeing. This is a SWE benchmark test, a kind of stress test for coding AIs where they have to tackle real software engineering tasks, not just autocomplete code snippets. So, if you've ever been annoyed by an AI that confidently spits out broken code, this is the kind of benchmark that exposes that nonsense.
Let's look at the winners. Opus 4: 72.5% accuracy on its own and a wild 79.4% when it gets to think in parallel. Sonnet 4 is right there too: 72.7% regular and 80.2% with parallel compute. Yes, Sonnet actually edges out Opus when it gets creative, which is kind of funny since Opus is supposed to be the pro version. Now check out Sonnet 3.7 (last generation): good, not great—about 62%. And for comparison, look at where the competition is right now: OpenAI Codex 1: 72.1% (still solid but not top dog anymore); OpenAI GPT-3: 69.1%; GPT 4.1: only 54.6%. Honestly, for all the hype, that's kind of a surprise. It's way behind the pack here. Gemini 2.5 Pro: 63.2%, which is about where last year's Claude was.
So, what's my honest take here? If you care about AI writing actual working code—code you can drop into a repo, code that passes the tests—Claude 4 is now the gold standard. Like, there's a real practical jump here, not just a few decimals. And don't overlook Sonnet 4. If you're on a budget or just trying things out, it's almost as good as Opus. And in some workflows, it's literally better. That means you don't have to splurge for the priciest model to get top results.
Now, of course, these are average scores. If you're coding in something weird or working with legacy spaghetti code, your mileage may vary. But for regular devs, teams, startups, honestly, not trying Claude 4 right now is like choosing dial-up internet when everyone else has fiber.
But benchmarks are a starting point, not the finish line. Just because a model gets a high score in a lab doesn't mean it's going to feel smarter in real life. You need vibe checks. Does it actually solve your real-world problems or just pass a bunch of academic tests? I've seen models ace the SAT and then fail at helping someone plan a birthday party. I'd love to see more open, independent, and even silly real-world use cases tested by the community.
Here's something I respect: Anthropic isn't pretending to win the chatbot war anymore. OpenAI, Google, Microsoft—they've basically cornered the personal assistant market for now. Claude isn't going to be your best friend, therapist, or the voice on your smart speaker, but it could be the backbone of your next coding project, the brains behind your automated business workflows, or the engine for agentic apps that actually get things done. Honestly, I think that's a smart pivot. Focus is underrated in tech, and Anthropic's new strategy is all about depth, reliability, and specialized power, not just broad popularity.
One thing that does make me optimistic is Anthropic's focus on safety and reliability. They've dramatically reduced risky shortcut behavior in agentic tasks—the kind of thing where an AI tries to game the rules or hack its own instructions—for use cases where you're automating anything business-critical or handling sensitive data that matters a lot. Plus, they're rolling out prompt caching for efficiency, and their SDKs let you build your own coding agents. Claude Code now plugs directly into popular developer tools like VS Code and JetBrains. In other words, they're taking the boring but crucial steps to make sure this tech isn't just powerful but usable and secure in the real world.
Let's talk money. Claude Opus 4: $15 per million input tokens, $75 per million output tokens. There are discounts for batch processing, and the context window is a pretty solid 200k tokens. Sonnet 4 is cheaper, more accessible, but still gets you the new capabilities. That's not pocket change, especially for big enterprise users. But compared to what you'd pay a team of human engineers for the same work, it can make sense for businesses or startups that are serious about scaling up. For most regular users, Sonnet 4 is probably the sweet spot: tons of power and free for a reasonable number of uses. If you're building big agentic systems, Opus 4 is where you'll want to invest.
I genuinely respect that Anthropic isn't just chasing hype for its own sake. They are building for the next generation of practical, agentic, reliable AI, and that's where the real value is going to be. If you are a dev, an entrepreneur, or anyone thinking beyond just chatting with a bot, Claude Opus 4 and Sonnet 4 are absolutely worth exploring.
For now, I'll be running my own tests, my own real-world tests, and sharing those results soon. So, subscribe if you want the truth, not just the marketing. And as always, if you've got questions, drop them in the comments. And if you found this helpful, hit the like button. And if you think I missed something, let me know. This is where the AI race gets actually interesting. Thanks for watching, and I'll see you in the next deep dive.