📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

I Tested Gemini vs Claude vs ChatGPT So You Don't Have To

Parker Prompts8:51

Transcription

For the past 2 weeks, I've had ChatGPT, Claude, and Gemini open side by side on my screen, running all three through the same prompts across writing, coding, reasoning, research, [music] and everything in between. And after running all three side by side, I can tell you that every single one of them has a category where it gets outperformed badly. And if you're paying $20 a month for the wrong one, you're getting worse results than someone using a competitor [music] for the same price. So, I'm going to show you exactly where each one wins, where each one loses, [music] and which one is actually worth your money based on what you use it for.

The first round is writing, because this is the task where you feel the gap between models faster than anywhere [music] else. I wrote one prompt, and I'm going to drop the exact same thing into all three models and compare what comes back. And instantly, ChatGPT 5.4 comes back with a structured, competent response. The information is there, and the logic tracks, [music] but the tone feels flat. Claude Opus 4.6 looks like a person wrote it. The tone matches my instruction, the phrasing is natural, the formatting stays clean without adding anything extra, and the overall voice feels like something I could use without changing much. [music] This is where Claude has been ahead for a while, and that gap hasn't closed. Gemini 3.1 Pro lands somewhere in the middle. The writing is clean, and the tone is solid. [music] This round definitely goes to Claude. If writing is your primary use case, whether that's emails, content, [music] client work, or anything where voice matters, Claude is still the clearest winner, and it's not particularly close.

Writing showed how each model communicates, but none of that matters if the model can't actually reason through a hard problem, and that's where round two gets interesting. By the way, I put together a free cheat sheet that breaks down exactly which model wins for which task, plus all the prompts from this video ready to paste. You can grab it in the description below.

The reasoning round is where I wanted to see how each model handles a messy, multi-layered problem with no obvious answer. So, I gave all three the same analytical prompt. So, ChatGPT 5.4 actually has a new feature called steerable thinking plans, where it shows you its reasoning up front before generating the full response, and you can adjust the direction mid output. The analysis is solid. It frames churn clearly, ranks the causes, and explains the reasoning. It also gives practical steps like segmenting churn and tracking activation, but it still stays at a fairly high level. Claude Opus 4.6 goes deeper. It's a solid and well-structured answer, especially in how it prioritizes onboarding and time to value. The diagnostic steps are practical and useful. That said, it runs a bit long and explains things that are already fairly obvious. It's good, just not very sharp. The extended thinking mode gives it time to reason through the problem before responding, and you can feel that extra depth in the output. The tradeoff is that extended thinking burns through tokens fast, [music] which is worth keeping in mind if you're on a usage-capped plan. In independent benchmark testing [music] that specifically measures how well a model solves problems it has never encountered before, Gemini scored 77.1%, which is over eight points higher than Claude's 68.8%, and you can feel that jump in the output. The reasoning is sharper, the connections between ideas are tighter, and it reaches conclusions that the other two missed entirely. This round goes to Gemini. The reasoning upgrade in 3.1 Pro is the single biggest performance jump any model made this cycle, and it shows in every analytical task I threw at it.

Thinking through a problem is one thing, but actually building something functional from a prompt is a completely different skill, and that's what round three puts to the test. I'm not a coder by trade, but I've personally built small projects in all three of these tools, so I know what good output should look like, and I can tell when something is cutting corners. For this round, I gave all three the same build, and right away, ChatGPT 5.4 spits out code quickly, and the output runs on the first try. For [music] quick scripts and simple prototypes, it gets the job done fast. But when independent teams run these models through real-world coding challenges involving multiple files and complex dependencies, ChatGPT consistently scores around 23% lower than the other two. Claude Opus 4.6 is a bit different, and I'll show you why. It explained that local storage wouldn't work in its environment, changed the approach, and generated a live Kanban board preview instead of giving raw code. So, rather than outputting code you can copy, it built and showed the result inside its own workspace. Where Gemini 3.1 Pro pulls ahead is on larger codebases. It's more consistent at handling very large context, which makes it better when you're working across big projects. This round goes to Gemini, with Claude close behind. Claude is strong, but in this example, it didn't even return raw code. It built inside its own environment instead. Gemini is more consistent for actual coding tasks, especially as projects get larger. ChatGPT is still behind both here.

Three rounds in, and every model has shown a clear strength. But the next round isn't really a fair fight, because one of these three doesn't even show up. Claude does not offer any image generation at all. If you need an image while working inside Claude, you're leaving the platform and opening something else. Anthropic hasn't announced any plans [music] to change that, so it's not something that's coming anytime soon. That leaves ChatGPT and Gemini, and this is where I expected ChatGPT to win easily based on how long it's had a built-in image [music] generator. But Gemini 3.1 Pro now has Nano Banana 2, and after testing both of them side by side, Nano Banana 2 really impressed me. ChatGPT's image generator is solid. [music] It handles text and images reliably, follows complex visual prompts, and produces results that are usable for social media, thumbnails, [music] and mockups. For a long time, it was the only real option built directly into a chatbot. But Nano Banana 2 on Gemini is producing photorealistic outputs that are on another level. The lighting is more natural, the detail on textures and skin is noticeably better, and the prompt adherence on complex scenes is tighter than what ChatGPT gives me. I tested both with the same prompts across a range of styles, >> [music] >> and Nano Banana 2 matched or beat ChatGPT's generator on nearly every one of them. The photorealism in particular is where the gap is most obvious, because Nano Banana 2 produces images that actually look like photographs, [music] whereas ChatGPT's outputs still carry a slight artificial feel even at their best. This round goes to Gemini. If image generation is part of your workflow, Gemini now gives you the best built-in [music] option of any chatbot on the market. And if you're on Claude, you'll still need a second tool no matter what.

The next round shifts into something completely different, which is how well each model handles serious research where you need accurate sources and structured analysis. I gave all three models an identical research prompt on a specific topic, and asked for a structured breakdown with cited sources. ChatGPT 5.4 has deep research mode, which fires off hundreds of secondary search queries across the web, consolidates what it finds, and delivers a structured report. It's [music] fast, it covers a lot of ground, and it surfaces current information well. Claude Opus 4.6 takes a different approach. It's strongest when you upload your own documents and ask it to work strictly within those sources. The answers stay grounded, the citations are traceable, and it doesn't wander outside the material you gave it. Claude even flags the parts that are still uncertain. Research mode is available on the Pro plan, but Claude's real advantage is when you're the one deciding what information the AI draws from, rather than letting it search on its own. Gemini 3.1 Pro has the deepest integration with Google's ecosystem. It can cross-reference massive document sets, connect to Google Scholar, and Gemini 3.1 Pro is now the model running inside Notebook LM, which means if you've been building notebooks with curated sources, Gemini can now reason across all of that material with significantly stronger analytical ability than what Notebook LM had before. For web-based research at scale, it covers the most ground of any model I tested. This one is a split, but if I have to pick a winner, Gemini takes it. The workspace connections, the Notebook LM integration, and the sheer scale of what it can pull from in a single query gives it an edge that the other two can't match right now.

Over the last five rounds, I've tested how each model writes, thinks, builds, creates images, and researches. The last round tests something that sounds simple, but is actually where a lot of models quietly fall apart, which is whether they can hold a massive amount of information at once without losing the thread. For this one, I took a large set of documents and uploaded them into all three models. Then I asked identical questions that require pulling specific details from different sections of the material. Things that you'd only get right if the model was actually holding the full document set in memory, rather than skimming. All three models now support a 1 million token context window, [music] which is roughly 750 pages of text in a single conversation. So, this round came down to recall quality and what each model can actually do inside that window. ChatGPT handled the full document set well and found details I expected it to miss. Claude's recall was the most precise of the three. The answers were tight, accurate, [music] and it rarely pulled in information I didn't ask for. And where Gemini pulls ahead is on what it lets you put inside that window. It handles multimodal inputs natively, which means you can drop video, audio, images, text, and even YouTube links into the same context, and ask questions across all of them at once. That's something neither of the other two can do in a single conversation. This round is close, but it goes to Gemini for the multimodal advantage.

So, after six rounds, the picture is clear. Claude wins on writing, Gemini wins on reasoning, coding, image generation, research, and long documents. ChatGPT offers the broadest feature set overall, but didn't take the top spot in any single category. There is no single model that takes every round. If you write or need your AI to follow instructions precisely and maintain a consistent voice, Claude is the right choice. If you work with large documents, need the strongest reasoning, or live inside Google's ecosystem, Gemini gives you more for less. And if you need the widest range of capabilities in one place, voice, computer use, and the biggest plugin library, ChatGPT covers the most ground, even though it's no longer the leader in any one area. The real answer for anyone doing serious work is to use two of them and match the tool to the [music] task. And if you want to get the most out of whichever model you choose, I just posted a full breakdown on how to use Gemini 3.1 Pro better than 99% of people. Thanks for watching, and I'll see you in the next one.