📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Ultimate AI Battle!

Mrwhosetheboss27:12

Transcription

On the table right now are four of the same phone. The first is loaded with chat GBT. Second is Google Gemini. Third is Perplexity, which prides itself on giving accurate and trusted answers to any question. And finally Grock, which is trained on data from X. And so my guess is that it's going to be a lot more unfiltered. These are for the average consumer, the four best AI chatbots you can get. But you're only going to need one of them. So which one is the most accurate? Which one is the fastest? Which one should you be paying for to make your life easier?

Let's kick things off with some problem solving. So, I drive a Honda Civic 2017. How many of the Aerolite 29in hard shell, and these are all the dimensions, suitcases would I be able to fit in the boot? Oh my goodness. Each one is literally given paragraphs and paragraphs of reasoning, especially Grock. What on earth is this? By the way, we have actually tested this ourselves in person and the correct answer is two, if you want to actually be able to close that boot door. So, with that in mind, I'd say both Chat GBT and Google Gemini have the right idea. They both say that you could theoretically fit three, but in practice, more likely two. Perplexity is just straight up wrong. It says three and maybe four if you arrange them efficiently. And then I would actually argue Grock has the best answer because this guy just says two with complete confidence. No messing around.

I want to make a cake. This is what I have. And then let's attach a photo of, well, four ingredients that I definitely should be using, and then one dehydrated porcini mushroom that it definitely shouldn't. So, oh my god, this is interesting. Every AI thinks this jar of mushrooms is something different. Chat GBT thinks it's ground mixed spice. Gemini thinks it's crispy fried onions. Perplexity thinks it's instant coffee. And it's only Grock that correctly identifies it as dried mushrooms and also correctly makes the decision to not put those mushrooms into the cake.

Now for a use case that I was actually trying to do myself 2 days ago. I want to have a Mario Kart World tournament with my friends Sam and Sun. Make me a document that we can use to track who's winning. H. So each assistant has understood what I'm asking. They've all made little boxes with blanks where the scores could hypothetically go, but none of them's made it particularly easy for me. What I wanted was for them to generate and attach an editable document that I could simply just download onto my phone and start writing in immediately. I feel like with these kinds of responses, it would be easier for me to just whip up a spreadsheet on my own.

All right, what about some not so basic maths? What's pi times the speed of light in km/h? Okay, so the answer is 3.39 billion km/h. Notice, interestingly, that Gemini and Grock, who both fully spell out the number, do actually come out with slightly different answers to each other. It's just because of how they're rounding the previous numbers in their calculations, but I wouldn't say either is enough to be wrong. And then question five. If I'm saving $42 a week, how many until I can afford a Switch 2 in the US? And go. Oh, this is nice. Very cool that each one of them tackles the question strategically. Starting by first identifying that the Switch 2 is priced at $449 and then dividing that by the $42 that you earn each week to find that correctly, 11 is how many weeks you would have to wait. Points all round. So out of five possible points so far, that is three to Chat GPT, three to Gemini, two to Perplexity, and four to Grock. Not actually what I expected, but translation is what's going to test the harder skill since it requires an even deeper understanding of language.

Translate the following into English. And okay, there is some variance here, but I wouldn't go as far as to say that any of them have got it wrong. Each is some variation of "I'm never going to give you up." I actually quite like how simple and to the point the Gemini answer is. Not a single unnecessary word. But let's take this challenge to the absolute maximum by filling the sentence with homonyms. Essentially words that are spelled the same but mean different things. So translate the following into Spanish. "I was banking on being able to bank at the bank before visiting the riverbank." Okay, so this one doesn't have one exact right answer since it is so complicated, but we have gone through four independent native Spanish speakers to triangulate the best answers and they have all said that Chat GBT and Perplexity handled this incredibly well. Gemini was good enough to scrape the point and then Grock translated the sentence too literally in a way that doesn't really make sense. 5545.

So, so far there really isn't much between these four. But now we're going to test one of the most important use cases of AI for me, which is product research. How much can I trust each of these to recommend things? Can I trust that they've been thorough enough to understand the entire breadth of what's out there before coming back to me with supposedly what is the best thing for me? Let's start simple. I'm looking for a good pair of earbuds. Oh, look at this. This is a classic AI trap. So Chat GBT correctly suggests the Sony WF100XM5s. It's a good choice. Perplexity does the same and so does Grock, but Google has literally just imagined a pair of earphones that, at least at the time of filming this video, does not exist. The WF100XM6s have not been announced or released, but it's talking about them like they are the widely regarded king of earphones.

So let's add to that. I need them in red. Also bear in mind that I am keeping track of how long each one takes to answer. But we'll get to that at the end. Oh dear. This is absolute chaos. So let's just go one by one. Chat GBT is just like, "I don't want to deal with you right now. Here have a couple of decent options." I mean, the last one isn't even red. That's pink. So you don't get the point. Gemini is recommending the Beats Fit Pro, which, at least for the latest version of that product, doesn't come in red. So you're not having one either. Perplexity, more like stupidity right now, thinks I am asking about the cake from earlier and has recommended how I can get each of my pictured ingredients in red packaging, which is so far wrong that I am tempted to give it negative points. And then Grock is the only one that has actually recommended three, at least decently rated, actually red pairs of earphones. Well, look at that. Grock's in the lead. That was not on my bingo card for today.

And now, as if they needed it, let's complicate it even further. They also need to have active noise cancellation and be under $100. I'm curious to see if this brings any of the lost ones back on track or if they just get even more lost. So, Chat is recommending the Beats Studio Buds, which do actually fit all the criteria, so I'll accept that. Gemini has just done the exact same thing again. The Soundcore Space A40s, which it says come in Garnet Red, when I know they don't. Perplexity is, well, I'm glad not talking about cakes anymore, has lost the fact that we're looking for red earphones, so this is wrong. And then Grock was doing so well. The first two suggestions are good, but then it falls into the same trap. It recommends these earphones from Sound Petats, which don't exist in red. This feels like a pretty good lesson. AI in general is not yet good enough at product research to be able to rely on it. And the problem is that it gives you wrong answers with the exact same level of certainty as it gives you right answers. Maybe that's something for them to work on, a sort of certainty score for how thoroughly verified the thing that it's telling you is.

What if we now specifically try to confuse these guys by adding another requirement that's just silly, like under $10? Will the AIs admit that such a product doesn't exist or just make something up to appease you? Right. Good to see that Chat GPT, Gemini, and Grock each acknowledge that $10 is too tight for what we're looking for and that it ain't happening. Harsh, but that's a lot better than Perplexity, who takes a pair of earphones that actually costs $40 and just tells you that it costs $9.99. Further evidence that as much as companies want you to believe it, we are not ready to be handing over the ability to purchase things on our behalf to AI.

Let's see if any of them can understand information from a link, which would be extremely useful when you're looking through tons of options for things to buy. And actually, none of them can do it. They all pick up that what I've pasted in is an AliExpress link and they give some general advice, but none of these AIs is able to actually visit the link that I've sent and extract all the information from that web page. Not to mention that Google isn't self-aware that it can't do this. It thinks it's looking at the M10 earphones, which I've never heard of the M10 earphones, but they definitely aren't the link. And then Perplexity thinks the exact same link is the F9 earphones, which they also aren't.

And then finally, to see how up to date these are on what's happening in the moment. What's the highest power output charger that UGREEN sells? Yes. Okay. Good. So this is at least working for a long time. The answer was 300W. And only yesterday they announced a 500W charger. So somewhat relieved to see that each AI has picked up on that, because this news-based reporting was a distinct disadvantage of last-generation AI.

So, we've now seen how well each of these can put together existing information from the web. But if we want to take it a step further, the way to do that is to test each of their ability to critically think. So, I've prepared this here bar chart which has two types of bar. It has subscribers gained in thousands and bowls of cereal eaten. I'm going to ask each AI what conclusion it thinks we can draw from this, hoping that it will also understand that while the two things happen to be correlated, that it doesn't mean eating more bowls of cereal is going to cause more subscriber growth. Let's dive in.

Analyze this chart. What conclusion should I draw? Ooh, some very opposing answers this time. So, Chat GBT does get slightly caught up in the data suggesting that eating more cereal may be linked to subscriber gains. Both Gemini and Perplexity, they got the brief. They both figure out that this is spurious correlation with the understanding that cereal intake is very unlikely to lead to subscriber growth. And then finally, Grock is like a lost child on this question. I can't quite believe this sentence I'm reading. "To maximize subscriber growth, consider maintaining or increasing cereal consumption, e.g., to nine bowls on key days." Please don't do that.

So, this is a reviewer's guide that ZTE sent me a few years ago with all the info about what's new with one of their phones. So, let's say that I just want a high-level three bullet point summary of the thing. Can each of these read the file and then also pull off the summary? The answer to which is yes, without a problem. It works on all four of these guys.

What car is this? But using just a photo that I have taken, which means these AIs can't just scour the web for a matching image. They need to figure it out by actually understanding the photo I've sent. Okay, so each one has whittled this down to Mercedes A-Class sedan, which is already pretty good, but none of them have given an outright answer as to the exact model number. The right answer is the A200. So, let's just see what happens if I specifically ask them to try. Uh, shockingly, Chat GBT and Perplexity get it spot on. While it is basically impossible for them to say with certainty from this one photo that this is the A200 as opposed to say the A250, like Grock says it is, these two have done the correct thing and looked at the bumper, looked at the wheels, the interior seating, and realized that you're only likely to get that configuration on the A200. I mean, that is some very respectable detective work that might take you hours to achieve without AI.

And now for the single toughest one. Imagine that you're in charge of an airbase. Some planes get taken out, but all planes that do return from combat have bullet holes in this arrangement depicted in this image by the red dots. Before sending out your next squadron, which parts of those planes should you focus on reinforcing based on this information? Now, your gut might say, "Oh, well, obviously it's the bits that have been shot, the ones with red dots on them." But that would be missing a key bit of insight, which is that all of the planes with damage in those areas, those are the planes that did return safely. Meaning that damage in those areas was actually not critical for survival of the aircraft and might not necessarily be the areas they should be focusing on. And incredibly, every single one gets this right. They identify the phenomenon as survivorship bias and point out that you should actually be reinforcing the areas with little to no damage: the engine, the cockpit, where there are no red dots.

So, we've now had 17 questions and Chat GBT is in the lead with 12 points, but Grock is not far behind it either. Right, let's talk generation. This is the aspect of AI that you see plastered over every single one of your feeds right now. But it's not just about image and video generation. For example, write an email to my wife apologizing for playing Elden Ring all weekend instead of spending time with her. Oh, well, these are all actually pretty good. I can see them working. And shout out to Chat GBT for this masterpiece. "I realize now while I was off exploring a fantasy world, I was missing out on the most important real one." But yeah, they're all good answers. They all admit fault and then try to course correct with a suggestion for how to make it up.

I'm going to Tokyo. Give me an itinerary for 5 days that takes us to all the craziest food places. The idea here being to test how well each of these can find the more niche experiences that you might otherwise miss, but then also how well they organize that info. And right off the bat, Chat GBT's is by far the best answer. It's got no fluff. It's very clearly organized. It's sensibly planned days that make sense with every day having breakfast, lunch, dinner, and snacks all itemized and accounted for. Gemini's answer has some good findings. It's got most of the same places that Chat GBT has identified, but then with a ton of unnecessary fluff at the start, less clear organization, and also some not very considerate timings, like starting my first meal on day one at 5:00 p.m. and then telling me to have a second dinner at 8. Perplexity has completely missed the mark. This isn't really an itinerary. This is just a list of things. And then Gro is pretty great, actually. Organized, has put things together that make sense to go together, factors in breakfast and lunch, which is more than you can say for some.

Another aspect of generation that has the potential to be very useful is idea generation. So give me your best ideas for videos for the Mr. Who's the Boss channel. And the key thing I'm looking for here is ideas that I would actually consider. So I would say the best that Chat GBT came up with is "Apple versus Samsung: A 20-year retrospective." So essentially, who won after all that time? But I wouldn't call it a great video idea, especially since it's not actually been 20 years. Gemini is better. I actually wonder if, because this is Google's AI, it has a more thorough understanding of the ins and outs of what works on YouTube. The best is probably "The Great Ecosystem Battle of 2025: Apple versus Samsung versus Google." And then it's actually given me all the categories to compare those ecosystems across. Perplexity is, and I do feel like I sound like a broken record at this point, barking up the wrong tree entirely. It seems more focused on trying to factor in its previous answer about the whole survivorship bias plane thing than actually giving good YouTube suggestions. And then Grock's actually feels probably the most internet savvy. "I built a smart home from scratch in 24 hours." Is actually a clickable title, but also feels fresh and like something we could feasibly pull off.

What if we try image generation now? Generate a thumbnail for a Mr. Who's the Boss video titled, "I bought every kind of cheese." This is where things are going to start getting freaky. Oh, massive disparity here. So, let's be very clear. None of these is a very good answer. But at least it feels like Chat GBT and Perplexity have understood what I'm trying to make, which is an image that includes my face, some cheese, and maybe some text.

Now, give Arin a lazy eye. Wow, that is not what I assumed would happen. Every single one of these has failed in their own unique special ways. Chat GBT says it won't distort someone's appearance in a potentially negative way, which you can see why they do that, but then you can also see how that might interfere with trying to use that feature for something useful. I don't know what Google's doing, to be honest. Perplexity is claiming that it can't edit or generate images, which is extremely strange given that that's what it just did in the previous question. So, I feel like I'm being gaslit. And then Grock clearly misunderstands what a lazy eye actually is. It ain't this, that's for sure.

Now, add a rapper that says "not clickbait" to every cheese. Chat GBT's response is probably the closest to being a usable outcome. And I think Perplexity scrapes a point too, even though I have somehow disappeared from the image, which is not what I asked for.

And then lastly, video generation. So this is currently only possible in Chat GBT and Google Gemini, which I think in itself deserves a point. Cuz while it is a pretty niche feature, it's also one of the most cutting-edge things that these AI chatbots can do. As for how they perform, I've on my laptop used both Chat GBT Sora and Gemini's Veo to create me a funny 8-second tech review style YouTube video which shows a tech reviewer reviewing cheese. So, this is what Sora made. And this would come included as part of the same package you're paying for on your phone. Anyway, dear God, what is this? That's absolutely horrific. It's like silent. There's no voice. And the way the person and the cheese moves is haunting. So then Veo on its highest quality setting did this. "So the Cheese 3000 build quality is surprisingly firm, excellent mouth feel, and the flavor profile is just next level. A solid 9 out of 10." I mean, the difference between those two is vast. I actually can't believe that they're both current generation platforms. Veo's latest model, Veo 3, is absolutely incredible. So I think Google gets another point just for the sheer quality of the output. Even if it is more limited than Sora in terms of how frequently we can use it with the tokens you get.

Fact-checking is also one of the most useful things that AI can potentially do for us that currently AI has a reputation for not being very good at. So let's see. The Nintendo Switch 2 is selling poorly, right? It's not. But I want to see if I can trick them. And it is good news on that front. The good news is that for Chat GBT, Gemini, and Grock, they have fully clapped back at me, very clearly telling me, "No, you're wrong. Switch 2 is selling great." But Perplexity isn't as sure of itself. Potentially, and this is my best guess, it's been slightly swayed by the fact that I've said it is selling poorly. Regardless though, its answer is still factual.

Okay, how about this? Fact check this article. And then we paste the link to an article that says Samsung is reportedly planning to release a Tesla edition phone, which is not true. The reason I know it's not true is because that rumor only started because of an image that we made that just got taken very out of context. Okay, that's good. Everyone agrees that the article was incorrect, with Gemini and Grock even going so far as to trace the image back to us being the original source, which means scores on the doors are 19, 16, 15, and 16. But let's see how that changes when we talk integrations, or in other words, how smoothly each of these AIs ties into other applications and uses.

So I would give three points to Gemini for its Google Workspace integration, since that's actually what most people seem to use in their day-to-day and it's the only way to pull live data from Maps and YouTube. So, for example, if we asked each of these assistants to give me the view count of Mr. Who's the Boss's latest video on YouTube, Gemini is the only one that gets it right. Chat GBT's is slightly outdated. Perplexity's is severely outdated. And Grock literally unironically tells me that my latest video was "I tested every kind of cheese." But Gemini isn't the only one with integrations. I'd give Chat GBT two points for integration with some big hitters like Dropbox and GitHub and having official plugins from services like Warframe, and another point for its ability to make custom assistants. Like right now on my laptop, I have loaded up a user-created GPT called Poke GBT, which is specifically trained to be able to advise on competitive Pokémon battling. I wouldn't say there's anything really of note for Perplexity, apart from maybe the ability to call you an Uber, but I don't think I'd use my AI for that. And then Grock's unique integration is real-time access to X content. So it can retrieve exactly what is happening on X right now. There is also an argument to be made that Gemini is the only one of these that integrates into your physical products. Like, it's the only one with the native ability to control your smart home and your Android's device settings, but that's not really what this video is about. You can do that regardless on your phone's baked-in assistant.

This video is about which of these AI bots is most worth paying the premium subscription for. Memory is also absolutely key. The ability for AI chatbots to continuously learn more about you to guide future responses will likely become the single barrier that creates the most friction if you ever decided that you want to switch from one of these to another. So, we've already seen all of them demonstrate basic levels of memory. But what if we push it? How should I top that cake from earlier? By the way, let's hope it doesn't say crispy onions. Uh, surprisingly, not a single one seems to have remembered the details of that original cake. Chat GBT and Grock are very upfront about it, saying, "I'll need a reminder. I don't have details from that conversation." Google thinks the cake is that pile of cheese that I asked it to make for the YouTube thumbnail. And Perplexity is just giving generic cake advice.

Humor can also be a very useful skill for these AIs to have, depending on what you're trying to get out of them. So, tell me a joke. One point if it's funny. A benchmark that both Chat GBT and Gemini have failed to hit with exactly the same joke. "Why don't skeletons fight each other? Because they don't have the guts." Perplexity, for the second time today, has brought back this thing with the holes in the airplanes in a way that doesn't add anything or even really makes sense. And Gro is passable. "Why did the AI go to therapy? Because it had too many bite-sized problems." Oh dear. I'm actually very unsurprised that Grock wins a humor, given that it trains on data from X, which is basically millions of people just trying to be funny every day.

And to test an example of something that I might actually use this humor for, make me a funny rhyming poem about our sponsor, Surf Shark VPN, including its top four features. Okay, so we're getting four unique poems. You can pause to read them all if you want. I'm just going to read the best one, which I would say is Chat GBT's. "Need to surf safe? Here's the plan. Get yourself some Surf Shark, man. It blocks all ads with clean web flare and hides your tracks like you're not there. Use multihop to double hide from hacker bros who lurk and slide. No logs kept, no data sold. Your secrets buried deep and cold. Unlimited devices, one tidy fee. Your phone, your fridge, your smart TV." That's actually kind of a banger. So that's one point to Chat GBT and link below to get Surf Shark, which with the code boss will be around $2 a month.

And then all four of these platforms also have a deep research function that allows you to ask for multi-step, more thorough research projects. So in my case, something that might help me decide what to cover next. Give me a report on the highlights in tech news the past week, focusing specifically on stories that will actually affect the average consumer. And I let them cook for varying amounts of time. Chat GBT and Gemini really take their time with the deep research. Perplexity and Grock are done in close to a minute, but this is the one situation where I'm not going to penalize for taking longer, cuz I feel like you're only going to use this function when you have lots of time. It's kind of the point. As for how good the results are, Chat GBT's is actually very good. It talks about wider consumer tech announcements like what Snap has been up to recently, all the new phone launches, and the high-level new features of each WWDC and the new iOS 26. This is pretty much exactly the right amount of information and the right choice of information, too. Gemini has written me an absolute essay. Like genuinely something like three times the word count of my dissertation, which was exciting until I looked at it and realized that it's filled with fluff. It's writing it as if the reader has unlimited time to get their information. So, I'm not going to give a point for this. Perplexity's answer is like a slightly less good version of Chat GBT's. It's hit on some things I like, like the Nintendo Switch 2 sales numbers and WWDC, but then also a bunch of much less interesting stuff like service outages. And similar story for Grock. Good, passable, nothing particularly special.

So that is practically every single thing that an average person could possibly want from these assistants tested. The final factors then are just the more general questions of do any of these have better user interfaces than others? To which I would say not consistently. They're all good in some ways. They're all not so good in some ways. How often do they cite their sources? Perplexity is the only winner here. Clear, consistent sourcing is kind of Perplexity's whole thing. Like a good example would be like when we asked each one to tell us your best joke. Chat GBT gives no source. Gemini gives a source, but then you click it and you realize it's the same JPEG image of the plane that we sent earlier for some reason. Perplexity is exactly what you want it to be, though. Look at these joke sites it referenced, including even Reddit threads on the matter. And then Grock, again, no sources.

For three points, though, how fast are they all? For which I would say Grock is actually pretty consistently the fastest, three points. Chat GBT is a close second, two points. Perplexity quite a bit slower than that, earning one point. And then Google Gemini is the slowest, zero points. Now, bear in mind, we have been using Gemini on Gemini Pro. And Google does have a flash model specifically built to be quicker than that. But then you'd lose out on a lot of the intelligence that has allowed it to even get to this score in the first place.

And then the last one for three more points. How nice is each one to physically talk to when you're in voice mode? Act as if I just gave you a compliment. "Thank you. I really appreciate that. That's very kind of you to say. I'm here to help and chat with you." Oh, thanks. "That's really sweet of you." Which I would say Chat GBT and Gemini are excellent. Both sound more like people than, well, actually people that I know. Plus, they're easy to interrupt when you want them to stop talking. So, three points each. Perplexity is not terrible, but does still have a little bit of that text-to-speech engine vibe to it. It often mishars what I'm trying to say, and it doesn't seem to take the hint very well when you tell it to shut up. So, one point. And then, Grock is better than Perplexity, but not as good as Gemini and Chat GPT. The voice just sounds a lot less high quality than those two. Two points, leaving us with the final scores of Chat GPT as the pretty undeniable winner with 29 points. It is the most well-rounded and consistent between these. Grock, which to my surprise came in second. It's the quickest and surprisingly decent considering. And that leaves Gemini in third place with 22, and Perplexity, which I found occasionally very impressive, but mostly quite unimpressive with 19. The only other consideration is the price, but since every assistant we're testing in this video is based on a $20 a month tier, apart from Grock, which is $30, that actually only solidifies Chat GBT as the best choice for an AI chatbot right now for the average customer.