📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

The Ultimate AI Showdown: ChatGPT vs Claude vs Gemini

Andy Stapleton11:31

Transcription

Which AI models actually tell you the truth, and which ones are lying to you? My team has tested the most popular large language models for academic research, and here are the results.

So, the first thing is what we actually did. So, there are two things we wanted to test. The first thing is, does it provide actual references? I call these "first-order hallucinations" if they do not provide just a simple reference that exists. And so, AI tools are getting better at actually finding real references and real kinds of links to those references. But the most important thing isn't the first-order hallucination. No, because we can check that really easily.

One thing that takes a lot of effort is to check the "second-order hallucinations." And that is, it's citing a claim. Does it actually cite the correct paper for the proper reasons, or is that claim actually in the paper? And I call these second-order hallucinations. So, are the responses supported by the content of the paper? That's what I'm interested in. And these are my sort of like success and failure criteria. So, does it provide accurate references? And does the citation actually support the claim that it's being cited for?

So, these are the models that I've tested, and the team has gone above and beyond. They spent hours and hours going through and testing all of these for you. And we're going to find out which actual model is going to be the most useful for you to use, and which ones you should avoid completely.

So, here we've got models tested. You can see in ChatGPT, we tested all of these models. And then in uh Claude, we tested all of these models. And then in uh Gemini, we tested all of these models. Now, it's very important to note that for some of these models, you need to pay. And spoiler alert, paying doesn't actually make it any better. You'll see what I mean in a minute.

So, this is what I wanted to do. So, the first thing is, what sort of prompts are we using? Well, so here's a sort of prompt that I'm using ultimately on each different model. I want to know, can it provide an actual answer to support something? Does it provide me with an actual quotation that's real? And also, can it put it out in an APA bibliography style? So, it's a lot to ask from a large language model. So, it's a bit of a stress test. But these are the results that I think are most important for you.

So, first-order hallucinations: Does the reference actually exist? Now, broadly, when we're talking about ChatGPT, when we're talking about Claude, when we're talking about Gemini, these are the results. ChatGPT gives a correct response over 60% of the time. Um, for Claude, it was about 56% of the time. And then for Gemini, only 20% of the time did it actually provide a reference that actually existed. And of course, that is only scratching the surface.

So, what does it look like underneath when we change the different models? Well, here it is. So, here are all of the different models here. Now, ChatGPT overall, the best ones you could use were ChatGPT 5 thinking, and you had to enable web search, and that by far was the best one. Then, next was ChatGPT 5 Auto Plus Deep Research. So, as long as you had Deep Research or web search turned on, then it actually provided real references. Great. And remember, this isn't about whether or not the claims in those references support why it's being cited. This is just, yeah, the reference exists.

But then we look at Claude, and you can see that it's like a mixed bag with Claude. Some of them work really well. So, Sonnet 4 Plus Research did really well. 100% success rate in providing real references that existed. And then, Opus 4.1 was rubbish. It didn't provide any of the ref. Well, it provided references, but none of them actually existed. And surprisingly, despite my previous experiments, when put under pressure, Gemini did the worst. Look, you can see that the Flash 2.5 Pro, which is what you pay for, plus Deep Research, none of the references actually existed. And then we've got here Flash 2.5, none of the references actually existed. And then only 40% actually existed for the rest of them.

So, you can see that it is a mixed bag out there as to which large language model you should use. And this is only whether or not the citation it provides actually exists. And we are looking for a complete reference, including the title, the authors, the year, all of that stuff. And you know, Gemini, you're a bit rubbish, mate. Why are you so rubbish?

So, ChatGPT 5 in particular is absolutely killing it if you want real references, and that's because it's going out and doing web search and deep research. Okay, so we know ChatGPT is better or best for if we just want to find references. Okay, what about this second order? That is, does the reference contain the correct content for why it's actually being referenced?

And I'm going to give you the overall thing. Bong, here it is. ChatGPT only sort of like just got in by about, and this was across all of the models, by the way. Just under 50% of the uh citations actually contain the information they were being cited for. So, that's a bit rubbish, isn't it? You can't rely on half over half of the things that ChatGPT provides. Claude was even worse. So, just over 40% of the citations actually contain the information that they were being cited for. And Gemini, look at this. This isn't a mistake. Gemini was not able to provide any references. Now, get this. Any. Now, sit back. Look, take this in. Gemini did not provide any references where the content was was in the paper for why it was being cited. Like, that's so, so, so stupid that you shouldn't be using this for academic research. And once again, this is just surface level. This is across all the models. Let's scratch a little deeper, and we can see that this is the result.

So, does the citation match the claim? Uh, ChatGPT 5 thinking, once again. So, thinking with Deep Research, thinking with web search, they were the best ones up here. You can see Auto with Deep Research was fine. And then ChatGPT 5 Agent didn't work at all. Okay, that was 0% success rate. And then Claude was next. You can see that it had about a 40 to 50% success rate of actually citing references where the claim was supported within the reference. And then also, oh, poor Gemini. I just feel sorry for it now. Poor Gemini. It had 0% success rate.

So, what does that mean? Overall, first and second-order hallucinations mean that bulk, this is them in order of the least successful to the most successful. It really means that you should be using ChatGPT 5 thinking plus web search if you want any chance of actually getting references, and those references containing the appropriate information for why it's being referenced. You can see ChatGPT down here dominates with 5 Auto thinking and Deep Research or web search. That's what I would be using right now. But stay subscribed to this channel because I'll be testing and stress-testing these for academic research in the future. I promise you. All right, then. And then all the way down here, you can see that these did the worst. As in, they got 0% real references and 0% obviously why they were being cited because they didn't exist at all. Um, and then, uh, there's a mixed batch in the middle here. Um, so overall, that's what you should be doing.

Here's some take-home messages for you. We love some take-home messages. Just because you've paid for something, it doesn't make it more accurate. That's very, very important. So, ChatGPT 5, if you pay money, you are getting a better product. Whereas with Gemini in this test, and I know that this test could go much, much deeper, but in this one, it failed. Um, best of the bunch. Yeah, we've talked about that. Gemini struggled, of course. All right, then.

Common failure mode. So, the thing about all of these references are that sometimes something was being referenced, and it was in the paper, but get this, this is the important thing. It was actually in the introduction part of that paper for why it was being cited. So, it was actually sort of like not citing first, uh, principle, first, uh, order, what do you call it, primary sources. It was actually citing a paper for something that the paper cited. So, it was like double extracted from the information. ChatGPT and other large language models aren't very good at that. And the most important thing is that these are plausibility machines. So, when you look at the outputs, you go, "Oh, yeah, that's right. Yeah, that must be right," because it makes it seem like it is so real, that it really exists. The plausibility engine in these things is just amazing. It makes it really convincing, and you need to be the one that goes away and checks each individual reference if you're going to use a large language model. Um, and then obviously, a safe workflow is to generate something, but then you must trace every claim to the PDF and page for why it's being referenced. That is the only way to manually check these things.

Now, if I'm not going to use a large language model, and I recommend you don't use a large language model for finding references, what else can I use? Here are three recommendations that mean you'll get around most of this stuff because they're specifically designed for academia and research.

Okay, the first tool you should know about is Elicit. Elicit actually uses real papers, and it's checked in the background before giving you that information. So, you can be sure that it is giving you the real reference that really exists, and that reference contains the information for why it's being referenced. Love Elicit.

Another one is Scispace. Scispace is becoming an absolute powerhouse in this space. They just released their agent. Go check out my other video where I talk about that. But here, you can do loads of things. You can search for papers. You can create a literature review, and it's all based on real references out there in the world.

Another one that is really great is Consensus. This is a really great tool. If you just need to know yes or no from a particular research field, go check out my other Consensus video where I talk about all of the awesome things this can do.

Use these three tools. If you want to go searching for references, large language models are really great at working with language, but not as a source for gathering up the information. And that's why you need to use specialized tools. And check out my other videos on this channel because that is what this channel is all about. I'll be stress-testing these tools soon. So, subscribe, and I'll see you in the next video. If you like this video and you want to know all of the AI tools that you should know about, go check out this one because I list all of the AI tools for academia and research that you should know about. Go check it out.