📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Can $100 ChatGPT and Claude Fable Solve PhD Math?

Easy Riders10:10

Transcription

So, in the last month, I've been testing the $100 a month ChatGPT Pro tier. And before that, I've been using the $200 a month model. And I've been pretty impressed by its ability to help me solve research math problems for my PhD, or at least its ability to kind of clarify details on things which I didn't fully understand.

But the thing is, $200 a month is actually a lot of money to be spending on a clanker, uh, which is basically just helping you do a job for which you're paid basically minimum wage for. So, when OpenAI released the Alert, the Pro subscription, I was really interested to see just how different it was given it was half the price. So today we're going to be exploring just that, the $100 a month GPT Pro subscription.

And we'll also be testing Anthropic's new Fable model, which is supposedly a big improvement from Opus, just to see whether it's any better at helping you solve math problems. And we'll also be comparing to the equivalent option offered by OpenAI, which is an extended thinking mode. Both of these go for about $20 a month, and overall, I think if you're kind of starting out, they're way better than the free, and I do genuinely recommend them if you do any kind of programming or like technical STEM work. I think that they're really, really useful.

We're going to see Clanker Logic. We're going to see clankers hallucinating and tripping out, and we're even going to see clankers that you pay for that you don't even get to use. So, join me today because we're really going to see it all.

Now, I would love to say that Claude Fable was incredibly good at research level maths, but if we go to the questions that I've been asking it, and I've been doing this all morning, actually, every time I seem to ask it a question, and funnily enough, this was the same on Opus 4.8 on high thinking mode. um I'll paste in essentially just a whole bunch of maths uh describing the problem and basically a question asking it what to do and what I found is that it will reach the max length for the message and then if you press continue it will then go on to think again and then you'll just run out of tokens for the afternoon. This is actually the third time I've tried this and all three times I've run out of tokens and it's like okay wait until this afternoon until you can ask the question again.

If we compare this to the same price model as you get from ChatGPT and I haven't had a huge look at the answer here but at least we're getting an answer. I mean, I'd rather get something than nothing. But nonetheless, hopefully we'll get lucky on this third attempt of Fable because so far I haven't even actually been able to see any maths that it's produced because it never seems to produce it. And here we are. So, it's happened again. I've now reached my 5-hour limit and it's going to reset in a couple of hours. So, if I want to, I can kind of get it to keep doing the same thing and never give an answer. When the limit does reset, I will press continue again. Obviously, I can't do that at the moment cuz I'm out of credits. But so far, I am yet to get a single answer from Claude's Fable. Uh, which is really kind of crazy actually that I'm really not getting anything. And so, of course, if you're army and about testing Claude's Fable with the $20 subscription, it might be worth taking into account the fact that you might not get a single answer back from it.

Naturally, this does mean that the answer that extended thinking mode gave was better. If I were really grasping at straws for Claude in this case, you could argue, I suppose, that it's better to not give an answer than to give an answer that's wrong. The fact that the model keeps asking for more tokens and keeps running out perhaps means that the model does understand that the problem is actually quite a challenging one and that it needs to allocate a lot of compute to it. It doesn't have access to it because the subscription tier is too low. I'm not so sure if I would be happy with that if I was paying for it, which I am. But at least I can like make a video of it and make some memes. So, it's kind of funny. Uh, I don't mind that much. It's only $20. Uh, which I won't get back, obviously, which is just going into the ether. So, I'm paying for compute that I'm not even like getting the output of. So, yeah, thanks a lot, Anthropic. That's that's awesome. The model's great, by the way. I just wish I could see an answer from it, honestly. I I just wish I could see an answer.

After waiting again, again, again, it then came back with a glitch, and when I pressed continue, it had erased the whole history of the attempts I'd given it. And it's kind of strange cuz Anthropic give you access to these models with that tier of subscription, but they don't seem to work with that tier of subscription. So, it makes you wonder why they've given you access to those models.

So, now on the fourth attempt, we did actually manage to finally get an answer for this question. And as you can see, one thing that's pretty nice about the answers that Anthropic gives is that it does give you a kind of way of reading the PDF like within the interface on the on the website. But bear in mind, this took me like a whole day of reprompting. So, this was at 2:40, this was at 3:00, this was at 9:00, and only then it gave me an actual answer, which included a whole bunch of different things. So, it had three Python scripts and a PDF file. And this is the PDF file that it returns, which is essentially just its answer. And I've had a look at this and to be honest, it sounds really, really strange. Like the words that it's using are just bizarre. I mean, look at this lemma one spectator reality. What is it talking about? Like I genuinely don't know. It's either a complete genius or just completely tripping out and has no idea what it's talking about.

If we compare it to the answer that ChatGPT had given, it makes a lot more algebraic sense. Like what it's doing makes more sense. You can actually follow it and it reads as if it's actual math. And here Anthropic states proposition without a proof which is essentially just restating something which was in the prompt itself but then doesn't actually go into the detail which I'd asked it to with regards to this statement. Again here it's restating something which was given in the prompt but it doesn't go into any explicit detail of the things that I had asked. And actually one of the things that I'd asked it to do was to find these R_n and K_n and you see that it actually kind of states that you could find out what they are but doesn't actually do it. So it says all of which are gamma/3F2 evaluable. So maybe you can write them in terms of gamma functions or these 3F2 hypergeometric functions, but it doesn't actually do it.

Overall, the answer that ChatGPT gave is more legible and I feel like I trust it more. And although I like the fact that Fable kind of gave back the three Python scripts so that you could, if you wanted to have a look at what it was actually trying, I don't really like the mathematics that it produced and I don't think that it's anywhere near as good as that of ChatGPT. On the one hand, it does seem like it understands a lot of things and it's definitely able to do maths, but some of the things that it's saying just seem very bizarre. I mean, I don't have a clue why it's calling this spectator reality.

Now, that's actually the day after when I asked this prompt. And I'm just going to show you guys cuz this was on extended thinking, but they seem to have actually change the setup for asking questions. It seems like extended thinking corresponds to high. Maybe heavy thinking is extra high. And then pro still has extended and standard. I'm definitely disappointed with Fable's answer given that this took me a whole afternoon of reprompting in order to get it. And in all honesty, the quality of the answer is pretty bad.

And this takes me onto a similar point that I found with the GPT Pro subscription. Now, the main thing that I found with the $100 tier of ChatGPT Pro is that the usage limits are just way, way more constraining. Whereas I'd never run out of credits with GPT Pro when paying £200 a month, I run out of credits all the time with GPT Pro while paying £100 a month. It's actually quite annoying because you're paying half the price, but you're actually getting a quarter of the usage. Because if you look on the OpenAI website, with the $100 a month tier, you get five times the token usage, whereas with the $200 a month tier, you get 20 times the token usage. So evidently, they differ by a factor of four.

So, as you can see here, this is a long conversation thread that I had with GPT Pro yesterday. And as you can see, I actually ended up running out of credits. So, if we go to pro, you see that my limits will reset on the 2nd of June, which means I now can't use GPT Pro for a few days. And what's annoying about this, which has happened multiple times in the month I've had this tier of ChatGPT, is that it will always freeze your credits for like a few days at a time.

If we compare this to Anthropic's Claude where I was paying for the $20 a month model, what I did find is that the token usage was really, really limited, but it would kind of cap your usage like twice a day. And what I found was if I was using a lot of tokens in the morning, it would still let me use it in the afternoon that same day as long as I waited for the token limit to reset. It would have been nice if that lower GPT Pro subscription, this $100 a month one, was a bit more similar to that where you'd reach your cap sooner, but you'd also have access to the model again in a period of time that's like shorter than 3 days.

Now, I can hear some of you saying you're just a shill for OpenAI. But wait a second, cuz I have also had some funny stories with the $100 ChatGPT subscription this month. Now, this example right here was pretty impressive when it comes to seeing clankers doing research level maths. Okay, so let me just show you what this $100 clanker did. So, I've been talking to it and having it come up with a whole bunch of crazy stuff, which to be honest, I'm starting to think is just a massive, massive hallucination, just some sort of huge AI trip.

So basically what I thought you know what why don't you write a Python script to test this formula. So it's trying to test this thing which on the left you can test with standard things. It's just like an integral. So you can numerically solve this integral. So you can get a whole bunch of data for what the left hand side is. And basically what we're trying to see is if the right hand side formula which you can see is like a sum of three different things is the right answer. So basically you see whether if you sum up all of those three things which you also try and compute numerically whether you get the same as the left hand side. That's kind of just how you test an equation numerically is you test both sides and see if it comes out as right. And that was one of the things that when I was learning Mathematica like ended up being super useful for like seeing if you've done algebra right and stuff like that.

So ChatGPT comes back and basically tests the sum of the three terms. And as you can see here it starts to look quite promising. So the first thing is basically the left hand side. So it's the real answer and it's getting for a particular set of values which you can test it for. And then it does it again with effectively the same method and gets the same answer. then you see the sum of the three terms and it's giving the same answer once again. So I was thinking damn that's pretty impressive that it's managing to do that right but no because as soon as you start looking into the script you realize that it's actually had this important caveat here where it says in this numerical implementation I3 is computed as the residual projector. Okay, let's see what that means. So it means it's taken the answer and then subtracted the two other answers from it.

Now, um, to explain how stupid this is, it's actually kind of crazy. It's basically just saying that if you plug in for this third term here, this like I3, you let it be equal to the left hand side and then minus the first two other terms. So, obviously, when you add I3, which is your answer, subtract these other two terms, you're obviously going to get your answer back. So here I've quickly just typed up the reasoning of effectively what the clanker has said and it really does amount to saying that a is equal to b + c + a - b - c residual projector.

Overall, I hope this has made it clear just how limited these models can be if you're actually trying to solve a difficult problem and how if you're not aware of what they're doing, it very quickly becomes AI slop and the answers will just massively degrade in quality.

I want to say a genuine huge thank you to everyone who comments in all of these videos. I've been getting comments from people because I know I haven't uploaded in a while. It's just really really busy. There was a conference in Paris, then I had this week in Paris. I also landed a full twist, which is like a thing that I've been trying to learn in gymnastics for like 2 years at this point. And honestly, it's just been super busy. But I do really appreciate all of the support that you guys have shown the channel. It's actually crazy. And I'm also going to be applying for my first postdoc. And so if that goes well, I'll definitely let you guys know about that as well. Otherwise, thank you guys so much for all of the support. You really are awesome. But I hope you enjoyed the video and have a good.