Transcription
Hey guys, what's the show? Today we're at Apple HQ Australia Gold Coast, famous for being next to one of the most popular stores over here in Rabbina Town Hall. We're here to unsettle the debate to find out how much faster is the M5 Max. A lot of YouTubers out there, I said it's 20% faster. Me, I predicted it could be two and a half times faster, but we already seen the M5 Pro. Is it two times faster than the M4 Max? Let's find out for certain. It's going in.
Now, >> previously when it was announced the M5, we went ahead and did some tests with the M1 Max and the M4 Max to see how well the delta is between that. So, we can actually predict what the M5 Max could be. And we went ahead and guessed it could be two and a half times faster. 13.25 seconds. So, again, it's very, very similar to the performance delta of 2.6. 86. But then we went ahead and got the M5 Pro. Let's check it out. See how much faster that could be with the M4 Max. We found out that the app prompt processing at very least the M5 Pro was two times faster than the M4 Max. It was even faster than the M3 Ultra prompt processing 9.37 seconds. So the M5 Pro even beats the M3 Ultra at prompt processing. Got to say great job Appel just by a titch all at prompt processing token generation it was slower but at prompt processing using the neural accelerators it was boom shak good.
So we revised our analysis and we said yo the M5 Max has to be four times faster than the M4 Max. But we didn't stop there. We went ahead. We went to Apple HQ to find out for certain. Went ahead and got the M5 Max and we did some benchmarking. went ahead and I'm gonna show you the results live here because I can't actually show you anything live because if I start running OBS on my M4 and this this by the way isn't an M5 Max. This is an M1 Max but you know it looks like the M5 M5 we did it in in the actual store. Full disclosure. So yeah, let's just jump in ahead and look at a test over here.
So I'm going to jump in and we got inference right here with the software update using neural acceleration. We got a Wikipedia article about Appel. Now this is a different kind of app. So we need just playing it live to see how it performs. Now I had to sacrifice myself and disable restrictions. So I have to disable restricted mode because the Wikipedia article had some restricted kind of words. So it's called printal control this application. So we're going to hit play. And as you can see prompt processing away 1,00 2,000. So 10,000 token plus prompt 10,33. We're doing coin 3.527B. If you want to play along at home, 20 3.5 27B the 4bit edition just straight out of the box. Boom. We're doing 22 to 20 almost 23 tokens a second. 100 tokens were made only 100. We're just keeping it simple. Check out the processing and boom. It says right there 15.65 seconds. So you want to see how the M4 Max does in comparison. So we're comparing just so you guys are aware 15.65 seconds. That prompting generation that's also going to be different but let's just stick to the promp processing 15.65. 65 seconds.
So, this is the M4 Max running. I've just got the video. I got my my phone here just filming the screen. Had all my applications, well, most of my applications closed. You can see it's definitely slower the increment. We're at 9,000. It's going to go to 10,000. Come on. Come on, baby. 10. Boom. And we're getting 15 tokens a second. So, definitely token generation was faster than the M5 Max compared to the M4 Max. But look at the prompt processing speed. And it says right there 45.98 seconds. So, I don't know if I'm a calculator, but 15 of 45. That's only a 3x improvement. Now, 3x, don't get me wrong, 3x is amazing. In one generation, they boosted the pro processing by 3x. But it's not that 4x figure that Apple was saying on their website. So, it's only 3x.
But we didn't stop there, cuz there's more. Now, check this out. You see how thorough we are on channel. Expect a like and a comment for this. So, you're going to see it. We didn't stop at just the Quinn 3.5. We went ahead and tested out Gemma. So, I'm going to switch it over here with the model selector. Boom. Select out Gemma. This is Gemma 4 26B A4B 4bit. Super fast lightning. It's loading. Now, the loading speed is is it was stuttering cuz I was running off this SSD which is slow. So, just forget about the loading speed. So, I paused it there. But, we're jumping in. We're doing prompt processing. Really fast. Lightning fast prompt processing. Just smashing out the tokens. And we think was it 85 tokens a second. We're about to see the big reveal. Prompt processor speed was 4.47 seconds. So if it's going to scale, then the M4 Max is only it's going to be three times slower.
But check this out. Check this out. And this is how much advantage I wanted to give the the M M5. I actually ran Infencer over here or the exact same model. I had Final Cut Pro in the background. And I was too lazy to close all my applications down again cuz I was like, "Oh, hold on a second test." So, we're processing the prompt. Boom. 6,000 7,000 8,000 9,000 10. Boom. Done. 93 tokens a second. So, in this model, the M4 Max was actually faster than the M5 Max. On this model, in Quen, the M5 Max was faster at token generation than the M4 Max. But in Jammer, maybe it's optimized. They need to find out exactly what's going on. 93 tokens. So, it's smashing it out on the token generation. Let's see the prompt processing. Boom. 8.79 seconds. It's like 8 seconds. So, it went from 4 seconds to 8 seconds. Yeah, that works out. 4 to 8. So, that's only twice as now. Only twice as fast. It's still fast, but it's not 4x.
So on one test, we got a 3x improvement in processing. And in another test, we only got 2x improvement in prompt processing. I'm totally totally uh I wish I mean maybe I don't wish but maybe if anyone has any benchmarks which are credible run these tests let me know what kind of speeds you guys are getting I will if you want you know just if anyone has an M4 and M5 let me know if it's actually 4x faster like appella says and some of the fanboys have said yeah it has to be appel has to be 4x or not in our test we got 2x and we got 3x X and I'm confused now what's going on. So, I'm wrong about everything. I thought it was going to be two and a half, which is kind of like landing in the middle. I like that.
But there is one more twist to the story and that is unfortunately the only MacBook Pro M5 Max that they had in the store was the 36 GB model. And looking into it, the 36 GB model is only the 32 core GPU. So, the fastest M5 Max is actually a 40 core GPU. So, it's 20% less GPU than the 40x version. So, pretty much if it's 2x the speed, you got to give it another 20% which comes to 2.4x 2.4x faster or 3.6x faster. Both ways, it's still not 4x faster surprisingly. Don't know what's going on there. Still not 4x, but it's still, you know, it's still a good amazing jump in performance. And uh I just wanted to share this video cuz some YouTubers out there said it was only 10% 20% faster. And I predicted it was going to be 2.5x faster. And we ended up landing that it's around two to 3.6x faster depending on the model you use and all this kind of extra stuff to use.
So did you find these experiments interesting? Because some guys out there with with the M5s, they said it was so fast they they watched this video before it was even made. Hope you guys found this video useful and enjoyed the show. right here. >> Hope you guys found this video useful and enjoyed the show. Can you again? Yeah, sure.