📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

After This, 16GB Feels Different

Alex Ziskind12:35

Transcription

You probably can't tell the difference between those two images, but the computer can. The image on the left is 14.6 megabytes and the one on the right is compressed to 1.8 megabytes. They look the same to me.

So, if you use compression on a given disc, you can fit a ton more images or videos or whatever it is you're compressing than non-compressed. And same goes for LLMs.

This right here is a Mac Mini with 16 gigs of memory. And this one is my daily driver with 128 gigs of memory. I can run pretty decently sized models on the big boy here or even on these with 512 GB of memory each. But on this one with 16, I have to be very picky about what I run.

For example, a really popular model set that just came out is Quen 3.5. This whole family of models is really performing well. Look how many downloads in the last month. And I really want to run this 9 billion parameter model on this Mac Mini. Well, guess what? The full thing is actually 19.3 GB. And you want to be able to fit the whole model inside the memory. And we just won't be able to do that here.

Well, this is the full BF-16 unquantized version. What does that mean? It's not compressed. That's why people started compressing different things here and there. And Turbo Quan helps with that.

Now, before Turbo Quan came about, we had different ways to handle compression. We had quantizations. So for example, BF16 is like a floatingoint 16bit models. The weights for the models are basically a bunch of numbers and their 16 bits. Not only do smaller quantizations take up less space on disk, but when you load those models up into memory, they decrease the memory requirement allowing you to run it on smaller hardware, but you still have this KV cache issue which also requires memory.

So, for example, here is uh the same model, 9 billion parameters, 8 bit, and that one is only 10 GB. We can go even smaller, 4bit. That one's 5.98 GB. Well, just like with those images I showed you, sometimes the models and their compressed versions don't vary too much in quality, but sometimes they do. For example, if you go too low here, let's take uh Bowski's quantizations. He's got all these different ones. Look, Q8 8 bit, 6 bit, 5 bit, 4bit, 3bit, 2bit. Now, I'm sure these work, but they're not going to give you the best results. Sometimes the LM gets into a loop where it just spits out garbage. Usually, quantization of four bits is probably the lowest you'd want to go.

And if we take a look at this one, it's 6 GB. Well, you'd say, "Oh, 6 GB fits no problem on this Mac Mini." But wait, what about when you actually run it? I'm using 77 out of 128 gigs on this machine. Jeez, that's a lot of gigs. I'm going to load up this model. And now I'm up to 84. Huh, that doesn't add up. That's because we need to reserve more memory for context and for cache. And this was only 4,000 context length, which is not very useful. This model supports way more than that. So, let's crank it up all the way. Reload. And now we're suddenly up to 92 GB. And that's without even running any prompts at all.

Hi. I know you all love that prompt, but let's see how that affects our memory. We're still good. 92. What if we send in a really long prompt? This prompt is uh about 17,000 tokens. Let's send that in. Look at that processing prompt. It's taking a little while. Even on this powerful M5 Max MacBook Pro. And look at our memory. We suddenly have a little bit of a spike there. And there it goes printing out the answer.

I'm going to paste in the same prompt again. And what happens here is the entire conversation is sent in again for processing, not just my prompt. But when the LLM is generating text, it doesn't reread the entire conversation from scratch for every token. Instead, it stores key value pairs, the KV cache. These are basically mathematical summaries of every token it's already seen. Think of it as like short-term memory. And it lives inside the memory alongside of the model weights. With every token, KV cache grows and fills up memory.

Now, quantization solves for the model weights and shrinks them down. Whereas, Turbo Quant, this new amazing thing, actually works on the KV cache and shrinks that down.

This is one of those free Wi-Fi everywhere situations, which sounds great until you remember that everyone else is on it, too. That's why Surf Shark gets switched on before I do anything else. One click, and my git pushes, SSH sessions, and package installs all go through an AES 256 encrypted tunnel. In the meantime, Clean Web quietly filters out all the usual garbage, ads, trackers, all that stuff you don't need. Surf Shark keeps no logs, and that's been independently audited, and their servers are RAM only, so a reboot wipes everything. The feature I like most is unlimited devices. This MacBook, my phone, and the rest of my gear are all covered without me thinking about it. SSH into the lab feels a lot less sketchy on public Wi-Fi. Go to surfshark.com/allexiskin and you'll get four extra months free, plus a 30-day money back guarantee. Stay private, stay productive, and now back to the video.

Now, Turboquat might not necessarily help you out on a system with a lot of memory, but it will help you out on a system with a little bit of memory. And the amount of noise this MacBook Pro is making is just crazy. Listen to that. That's crazy. You're done with your job. What are you doing? Chill out. Ah, cheesy jokes.

You might have heard about Turbo Quant recently. It's making some waves, but whether it's going to be a thing or not, we'll see. So far, the experiments are promising, and you can check out the official Google research paper there. I did my own experiments, though. Llama CPP, a very popular project. It lets you run LLMs locally. It's pretty cool. They've not rolled in TurboQuant just yet. So far, it's a community effort. Here's one very popular one on GitHub by the Tom Turboquant Plus. You can go grab this, run it yourself. Basically, it's a fork of Llama CPP that implements TurboQuant.

Now, my initial tests with this were pretty bad. Um, I tried it out on the M5 Max. I tried it out on the M4 Mac Mini and the KV cache space savings were present already. So, that's good. Here's the context depth scaling. I went all the way up to 32K. The problem that I saw was that both prefill speed and decode speed suffered a lot. Also very model dependent. For example, I started out with Quen 2.5, an older model. Then I did Quen 3 8B. And they all kind of showed similar results as far as Turbo 3 and Turbo 4. There's three different variants of Turbo Guantan, Turbo 2, Turbo 3, and Turbo 4. Turbo 2 is the most aggressive one that's supposed to squash it four times. Turbo 3 about 2 and a half times and Turbo 4 about 1.9 times.

Then I ran Quent 3.535B, which is a mixture of experts model, not a dense model. And I ran it on this one. It's a 34 gig model. Obviously, it's not going to fit on the Mac Mini. And that performed the best as far as squashing that KV and not losing too much on the speeds of decode and prefill. However, it still was slower. I was like, "Okay, well, I I suppose it's worth the penalty if you're going to be getting such savings in compression, but it wasn't over."

Tom Turney worked on it a little bit more and gave me some hints. The way I was running it was not exactly the best way to do it. I was running it symmetrically. So, K and V can actually be turboed separately. When I ran it, I applied turbo 4 or turbo 3 to both the K and the V equally. That's called symmetric when I applied it to both. And Tom suggested I use Q8 for K and then the turbo one for the V part, which would be considered an asymmetric approach. And then if you want more aggressive, you'd still keep Q8 for the K and you would use Turbo 3 for the V part.

On the Mac Mini, loading up the Q8 version with 131,000 context window. just crashes. However, Turbo 3 runs 131K context comfortably with 3.6 GB to spare. Same model, same machine. Turbo gives you two times more usable context. And I tried different context length, 32K, 65, 131. And at each level, we saw a huge difference in the number of gigabytes that were chomped up. And this chart shows a little bit of a breakdown between the model weights. And they're about the same. Well, they are the same actually right here between the turbo run and the Q8 run. But here's that pesky KV cache that grows so much when you're running it at Q8. And here's turbo 3 KV cache. Much much smaller, leaving us with some extra headroom. Nice. So, it kind of does that job. What about the other thing? The speed. Well, we recovered that, too.

Hold on a minute. Before I get into that, I wanted to do a little bit of a needle in a haystack because not only do you have to worry about the memory, but is turbo quant going to affect the quality of the output? One test for such a thing is needle in a hay stack. There's a lot of different tests, but that's one of them. That's an easy one to understand. You have a bunch of text and somewhere inside of that text, you put a little secret and then you ask the model to find it and you test different lengths of that full text. So that's the context length. 1K all the way up to 32K. 32,000 tokens.

Initially, this was a total disaster. Look at this. So three out of three at the top means we hid three secrets and found three secrets. That's 100%. This is on the Mac Mini, by the way. Quantization of eight got all of them. Turbo Turbo, which means we had symmetric. Turbo for K and Turbo for V got 100%. But then we had a huge drop here. Terrible performance. Turbo 3 got one out of three and Turbo 2 got one out of three. Furthermore, if we look at Turbo 3 and Turbo 2 over here for larger contexts, 8K and 16K, they got nothing. They got zero. This was before I listened to Tom's suggestion to do it asymmetrically. And when I switched to asymmetric, look at this. Three out of three for all of them all the way across the board. Beautiful. So, this shows the quality of the result is actually good when we're using Turbo Quant. at different levels of turbo quantization at different context lengths, too.

Okay, back to speed. There was a surprise. Now, the M4 did okay. Um, at short context lengths, turbo doesn't do that well here. Uh, it slows down quite a bit. We're about 1 to 4% difference in the speed for decode here. But on the M5 Max, this is where we saw a huge difference. Check out the decode speed here. So on Q8, the full baseline non-turbo quant, we dropped down from about 54 tokens per second to about 37 tokens per second, going from a depth of zero to a depth of 8K for context depth. And then we went up a little bit more for 32K. We went up to 44. But Turbo Quant stayed relatively flat all the way across the board. And this isn't some weird glitch. I ran this many times, so this is an average. This is pretty cool and it's pretty promising. I wish I saw the same kind of curve on the Mac Mini.

Now, the reason for that is here on the Mac Mini, we were computebound. Reading from KB cache was not the bottleneck here, but the model's matrix multiplications are. So, if we speed that up, then we might see a similar curve to what we saw on the M5 Max. And that could mean that when the M5 Mac mag minis drop, even if they have 16 gigs, which they probably will, they will have a significant boost and benefit even more from Turbo Quant.

So, the takeaway is that every model is going to behave a little bit differently, as usually is the case, but some models perform really poorly. However, the good news is the latest Quen 3.5 models behave really well and respond really nicely to Turbo Quant on the Apple hardware.

Now, you could try this fork out yourself or you can wait until this is rolled into Llama CPP and other tools. I think VLM is also working on it. And if Llama CPP gets it, then tools like LM Studio are going to get it. Go ahead and try it on your machine. Let me know what you get. Do you see good results from this? Also, leave a comment with what models you used as well. Thanks for watching and I'll see you next time.