📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

I Replaced My 4x3090 GPU Rig With a Mac Studio, can it run GLM-5.2 Local AI faster?

Tech-Practice5:54

Transcription

Last video I ran the GLM 5.2 model on four SD90 GPU. They have loud fans, a rig that likes a crypto mine. And this is the actual rig I'm running the model on this time. That's it. Just a one box which is just a Mac Studio that you can put on your desk, no discrete GPU. So today let's see the performance of running the GLM 5.2 on Mac Studio.

Before we get started, I want to quickly mention that if you want to better manage the big AI model files using the MTFS drive to use that drive for Mac, you will need some software. Recently, I found that there's a software called iBoysoft MTFS for Mac. That's really helpful to use that. You just Google that and go to the website. You can directly download that for free. How to use it? It's very simple. Connect your MTFS drive to your Mac using the USBC and then you can enable the driver. Make sure that you find the MTFS drive and then you can enable that. After that, you are able to read and write the files on Mac very, very easy. I highly recommend that.

The GLM 5.2 model was released on June the 17th. There has been various benchmarking that compare it with other top language models and it appears that this is at a serious level. The benchmark shows that it is comparable with the top closed models including GPT 5.5 and the Ops model from the cloud team. The best thing about it is that it's open source in MIT license. So anyone is able to download it and run it if you have the right hardware.

We will be using the GGUF quantizations because they save lots of space. However, because it's a really big model, size is still really big as you can see from this guide. So it really depends on the total memory including the RAM and the VRAM or unified memory. So even at one bit, it requires at least 223 GB. At 4-bit, it requires 372 to 475 GB. So for the Mac Studio, you will need at least 256 GB RAM in order to run the one bit. To run higher precision, you will need at least 512 GB of the unified RAM.

Looking at the 512 GB unified RAM, we are looking at the M3 Ultra. This is great running the GLM 5.2. The speed can go up to 24 tokens per second and the pre-fill is at 150 tokens per second. On the other hand, for the 256 GB version, it is still be able to running that. However, only at one bit, speed is at about 9 tokens per second.

Now, let's look at a real-time demo using the LM Studio to run the GLM 5.2 on a 512 GB M3 Max Studio. We are using some examples to see how the generating works. Because it's a China model, we use the Chinese to test. We see the text generating is working really well. So the speed comes out to be around 17.45 tokens per second. We see that it's working really well.

People have doubts that how useful is the one bit. There is an interesting comparison by Onslows. They compared the one bit GLM 5.2 with the Ops 4.8 and the GPT 5.5 using the same prompt and then you can see that generated some really working great animations by all three of them. I think the one bit works quite well in this case.

To conclude this video, I hope you find this video useful. I think it's amazing that one Mac is enough to run the 744 billion parameter model locally. However, I think the limitation is still the amount of the unified RAM. And right now, Apple has increased the price due to the price of the RAM. However, if you have one of the Mac, I think you definitely should give it a try and please post your numbers in the comments. Thank you for watching. Please give it a thumb up and share it. Please subscribe to the channel for future content. Thank you for your support. Goodbye.