📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Apple JUST Dropped a Game-Changer

Alex Ziskind22:35

Transcription

Every Mac cluster I've built in the past has had the same painful ending, considerably worse. Oh, come on. Until this one.

I built a Mac Mini cluster or Mac Studio cluster. And yeah, it works. You can run bigger local AI models, bigger LLMs, and bigger models usually means more knowledge, smarter answers. Big models also require a lot of RAM. Combining several machines together into a cluster means your models can share all that RAM or unified memory as it's known in Apple Silicon World. But there's always been a trade-off. The more machines you add, the slower things get. Well, that ends today. What I'm about to show you changes everything. More machines and it actually gets faster. We can do this, folks. And you don't even need Formax Studios with 512 GB each for a total of 2 TB of memory. You can even cluster on baby Macs that cost 500 bucks each. Aren't they cute? Network Chuck, I hope you still have your Mac Studios because you're going to like this.

As it turns out, Exo, that's the tool that I showed before on my Mac Mini cluster, the one that's really easy to set up. Then they disappeared for months, and I thought the project was abandoned, and so did everybody else. Well, they've been tweeting, "Linear scaling achieved. Deepseek with 671 billion parameters at its native 8bit is over 700 GB in size. And they've been teasing this all summer." Well, finally, they've reached version 1.0. But in order to get that linear kind of scaling, sorry the green line is not actually linear, but you get the idea. It's not just Exo. There are three levels of technologies that had to align, like the stars aligning in order for us to get here. And yet somehow version 1.0 of Exo got even simpler. All it is is just an installer that you install on every machine you want as part of your cluster.

By the way, this video is still not sponsored by Exo, even though my last video about it did really, really well. So hang on a moment while I pay the bills. True story. My first month in a security operations center, I was terrified of touching production. I just needed a place to practice legally and safely. That's what Try Hackme gives you. This is Try Hackme. Browserbased cyber security training. No installs, no setup. Click and you're inside real labs. Let me show you a beginner friendly OSEN challenge. In the OSET room, you get a single image and one question. How much can you discover from just one photo file? I'll open the virtual machine, grab the image, and run. Exit tool image.jpeg that pulls metadata including a username and now we pivot. I'll search that handle across public sites. Found the avatar. It's a cat. The city tracing the profile footprint points to London. Personal email on the developer footprint. I find o woodflint atgmail.com. All right, I'll stop there. There are more questions in the room and you can finish the rest yourself inside tryh hackme. Why I like this for devs and beginners. You learn by doing. structured paths, scenarios modeled after real attacks and defenses, gamified progress, a massive community, and Echo, their AI tutor to coach you through. Jump in with the link below. Start with the Osent room and see how far you get. Many rooms are free, and with my code, Ziskin 25, you'll get 25% off the annual plan if you want full access. Link down below.

To kick things off, I'm just going to start Exo on all four machines. Now, let's keep an eye on things with activity monitor. Memory usage and GPU usage are two things we're mostly interested in here because memory is really involved in machine learning tasks. And the GPU usage is something that's going to let us know that these processes are happening on the GPU, which is where we want them, not on the CPU. So, things are fast. By the way, for M3 Ultra Mac Studios, I hooked them all up. They have to be hooked up using the mesh networking through Thunderbolt. I'll explain that in a bit. And together they're all using only 66 watts of power. That's crazy.

Everything in Exo 1.0 is kind of UI driven even though they do have API access. You can see Exo Hangout here in the menu bar and the topology of your four machines or however many machines you're using, 2, 3, 4. And then you got this dashboard button that takes you well, you guessed it, to the new Exo dashboard. This shows you the topology, chat history. I'm going to delete my hi's and how areas over here. And here is where you get to select the models you want. You can do DeepS 3.1 8bit, which I'm going to show you. Kimmy K2 thinking, Quen Coder 480. These are all large models, but they shard slightly differently. Shard. What is that? That sounds gross. Oh, sharding is where you separate the model out into different parts. So, you can run certain parts of the model on one machine, other parts of the model on the other machine. And in the old way of doing things, you would use pipeline charting, but now you can use tensor parallelism.

"Previously, the way that Exo split up models was something called pipeline parallel. Each stage still runs sequentially."

This is Alex Chima. He's one of the founders of Exo. "So you don't actually get any speed up because only one device is active at a time. Now with low latency RDMA, that enables different kinds of parallelism, namely tensor parallelism, which enables you to actually get true parallelism. Instead of just running in stages, now you split up each layer into separate pieces of computation that happen in parallel."

Oh yeah. Let's kick things off with this quen 235 billion parameter model. And we'll do MLX RDMA and we'll use all four nodes. That's going to be our topology. I'm going to launch this and keep an eye on that memory. The memory is going up on every one of these machines because it's being loaded into memory and that takes a little bit of time because it's a large model. You can see the memory pressure is getting a little bit higher, but come on. This is ridiculous. Each one of these machines has 512 GB. We're going to load bigger models. And it's ready. 84 GB on this machine, 82 on this one, 81 on this one, and 81 on this one. How is it splitting up the model? Well, it's using a new technology from Apple where it can split up the model using tensor parallelism and it can take different layers and put them on the different machines. This really optimizes inference.

By the way, this little interface pretty cool. You can delete the instance, but you can also start multiple instances. And since each one of these machines has 512 gigs of RAM, I can start multiple instances of Quen 235 billion or other models if I want to. Let's ask it something. Hello. I know, I know I have a lame prompt. What can I tell you? That's not what this video is about. This video is about the technology behind this and not necessarily how good is your prompt. That's for other videos. We get 37 tokens per second here. Remember that number because now I'm going to delete this instance and I'm going to start up this model one more time on only one node. Oh yeah, it'll run on one node. There's plenty of room. Hello. Hello. How can I assist you today? 30 tokens per second. 30 on one node, 37 on four nodes. And this is why tensor parallelism is fantastic. It's actually a bad example of this because this model is or mixture of experts, which means, yeah, it's 235 billion parameters, but when it's actually loaded and you're using it, you're only going to have a small subset of those parameters that are activated. And these types of models don't typically shard very well or split up across different machines, as well as dense models. dense models split up really well. I'm going to show you that in a moment.

Wait a minute, Alex. You mentioned RDMA. Do you want to use RDNA? No. No, not DNA. RDMMA. It's remote direct memory access. Oh, so Apple has done something incredible here. Do you remember how blown away we were by Thunderbolt 5? Yeah, it just came out like less than a year ago. We started getting our initial devices and then Mac Studios had it. M4 Max, M4 Pro got it and it was already super impressive at how fast it can transfer files or drive 8K monitors. And Thunderbolt can be used to network machines together, which is exactly how you've seen me do it in my Mac Mini cluster video and my Mac Studio cluster video from a few months ago. And I thought, okay, well, that's kind of limiting and that's why it was a little bit slower. But Apple didn't just sit there, release this technology, and there you go. You're done. No, they decided, "Oh, I think we can use the same exact hardware and make it better in software." With Mac OS Tahoe, say what you will about Mac OS Taho's UI design, but with 26.2, they really cooked. They've enabled RDMA over Thunderbolt, which means communication between machines can now be 10 times faster. And that, my friends, is what eliminates our bottleneck. Now, there's no reason not to have 50 Mac Studios hooked up together like Markiplier. Gasp at this. Jealous. Jealous in my big old bathroom. Yep. That's a big one, huh? Oh, yeah.

"So, the hardware was all set. There's one more missing piece of the puzzle that needed to be solved to have the stars aligned and everything working, and that's MLX."

MLX is an array framework designed for efficient and flexible machine learning research on Apple Silicon. It's specifically designed and optimized for Apple Silicon. I've shown it many times on the channel before. If you're using LM Studio, you can get GGUF models or MLX models. GGUF models go with Llama CPP, a very, very popular tool for running inference. In other words, running models and generating text. Olama is built on Llama CPP. LM Studio is built on Llama CPP. But LM Studio also has access to MLX. And MLX is kind of like CUDA on Nvidia hardware. It's specifically designed to take advantage of the hardware. And it's actually more performant and faster. It only works on Apple Silicon. It's not crossplatform like Llama CPP, which is Llama CPP's advantage. It'll work on Linux. It'll work on Windows. It'll work on Mac. But MLX is faster on Apple Silicon. And last time when I did the Mac Studio cluster demos for you, I used MLX distributed, which is this way to establish communication between multiple machines to run MLX models on multiple machines as a cluster. And at that time, MLX was using regular Thunderbolt communication networking, which was slow. So even though we ran large models, we saw a big degradation in speed, and I couldn't even get the really huge models to run. Well, MLX has incorporated the new RDMA functionality into its library, and you can run it directly if you want to. I got to tell you, it's not an easy setup. Similar to my previous video, it took me a while to set up because there's a lot of parts that you have to install. These are the kinds of parts that Exo eliminates for you with the one installer. It basically automates everything. But if you wanted to run MLX distributed yourself, you can.

Check this out. Developers are going to be interested in this one. Devstral 2 123 billion parameter model. Devstral large is now considered to be the best open weights model for developers. It does really well on software developer benchmarks. Let's run it on MLX distributed. Now I had to write a bunch a bunch of scripts in order to orchestrate all this stuff so it's easy to run from one machine. First of all, you have to enable passwordless SSH between all the machines. You have to make sure that RGMA networking is hooked up through Thunderbolt in a mesh pattern like I showed you. but also they have to be on your LAN because they have to communicate through the root nodes IP address. Then being able to copy and distribute all the models across all the machines and all the scripts across the machines is something you have to write yourself. By the way, I may just put all this stuff in my GitHub. Let me know if you're interested down below. But usually Anie from the MLX team uh creates his own gist and puts all the instructions there. So I may not have to check out Anie's Twitter. He's probably going to tweet about it pretty soon if he hasn't already. I actually link to him down below.

Let's take a look at one machine right now. I'm going to kick off MLX generate command. And I'm using Devstral 123B instruct 4bit. All right, 4bit. I'm going to try 6bit in a bit. That's sounds weird. Um, but you'll see it's going to make a big difference. And there it goes. Write 200 words about Thunderbolt RDMA. This is happening only on one machine right now on this one. And it's not super fast, but it is a really large, dense model. This is a dense model, one that's supposed to shard really well across machines. Hint hint, we're going to see that. So, right now, we're getting 9.2 tokens per second. H, I like it a little bit faster. So, I'm going to run this on all four machines. And there we go. Keep an eye on the memory and the activity monitor here for the GPU history. We've got about 36 gigs on each machine here. There's the GPU history. You can imagine that's happening on the first machine as well. And it is using the GPU really nicely. Uh pretty much 100% of that from what I can tell from the chart. Ah look at that. 22 tokens per second. We've more than double the throughput of tokens per second when using more machines. That's what RDMA gets you.

Now, not only dense models get you faster performance when charting with tensor parallelism, but also this was quantized down to four bits, which means the full model was perhaps 16 bits and some information was thrown away. That's what quantization is, so that it can run on smaller hardware. Typically, that's what you do when you quantize things. If you haven't seen Julia Turk's channel, she has a really nice video on quantization. I'll link to it down below. And you can quantize to eight bits, six bits, four bits, three bits. Each time you quantize, you're losing a little bit of the information. So you might not get as best results. So what if you take Devstro and quantize it less? Let's say you quantize it only to six bits instead of four. Well, when you do that, you actually get very different performance. In fact, it's better. All right, let's run 6bit on one machine. This on one machine should be a little bit slower than 4 bits or it might be about the same, right? 200 words about Thunderbolt RDMMA in 6 bits. Yeah, we're getting 6.4 tokens per second here because this is a bigger model running on one machine. But what if we run this model on all the machines? Boom. So, we're using a bit more memory here on each of the machines. We're up to 40 41 GB instead of 30. and we're using the GPU to its fullest as you can tell by the GPU history chart and we're getting 17 tokens per second here. So here we're seeing that the sixbit gives us almost three times faster results using multiple machines.

Let's take a little detour for a second and take a look at a small model running on one machine. But I want to show you the differences between GGUF and MLX. I got Quen 34B tiny model. They're both four bits. One is GGUF and one is MLX. Write a paragraph. Pretty fast. 131 tokens per second. Which one was that? Can you take a guess? Let's start a new chat. Select the MLX version this time. Write a paragraph. 168 tokens per second. So, you can see for the exact same model, the MLX version is way faster on Apple Silicon. By the way, this is LM Studio running and it has the ability to run both. And LM Studio is pretty easy, but it does not have the ability to cluster. If we take this Quen 34B and we cluster it using Llama CPP RPC, which is the way that Llama CPP can cluster uh your models and yeah, it's possible, but it's also a little bit difficult to set up for you. I've done it all and I've been testing. I haven't had that much time with these. So, a lot more testing to come for sure. But here's what I got. We are down to 4546 tokens per second for the same model across the four nodes. Pretty significant drop from 131. That's where it goes down. That's because Llama CPP RPC is not using RDMA and tensor parallelism. And Jeff Gilling's been doing a lot more testing with Llama CPP RPC as well as Donato Capitelli who's been creating toolboxes for clustering on framework desktops. And I've ran the Quen 3 coder 480 billion parameter model. And across the four framework desktop nodes, if we're clustering this, we're getting seven tokens per second, which is not great. It is possible though, which is really cool. It's a 480 billion parameter quen coder model. By the way, I ran this for 128 prompt size, 256 all the way up to 16,000. I actually do include a 36,000 token prompt, but the framework cluster did fail that one. What is nice seeing here is that it's pretty consistent across all the different prompts and we're getting about seven tokens per second.

And I want to see what we get with Exo clusters up. Quen coder 480 billion parameter model. We're going to go to tensor with MLX RDMA and run it on all four. Launch. I love seeing the memory on all four machines go up at the same time. It's so cool. I wish I could play with this cluster more, but Apple loan this to me and they want it back. So, I have to do all my tests as soon as possible before I have to give it back. Write a paragraph. Boom. Look at that thing go. And we got 40 tokens per second here on four nodes. 40 tokens per second for a 480 billion parameter model is pretty nice. All that's left is hooking this up to your code editor and I made videos about that. I got member videos for that. By the way, if you are a member of the channel, thank you so much. Members of the channel get extra videos and I thank you for your support. But wow, this is really cool. Seven tokens per second framework desktops of course, so a little bit different. 40 tokens per second. But I've already showed you one node on Mac Studio versus four node improvements. So you know that this is what it's capable of. But speaking of even bigger models, let's see what else we can do.

I'm going to delete this instance. I'm going to launch this Kimmy K2 instruct which is 578 GB in size. Also, Tensor RDMMA launch. This one might take a little bit longer to load. We're at about 115 GB of memory usage on each machine right now. 130 135. It just keeps going. By the way, Kimmy K2 thinking has 256,000 context size. So yeah, my little hello world prompts are not nearly stressing the system enough, but MLX distributed adjusts dynamically based on the size of the prompt. So you can see it using 150 GB of memory. But if I give it a bigger prompt, it's going to use more. And check this out. It loaded and it's ready for me. And the memory pressure went down. Yet the memory is still at 150 GB on each machine. Hello. Boom. And it just says hello to me. I think it's mocking me. It's It says, "You gave me a short prompt, so I'm going to give you a short answer." It could be either really smart or really dumb. I don't know which one. Fine. You know what? Here is my architecture prompt. Design a scalable web application architecture for an e-commerce platform. This one usually gives me a pretty long answer. It's a pretty decently sized prompt. Boom. And there it goes. It's designing the architecture. It's going pretty decently fast. Tokens per second right now is 34.8. 8 and it actually adjusts the display on the fly for me. 34.6. You can see the memory has gone up just a tiny bit by a couple of gigabytes on each machine. And the GPU usage for this model is not 100% which is interesting. It's uh towering maybe about 90% there. I wish these charts had a better display, but that's all we get for the GPU history. 33.8 tokens is the final answer. Now, if you don't know this, Kim K2 has a trillion total parameters. We just ran a trillion parameter model. It's a mixture of experts, so it's not using all the parameters at the same time, but still 30 something tokens per second. Wow.

All right, time for the big boss. Deepseek v3.1 8bit. This is not even 4bit, folks. This is 8bit and it's the original training size, like the full thing. I want you right now to go in the comments and put down your guess for how fast this is going to be or if it'll even launch. Let's see. 120 GB on each machine and still going. 150 170 folks. We reached 200 GB on each machine. It's settled down at 183 per machine. That's a lot of gigabytes for every machine. I admit this is kind of nice. It's kind of luxurious. But it actually loaded. Let's paste in our prompt here. Deepseek V3. One of the highest ranking software developer models out there that's open ways right now. I gotta say though, this is the highest I've ever seen Mac Studios go. We're at 59 or 512 watts being used by this cluster running Deep Seek. Oh, wow. That that is uh no longer sipping. It's now drinking. And there it goes. Now, it's not as fast as Kim K2, but it's plenty fast. I mean, that's uh way faster than I can read. We're going at 25 tokens per second. And for such a giant model, that's actually really nice.

Now, a lot of these large models are now mixture of experts models. I do want to try a very large dense model. But right now, we need to have support at the MLX distributed level. We need to have support at the exo level in order for these models to be available, to be surfaced, and for us to be able to use them. But luckily, it's not super difficult to make the models available. We just probably need to open an issue or something.

Now, instead of running this $50,000 cluster, you could run EXO on tiny little M4 Mac Minis and still get the benefit of the easy setup and automatic discovery and combining unified memory to run larger models. However, the M4 Mac Minis don't have Thunderbolt 5. They have Thunderbolt 4, which does not support RDMA yet. Apple, please fix that. Or give us Thunderbolt 5 and M5 Mac Minis, please. To get RDMA support, you need M4 Pro or higher chips so that you can get Thunderbolt 5.

So, check out the new Exo 1.0. Check out MLX distributed and check out Apple's Mac OS 26.2 which is going to have this functionality enabled in software. That's crazy. I hope you enjoy this video. It was a lot of fun to do. If you want to see more details on MLX distributed and how I ran that, watch this video over here. Thanks for watching. I'll see you next time.