📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Why you NEED to be running local AI models (FULL beginners guide)

Alex Finn21:27

Transcription

I'm about to show you the future of AI, AI agents, and OpenClaw. Over the past two months, I've spent over $50,000 to use, test, and learn about local AI models. What I learned, I think, can dramatically change your life and save you tons of money, even if you're on a cheap computer. You don't need to buy Mac Studios like me.

In this video, I will cover everything local AI models. I'll cover what computers you need, what local AI models even are, which models you can run, what use cases you can do today, and how this lets you use Open Claw completely for free. I'll also show you a glimpse of the future that I am 100% confident is what's going to happen. By the end of this video, you'll be a local AI master and you'll be running your own local super intelligence on your computer. So, let's lock in and get into it.

So, this might be my most important video yet. I'm so excited to take you through what I've learned over the last couple months, even if you have no idea what local AI models are. You're going to get so much out of this video. So, let's start off with why local AI is so important and why you need to be using it.

This is what you're probably doing today. If you're on Chad GBT or Claude or using any of the AI frontier models that everyone knows about, you're using a cloud model. That means you're using AI that are running on these big servers that might be underground or one day in space because of Elon or might be on an island. But because you're using these models that are running on these servers, that means a lot of different things.

One, it's expensive. You're paying for every token you use. Every time you send a prompt to these servers, it's doing a bunch of calculations and they're charging you for each one of those calculations. These are where these massive API bills are coming from. These are where the $200 month subscription plans are coming from and all your API usage. It's very expensive to be running AI models on these billions of dollars of chips.

It also has many other downsides. Zero control. A lot of people complain all the time. They feel like their AI models are getting stupider. In reality, they probably are. These AI companies are constantly dialing the knobs and changing things to try to save money. You have zero control over the AI models running on these servers. You have zero privacy. Every message you send to Chat GPT or Claude or Gemini or any AI model you're using on the cloud, those employees can read those logs. Nothing you say is private and secure. Every question you ask about your health or maybe if you're a sicko and you have your own AI girlfriends, they can read all of those messages you're sending. It's also laggy. You need to be connected to the internet. If you don't have great internet, it could take a while to get your prompt sent there and sent back. So, there's a high latency. And on top of that, it's just not scalable. A lot of people have been learning this lately with OpenClaw. Maybe you connect it to Opus 46 API. You send a bunch of prompts. You look at your API bill and whoops, you spent $1,000 over the last day cuz you sent a bunch of prompts. It's not scalable at all. And if you want super intelligence working for you 24/7, it's going to cost you millions of dollars.

But with all that being said, you do get one benefit, which is you get front tier AI. You're getting the best AI models. They're running on these servers and you're getting the best performance. That is probably what you're used to today. But where I strongly strongly believe the future is going and what I actually believe you will be doing in the next 12 months is you will be using local AI models.

What are local AI models? These are AI models. Instead of all these complex multiplication equations happening on servers across the world, they're happening locally on the Mac Mini on your desk or the Mac Studio or the old dusty Lenovo laptop, whatever you're using. The models run locally and that has a tremendous amount of benefits.

First of all, it's completely free, right? You're not paying for tokens. It is just the cost of the electricity going into the computer you have plugged into the wall. It's fully customizable. If you want to take a local model and make it sound like you or make it rap like Kendrick Lamar, you can do that. They are fully customizable. It's also fully secure and private. So, every message you're sending to your local AI running on your computer on your desk stays on your computer. It does not go to the internet. Nobody can read your prompts or your messages back and forth. If you want to get freaky deicky and make your own AI boyfriend or girlfriend, you can do that and no one will read those messages. Not that I would know anything about that, but also zero latency. There is no messages going to the internet. It's all staying on your device. So, you literally can unplug this from the internet and the AI would still work. You can be on an airplane vibe coding to your heart's content and it doesn't matter because there's no internet. It's all local on your computer.

And here's the best part. Here's why I'm bought it. And here's why I think everyone will be using local AI in the next 12 months. Because it's local. Because it's free, it is extremely scalable, which means you can have AI doing work for you 24/7, 365. I have right now, and I'm going to demo this later in the video, so make sure to stick around for this. I have right now four local AI models, doing things for me continuously, going on the internet and scraping websites, writing me content, writing me newsletters, writing code for me, just doing things at all times of the day. It's like I have multiple employees working for me. This is an advantage I have over all of my competition because they are not running local AI models. And if you do the things I'm about to show you in this video, you will have the same crazy advantage over everyone else in the marketplace as well. That's why it's super critical stick here till the end.

Now, the one downside, what's the one downside to all of this? Local models aren't quite as smart as the frontier models. I'd say they're about 6 months behind at all times. So, so 6 months ago was like Opus 45, Sonnet 45, around that realm. The local models are about there. Now, if you think back to 6 months ago when Opus 45 came out, it absolutely blew people's mind. So, we're still we're at that point when it comes to local models. So, it's still really really strong.

So, that brings us to our next point, which is what computers do you need? Do you need to run out and buy $50,000 worth of Mac Studios and DJX Sparks like me? Well, the answer to that is no. You can literally run local models on any computer you have. So, you have an old crappy laptop in your closet from like college or something, you can take that out and run local models. If you have the new $600 Mac Mini that everyone was running out and buying a few months ago, you can run local models on that. That was a very good purchase.

Now, are the models you're running on these cheaper, smaller machines going to be Opus 45 level? Well, no. But there's still use cases you can run. You can still do things like memory management for your open claw. Having a very small local model, deciding which memories get loaded into context for your OpenClaw or your AI agent or whatever is still a really powerful use case that you can run on your $600 Mac Mini. And I'll go through the exact models you should be downloading for each device in a second. But even if you have these old dusty computers, you can still run local models. And also, as a side note here, even if that use case doesn't interest you, the the fact that you're downloading and using local models, which I'll show you how to do in a second as well, will teach you so much about this incredible technology AI. you will be learning and tinkering with this technology which if you're watching my channel you probably agree with me is the most important technology in the history of our species. So you're getting that advantage as well just the pure education part.

Now we take it a step further and get a little bit more expensive. You have the devices like the DJX Spark and the Mac Studio. This is the devices people are typically going for when they want to spend a little bit more money. And now you're getting Opus 4, Opus 4 5 level intelligence running. Now, there are different advantages and disadvantages to going with a Mac Studio versus a DJX Spark. The advantage to going with an Apple device is they have what's called unified memory, which basically means you can use any of the memory on the computer for processing. So, you can use all the memory to load full AI models. With Nvidia chips, you can only load it into the VRAM. So, the memory specifically for the GPU.

Now, there are advantages and disadvantages to both. The additional memory in the Mac will allow you to run much bigger models, which means you can get much smarter intelligence. The downside is is you lose a lot of the speed benefits you get from the VRAM. So, it's going to go a little bit slower, but you do get bigger intelligence. There are also some amazing developer tools available just for Nvidia chips that are not available for Mac chips, although I think that will probably change in the future that allows you to do things like fine-tune your own models. Train Loras. For those who don't know what Loras are, they're basically plugins for your AI models. It's like when Neo had the USB and put it in the back of his head in the Matrix and he just learned something new. That's basically what Loras are, are those USB sticks. You plug in a USB stick in your AI model. It teaches it something new. And then you can also do auto research much better on Nvidia. For those who don't know what auto research is, it's basically this new framework that allows AI models to improve themselves. It runs very well on Nvidia. Although they're coming out with this functionality for Mac that doesn't perform quite as good, but still pretty decent.

So, if you're more of a tinkerer, I'd probably say the DJX Spark is better. If you just want to load really good super intelligence locally that you can use for free, the Mac Studio might be better for you. The Nvidia DJX Spark I think is like $4,800 now. I got mine for $4,000. I think they just raised the price. That's still a pretty good price, $4,800. The Mac Studio scales. If you want to go baller like me, I maxed mine out. I actually have three maxed out Mac Studios here. They're $10,000 for the 512 GB, but you can spend like $4,000 and still get a very solid Mac studio that will give you a ton of performance. And then for the big ballers out there, Nvidia just announced the DJX Station, which is going to be a $100,000 computer, which is going to have like 750 gigabytes of memory that you can use for AI models. This is going to allow you basically frontier level intelligence with massive multi- aent workflows where you have multiple models going, doing different things on your computer all at once, and you can train your own models and do that all simultaneously. You'll basically have an entire AI research lab on your desk. It's going to be $100,000. The good news is though is the cost of intelligence is constantly coming down as well as the cost of hardware. So I believe in the next 12 months or so you might have this amount of power in this form factor. So we are getting close.

The point of all of this though is to say you don't need to spend $100,000. You can have the $600 Mac Mini and be running models and doing really, really cool things. But let's do this real quick. Let's talk about which models you should be using and then I'll go into the juicy part, the fun part is what the hell do you do with these models anyway? How do they plug in OpenClaw? How do they save you money? How do they do really amazing things? I'll show you my setup and all my use cases as well, but I just wanted to make sure we're on the same page of why local models and what you need to do to run them.

Now, let's talk about which local models you can run. There are three I use that I pay attention to that I think are great. Quen 3.5 is excellent. The reason why it's excellent is one, it's just smart, but two, they have versions for every computer you have. If you have just a $600 Mac Mini, they have like a 9 billion parameter model that you can run on like the 16 gigabyte Mac Mini. They also go much bigger. They have a 29 billion version, even bigger version than that. This is my daily driver. I'm able to run it on my Mac Studio that I have here. But even if you're on a Mac Mini, you can run Quen 3.5. Nvidia just entered the opensource race with Neotron 3, which is also an excellent model. It just came out. I'm testing it. It's very, very strong and very good. They also have versions for different devices. So, I would recommend trying them out, testing them against each other. And then there's Miniax 2.5. Miniax just came out with 27, but it's not open source. But 2.5 is still a really, really good model. It's lightweight and it's extremely quick. It might not be quite as intelligent as Quen 3.5, but the performance is amazing if you want really quick tasks done. I'm super excited about Neotron 3 to be quite honest with you, just because it is basically the only American company going hard to open source. Meta is doing it as well, but I don't know, they haven't really had great results, but Nvidia is doing a really good job. So, I'm excited for Nvidia to enter the market for open source models. But these are the three I would focus on if I were you.

And if you want to download any of these, load them in, the best website for doing that is HuggingFace. So, this is HuggingFace. This is basically a website that hosts all the different open- source models out there. You don't even really need to interface it to be quite honest. If you want to load a local model onto your device, just go to your OpenClaw and say, "Hey, based on the computer or hardware I have, what would be the best local models on HuggingFace to load? I'm thinking about maybe Quen or Neotron. Take a look at Hugging Face, see what the latest models are. Look at my device and see which models we can load." If you just reverse prompt that to your OpenClaw, it will find the best models for you and the hardware you have. But HuggingFace is an incredible site. I'll leave a link for it down below, so you can go check it out, browse the models, and load one up onto whatever device you have currently.

And just as a quick demo, this app here is called LM Studio. It's completely free. I have it running here on my Mac. This allows you to load in those AI models from Hugging Face and start to use them just like you would a normal chatbot, right? I I have Mini Max 25 loaded right here. I said, "Hey, how are you?" And I can even say write me code for a snake game. And I can hit enter. And you can see right now it is going to go, you can see the thinking all this is happening locally on the Mac Studio on my desk and it is going to go ahead and write the code for the game. Those are pretty good speeds. A lot of people crap on the speed of like Mac and Apple silicon, but I don't know. This looks pretty good for me. And boom, there you go. It's all done. I have an entire game written in that was like real time in like 15 seconds.

Now, let me take it a step further and show you how I'm plugging this into OpenClaw. What you see here is my software factory. This is where my OpenClaw and all its sub agents are doing work and building software 24/7 365. And what you can see here are a couple things. Charlie is my main coding agent. that Charlie is powered completely by Quen 3.5 and it is coding 24 hours a day 7 days a week completely for free. Now Charlie is being managed by Ralph. Ralph is like the software developer manager of this whole operation. Ralph is being powered by Chad GPT still. Ralph is the manager kind of orchestrator of the entire software factory. This is what I believe is the best approach at the moment and what you probably want to go with no matter which device you're on. You want to go with the brain muscles model where the brain the one making all the decisions is the cloud models. You can go with Chad GPT. You can go with one of the cheaper cloud plans like Miniax if you want. And then the muscles, the ones actually taking orders from the cloud agents and doing the work, writing the code, doing the research online are your local models. So again, for instance, Charlie is Quen 3.5 running locally on my Macs to just writing code at all times. Right?

To show you another example, I have a local model running on Miniaax that's constantly scraping the web for business opportunities and gets sent to me here in my chat right in Telegram. Opportunity scanner flagged a heat 8 signal AI app reality check testing for vibecoded products. My local AI model is just going on the internet literally 24/7 365 looking for business opportunities that I can capitalize on. I basically have a scout going and finding ways for me to make money and it cost me zero to do because it is a local AI model doing it. If I were to try to do this with Chad GPT or Claude today, I would literally be spending thousands of dollars a day to do this. But because it's a local model, it's fully scalable. And the reason why I'm showing you all this is to show you what's possible. If you leverage local models, you can have these workers, these employees going out and finding you opportunities and writing you code and doing you research. You basically have your own digital company.

And this is just the work side. If we talk about the cost side, this has driven down my cost a ton. I get the question a lot, you must be spend thousands of dollars a month on all these things you're showing us. No, I'm probably not spending any more money than you because the cloud models I'm using are just doing the orchestrating. They're just doing the thinking. All the execution is done locally for me.

Now, if you want to get even spicier and take it to another level, if you have good hardware and you connect maybe a couple Mac Studios or a couple DJX Sparks, you can load a really big heavy intelligent model and have that orchestrate everything and have this entire digital society, digital company running 100% locally, 100% for free, just the cost of your electricity. That's possible if you want to do it right now. And if you think about it, if you do that, yeah, maybe you'll have to spend $20,000 on Mac Studios, but at the same time, it's going to eventually pay for itself. If you're a power user like me, you probably spend a couple thousand dollar a month on API credits. But if you're doing it locally, you're saving money on those tokens.

Although, I will say this, I don't believe this is about cost savings. Out of all the reasons I showed you earlier to go local, I actually think cost savings is one of the least reasons why to do it. I think the power of running models 24/7 365 is the reason to do it because by running them 24/7 that just unlocks so many use cases you could have never done before with the cloud models. So I believe right now at the time of this filming the hybrid approach is the way to go. You use a cloud model like Opus or Chad GBT whatever you want to use as the cloud orchestrator as the brain that's running your open claw and then you use your local models as the tools as the muscles to your brain right you use Quen to do the coding maybe use Neotron to do the research Miniax to do the writing your brain cloud model is passing off these tasks to the muscles to do the work for you.

Don't need to run out right now and buy $20,000 of Max Studios. You don't need to be like me. You can start cheap. You can start with your Mac Mini and maybe you hand off one small use case to a local model. Then once you do that, you think about, okay, what else can I hand off next? If you think of a really good use case in your workflow that you can hand off next, then maybe upgrade. Then maybe you go out and buy a DGX Spark. You utilize that. You max that out. You start seeing an ROI. Then maybe you go out and with your savings, you buy a second DJX Spark. I believe that is the best way to be doing this. I believe that's the most scalable way to do this. I see a lot of people running out buying like $40,000 of computers and they have no idea what to do with it. I believe the most responsible way to handle this is start with what you have. Learn about the technology. Start with one or two small use cases, then slowly scale up and layer on to your local hardware.

I'll also say this, the M5 Ultra, which I predict probably comes out by June this year, is going to be incredible. The M5 we're seeing on the MacBook Pros are amazing. They're like five times faster with AI processing than the M4. They are a revolution for Apple silicon. So, I would also at this point see what the M5 Ultra looks like when that comes out with the next Mac Studio, but that's going to be incredible. I'll probably end up buying the $100,000 DJX station just because I'm about that life and want to get psychotic. So, make sure to subscribe and turn on notifications for demos of that when that comes out. If you learned anything at all, make sure to leave a like down below. I also do weekly live boot camps on local AI models in the Vibe Coding Academy. You can join, ask me questions, show off your use cases to the other, 1300 people in the community. Make sure to join that link is down below. It's the number one AI community on planet Earth.

What you see here is all I do. I spend 24 hours a day playing around with this stuff, trying new things, trying to find incredible new use cases. So, if you enjoyed it, subscribe, turn on notifications. You'll get these videos the moment they drop. I'm so grateful you would watch these videos. I'm so grateful you'd spend time listening to me talk for like 20 minutes. I couldn't do it. So, thank you so much for watching and I'll see you in the next.