Transcription
We need to talk about hardware and, um, how to build your local AI. Now, today's video, I want to share what I learned. I've been using AI for the last four years since ChatGPT came out, and I've been on my journey to build, uh, one AI that I can trust. And that means moving from the dependency of those corporations to they give us those AI. They basically spy on us. Okay. They just train on our data. They look at our data and so on. So we have no sovereignty. The opposite of sovereignty is to build something that I, in somehow, I trust, works for me, works locally, 24/7.
Now, I don't want to give you into, put you in this to this trap of thinking, uh, this binary idea. Okay. Uh, 100% on the cloud for dependency and 100% locally. Because in this moment in history, the strategy is different. The strategy is to use the best AI you can to build what you need. Then the future is going to change really fast. So let's leverage the ability to, you know, the the opportunity to use those, uh, frontier model AIs at a price that are actually cheap. They're going to change the price. They're going to go up. They're going to do other things. So, it's a journey. That's what I think is important to understand. It's a journey towards sovereignty. It's a journey towards building your hardware and knowledge. So, you have total control of your system.
Now, in my journey, I'm going to explain my journey. Okay? But in my journey, I end up having so much hardware. Okay? There is a series of computers that I own. There is another one, I mean, here that is on. You can see there's another computer down there. There is this laptop, and I also run local AI on my mobile phone. So, a lot of different things. Let me explain how everything started. Okay. I'm going to try to make it fast so you don't lost. I don't get lost.
So, everything started from this laptop. Okay. This is my. It's pretty old now, but it's still an M1, 16 GB of RAM MacBook Air. Love it. Done a lot of work on this machine. And at the moment, it syncs, of course, to the cloud. And the fact that it syncs to the cloud means that all the data that is here is also on this computer. This is another Mac Mini. It's a Mac Mini with 32 GB of RAM, M4, 1 TB of hard drive. And then you say, but why you have another one? And why you have this one? Why you have this one? Okay. So, I have this other one for another reason. Okay. When, uh, Claude came out, okay, beginning of the this year, when it came out, the idea of running Claude on your personal machine would be madness. Okay. It's a new agent that we don't know what it can do. We know there is a lot of, uh, virus and so on. Okay. Prompt injection going on. And there is a kind of nightmare. You, I wouldn't, and you obviously, you didn't give to that agent access to your most important data. So this computer contains, contain most, my most important data, synced, like I said, with my laptop, with the cloud and so on. Okay. And because of, uh, I care about my client, my work, and the things that I'm doing, I didn't have give access to Claude to this machine. But what I did when Claude came out, before even buying this, I created a virtual machine, basically a computer within a computer within this Mac Mini. So I could test Claude. I installed it there. And what I did in the beginning. Okay. The first few days was insane. Okay. Finally, I had an agentic AI that I was trying to build myself for almost a year. I mean, at the time, it was nine months, okay, that I've been working on building my agentic AI, but I couldn't manage to do it. I had all the architecture written down, so I knew what I wanted. But I couldn't build it. Then Claude came out. I put it on this machine as a virtual system, meaning sharing resources. Okay. And it was still running really well. And in a few days, I built my personal AI. Okay. I modified, basically, Claude to do what I wanted. I had designed a really much better memory system and, uh, advanced system that would change the awareness level of the AI. Uh, I, I created this logician that is the one that controls in a deterministic way that what the AI is doing is actually happening. So AI works in the probabilistic world. Okay. But therefore, we can't trust the things that actually happen, that does, that does what it says it's doing. But we can have a logician outside that checks and forces this to behave in a specific way. Plus, I had a lot of other elements which, in enough, really in, in days, I implemented them. That is the, what, what I discovered with this AI. Okay.
Then I said, okay, this system is the future. I love it. This is amazing. I need to give a proper computer to this. Okay. Because it was taking resources from my machine that I used for work, and, uh, and the phone also had less resources that you could have. That's where this, uh, come into play. This is, I found this on, on discount, crazy enough. Okay. So, at the moment, it's even hard to, to buy them. This is the basic model of, uh, Mac Mini. So, on the basic model, the 16 GB of RAM, M4, uh, 256 GB of hard drive. Okay. Tiny. I paid $150 less than the, the price online. Insane. Okay. It's such a good deal. So I couldn't resist. I bought it. And, and now everything that to do with AI runs in here. Okay. So the software actually runs here. Hermes is in here. Claude is in here. Uh, Piperclip is in here. OpenCode is in here. Okay. CodeX is in here. Okay. When I used to work with, uh, cloud code, it would run on this machine. Okay. Everything was in this machine. And it can connect to the cloud and therefore use those cloud AI systems. But then I wanted something that would run locally. And while you can run something on this machine, it's not going to be powerful enough to actually be really productive. It's good for chatting, but not for high-level thinking or high-level coding.
And then I got this. Okay. This is a, a micro supercomputer. Okay. It's an incredible machine. This has 128 GB of RAM. Uh, has the same hardware of the DGX A100. And, uh, for those that are curious about the connection, you have all those USB card, USB, uh, input and, of course, Ethernet and those ports that are used to connect more than one of those machines together. I think up to four is possible to connect them. Incredible.
Now, are those hardware the best hardware you can buy? No. That is the thing, okay? You buy what you can afford. That's, that's the, that's the truth about life, okay? You can only buy what you can afford. What would be a better solution? And it's a weird thing to understand, okay? Because a better solution is an AI that is faster on one side, but is also stable on the other. So those hardware are extremely stable. Okay. Because AI at the moment has been developed over Nvidia systems. Okay. This one being designed by Nvidia is extremely stable. Really good. Extremely happy about this machine. Has only one problem. The speed of the RAM is not that great. Therefore, the token generation, the token per second generation when you chat with AI, for example, it can go really low. It depends on the models, depends on the size of the models, depends on the architecture of the model, but it can go slow, slow to a level that is frustrating. So having a higher, uh, speed in RAM, you would have a better interaction, basically.
Now, is it possible to run on this machine an AI that is pretty good? And is the, is the, what I use on my Hermes, which is pretty good and pretty fast? So, what is the solution automated on this machine? The best solution for me is the Qwen 3.6 35 billion with three billion active parameters at the time that allows the speed basically of a three billion parameter having the intelligence of a 35 billion model. Now, there are better models. Qwen 3 does the 27 billion parameter, which is smaller than the 35 billion, but that one is, uh, uses all the 27 billion parameters all the time. Okay. So meaning it needs to do much more higher compute. Therefore, you have higher intelligence, but also a much slower, uh, token per second. So talking about token for for a second, for those that are curious, on the Qwen three billion, a three, uh, billion, say with three billion active, we have around 70 tokens per second, which is pretty good. Okay. When the context window starts to grow, you slow down a little bit in a linear way. So, really pleasant way, and, uh, 70 tokens per second, even more sometimes. It's really good. Okay. You don't feel any slowness. It's, it's really fast. It's like working with those AIs from the cloud.
Now, the 27 billion instead is around 10 tokens per second. You can optimize, find ways, maybe you can go a little bit faster. And the, the truth here is, they are, those software getting better and better, so they're optimizing more and more, so they're getting faster by themselves. But that is the speed. Okay. 10 tokens per second, it's in somehow usable. I'm not saying that it's not usable, but it's frustrating. Okay. It's not, it's not how I like to work personally. So speed is one of the elements that is important because it's, it's, it needs to match, in, in some form, the speed of my brain. Okay. I go fast. So stay there. Look at this thing working slowly. It doesn't give me joy. Let's say the, the other model is much faster. Love it. Gemma 4 also is a really good model. There is the 27 billion with 4 billion active parameters. Little bit slower, on the 50 tokens per second is a good solution, but at the moment, Qwen 3.6 35 billion is my go-to for this machine.
Another thing to consider is that you can run more than one instance of that model. Okay. And, according to the size of the context windows, you decide how many instances you can have. So you can have really a lot of them. Let's even, for example, I think you should be able to go into the 20s with that model, 20 in parallel, but the context window is going to be small. So for a context window of 200,000 tokens, you can have, uh, two or three running at the same time in parallel, which is a lot. So you can have, for example, one for Hermes and one for the OpenCode. So to write code 24/7 while you chat with Hermes and use Hermes to control OpenCode and so on. So you can have this advanced system working simultaneously. Love it. The quality is pretty good.
Now, I want to give you another reflection. When I bought this, Gemma 4 didn't exist. Qwen 3.6 didn't exist. Okay? When I bought this one, I paid $3,200. Okay? Which is approximately, I think, $4,000. Now, it costs almost $1,000 more. Prices are going up for this machine. And they are not getting old in the sense. Yeah. Yeah. But now you can do, uh, the new things that come out cannot run on the, no, actually, the new models that came out, they give more power to this machine because now they run even faster because the models are optimized, they are even more intelligent. So it's such an interesting scenario. So it's the, the value of those machines is not going down at the moment. And, uh, I don't know what is going to be the future like in the sense, 128 GB of RAM is start, is, is quite a lot. And what is possible to do now through this more advanced models, they are getting, they're solving some architectural problems. So by doing so, what they do, they optimize, so they become more intelligent even within the same amount of RAM, even with the same amount of, uh, parameters. That is a really interesting, uh, system. So if you see, for example, the DeepSeek version 4, they added some specific architecture inside which allows that model to be extremely intelligent for what it is.
So now, okay, we can match with those, uh, frontier models like, uh, GPT 5.5 or Code, I can't remember now, 4.7. So we can match those ones on local machines. But in some most scenarios, actually, we don't need to match it. In some other scenarios, it makes a difference. Now, so what is the strategy? I don't want you to go into the binary thinking of, okay, I can, because I can't have total local AI. Okay, I'm going to go for the 100% cloud. Or because I hate the cloud, those AI corporations. I'm going to go 100% local. I think those, this is a source of trap. Okay. We need to understand the, the moment in history. In this moment in history, the local AI is powerful, but not perfect. The cloud AI is, it's, it's powerful, let's say, but everything is not perfect as well. The, the business model behind this is, is the problem. You know, the lack of transparency, the, the training on your data, the controlling on your data, the profiling of who you are, all those things is unacceptable from my point of view. The, the illegal way of how they train their models, it's, it's all, it's all no good.
So, but what is the scenario for us? So what shall we do now? The answer for me is simple. Okay. Get a cloud AI which is convenient for you to do high-level work. Okay. For example, building the architecture of your software. You build architecture with this high-level, so you know what to do. And then build most of the software with your local AI, whatever it is, okay, whatever you have. Use the local model to build it. Then, if it doesn't work, if there is a bug, if you want to review the code, you give it to the bigger one again. Okay. So the bigger one just works every now and then. But what, when it does, it does really high-level, uh, work. So stress testing, there is a problem, a bug, make it fixed by the bigger model. Okay. And there is even the architecture. Okay. Building the architecture in a way that is more modular. So at the size that is that your local model can deal with. So let's say a code is huge. Okay. Therefore, needs a a massive context window to look at it. But if you cut this code in blocks, okay, in, in parts, then each part is kind of self-sufficient. That self-sufficient part can be looked by your local model and understand and fix it. So start to think in those terms as a mixed scenario.
Then, of course, if you have more money, what would be the best solution? More money with, with what we have available at the moment is not even those scenario. And it's not even, from my point of view, Mac Studio. Mac Studio has this beautiful thing that for you get a lot of RAM for a lot of money, but nothing compared if you buy like the RTX 4090s. Okay. This card has 32 GB of, uh, of RAM each card. So to get to 512 GB of RAM, you need a lot of them. Okay. To a level that is probably impossible. I think you can maybe run up to four of them in parallel. Uh, so 128 GB of RAM, which is good, but they are pretty expensive. Okay. So think about $4,000, $5,000 probably each. So, four of them is, uh, $20,000 plus the computer, plus the RAM, plus the CPU, plus the case, plus the electricity, which is insane. Those, those, uh, hardware burn a lot of electricity. And you still have 128 GB of RAM. Think about the 512 from the Mac Studio.
Now, I want to share this my thinking about this. I don't think you need 512. Why you don't? Even though you think, okay, I can load a massive model. The problem here is the speed at which the model will run. And if the speed is not good enough, you're not going to use it. So a sweet spot is 128, 128 GB of RAM like this one, because you can load a model which is big enough, but still have a large context window and still have speed enough for actually using it. So if you have a, like, for example, at the moment, you can load on this large Mac, what's the name, Mac Studio, you can load Mixtral 250, uh, billion parameters, you should be able to load it, but quantized, okay? And then, uh, it's going to run slow. What's the point? There's really not the point. Okay? So it's better to have something that is balanced correctly to the speed. For this is my thinking. Okay. You might, if you have a 512, tell me what's your story. But my thinking here is, uh, something built on Nvidia at the moment is better because it's more solid. Okay. It's true that in somehow, even though they are really expensive, it's true that the Mac processor with the unified memory, you have a lot of memory, really fast memory, but you have a less solid infrastructure and not fast enough to actually justify the amount of RAM. So for me, the best solution is the RTX 4090. You can start with one, okay, and you can run the 27 billion with a decent, not full size, but decent context window and, and then you're happy. Okay. You are really in a good position. Decard $5,000 plus the computer, you go into the, uh, $8,000, $10,000 machine. A lot of money, around four, five, maybe now that the price is up. Okay. You see, it's so complicated making the right decision. The right decision is connected to your money.
You can have, uh, something like this and run cloud computer. Okay. Maybe a Whisper here to to transcribe your audio messages to, uh, text. And having Hermes running here is fine. You can do it. I think those machines are $3, $400. And, uh, Mac Mini is the next step. Okay. Much better. A little bit more expensive, but much, much better. I wouldn't advise personally to use a laptop because, like I said, I want the system to work 24/7 and, uh, a laptop is not ideal to keep it running 24/7. That doesn't make much sense. It's too expensive. So, it's better. In my case, I have this Mac M1 version, 16 GB of RAM. It's still doing a lot of work. Fine. Okay. I can connect this one through the cloud to this machine and using AI. Okay. And them together, it costs probably less than the Mac Pro 128 GB of RAM. But here you have something that runs 24/7. So that is my, my thinking.
Now, I want to conclude by saying a few things. Okay. First of all, like and subscribe. Help me build this channel. Second, it's also for a reason. Okay. Also for a reason when this channel, I mean, this channel is already monetized. Monetization means I make actually money from this, uh, YouTube channel, but so little that it doesn't make any difference. Okay. But growing this channel means I can afford now getting the RTX 4090. Okay. The latest, uh, uh, card from Nvidia. Therefore, I can, we can try more and learn from it. I'm going to share everything, all the experiences. So it would be nice to get to the position.
Uh, another thing that I want to say, every day of the week, we have inside the community, we have these, uh, calls. Okay. Every day at 8:00 PM Central European Summer Time. We have those. Now that it's summer, it's summertime. Uh, so 8:00 PM, we have those, uh, community calls. Those community calls, we discuss a lot of things about AI. Every Tuesday, there is this call that is the academy. Academy means is a place to learn. This Tuesday is going to be about, basically, tomorrow. This Tuesday is going to be about V-coding, a sort of a masterclass of V-coding. I'm going to share how I've been V-coding. There's going to be other people, professional coders and amateur coders like, like you and I, okay? That I'm going to do sharing their experience. Now, I said you and I, maybe you are professional coders. So, just like me, okay, just like me, I'm not a professional coder. So, I, they're going to share their experience. You can share your experience. We can learn from each other. The thing here is, there is a lot of people that has been V-coding amazing stuff without any previous training, and therefore they learn by doing. Learn by doing means a lot of mistakes, and then you fix those mistakes. There are some mistakes that they, that we are making, making as a V-coder, that doesn't make, uh, sense. Okay. But we can't understand them. So working with other people, we can see each other how we do, how we go around some problems, and see professionals what they do. Now, if you're a professional coder, I tell you this thing because you have been working as a professional coder, you are missing out from actual crazy V-coding because as V-coders, we don't understand completely the limits of the machine. Okay? So we ask things maybe are crazy to do, and they actually do it for you. Say, this is not going to happen. You don't even ask, you don't even try. You, you start thinking, I need to stay in control, and so you delegate only part of the thing. But we, as a V-coder, we have to delegate everything. Okay? We can't write any single line. We can't understand any single problem. We can't audit any single line of the code, and we can't fix it. So for us, it's 100% delegation. What is changed is how we need to think about this. Okay? And because we've been 100% delegating to this AI to do work, first of all, we do a mess. Okay? For sure. But also, we achieve things through V-coding that you maybe never thought was possible because we had only that option, V-code or nothing. Okay? So it's two worlds that are different, two different approaches that if they meet, they can learn from each other. So, I really invite you to join the community. Link in the, in the description. It's a Discord server. You join there. There's going to be the link shared every day. We have calls every day, okay? I mean, five days a week, Monday to Friday. Each one has a different theme, but you can go in the Discord and read about those calls. Tuesday is about this. Tuesday is about V-coding. It's about the academy. So, I'm really looking forward to hear your story, what you're doing, what's your hardware, what you're building, and, um, how you V-code. Thanks for staying to the end. Like and subscribe, and, uh, see soon. Ciao. Watch the next video.