📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Your $20 AI Subscription is a Trap. (The Local Solution)

Manolo Remiddi20:36

Transcription

We need to talk about local AI and what is possible to run on these kind of machines. Now I discovered something really important I want to share with you because it's kind of surprising and disappointing at the same time.

Now before we start I want to share you one one thing why I'm getting into this local AI. So I been using AI since JTP came out. Okay, 3 four years time and more and more I've been adding AI to my work. Now I can't even think about running a business without the use of AI. Okay, what's happened to me that I use AI to augment my ability. Now I can do more. I can do faster. I can do better and that is already a lot. But then I started doing things that I couldn't do before. So now I can do things like coding for example that I couldn't do before and this open the amount of possibility that I have. My feeling was like okay now that I can't I can't stay with AI. I need it. Okay, of course I I I can adapt and so on. But the idea of working with AI is not an idea that I'd like to address. Okay, a director that I like to do. Therefore, I felt like I need to have at least a simple AI, a basic model that I can control it. It's mine. It's always going to be there. I don't want to leave. I I don't want to live without. So because I have this now I can sleep much better. So this is the my personal reason but there are even better reason that just sleep at night. Okay. And the direction of using local AI is pretty simple. Okay. If you want to really learn how AI works, you need a machine like this one. Nothing else. A machine like this one. If you want total uh privacy, you need to run AI. locally. If you want the total sovereignty, you need a local AI.

Now, let's go back to what I discovered. Actually, before we do this, okay, I need your help. I need you to help me grow this channel. Okay? So, you need to like, subscribe, watch the video until the end, and leave a comment. Let me know uh what kind of a model you use, if you have a machine like this one or something else. I'm going to share some data during this video so you can also tell me if this data match your experiences because there is something really kind of upsetting about what I discovered before for those people that don't know what this machine here is. This is Asus GX10. Okay, this is the equivalent because inside is exactly the same of the Nvidia DGX Spark. Those machines are like of a micro superco computer. Okay, inside have the GB10 processor. This is an incredible uh system to run AI for training AI to to you can do everything with this one. Okay, the problem with the only problem those machine have apart of being expensive. The only problem is the RAM is low. Okay, the therefore inference is a little bit slow. So in the sense a MacBook M5 Pro uh Max with 128 GB of RAM run cost more than double of this and but run at almost double of this speed. Okay. So there is uh a difference there.

Now, what I learned yesterday, I was watching this video and was this person testing this uh ladies uh MacBook Pro M5 Max 128 GB of RAM with Gemma 4 26 billion. This Gemma for 26 billion is the AI model that I use on this machine to run Hermes. Okay, so I've been using Hermes with their model since it came out. Okay, now I think it's a month and a half. So I had a lot of experience with it and I was pretty pleased. So I was really curious to see what this person would find out through these benchmarks especially because when I bought this you never know. I think I thought like okay maybe I should spend more and get the MacBook Pro or or should I go for something like this? So I was really curious. So, I watched this video and uh I'm going to share in the description uh a Substack article where I'm going to share all the data that I present here plus the link to this uh video. Okay, which you should like and subscribe uh because the data is really good. So, but it it only shared this data on the MacBook Pro and then I did the test on here. So, now we can compare.

So what came out is uh what was this benchmark it was doing. So the benchmark was about not just what token per second can the model work okay but also what kind of to that model can work at different prompt sides and if the quality of the output from the prompt is actually valuable or not. Okay. Is there an error? This is hallucination or whatever. Okay. So that is also the risk because larger prompt could lead to hallucination and problem and error. So this person run this benchmark and uh of course like I said double of the speed 115 second token per second on the Gemma 4 26 billion. Here we have around 55 60 token per second sometimes upper sometimes lower but around that number on this machine. So almost double the speed. But something really upsetting happened. So on the MacBook Pro M5 Max, the 128 GB of RAM, the when the prompt size reached 32,000 tokens, the output has error. Okay, that is 32,000 token. It's it's big but it's not huge. Okay, think about that. The context window of that we usually have is around 200,000 tokens and on Gemma 4 should be able to manage 262,000 token contest window 32,000 it's not big at all but at that size it would create error hallucination and so on. So 60,000 token was the uh on this benchmark was the level where this AI would work at a decent speed, acceptable speed and u without errors. So I was like wait a second I never saw this problem with my eye. It was me I mean with the same model running on this machine. So I was really curious at at that point to understand is it me not realizing that the is telling me you know hallucination and just say yeah yeah yeah all good or is actually is better on this machine. So I wanted to test it and I run some preliminary test and then I'm on the article going to add all the details that I have. Okay. Now I wanted to understand is do I have this problem or not because I never experienced through airmes to have this kind of problem. Now what I found out is this machine can run without any error even a 261,000 tokens compared to the 16,000 token on the MacBook. That is huge different. But I want to tell you more.

So how does it work? Okay. Then the the context window is where we work with AI and uh 32,000 is just a part a part of it. So you imported you inject that prompt with within the AI then you send it to the does the process and and send you back. Now the context window keep growing when you use AI with agents like Hermes and open claw and so on what you have you you create a lot of uh tokens. Okay, because you give task to those AI. So they need to think, they need to do things, go tools and so on. Okay. So they you burn a lot of lot of token. So 16,000 token is nothing. It's really nothing. Now how because you start using a lot. What's happening with this this contest window? The contest window usually is set to 200,000 token. Okay, Open close by default. Put it at the sides. What is happening when you are around 170 you have uh 170 180 the system go into compaction. What is compaction is basically taking the raw data summarizing and push it back and removing the old data. In this way you you now have again a lot of space that comp the quality of compaction is uh is extremely important because that summary contain the elements that should allow AI to carry on working. So a lot of people have been experiencing that open crow for example would give you such a bad summary that the system just would like get lost forgetting important things like a suffering of amnesia that was connected to the compaction system. This comparison system is used by any kind of uh chp cloud gemini everybody using this idea. Okay, because nobody wants to go to a huge uh contest window because it become really expensive. So this is how we save and work.

Now the experience of why I'm saying this because when you use with Gemma 4 with log AI the size of the contest window the side of those um the number of those tokens change the speed of which you receive the output. So at 192,000 token okay which probably is where the compaction is going to for sure be done the on this machine you would have 20 token per second. Okay, which is getting into the frustration level. Why I'm saying this thing? Because even in that video that I'm going to share, it's this this person say over 30 second 30 token per second you start uh working properly and you start feeling the the frustration. So 20 token per second is in the frustration area. The truth here is you need to remember that you start that works around 60 token per second and then over time longer it gets this contest window longer guys this prompt this back and forth and bigger it's going to be and therefore it's going to slow down more you use it to a level to reach frustration level then it's going to get compacted okay so summarized and go back into going fast enough so there is this experience that is not pleasant. Now if the MacBook would be able actually to work until that kind of size then the experience would be better because like I said the starting point is double of the speed. So let's say that this the speed maintained the double means that at around 200,000 token you still have 40 token per second. So the experience on those mech potentially would be better but in practice they cannot be used. Okay that is the I don't know extremely frustration for me. So in somehow I'm super glad I went with this machine and again those machine also allow you to do much more experimentation much more learning they are much better for AI. Okay. So if you want to learn get better if you want to work in uh with with you know in somehow a system that is more solid those are better.

Now there is always uh something that probably you could answer in the comment below. It's this is depend of the model. Okay, this gem for create this problem at that kind of uh size of prompt different size create different problem. Okay, different pro different model create different problems. So let me know what you're using. Maybe an alternative to this one. There is the uh quen 3.5 I think 35 billion parameter that that one is also fast and uh could be maybe a better option than the Gemma. I don't know how it works on your system. For me, I'm really pleased about the the Gemma.

What I want to to share now is something another that is also really important because the there is also something that we need to look at the bigger picture. Okay. Why we are using this logo model? Why when for $20 a month you have extremely powerful solution. So it's really sometimes really hard to justify uh something like this. Okay. In my case I sleep better. I'm so much happier to have this. I can learn, I can experiment, I can do so many things. So I'm in my case, I'm pretty pleased that I bought this machine. And like I said, I bought it this in a scenario where was uh on discount. Now it's $1,000 more expensive than when I bought it. And this makes us think also about one thing. Okay, the demand of token is increasing. Okay, the price of those token is going to be increasing. there is more demand than supply. So to meet the demand they are still investing a lot of billions to to build the data center and so on. An entropic just managed to secure from a deal with Elon Musk a lot of compute but still it's not going to be enough. Okay, they buying even from Amazon even from Google so they keep buying. Everybody are really in need of a lot of compute. This is happening because corporation understood that they need to implement AI correctly within the system because as you and I can produce more and faster and do things that we couldn't do before the same thing is happening with those corporation. So they can do more faster and better and uh and therefore they want it now implementing is is difficult and takes time but we are in an acceleration environment. There is this demand that is growing. how a corporate work it works in the sense they don't just do uh a monthly subscription. Okay, they don't work in this way. They need to be sure that they if they do this investment they have enough token to actually run it. So they they do a kind of a different kind of deal. They want to secure a specific number of token over a year for a series of year years. Okay? Because that is makes them in a scenario where AI is stable and it's not going to disappear. So they're going to buy pre-by all those tokens. Now because of this high demand, the price of token going to go up. Okay. At moment today, we have the luxury of having a $20 subscription that gives a to us the most powerful AI models. The, you know, a lot of tokens to use and basically almost zero friction. Okay, zero privacy, zero sovereignty, but really minimum friction to have something incredible. This is something that we have today. Okay, the price of RAM is going up, the price of electricity is going up, the scarcity of those things is is evident now. Even though it's hard to justify spending so much money for those kind of machine, it feels good and secure to have at least something that you know this is going to be with me for a several years where at least I know that I can build something with if can run here my business my you know whatever I'm building is secure because at the moment it's $20. What's happen when the minimum to use those frontier model is going to be $500 a month, $1,000 a month. We go we are going to go into the a level there. It's it's going to be really expensive. So for us, they're going to offer smaller model at a cheaper price, but for the Frontier model, it's going to be expens more and more expensive. We see this already with meters, okay? They don't want to even share it with us. But when they're going to share it, they already said it's going to be really expensive. So that is the direction okay they care more about corporation they care more about you know selling the token a higher price and so on. So this is what is happening, what is building. I'm not saying that now you need to run and buy, you know, spend thousands and thousands on buying, you know, those super powerful computer. But if you're like me, okay, a sort of digital prepper, you want to have a a little bit of security.

Now I want to give you this little security without spending so much money. So the the easiest thing also as a good way to start and learning understanding gemma 4 came out with four different kind of models and uh there are 2 billion 4 billion 26 billion and 31 billion parameters okay the one that I'm focusing on is in the 26 billion because even though is 26 billion is activate four billion parameter per time so it has the speed of a 4 billion model Okay. While have the intelligence of the 26 billion one the 31 billion doesn't have this functionality. So you need to load the all at once. So the speed is really really slow but it's it's better let's say higher intelligence. It's equivalent of sonet. We are at that kind of level more or less. So the thing here the smaller one instead they can run this the two billion parameter can run on a mobile phone. Okay. So you can run it log on your mobile phone and test it and start learning. The four billions one probably can run on um any kind kind of computer. So the last two you can start learning. So how do you run those last two with minimum friction? Okay. So the easiest way is to use LM Studio. It's an app. You download it. It's free. You download it. when you open it tells you uh which model you want to download. Okay, start from searching for the model. The model you want to download is going to be share what is uh it's going to analyze your system and and tell you look those are the model that are good for you. So you will see probably at the top level you will going to see this gemma 4 and if you can run the four billion go with that one if not do the two billions depend of your system the token per second is going to change. Okay, some is going to be slow, some is going to be faster than enough. So it's depend of your system. uh for sure you can start understanding this and it's not difficult it's easy everybody can do it and uh you will see the beauty of of working with AI locally that you don't need internet you don't need anything and you can do more or less almost yeah more or less the basic thing that you do with chajp okay free and uh and it's I don't know I find it amazing now have in mind that those are much smaller less int intelligence and uh and so on but amazing.

Okay, so thanks for staying to the end. Like and subscribe. There is a community that I'd like you to join which is my our community is the agumentatism community where people like uh me yourself meeting discuss AI discuss local model we building together we are creating resonant OS we are building the resonant DAO which is a decentralized autonomous organization where we can own it and work together to build the future that those corporation are destroying. So be part of the movement, be the the rebel that is building the they are going to build the alternatives. And uh watch the next video. Ciao.