📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

How I Use OpenClaw for 95% Cheaper (Feels Illegal)

The AI Growth Lab with Tom19:02

Transcription

People are spending hundreds, even thousands of dollars every week on OpenClaw because they're using it completely wrong. They buy a Mac Mini for 600 bucks, they install OpenClaw, they run it for a week, and then they get their tokens drained, and they shut the whole thing down. And that's pretty insane to me because Nvidia's CEO Jensen Huang just called OpenClaw the new computer and the most important release of software probably ever.

You see, the problem is not OpenClaw. People are just using it wrong. And in this video, I'm going to show you the exact setup I use to run it for dirt cheap. Specific models, the settings, and techniques that can cut your build by up to 95%.

You see, OpenClaw has two costs. The hosting, which is the machine that it runs on, and the API cost, which is essentially the fuel. So, every time your agent thinks, replies, or checks your email, it's burning tokens. And most people are putting premium race fuel in a car that they're driving to the grocery store. So if you picked Frontier models like Claude Opus 4.6 as your default, you're going to be paying what, $15 bucks per million input tokens for a heartbeat check that a simple 10-cent model could handle.

The other killer is context accumulation. Every message you send is going to include all of the previous messages in that chat for that agent. So by message 50, you're paying for the weight of every conversation stacked on top of each other. And this is the single biggest hidden cost that most people don't even know is happening.

So let me put some real numbers on this. Reddit is full of people sharing their OpenClaw bills. Unoptimized setups running Opus as the default model. I mean, you're going to be looking at what, $300 to $600 a month. And one person on Reddit even reported $3,600 in a single month spent on API credits. But people running the setup that I'm about to show you probably get that down to $6 to $25 a month. The same tool, the same features, just configured properly. So let me show you how to fix it. Starting with the biggest lever.

Okay, so first thing we're going to do is install Open Router. There is no point in using Opus 4.6 for absolutely every single action inside of OpenClaw. So I'm going to open up my OpenClaw instance. I do use Hostinger. This is a a VPS with a Docker deployment. Uh, it's a super easy one-click deployment and I can just simply open this, uh, enter in my gateway token. Boom. I'm going to be logged in.

Now, I'm going to assume that you've already got OpenClaw set up and running on your computer. Now, if you don't, I've put a full video on my YouTube channel. It's going to be linked in the description of not just how you set up OpenClaw, but how you set it up securely. Now, my choice is using a VPS. And you can install it on your computer or VPS. Either way, you do need to go through the security process, uh, to kind of harden the access because you know that's one of the risks of using OpenClaw.

Okay, so I'm going to assume that you've already got this set up and you're already connected with like a basic install with Telegram connected. Okay, so we're going to head over to Telegram and we're going to start chatting to OpenClaw to get it to install everything that we need it to do to reduce the token usage.

If we take a look at Open Router, we can find out some of these models and how much they're typically going to cost, right? So, one of the most recent models that is super cost-effective is the Minimax M2.7. So, we're probably going to hook that up. Deepseek V3.2 is a good one. Which other one did I want to take? There's a, there's a Kimmy Kimmy model. I think it's 2.5. Yeah, 2.5. So, we can add any model that we want here from Open Router. And if you just take a look at Anthropic Opus 4.6, you can see it's $5 per million input tokens and $25 per million output tokens, which is considerably more expensive than, let's say, Minimax, which is $0.30 per million input tokens and $1.20 per million output tokens. So that's a 10x difference.

So we're in Telegram and I'm going to ask my bot, this is just a fresh install for me. I'm going to ask my bot to set up Open Router. "Hey, can you help me set up Open Router? I'm going to drop my API key below. Um, I want to be able to configure different models so that we can reduce the cost of using OpenClaw."

Okay, so it's working on configuring Open Router. Now I know that putting your API key in this chat is not the most optimal solution, but because of the security protocol that I've got set up, I, I'm not worrying about it. And in any case, I can go ahead and just delete that message after it's set this whole thing up. So, what I'm going to do is I'm actually going to get the model names of the models that I want it to specifically add to my Open Router, uh, sorry, my OpenClaw instance. So, I'm just going to copy the Minimax. We're going to go for Deepseek and then we're going to go for, uh, let's do Kimmy. So, it's just setting all of this up for me and I'll be back when we're up and running.

Okay, so that's completed. Just took maybe a minute or so. It's added these default models. So let's just check if we can actually take a look at the models that it's we've got here. I think it's maybe models. Okay, so it's brought back the providers. So if you click Open Router, you can see we've now got access to these three models. So the first way you're going to be able to reduce your cost is by manually switching to one of these three models. So I'm going to select the Minimax M2.7. "This model will be used for next message." Let's just check which model are you using right now. Okay, so it's saying that it needs me to adjust the think, which is bizarre. So let's stick this at medium and then ask the question again. Okay, you can see it's using Minimax M2.7. Excellent.

So there is another level that we can go here. Open Router has an auto feature. So we can route prompts to Open Router and it'll automatically determine which model to use based on the complexity of the prompt and the type of tasks that you're asking OpenClaw to do inside of the prompt. So whether it's a coding task, it's going to select a better model for that. Or if it's just a really simple request, we're going to go for something that's super cheap. So let me show you how to set that up.

Let's ask OpenClaw to set this up for us. "Hey, can you set up the Open Router auto mode? So that when a prompt comes into OpenClaw, it then there is a mode in Open Router called auto that will allow Open Router to select the model based on the complexity of the prompt. Can you help set this up?"

So you can see here in the documentation in Open Router, which, uh, I can put a link to below as well. It's just in the docs, you can Google this or whatever. You can see "Using Auto Mode for Cost Optimization." So, "OpenClaw agents perform many different types of actions from simple heartbeat processing to complex reasoning tasks." And the whole point of this video is the fact that you don't need to use a powerful model for every action because it's just going to waste money on tasks that don't require these advanced capabilities. So, the Open Router auto model, which automatically selects the most cost-effective model based on your prompt, is ideal for OpenClaw, right? So, it's already in the docs. That's kind of that's how I found it. Um, so let's see if OpenClaw is going to detect that. Cool.

So you can see here, "Got everything I need. Adding Open Router auto now." So this is basically another model essentially. You can see how it says "openrouter/minimax." We've got "openrouter/auto." And now it's telling us exactly what it does. So now Open Router auto is the new default. So whenever we put in a prompt, it's going to automatically reroute that to the most cost-effective model. Sounds pretty cool, right?

So there are some other things that we can do and to reduce the cost even further. But before we get into that next step, I want to tell you about today's sponsor of this video, which is Hostinger, who I've worked with for quite a while now, and I use their VPS for anything that I want to run in the cloud, including OpenClaw. And they've just released a managed version of OpenClaw which takes away all of the headache of setting up OpenClaw on a VPS and having to go through the whole security procedure which I've detailed in one of my videos. So instead you get an instant setup. It's like a one-click install, zero maintenance. So you've got all your security handled, your updates, your backups, and your agents are going to run 24/7. Right? If you had this installed on your computer, your agent is only going to run when your computer is awake and open and online, right? And there's a ton of built-in tools, different AI models, web browsing, agentic mailbox already built in with your OpenClaw instance, right? So, I think this is one of the easiest ways to get into OpenClaw. And yeah, you're not going to have to spend $600, $800 on a brand new Mac Mini. This is going to cost you literally $6 or $9 a month and I'm going to give you an extra 10% off. There's going to be a discount code and a link below in the description if you want to sign up.

Now, these are obviously two different options. One's a managed OpenClaw instance. Another one is OpenClaw on a VPS, right? So, if you were to choose the managed plan, you know, for 2 years, you're looking at $150 bucks, right? And this includes some additional extras here, which you don't have to opt for, but you can if you want. And obviously you can go for a smaller plan as well. It's totally up to you, but it's going to be way more cost-effective than buying a Mac Mini and doing that whole setup over there. And the setup is just super easy. One-click install. You choose where you like to chat. So whether it's Telegram or WhatsApp, and then it's basically ready to use. And with everything that I'm sharing you with you in this video, you're just going to be able to use it for literally pennies every single day. So, in my opinion, it's the fastest and easiest way to get started with OpenClaw. If you're just curious about it and you're not really willing to drop like five, six, seven, $800 on a Mac Mini and just want to try this out and see what it's like, see what all the hype is about, then this is going to be the quickest way to get rolling. So, anyway, with that being said, let's get back into the rest of these steps.

Okay. So, what I want to do next, if you've got any heartbeats running, we need to optimize the model that we're using for the heartbeat and a few other things. So, first of all, I'm going to ask, "Do we have any heartbeat? Do we have any heartbeats running? And if not, can we just set one up real quick?"

Okay, so I've just set up a heartbeat. Now, we're going to optimize this and make sure that we can reduce the token usage for every heartbeat. Now, it's giving me every 30 minutes, which I think is not really necessary. So that's something that you have to consider on on your own. How often do you want the heartbeat to run? And what's like the minimum viable? Is daily okay or like twice a day? For this context, I'm going to go for twice a day. "Can you set the heartbeat just to run at twice a day? Let's say 9:00 a.m. and 6 p.m. And can you select the Minimax model to run all the heartbeats?" So that's something that we want to mention. We want to set the cheapest model for these heartbeats as well.

Now you can see here it's using light context mode, which was one of the suggestions I was going to make. So if you don't know if it's using the light context mode or not, then you can definitely tell it to incorporate that as well. "Can you make sure light context mode is on?" We also want to set the isolated session to true. And then we want to set the active hours. Now in this case I've told it to run at 9:00 a.m. and 6:00 p.m., which is going to be the time I'm going to be active. But "set active hours so it only runs hobbies while you're actually awake. So I'm awake 9:00 a.m. to 9:00 p.m." Boom.

Okay, so you can see here, "Hobby is running on an interval, not a clock." So it's taken the prompt that I've given it and it's created two cron jobs instead. This could be definitely an option for you if you don't want to run something every 30 minutes or every hour and you just want to run it at a certain time every day, then setting up a cron job is going to be way more efficient. Now, it's still going to load the context from the heartbeat.md, but it's going to run it at specific times every day instead of on an interval. That makes sense. And you can see it's selected the model "openrouter/minimax M2.7." So, that's reducing the cost for all of these check-ins.

One of the other things is your context window, right? And compacting your sessions before continuing. Now, if you're just going to use this one chat window here in Telegram, then it's going to be important for you to at the end of a session say, "Hey, can you save all the relevant key points that we've discussed in this conversation?" So, we want to make OpenClaw save any important stuff that we've had in this conversation so far before we end up compacting the session, right? And if we use the command "compact," then it's going to compact the session and give us a lot more context space and it's going to use up less tokens. You see, if you're using, let's say, 50% 70% of your context window, then you're going to be paying more tokens than if you got a fresh session because it's going to send the entire context through every single time. Does that make sense? So the longer your conversation, you're sending you're basically having the AI like read the entire conversation every time you're sending a message, which is basically going to take up too much tokens.

Cool. So we're going to run the "compact" command here. And now it's going to compact the entire context. Now this isn't hasn't been a long chat. This is fresh and it's also just a fresh install. So there's not a lot really going on in this instance. But if you did spend like half a day, a full day going backwards and forwards, this is something that you would absolutely need to do. And whether you pull stuff into an memory.md file or whether you update your identity or your soul, these are all things that is worth considering. Now, I do have a memory system that we're going to get on to in just a second.

There is something else that you can ask your OpenClaw instance to do. You can come in here and set max tokens so you don't load of runaway responses. So that's what we're going to do here. here. And you can see, look, it's compacted from 55K to 23K. Okay, nice. So, I'm going to just pop in this command here. And well, it's not a command, just a piece of text. "Set max output tokens to 248 in in the config. This prevents runaway long responses that burn through your output token budget." We're going to ask it to do that. This is basically going to rain in the length of the responses that you get back. So, your output tokens.

Okay. The next thing is QMD, which is Quick Markdown Search. This is the GitHub repo here. What I'm going to do is just ask OpenClaw to install this. "Hey, can you install QMD from this GitHub repo?" And let's see if it's going to be able to execute on this.

So, it's gone ahead and installed QMD, which is basically a local search engine for markdown notes. It turns your markdown notes into like mini vectors essentially with reranking and all that good stuff. Though is given us some commands here that are available. We can actually use a QMD search if you need. I think it's post Q and Okay, that's not you probably just enter those. So, you could use these as search commands.

Now, one thing that I did see in the GitHub repo was to add the agent integrations. So, I'm just going to pop that into Telegram. "Hey, can you double check this? You need to add it to the agents.md file." So I should have mentioned that earlier, but it should figure out what it's doing. So it's just going to add this to the agents.md file. So it knows to check the QMD search whenever you're maybe asking for something. So the reason why QMD is useful is once you start building up a large range of files, different markdown files, then it's going to take up more tokens to search through those files. So, this is a search method that's going to take much less tokens if you're wanting to reference uh different markdown files or you're trying to use different agents and and different context. That make sense?

So, you can see here that it's currently working, which is great. We've got the memory search tool. Okay, it looks like that that's not functioning. We can go in and fix that if we need to. Let's go with option two. So, you can see we've added that to the agents.md file. So before answering any anything about prior work and decisions, dates, people, preferences or to-dos, it's going to run the QMD search. Uh, it's going to run a query. Cool. So that's going to reduce a lot of token usage when it comes to searching for files. Cool.

So that is everything. So the keys from this video to install Open Router to select the heartbeat model as a really cheap model. I'm using Open Router also to autoroute my prompts to the the cheapest and most cost-effective models that are available. Now you may find that you're not getting the performance you want from that so that you can select your models manually and you can also select them for each of the agents that you build as well. So you might have one agent that is specifically for coding and you give them the Opus 4.6 model, right? But another model might be on research where you use Kimmy K 2.5. So you can choose whether you want to go the auto route or the manual route. And then we've got the QMD search, which is another big part of this process and you should be able to reduce your cost down by 90-95%.

In any case, one of the best ways to reduce your cost is to use a VPS. I use Hostinger like I mentioned in this video. You can grab their managed OpenClaw instance for just $6 a month. There's going to be a 10% off discount code below in the description as well.

Okay, so let me put all of this into perspective. For an optimized OpenClaw, you're going to be looking $300 to $600 a month in API spend. Some people even more. Now, with model routing, a VPS, optimized heartbeats, regular compacting, and QMD, you're going to get that to under $25 a month. And that's up to a 95% reduction.

So, here's my honest take. If you're spending hundreds of dollars a week on OpenClaw and you're not building anything that either saves you money or makes you more money, then what are you really doing? The people who are actually winning with AI automation are the ones who learn to build these workflows to solve expensive problems. And that's what I teach inside of my AI Operator Mentorship Program. It's 16 weeks and you learn tools like n8n, Claude Code, and you learn to think in systems and build production-ready automations that actually work. So, if you are interested in learning more about it, you can click the first link in the description.

All right, guys. Thanks for watching. I'll see you in the next one.