Transcription
When you combine Hermes Agent with tools like Deepseek, you unlock capabilities that 99% of people don't even realize exist. And I'm going to show you exactly how to connect Hermes with the world's most powerful models, which means that you can use Hermes from $0. Build as much as you want to with no limitations or rate limits and my new system for using Hermes most powerful feature that will make you 10 times more productive even if you're a complete beginner.
And if you're new, I'm Jack. I bought and sold my last tech startup with gazillions of customers. Now I'm building my own AI businesses and I just shared the stuff that actually works. So if you haven't already, grab that beautiful coffee and let's dive straight in. Beautiful.
So let's talk about Hermes and exactly how we leverage this with Deepseek as I'm sure you know if this Hermes thing is new to you and you know how to set it up. I'll put a link on screen somewhere so you can get it all rocking and rolling. And now officially we can actually jump in and have a look at the Hermes plus Deepseek. And I'm not just talking about Deepseek. Although Deepseek is going to be doing some heavy lifting for us, but you've got to see why and why I haven't seen many people talk about this strategy I'm going to cover here.
So first thing we have to understand though before we even take any step is that software now costs less than minimum wage. So essentially the balance has shifted such that we can employ these AIs that are as good if not better than junior developers to do things for us over time. So the question simply becomes for us, how many minutes of the day do we have basically AI systems working for us building, improving things whilst we're going about our day-to-day things using Hermes.
Now we talk a lot about building no-code software with Claude code. We talk about Hermes. Remember, they're two very different things. Claude code lives inside repos. It's got a tight tool loop, session-bound, and it's built for codebases. Hermes lives across our entire life. It's persistent. So it learns from every task that we give. It is self-evolving in that sense. The more stuff we say, the better it gets. It schedules background jobs and it builds essentially the idea of Hermes is a deep model understanding of who you are and the better it knows you, the better it can help you with your life. And the whole point of systems like this connect code to Hermes quite nicely.
So let's talk about the reflection of the pyro. Okay, so here's a learning loop. Every task teaches Hermes essentially who you are and that's how it actually gets better, understands more things like this. Hermes itself works whilst you sleep. So we're going to give it something in this video that shows you how you can leverage Deepseek overnight to do some pretty exceptional stuff and why that's valuable.
So the idea with this is that effectively we pick our brain, we pick a frontier model, the most intelligent and effective model to be what we'd call the conductor, the organizer. I have personally found in my experience Claude Opus 4.7 is the best model for that. And the idea here then is that what Hermes is can hot-swap between every single frontier model and use the right model for the right job. Specifically, I'm going to show why Deepseek V4 is so powerful in this setup. I've seen a lot of interesting ideas about running it on your computer and that is really fun. Sometimes you'll find that your literal like MacBook will like melt through the table because like the amount of effort it requires. So, there's some different trade-offs to be aware of here. And we can even see here just how powerful Deepseek V4 is.
Now, the message from this graph, because we don't want to date a benchmark queen, like I always say, we've seen a Tinder profile. We got to take her on a date first, see what's up. It's not trying to say that Deepseek V4 is better than Opus 4.7. Of course, it isn't. But the question is, would you pay 1% of the price for 95% of the value? And if we can truly get like 100x output and we don't need that full max-level redlining brain power, what could we realistically accomplish? That's kind of the question here. And you can see just how comparable they are to each other and how powerful and how well Deepseek V4 has actually physically done here. And this is the benchmarks just for your reference. And so the idea is it's 100 times cheaper and we get the same job done overnight, which is freaking fantastic. And you can see here $75 per million tokens out versus immediately 87 cents. I don't know what you can buy with 87 cents. You can't even buy a hamburger these days with that amount of money.
And so the idea here is that we're going to tag in OpenRouter. OpenRouter is overpowered because once we give it access to OpenRouter, we can see our usage. We can track our usage in a beautiful dashboard. And on top of that, we can access all of the models and switch it dynamically whenever we want to without the need to have and control a thousand different keys and track usage over here and track usage over there. Another it's very powerful. It's one key. It unlocks everything.
And this brings us nicely onto the idea of the multi-brain model system being that essentially every model has its own strengths and weaknesses. And we can build systems that bring in the best model for the particular job. Meaning it runs 24/7 overnight.
So first of all, two models that you absolutely should have in your Hermes agent. Number one is going to be OpenAI's ChatGPT because for your $20 subscription, you effectively get to use your ChatGPT, which is going to be insane amounts of value and gives you access to ChatGPT 5.5. The second is the Gemini CLI and with just simply an email address and a Google account, we can use Gemini.
So this is an example of leveraging a model for a specific thing. Then I'm going to show you how Deepseek comes to this and makes it incredible. So check this out. I can say, for example, "Hey there, I would like you to use the Gemini CLI to go ahead and look at Jack Roberts's last video and give me a breakdown analysis visually of what you see in the first 10 seconds and send that one off." And CLI just stands for command line interface. It's just a very quick and the easiest way to connect to any service. If you don't have the Gemini CLI installed, all I want to do is come over to this GitHub repo right here. Click on code, click on copy of that code, and then head over to your language model of choice. Could be the code app, CodeX, or Anti-gravity. And so, for example, if I'm here, I'm just going to give this command, which is, "Hey, I'd like you to install the Gemini CLI onto my computer." Okay? And all it's going to do is go ahead and grab that GitHub repo. And as you can see, I've already gotten mine installed. And this CLI works exactly the same way with GitHub, with OpenAI, with Vercel. It's incredible.
Now, let's come back over and see how he's gotten on. So, check this out. It's user CLI. Fantastic. It's got my YouTube channel and it's literally broken down visually, all the stuff it can do because Gemini is so powerful at breaking down video. It is the multimodal model and we can now basically tag in and this one cool thing about Hermes, it can bring in any model it wanted to and this is running on my local computer, right? And so essentially I have the Gemini CLI on there. So, if I say, "Hey, use it," it can literally use it for us right there directly. And if it was on hosted somewhere, I would need to basically install it. But because it's here, it's fantastic. And look, it's breaking down the way that I move in my first 10 seconds. But this is just the beginning. It actually gets way crazier than this.
Now, I've been playing with loads of different models. Now, one of the techniques I found is called the council or the triad. And the idea is we have a super intelligent model and we bring Deepseek V4 to do an insane amount of heavy lifting. Then we have a super intelligent model that reviews and delegates. But we need to build it in the proper system. Now, there's a couple of things that you need to know about OpenRouter to get the most out of this model. Some might actually surprise you. These here are expressions that we can add to the end of any model that we're doing and effectively does some really interesting and useful things that are going to help us with Hermes.
First of all is Nitra. So, this can append to any model and it auto-routes to the fastest provider at that moment. For example, Anthropic/Claude Opus Nitra. Fantastic. You've got Exacto, which is a little Italian, but it's very fantastic or a sat-only providers rigorously certified for tool-calling accuracy. Freaking really awesome, right? Because if we have these like systems and models doing agentic things for us, in other words, they need to tool call, they need to check databases, you know, not every model is great at tool calling. So we only want to grab the ones that are great at doing that thing. Another cool thing we have on the smart routing side is OpenRouter Auto. So this picks the best model for your prompt. They're not diamond, no extra fee. That's pretty cool, right? That's pretty fantastic. It can do that. Then you got basically bringing your own keys, which is fantastic. So, for example, if we're using something that's getting rate-limited, like a Deepseek V4, we can actually bring our Deepseek key into OpenRouter just to save us that time and make the whole process easier. We've got fallbacks. And the last one is zero completion. So, you're never charged for blank or error responses across their customer base. That saves almost $20,000 a week, which is uh pretty handy, we might say.
So let's talk about the triad and how Deepseek actually physically fits into this. The idea is three different models, one verdict, no single brain, no brain isolation. Okay, so if you think of it like this, this is the general um strategy. This is very well reflected in research. Uh I think the triad sounds a lot cooler. I think it needs a little bit of a rebrand. Triad sounds cool to me. The idea here is we have plan, we have execute and critique. And I've often found genuinely speaking now, I will never ship anything. I will never ship anything unless it is severely and brutally critiqued. I actually use the word brutally critiqued because I wanted to be as critical as possible. Trying to criticize is a skill set in of itself. The idea here is that we have Claude Opus 4.7, which at the moment is the king ruthlessly okay planning. Now when it's a plan, the Deepseek, the giant whale, okay, that is like 1/100th of the cost for 95% of the actual performance is doing all the heavy work. Some say it's Deepseek labor. I don't know. Call it what you want to, but he's working hard for us day and night. And this is the Deepseek that can churn in the background for 24 hours while we're sipping our lattes and enjoying our beautiful days and spending time with our families. And then we have a critic model which is just going to essentially pick that apart and make sure that it's correct. And again, then we have a planning model and it works in a beautiful circle like that. So, for example, we have Claude Opus that can decompose a task, write the brief, and execute the workflow. Deepseek V4 is going to grind through the plan overnight. It's cheap enough to retry often if it doesn't get it right. And then we can bring in a different model for the critic. It doesn't have to be Gemini 3. It could be any model that you want to. I typically like to not have it being Opus. And I like to just get a slightly different flavor, a different scoop of ice cream, if you will, from a different very capable model. Most likely ChatGPT 5.5, but you can tag in Gemini if you want to.
And to do this, I'm going to be using what I call the Pantheon. So last video I showed you how you build up this entire beautiful Hermes operating system that effectively is a beautiful dashboard that basically allows you to connect Hermes to your Claude code operating system because we do coding on our computer, right? This gives you an overview of your spend, your costs. It dreams for you overnight. So based on your entire chat history with Claude, your usage, how you're using ChatGPT and Claude, and every model on your computer, this will give you dynamic feedback and suggestions. It is auto-dreaming. You can mark this off as done, go through these things, and it's incredible. What's really powerful here is we can connect this to Hermes, which basically means that Hermes has access to all the data and everything you're doing with coding. It shows you all your skills and all of your fantastic memory systems.
Now, in Hermes itself, one of the really interesting things that we can do with the Hermes agent here is actually connect it to our memory systems and we can connect it to everything we're doing. I'll put a link on screen for that full guide breakdown if you want to check that one out. You might find that one super helpful. So, this will be a link in the description. I referenced that video earlier so you get it. Obviously, we can chat to Hermes in the chat if we want to. It's just helpful to have it all in one dashboard. But what I want to look at here is a Pantheon. Now, you can do this just in Hermes chat. I like to do this because I like to have a visual look at everything that I'm particularly designing. I just find this way easier to do.
So, I'm going to add in a persona. Let's call this one something like Orpheus. That sounds fantastic. And we're going to give it a job. Right. Deeply reason on any topic. Okay. So, this is going to be a very powerful one. And I pulled together for you a template for this triad system that effectively breaks down the flow so that actually Hermes understands how it works. And it breaks down into three separate prompts. We have Opus, the conductor, who's the conductor of the Hermes triad. We have Deepseek, the worker. You're the worker. Here's a loop. You read the brief. You identify three to five angles listed by the conductor. And then we have the GPT 5.5, who is the critic that looks down and it kind of assesses it dynamically. And all you can do is literally come down here, grab the flow like so, copy all the stuff. Obviously, if you're using the Hermes dashboard, awesome. If not, you can just paste this into Hermes. And so, I've just pasted mine in here. And here, I'm going to add a little bit of description, which basically explains a few lines on what this persona is for and when to summon them. This is for when I want to go very deep on a topic. We're going to leverage Opus as a conductor. Deepseek is an extremely deep workhorse that can work for hours on a topic. And then we're going to have ChatGPT to review everything. This is for extremely powerful deep work. Awesome. And then we're happy with that. Again, you can just give this information that you want to to Hermes. I personally find the system really helpful. I like visually seeing everything. I think it's fantastic. Then pick the model that you want orchestrating it. So for us, it's Opus 4.7. And then we just click on create Orpheus. And then when that's done, your dashboard will reflect this. Of course, I'll just give it to your guy in Hermes. And I can click on Orpheus. And let's have a look at what they're doing. And we can see we've got the flow. So if at any point I just want to come in and amend it, I can do that, which is really cool. And just makes things a lot easier, I think, which is fantastic. But we're good.
So now we just want to sync these. So I'm just going to come down here. Here I'm going to cap this syncing prompt and show over to Hermes. And I'm just going to come up to the top of the screen and drop this bad boy in here so we can have a little bit of a conversation. Now, obviously, it's great. Again, I just like to visually see this. And once we've done that, what we need to do then is connect OpenRouter. And to do that, we've got to use the terminal. So, for example, we can do that on our computer or we can do Anti-gravity. If I do Control Spacebar and I just type in T for terminal, we have this guy pop up. Now, how we actually install this, many different ways. Easiest way to do this, if you haven't already, I'm going to grab the Hermes setup button real quick. Actually, I want the model set button. You're going to do Command Space, type in terminal, and it will appear, and we just enter in Hermes space setup space model. And then from here, you can basically select all the ones that you've got. And what you're going to find on here is OpenRouter. So you're going to press Spacebar to select it like so. And then I've already got mine in, so I'm just going to keep that one. But here you'll be prompted for an API key, and we'll just grab that from OpenRouter, which is this website here. And again, it shows you all the different performance, which is really interesting actually, see where the central gravity goes. But then you come over to OpenRouter here, click on Get an API key. Then you simply come over to New Key and there you can create it. You also have this B on the left. If I click on Bring, basically bring your own keys. I can search all the providers. So I can come down here and I'm going to type in Deepseek. I can configure this real quick. Beautiful. And I can add in any API keys I want to. So we're going to add a priority API key. And I recommend this for Deepseek just so you never hit any rate limits. Then you just head over to Deepseek API key into Google or Brave, and you're looking for the Deepseek platform. Again, you just sign in, drop $5, $10, as much as you want to into Deepseek. You are going to be using it a lot. So, I'd throw some good money. You never go crazy obviously, but put some money in there. And then you literally copy and paste that key here. And then you have essentially your own keys for Deepseek if you want to, so you never get rate-limited just in case.
Then essentially at this point, we can do anything you want to. Again, we could chat in Telegram, but I could come up to Hermes and be like, "Hey, the dude. I'd like to use Orpheus to solve a problem for me. Could you just confirm you know what Orpheus is and let me know if there's anything that you need from me?" Send that one off like that. Obviously, what I'm using speech-to-text is Glido. It enables me to yap and get a lot more yapping done in the time, which is fantastic. And just like that, we have a response. So I know Orpheus. Orpheus is your deep work triad persona. Opus sets, acts as a coordinator, first interrogating the problem. Five to ten clarifying questions. Then it writes one tight, one-page brief for a worker/critic loop. The intended flow goal is you give it a goal. Orpheus asks clarifying questions. They write a brief. The worker attacks it from multiple angles. The critic tears it apart until it's shippable. And then Orpheus validates the final artifact. And what I need from you is the problem you want solved, your desired output format, how deep do you want me to go, deep dive or just work for hours mode, and any constraints, budget, tools, deadline, audience, tone, things to avoid, and what success looks like.
Little hack for you as well, guys. You'll know for example that Hermes itself saves information in its agent memory and its soul. What I would love you to do for a public scale for you as well, and I'll put a link for it down below so you can grab it. And so this is the soul. Effectively come down and just replace this. This is exactly how to run it and how to use it. It's stuff on your identity, your mission, and goals. What's your goal this year? I want to hit $20,000 a month in my SaaS. I want to double my investment portfolio. I want to acquire seven companies. Whatever it is, whether you maybe you're a billionaire, you want to be a trillionaire. Whatever the thing is, make sure Hermes knows what it is. Explain your business, your revenue, your runway, bank payments. Um, any key information you wanted to know, key metrics that Hermes needs to check and be aware of. Uh, voice and communication details, how exactly you would like to have it speak to you. Short by default, one question at a time, how it needs to write to you, the rhythm, all this detail. Take a look at this. If there's anything else you want, you can add it in there. But honestly, these are just a really great questions. You can just feed it this information about yourself. Then Hermes is going to have that fantastic context. And then that will appear in your soul.md. And you can even say, "Hey there, I want you to add this to my soul.md. Go hands-free mode." And then just yap to your heart's content. And literally that will cover everything.
"Hey there, I'd like to use my Orpheus skill. Could you just confirm to me anything that you need to know about that persona before I give you a task, please?"
This is cool. Orpheus persona. What I need to ask for a task. Opus is a conductor. Deepseek is the worker. GPT 5.5 is the critic. We need a desired outcome, success criteria. Let's just tell it. "Hey there, my desired outcome is to know I would love to sell websites and AI services to businesses. My outcome is to know which niche should I personally choose. Success criteria is a list of the top three niches that have effectively high margin. Uh maybe typically unsexy and untargeted but have a high need for AI and automation services. Um could you do a quick one for me? Maybe like 10 minutes. I didn't answer pretty sharply on this. Um, the audience is for me. Uh constraint is just keep it nice and snappy and short. Uh not too short obviously. Uh like a good good amount of detail. I want emojis per thing. Um and then known assumptions. What do I already believe? Uh I well I have some beliefs that businesses like roofers and pool cleaning companies would be an excellent stop based on my experience and output format. Yeah, just give it me in emojis and breakdowns. I know you're set to interrogate me first, but yeah, you can ask me a couple cloud crane questions if you like to."
So you can already see this process and how valuable this is. I send this off now. Imagine actually setting any important decision through this. Like if you just ask Claude directly, I've actually found it myself. It just agrees with you. It just sometimes just agrees with you for no reason. I'm like, "Dude, you're just agreeing with everything I'm saying. This isn't good. We need to you have to back build in interrogators." And the beauty of this triad strategy here, guys, is the fact that we've got different models. We have critique. So, think of the number of loops. The whole idea of progress, right? WD40 is like effectively solves like 90% of household elements, right? The only reason it got that is they had a very quick improvement loop. Like WD40 was the 40th version that actually worked, hence the name. And think about how many incremental improvements that you get when you have a critic and review agent running around like this. And then we have obviously Opus 4.7 that's setting the strategy. It's just going to make you unbelievably effective. Like it's insane.
So when I scroll down, I'm effectively hearing what it's understanding. It's picking the top three niches, which is great. Success criteria. Click quick clarifying questions. Cool. So geography is going to be worldwide. Actually, let's do a let's do Texas, please. Office shape. Um, I'll let you tell me what you think is going to be most effective for that price point. I'm going to guess again that could be, you know, one to 15K. That's absolutely fine. Uh, sales motion, I'll let you lead on what you think the best one is going to be and to basically guide my decisions, guide my thinking on it. I think that'd be fantastic. And then, yeah, go ahead and let me know what the output is. Now, I've just used this as a random example, but you get the idea of how the system could actually work. One little tip that might save you as well is add in fallbacks. So, for example, if tokens ever run out, default to XYZ instead. So, it can physically do that for you.
Beautiful. So, now I come down. Look at this response here. I got fire, water, mold restoration. It's saying why it wins. Emergency leads can be insanely timesensitive. One job can be $3,000 to $50,000. Thoughts. Guys, if you're not working with these companies after this video, we need to go ahead and do that right now. Emergency lead capture is really cool. It's giving you pricing ideology. It's giving you um, you know, foundation repair, drainage, waterproofing, and all the reasons why. And what I could do is basically come down and say, "This is awesome. Could you just explain to me your thinking behind this, what each model did and how you arrived at this conclusion, please, so I can best understand that." And so let's see what it did here. And I just do this to show it's working. I feel like a math tutor now. And Hermes is my little student telling me everything it's done. So we've got Orpheus. And but Opus was the conductor, set the frame. DC worked. GT was a critic. And then this gave the final synthesis. So the point is to avoid one model just by being the answer. So what is a good niche? Instead of asking that, you ask which Texas local service niche is most likely to buy a high margin website blah blah blah. So it judged this on this, which is cool, but Deepseek worker scored basically scored the market brutally. Went through everything and just kind of explains its thought process, which is cool. But you can set it up to do these leaps overnight whilst you're sleeping. You can even use free models if you want to, but I would generally just advise against that just because free is not free as we say in the USA. The point here is that like for a fraction of the cost that you get with Deepseek, your quality differential is insane. At the end of the day, if it runs for a million years and the answer is still garbage, what use is that to you? So, I like to go for that ratio. Well, I personally use Opus 4.7 a lot. But for massive stuff, using Deepseek is insane because you get like I'd say 95% the value for like 1% of the cost. And so then we're building out this Pantheon of special skills with this particular triad skill. Now, the idea here is that Hermes plus Deepseek with the whole infrastructure we've got is an agent that grows with you. But it does lead us on to one final question, and that's how to get Hermes to its maximum potential, which we're going to learn in this video, right?