Transcription
Clot code is the best AI coding tool right now, but it only works with anthropic models. What if I told you you can run it with GPT, Gemini, Llama, Deepseek, and even completely free models? Two terminals. Any model. I built this and it's [music] called any model.
I've just created any model proxy for OpenAI GPT 5.4 for model so that all the requests I redirected to open router model and when we actually run any model any request is redirected in fact to open router. Any model is an npm proxy. It sits between the client and your model provider like open router or lama or any other open AAI compatible one. It strips incompatible fields, handles retries, translates formats and has no dependencies. Fully open source and free to use. You set the model on the proxy and the client just connects. That's the diagram that visualizes the process.
Let me show you this in action with a deepseek preset first which is basically redirecting us to deepsek R1 05 to8 model in open router I want to have a clear command here clean window and what we're going to do is to run proxy with a deepseek this is a keyword like preset that we are Usually you have to update to the latest version when there are some changes. We have it running on the port 1990. So on the right, let's go and run our npx any model.
Why open router selected as a default option for proxy? because it has more than 200 actually 600 models already. It contains open- source proprietary open weight models. One of the latest one which I recommend looking into is Gemma 4 from Google Deepline which was released just today but already got some real usage. So if we click on Google Gemma 4 model, we could see the exact name of this model. And in order to use it, we have to create an API key. So I'm going to settings, API keys. and you need to create one. I already have one. So, let me show you how I how to apply it. I'm going to stop my proxy for a while. And uh just if you execute export and then enter open router API key and then command V or control V depending your operational system. entering this key. It's going to save it as envoir. I already have one. So, this proves it. So, I'm going to restart uh the proxy again. And then I'm going to give DeepS something to think about. I'm using Whisper Flow to dictate it faster than I type. What's 1 to7 multiplied by 389? Think step by step. We can see it start thinking and that's the results produced. Looks amazing.
If you see it helpful for you, please like and subscribe. It will help me to build a lot more powerful and detailed tutorials for you going forward. I actually prefer VS Code with this terminal more. And here's we've got the empty folder so far, but for the sake of the demo, let's stick to this Mac terminals going forward here. Let's build something great leveraging codeex preset. So, I'm going to change the proxy to codeex, install the latest version, and let's let's just stop it and run it again. Vix any model. This time I suggest to build a small project but instead of just wipe coding. Let's use specwave which is the open source project that I've created in order to avoid this mess that you could get easily when running huge projects with many product increments which are not tracked with specifications. So basically specwave is about bringing the structure and I have a separate video for it. You could watch a small demo available on the page itself. There is a proper documentation which you could dive into. But um I also wanted to touch base briefly on the verified skill project which is the registry of more than 100,0ands of secure verified skills. So basically the main idea is to scan them with known patterns and avoid vulnerabilities when using or installing those skills. So there are millions of skills scanned already and you could read more about it on the documentation page.
So let's just take a brief recap what just happened. We could see that there is a back end, a node HTTP server. There is a front end, animations, expressions. And if we look into the network tab, after I press the button compute, indeed it comes to the server. That's fantastic. Codeex model is relatively cheap via open router, but we still have to pay. of course way less than uh with the uh pro plans of chat GPT openai or cloud code.
Another thing that I wanted to show is that we could run cloth as well connecting our pro or max subscription. So I'm going to call any model clot and I just stopped proxy because in fact uh we could just connect to our local Opus 4.6 six model with my max plan. You could see configuration changes usage. That was a tough week with some extra usage as well. Although I have actually two [music] max plants and let's give it a shot. too easy for cloth code. But let's see if we could use something completely free. I would suggest to look into Neatron, which is quite a popular solution and works pretty well from Nvidia. All right, just going to use the preset. And that's the actual model that we run. Connecting to the proxy. What is the capital of France? Here we go. Amazing. You could run as many terminals as you want. What is the capital of United States? And moreover, you could run as many proxies as you want. Let's just use another window and run. So this time we could run something on port 9092. And here we go. I'm just great limited because I I just did too many things with quen 3 coder which is by the way great fully open source and free model. Let me use latest gem of 4 this time. Perfect. Worked like a magic.
And as you might know, agent swarm is one of the most powerful features in agentic programming. You could basically hire [snorts] any number of experts to work on specific problem. Let's just leverage specwave in order to see it in action. As we've built a beautiful web calculator, let's explore some u potential like with brainstorming on the future ideas. Maybe building a scientific calculator further on and do some kind of a code review on what we've got here. Optionally, for better visualization, you could install T-Max. I already have it, so I just going to run it. This is the number of window that's created in T-Max. And then let's do npx any model. [clears throat] And then let's just leverage the skill available from a spec weave called team lead. Analyze thoroughly from the security perspective, performance, implementation, our solution of web calculator and also brainstorm the ideas for potentially new features and provide a decision matrix uh from yeah different points of critiques, advocate and pragmatists. Let's give it a go. And that's what we see here. We already got brainstorm advocate pragmatist and critic created though could figure out it on his own. Whom are those experts to best leverage and which skills to use? That's the beauty we see here. Agents are spawned. The default mode to run any model is don't ask which means similar to uh dangerous escape permissions. Once the work of agents is completed, the orchestrator, the team leader consolidates the whole information and provides the report. Let's take a look at this. There is a recommended plan to harden in the sprint [music] first with security headers rate limits tests [music] first ROI [snorts] feature pack then controlled expansion. That's so beautiful. So those are the full critic advocate pragmatism [music] inputs and there is a consolidated decision view what to do as a P 0, P1, P2 and the whole feature decision matrix. I love it.
But what if you want to go fully offline? No internet, no API keys. Let's explore Lama offline. You'll need to install O Lama locally. I already have it and this is the list of model that I have locally available. In order to grab a new model, you could do something like Llama Pool Gemma 3N and then it starts to grab it. I just cancel it. Let me run our proxy with any model proxy or llama and pass the model. This time it's going to be this llama 3.1 8 billion parameters. Okay, as usual, let's switch to this and run it. something simple. It takes some time to initialize the model. So, we have to wait. During execution, I faced that I need to release a fix with this 35 version. Restarted it and basically got this weird message uh that I never seen before. But uh anyway, that might be some specific of this quite a small llama 3.1 model which is quite outdated. So what I'm going to do is doing the addition. Let's see what we've got. Terrific. That's it. We finally got it running fully on local.
Any model is free, open source, MIT licensed. You just need to get open router key or run it fully locally. That's it. And any model defaf is here full of docs. Subscribe if you want more tools like this. I'm shipping weekly. Happy building.