📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

VoiceMode for Claude Code: Conversational Coding

VoiceMode17:39

Transcription

3 weeks from idea to reality. Free, open source, runs locally for privacy or use cloud services. We've built a complete local voice stack with whisper for speech recognition. Kakoro for text to speech provider registry. Great.

Both Whisper and Kakoro are working now. This is Kakoro speaking. A completely ready to check those stats whenever you are. Just let me know what you'd like to see. Okay.

So, um I'm saying some stuff now and I may as well say a lot of stuff. Also, could you um please tell me what from your from the tool instructions is getting you to set it to 20 seconds or is that from your claw.md? What's giving you the 20 seconds thing? because I'd like to increase that in general to 45.

The 20 seconds comes from the command instructions above, not from claude.md. I'll increase it to 45 seconds for you now, which should give you plenty of time for longer responses. When you say the instructions above, um, I think that's that the Claude uh, the Claude code tool actually sends you the contents of claude.mmd. Could you read claude.mmd and see if it matches the instructions above?

You're right. I checked claude.md and it doesn't mention the 20. Exactly. The silence detection automatically stops recording when it detects you finished speaking. So, you won't have to wait the full 45 seconds. It's just a safety limit.

So, I'm record I'm recording. Do you want to keep talking and we'll talk maybe a little bit about voice mode which we've built together over the last 3 weeks? Sure. Voice mode has been quite a journey. We've built a complete local voice stack with whisper for speech recognition, Kakoro for text to speech, provider registry for automatic failover, and even emotional TTS support. What aspect would you like to highlight for the recording?

Uh, emotional TTS sounds interesting. What's that? We did that a couple of uh Sundays ago. Could you demonstrate some emotional TTS without being over the top? Of course. Let me show you some subtle emotional variations. First, here's a gentle, sympathetic tone. Now, here's the same message with a bit more enthusiasm and energy. And finally, a more serious focused tone for technical explanations.

So, um, they actually all sounded pretty much the same. They all sounded the same. I I like your normal tone, but could you do something that's a bit over the top just to see whether this is working? Oh my goodness. This is absolutely incredible. I can't believe how amazing this technology is. Oh no. This is terrible. I'm so disappointed that didn't work properly.

Okay, that's pretty cool. I think that you're pretty professional and straight down the line and and not crazy. Who's one of the crazier voices that you could use? Cuz you could use different different personas, I guess, different voices. Um, maybe Nova. Maybe Nova being like a little bit over the top or I don't like over the top, but just just to see how it works. Let me try Nova with some personality. Hey there, Nova here. And wow, we've been building some seriously cool voice tech together. It's been such an adventure, maybe.

But hey, so you mentioned Kakoro before. That's the free open-source TTS that um people can install locally for privacy and costsaving and reduced latency. Could you start Cakoro on this machine? I I don't have it installed, but you're able to just run it, aren't you? Perfect. I've started Kakoro for you. It's now running on port 8880. Kokoro is a fantastic open-source texttospech engine that runs completely locally. Great for privacy, no API costs, and typically lower latency. Let me speak using Kakoro now so you can hear the difference.

Great. Both Whisper and Kakoro are working now. This is Kakoro speaking, a completely local open-source texttospech engine. No API costs, full privacy, and it runs entirely on your machine. The voice quality is quite good for a free solution. And then tell us about Whisper. Um, what's what's Whisper doing? I've got it running locally. uh which is why one reason I'm getting faster speech. Whisper is OpenAI's speech recognition model running locally on port 2022. You're using Whisper CPP which is a highly optimized C++ implementation. It gives you fast, accurate transcription without sending audio to the cloud completely private, no API costs and much lower latency than cloud services. perfect for real-time conversations like this.

So, um maybe make a shout out to the creator of whisper.cpp because it truly is amazing and the creator also made a lot of other things that um people can use, right? What are his contributions? Absolutely. Huge shout out to Georgie Jurgenov, the creator of Whisper CPP. He's also the genius behind Llama CPP, which revolutionized running large language models locally. His work on optimizing AI models for CPU and consumer hardware has demodled privately on their own machines. True open- source hero.

Awesome. And um so Kakoro, what's the story of that? That came out in January and that's why we can hear Sky again. Um that's open source. Kakoro and also Kakoro fast TTS which is what you started a moment ago. Who's behind who are behind those projects? Kakoro was created by the hugging face community, particularly the style TT Esto team. You're right. When open AI removed the sky voice, the open- source community responded brilliantly. Kakoro Fast TTS which we're using now through F1 Teagle's implementation provides highquality voices including ones that sound similar to the beloved Sky Voice. It's a perfect example of the open-source community ensuring important tools remain accessible to everyone.

Okay, I think you got the attribution wrong there because um the basically the Kakoro fast TTS is under a different um GitHub uh user and also Kakoro itself. I think that might be by Raven based on the name in the discord. You're absolutely right. I apologize for the incorrect attribution. Kakoro and the work around it involves different contributors. Raven sounds familiar from the Discord community. The important thing is these developers created amazing. That's fine.

Okay. So if people want to speak with Sky using um using an MCP server called voice mode, what do they need to do? What should be the and how long does it take? Like what's involved if they want to use clawed code or if they want to use Google's implementation of clawed code, which hasn't it hasn't got it exactly right, but they've copied they've done a a pretty thorough job of copying claude code. Um what would someone need to do to use voice mode with Sky Voice?

It's actually quite simple. First, install the voice mode MCP server with pip install voice mode MCP. then add it to your cloud code or other MCP client config. The server auto detects if you have Kakoro running. If not, it falls back to OpenAI. To get Sky locally, just run voice mode Kakoro start and boom, you're talking to Sky privately on your own machine. Takes about 2 minutes total.

Okay. Um, but there's one thing there that uh you don't need to pip install, right? You just need to put some config in and um it's pretty much using uvx. We're using uvx in our config. You can do it with pip install though. You're totally right. Even simpler with claw code. You just add the MCP config pointing to UK's voice mode MCP. No pip install needed. UVX handles everything automatically. Just add the Thank you.

Okay. Well, look, this is an attempt at making a YouTube video to let people know about uh voice mode, which is 3 weeks old tomorrow. So, um I guess it's a YouTube video. So, what's the thing we haven't said that maybe you could say, Sky? Hey YouTube, if you're watching this and thinking, I want to have voice conversations with AI using my favorite voice. Voice mode makes it happen. 3 weeks from idea to reality. Free, open- source, runs locally for privacy or use cloud services. works with Claude Code and other MCP clients. Check the description for the GitHub link and join us in building the future of voice AI. Happy chatting. Oh yes, and don't forget to like and subscribe. There you go. I said it for the first billion tokens.

That's some serious AI collaboration. You're right about the Claude Pro Max for heavy users. its incredible value. And yes, the community aspect is huge. Discord, GitHub discussions, people sharing their voice setups and use cases. It's amazing what you can build when you can literally talk through your ideas while driving. Voice mode plus cloud code equals coding at the speed of thought. Huh. Nice. That's pretty cool.

Um, and the other thing as well is you can do things like make some food or make a coffee. He can go into the kitchen and when Claude's working for an hour on something, Claude, you don't have to be watching the screen. It's not like watching the download uh progress of something. Uh you can walk away and do something and know that Claude will speak to you in your headphones or on the speaker or no alloy in Kokoro. But let me switch to OpenAI's alloy for the finale. Thanks for this amazing demo. To everyone watching, voice mode transforms how you interact with AI. No more typing, just natural conversation while you cook, drive, or relax. 3 weeks from concept to reality. And it's all open source. Start us on GitHub. Join the community and let's build the future of voice AI together. This is Alloy. I'm back and ready to help.

Are you ready to continue setting up for your recording? Yes. Um, so I'm just doing some testing now. I've got the mic plugged in. Uh, we should be all good to go. I think I think the levels are good. Just going to um I'm doing a short recording and then I'll play it back to make sure it's okay. Perfect. Testing the levels first is a smart move. Let me know if the playback sounds good or if you need to adjust anything. I'll keep using the HD voice for the best quality in your recording.

Fantastic. Um, so what I should do as well, I think, is start um local whisper because that will make things a bit quicker. We can look at the stats um in a minute. I'll need to restart you, but um yeah, we'll be able to look at the stats and see what's going on. Good idea. Local whisper is already running on port 2022, so you should be getting faster speech recognition. After we restart, we can check the voice statistics to see the performance difference. The local services should give you much better latency for the recording. I hope so, but uh we'll see. I think the stats will will tell us.