Transcription
Hello, my name is Evgeny Kovalenko, and Ken from Alibaba has released a TTS model, that is, text-to-speech. This is a model that can generate speech from text, audio from text, yes, and on their official page it is written that this is a series of powerful speech generation models, which, well, again, was released by N. They offer quite strong support in terms of voice cloning, voice design, ultra-high quality human-like speech, and voice control with natural human language. Below are the models they have and their features. Here are the principles of operation. Below are generation examples. Here we see Chinese and English. Chinese is understandable because it is a Chinese model. English, in principle, is the international language, so to speak, yes, the language on which everything is translated and spoken everywhere. For now. I will also try to write something, for example, in Russian in the model that I deploy. Let's see if it can read it, generate something. Here you can listen to the text, for example. Let's listen to an example. Well, quite, quite cool. Okay, now let's move on to installing Quen 3 on the computer, to downloading and installing. I will use Compi for this. It's very convenient, you can work with it. I will show from A to Z what is required, what needs to be downloaded and installed. And if you already have Kofei or some steps are already done, you can just scroll forward. Before downloading and installing Kofei, we need Git and we need Python. What do we do? We go to our search engine and type Download Git. Download Python. We type it like this: Download Git. The site appears. We go to it, select our operating system. I use Windows. In my case, it's Windows. Git for Windows X64 Setup. We click. It downloads. This is what the icon looks like. It weighs 63 MB. We launch it. We click yes. Next. Next. Next. I don't need a Start Menu Folder. Next. Next. Next. Next. Next. Next. Well, in principle, just click Next until the end. Click Install and wait for the installation to complete. After we install, we can uncheck Felistas Notes so that the site with updates related to GIT does not open. Click Finish. Next, we type Download Python. We download Python. Python has been downloaded. This is what the icon looks like, 28 MB. We launch it. We set use admin privilege. Add to path and install now. Click close. Next, we type in the search engine download conf UI. We go here and from here we scroll down to direct download. Scroll down, down, down, down, down. Here it is, direct download. We download, just click here and that's it, it should start downloading. Comfi has been downloaded. It looks like this. 1.7 GB, but keep in mind that besides this 1.7, you also need to consider the space where it will be unpacked. That is, it will unpack to about seven, so in total 1.7 + 7 is almost nine, right? You can unpack it wherever it's convenient. For convenience, I will just do it here in downloads, yes, I click all extract. And wait for it to unpack. Comfi has been unpacked. What is its size? 3.5. Strange, it used to be 6.8. Anyway, it doesn't matter. They optimized it, but we are not launching it yet. We need to type comfui manager in the search engine. Here it is. Go here. Ah, and we really need it. We will need it very much. We will install it, add it this way. We will add it this way. We will go to Custom Notes. We return to Confi, confui customes. And here at the top, we type cmd. Press Enter. We return here to our confui manager. Yes. We take Git Clone. This. Paste it, press Enter. This way, the manager button will be added to Kofei now, but we are not stopping there either. That is, we stay in the Custom Notes folder, return to Confants. And we need to clone this repository. Click here on code. The name appears. Click copy. Return to the address bar. Type Git Clone. Then paste what I copied. What we copied. Confon TTS. Press Enter. And this way, we should have the iManager Confui Manager button and confui quents. Now we go in. There are three types of launch. Run CPU, run Nvidia GPU, run Nvidia GPU Fast FP16. This, as far as I know, should be launched by people who have very powerful graphics cards. What is a very powerful graphics card? It's not Nvidia RTX 4090, 5090, 6090, I don't know what you want. It should be more powerful than that. People with Nvidia graphics cards launch this, and they want to do generation using the graphics card, that is, use the power of the graphics card. Run CPU is launched by people who do not have Nvidia graphics cards and who want to use the central processor of their computer or laptop for generation. And of course, Nvidia GPU, yes, GPU and CPU give different generation speeds. And in general, the power of your hardware determines the generation speed. That is, the more powerful your hardware, the more powerful your graphics card, the faster the generations will be. We launch Run Nvidia GPU. Comfi has launched. Here is our working environment, yes? This is it. And here we have, we click on the Menu button, browse templates, and scroll all the way down. At the very bottom, confui quen tts has appeared, yes? This is the custom note that we installed. Click on it. We have the option, we have two working environments. One is example, which is, well, like a working template. The second is multi-check dialogue. That is, here you can build, generate a dialogue between two people. Yes, here is such a working environment, this is the standard one. Let's use the standard working environment. Click on it. And this is what the working environment looks like. We have voice clone. Here we clone voices. I have already cloned something. Here we have custom voice, that is, there are already certain voices. We select the language here, we select the person who should use this language, to know which speaker, yes, from these, from the ones suggested in the dropdown list, speaks which of these languages. 1 2 3 4 5 6 7 8 9 10. There are 10 languages here, yes? So, to know which language is spoken by whom, you go to Quen, and here they write, yes, that Siren, Uncle, Vivin speak Chinese, Aiden, Ryan speak English, Ansohi, Korean, Chinese dialect. Uh, it doesn't say who speaks Italian, Russian, but perhaps those who speak English, maybe they can speak Russian, but I don't see it here yet, yes, from the list, from what is indicated on the site. And third. Third. Here we create a voice. How do we create a voice? In principle, you can also scroll up, uh, look at an example of how a voice is created here, yes? That is, we see here gender, that is, sex, male, pitch, how it should sound, speed, volume, age, and so on. Based on all these parameters, you can enter all these parameters. tone, personality, yes, personality, that is, yes, confident and performative, that is, confident and, how to say, presentable, perhaps. Here. So, based on this hint, you can, in principle, create. There is a female example, there is a male example. That is, in this, you can create voice design, which is creating a voice based on what you enter. Custom boys, you can simply use these voices right away, or perhaps something is downloaded separately for Russian, for Italian on their website. I haven't checked. And clone, to clone a voice, we have it in the TTS on the Kofei TTS repository, yes, here at the bottom there is a small hint on how, how to clone a voice, yes, use clean audio without noise, 5-15 seconds. Yes, and it is desirable, of course, to try to cut out a fifteen-second segment, yes. Uh, VRM BF16 is selected in the list. And here, yes, regarding this, I wanted to say, it's good that it's written here. It reminded me that when you launch the Workflow and click Run, uh, possibly, in general, by default, all models should be downloaded to the correct place, downloaded, and you will generate a voice from them. But if this doesn't happen, as it didn't happen for me, I tried for a very long time to set it up, to achieve this, you need to download manually, that is, you need to go to Hugging Face. Here they have PNTTS. Go here and download all these models and put them in that folder in customes. But if you click and it downloads, then you will be very lucky, yes, more than me. So that each of these nodes doesn't interfere, we should, for example, if we are cloning a voice, engaged in voice cloning, we should select these, holding down Ctrl, selecting with the mouse our nodes and right-click BYPs. That is, we disable them so that they don't interfere. Here we paste, click upload. And here we paste the segment. I uploaded Sam Altman's voice. This is, how to say, the CEO of this ChatGPT Open AI, yes? Here we have a segment. We can listen to how it sounds. Here at the top, we need to type the text that he will read to us. That is, I wrote that I am a trickster, yes? Well, and with Sam Altman's voice to say that I am a trickster, yes, such a, how to say, I don't even know how to translate this word, who wants to deceive you, yes, such a swindler, such a cunning person. Here. Now, and here we select the model. There is a model selection. I chose 0.6 billion parameters, that is, 600 million parameters, and 1.7 billion parameters. I chose 1.7 for better generation. Device, I have BF16 set, as it was written on the repository. Select BF16 here to save an incredible amount of memory with minimal quality loss. The language that should be generated. Again, there is Russian, Portuguese, Italian here. Possibly, yes, I will try to generate Sam Altman in Russian now, yes. And below here is reference audio text. Here you need to do a transcription or transcribation of the audio if X-vector is not enabled. If X-vector is enabled, then transcription is not needed. In what cases is it needed? When, for example, you have a voice, but it is not of sufficient quality, not clearly audible. Or in any case, you can disable it and write it down so that the generation is better overall. I enabled it because, well, just for demonstration, it was faster for me. And further, you can also play with all sorts of settings here. I left them at default. I clicked Run. After which, this replica was generated. Let's listen to it. It sounds very cool. It's very similar to Sam Altman, really, like. Yes, true, in the first one he speaks very fast. Here, in principle, he observes punctuation, yes, that is, if you don't put dots, commas, dashes, and so on, if you write everything in continuous text, it will read like this, yes, but since there is a comma here, he makes a pause, as he should. Now let's try to switch to Russian and try to make Sam Altman speak Russian. So, I am just a swindler who wants to deceive you, and click Run. The voice is ready. 4 seconds of replica. Let's listen to how well he manages to speak Russian with a cloned voice, so that Sam Altman's voice remains. >> So, I am just a swindler who wants to deceive you. Well, it doesn't sound like Sam Altman at all, to be honest, but overall it's an unrealistically high-quality generation. Well, I liked it. I don't know, very high quality, read very coolly. There are no unpleasantries, artifacts, any irritation. A simple, neutral voice that can, in principle, be listened to. Next, we move here. Custom Voice. Here we just without Yes, by the way, about cloning. You upload a segment from 5 to 15, as it was written, in MP3 format. If you don't have MP3 format, you can type in the search engine convert, for example, if it's video, yes, MP4 to MP3, and you format a piece of video into MP3. Well, and as for video, you can download it, in principle, also download video from YouTube or wherever you will download it from, or record it on something. Yes, I hold down Ctrl, select this, these nodes. Right-click, Bypass. Let's disable this. And, uh, our generation, we can click, in principle, right-click on it, Save Audio As, or go to the ConfI folder, go to Confy UI, look for the Outputs folder. Look for output. And here we have audio. And here our generated audios are stored. Yes. Here. They are saved. Or we can simply click and click save as and save the audio wherever we need it. Well, in principle, they are saved, nothing to worry about. And here in custom voice, we hold down Ctrl, select all this, right-click, click Bypass. Here we select speakers already. We have speakers here, yes? Click on the list. And we can see the speakers on the Quen website, yes? Here we have Aiden or Ryan, they speak English, for example, yes? And we can, uh, choose a person, for example, Aiden or Ryan. Ryan, yes. We can enable English. Uh, here below it describes, uh, how it describes here, what kind of person this is, what his story is, what his character is, what he does, and so on. Here you enter the text and, in principle, that's all. That is, with these voices that are here, standard ones, you can click run and it will generate the text for you. Now we hold down Ctrl, select this, click Bypass, and Voice Design, yes, this is a feature that was added to 11 Labs not so long ago. The Chinese have released it into the public domain. Here we write what we want our voice to say, which we will create. Below, we describe completely what kind of voice it should be. Here in the example, the character's name, his profession, what he does, some of his physical characteristics, some of his psychological characteristics, character, some backstory, history is written. You can write like this, try it, or you can go to Quen, and here they have, scroll, search. Here they have an example, yes, pitch, how the voice should sound, speed, and so on and so forth. That is, you return, uh, paste it, write it down. That is, there is even a choice here. For example, instead of English, you can create a Russian-speaking character. I don't know, I haven't tried to write in Russian here, but you can try to write in Russian. In general, I recommend writing in English, or perhaps translating into Chinese, but you can try to do it in Russian. That is, you describe the character, you describe his text, you click, and below your character is generated, which you can download. Well, that's basically it. This is Quen 3 TTS, which can be used through Kofei, it can be done with a different interface, deployed. There is also a version, an opportunity to try it in an online demo. I will leave all the links in the description under this video related to Quen 3 TTS. And that's all. Thank you for watching. You can support the video with a thumbs up, subscribe to the channel, and click on the bell to not miss new videos on this channel. Thank you. Good luck.