📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

F5-TTS! They DID IT! Perfect voice clone with Emotion with a 10-second sample!

Bob Doyle Media18:17

Transcription

Okay, this is a stop the presses kind of moment. What I'm about to show you is so freaking totally awesome. It's not perfect, it's not the end all be all, but it is probably one of the most impressive voice conversion/cloning text-to-speech solutions I have ever seen. It is absolutely free, it does so many cool things, and I cannot wait to show them to you.

Welcome back to the channel, where we discuss the creative uses of AI. And today, I am sharing with you F5 Text to Speech. Now, you may also see it as E2 F5 Text to Speech, but we're going to be focusing mostly on the F5 part because that's the part that's so awesome. This is another text-to-speech technology that clones your voice or the voice of whoever you put in there, but it only requires about 10 seconds of that audio, and it does an amazing job with that 10 seconds.

Now, over the course of the past couple of years, I've seen lots of solutions that promise voice cloning with 3 to 5 seconds, and you know what? No, they didn't work. They were maybe an approximation, but this right here is approaching, in my mind, 11 Labs quality in terms of vocal reproduction. Now, the conversational style and all of that still has some work to do, but where this really excels is the communication of emotion and the ability to mix emotions in one output that you generate. And it also has a podcast generation feature, and we all know how popular that has been recently. So let's go through all of this as efficiently as we possibly can.

I'm going to leave a link to this page here, which tells you what F5 Text to Speech is all about. Charlie Brown. And also, as you scroll down this page, you will see things like that, which I'm sure you can all make sense of immediately. I know I did. And then lots and lots of examples. We're not going to go through these right now because we're going to play with it itself. I will go ahead and take note that you see two languages represented here. We see English, and we have Chinese. Right now, those are the two domains that this works in, but wow, does it do a great job when it's taking English text and translating it to Chinese. So when they add more models, boy, this is going to be amazing.

At the very top of this page here, you see a link to code, and this is one place you can download this code on your computer if you'd like. And we're going to get right to that in just a second. However, down here, there are a couple of places where you can play with this without having to install it anywhere. If you don't have a Hugging Face account, you may not be able to get too far with this, but it allowed me to play with it about five times, and it did a pretty good job in terms of speed as well. So at least for me, the GPU they provided for me was at least as fast as my 3090, which is what we're going to be using today.

Now, if you know all about installing code from GitHub and all that, go right ahead, do your thing, and we'll see you on the other side. If you don't, there's a much easier way. In fact, when I heard about this technology, which was just a couple of days ago in the comment section of a recently posted voice conversion video, I immediately wondered if Pinocchio had it. Now, Pinocchio, if you don't know, is an app that allows you to easily install a lot of these AI applications without having to worry about all the configuration and the environment and being an expert on Python and all these other things. From this one interface, you can download a lot of popular AI software, but it starts with getting Pinocchio on your computer.

When you come to the page here, which of course I've left a link to, click on download, and you'll see that you can download this for Windows, Mac, and Linux. I won't walk you through every step of the Pinocchio installation process because I've done it before and it's pretty simple. You download the program, it tells you right here, this window's going to pop up, just click "Run anyway." Pinocchio has been around for quite a while. If it's your first time running Pinocchio, it will look like this, except with nothing here. You'll have this little interface here, and what you'll want to do is go over to Discover, and this is where you can search for and install any of these AI apps you're looking for. And as it happens, this E2 F5 TTS is right there on the top, and all you do is you click on this, you click on download, and follow the instructions.

Now, there's lots of things that could come up if you're installing Pinocchio for your first time. There's most likely a screen going to come up and say you need to download CUDA, you need to download this, but it will do all of that for you. You just need to click through the process until everything is all installed and you're ready to run the software. So once it's all installed, you'll have it here on your list, and you'll just click it and click on Start. And after just a minute or two, the interface will pop up inside of this window.

So, full disclosure here, I have installed this on two different computers. One has the Nvidia RTX 3090, which is what we're going to demonstrate with, and the other has the Nvidia 2070 Super, which is not nearly as powerful, and it shows. So although this software will run on that lesser card, it takes a lot of time, and some of the things I tried choked all together. But in theory, it should work on an 8 to 10 gig card from what I understand. So don't let my experience stop you from at least trying it if this is something you're interested in.

Once this interface loads up inside of Pinocchio, I actually like to pop it out into a proper tab in a browser. Just much more space, and I just like it that way. You can do what you want. Now, this is pretty simple. We have three main tabs that we're going to be working with today: the Text to Speech, the Podcast, which is cool, and then Multistyle, which is what's going to allow us to mix various emotional styles within one file of output. It only takes about 10 to 15 seconds of audio, and in fact, it won't let you work with any more than 15 seconds of audio at a time to create this model.

I've got a folder here with a few files I've already pre-recorded, and I'm going to drag "regular" over here. Just let you hear what it sounds like. Okay, this is just some regular reference audio where I'm not really showing any particular emotion other than just, "Hey, it's kind of nice to be here," but I'm not overly excited about anything. I'm not angry, I'm not sad, but I'm not showing great, uh, you know, expression of joy or anything either. Now, it has an idea of my voice. It also has an idea of the pace at which I speak, which is very important because another very strong feature of this solution is that it does pick up on your pacing and some of your vocal affectations, which is extremely cool. And that includes dialect, accents, things like that.

So now let's give this thing some text to say. "Hi, I'm Bob Doyle and I run this place. I'm going to want to see some identification, or I'm going to have to ask you to sit over there in the corner." Now, in terms of additional settings, we have a few. We have reference text, which is the transcription of what this is up here. Okay, this is just some regular ref. So if I had the transcription, I could paste it in here, and then it would know exactly how the words in the audio file match up with what's being typed. However, if you leave it blank, it's going to use the Whisper technology that we talked about on a recent video to do the transcription, and unless you're having real problems here, I would go ahead and trust that transcription. For my purposes, it's done great.

You don't see it, which I'd like to be able to do, but what it creates automatically has never caused me a problem, and I played with this a lot. The next line says, "The model tends to produce silences, especially on longer audio, and we can manually remove silences if needed. Note that this is an experimental feature and may produce strange results," which I've had it do. I generally don't have it remove the silences because I'm going to be editing this later if I'm going to use it. I can remove the silences there, and I don't want any weird behavior. PS, sometimes you'll get weird behavior anyway.

The speed, self-explanatory. How fast or slow is this speaking going to be? With crossfade durations, I've not yet seen a difference with anything I do, whether I slide this back and forth or not. So I just leave it where it is. You can play with it and see what you get, but that's really all you need to do. Choosing the TTS model, as I said, we're going to be working mostly today with the F5 TTS. It's just a better model, it's smoother, it sounds better. When I tried to run the E2 TTS model on my 2070 machine, it just choked. I don't know if that's a thing, but here we'll go ahead and listen to the differences.

Let's start with the F5 TTS. So the first time you do it, it's got to build that little model, but even so, it's still pretty quick. In fact, there it's already done. So let's listen. "Hi, I'm Bob Doyle and I run this place. I'm going to want to see some identification, or I'm going to have to ask you to sit over there in the corner." So for 15 seconds of audio, that's a pretty good representation.

Let's choose the E2 TTS model and see what happens here. "Hi, I'm Bob Doyle and I run this place. I'm going to want to see some identification, or I'm going to have to ask you to sit over there in the corner." The strong points of the E2 model are its conversational style, and I think you heard the difference there. That sounded really, really natural. Not that the other one was stilted necessarily, but I get far more robotic or unnatural sounding voices sometimes with the F5. Sometimes in that case, I thought the E2 was a little bit more fun to listen to than the F5.

Let's clear this out. Let's bring in another voice and let's listen to the pacing of this through a very tortuous, circuitous route. The book came to me when I was in Australia making a film with Matthew McConaughey, and, um, and I read it, and then so lots of pauses, the word, uh, the whole thing. For this, I'm going to paste in the text of a poem, "The Road Less Traveled." Now, it's pretty long, so it will take a little bit longer for this to convert, but let's see what this does.

Taking the voice quality and the pacing and every other little affectation we heard in there into consideration, I'm going to go back to the F5 model and click on synthesize. I'm watching the progress down here to see how long that takes because this is, like I said, a pretty long file that it's going to be creating. We're at about 25 seconds now, and it says 10% of the way through, although that I don't trust that necessarily all the time. It skips around a good bit. You just can't trust progress bars. You just, you just can't. Like when you're doing a Windows update. All right, forget it. I'm not going there. All right, 50% of the way through. It took right at 92 seconds for this generation with this model based on that clip we gave it. Let's listen to it.

"Two roads diverged in a yellow wood, and sorry I could not travel both and be one traveler, long I stood and looked down one as far as I could to where it bent in the undergrowth. Then took the other, as just as fair, and having perhaps the better claim, because it was grassy and wanted wear, though as for that the passing there had worn them really about the same. And both that morning equally lay in leaves no step had trodden black. I kept the first for another day, yet knowing how way leads on to way, I doubted if I should ever come back. I shall be telling this with a sigh somewhere ages and ages hence. Two roads diverged in a wood, and I, I took the one less traveled by, and that has made all the difference."

That is, that's crazy impressive to me with 15 seconds, and it caught all the little nuances. I just never get tired of when it picks up on the natural rhythm of a sample, and for it to do that in 15 freaking seconds, that's insane. But you want to hear it more insane? Let's hear it in Chinese. Now, we won't do the whole thing. I'll just take this first paragraph here, copy it, paste it over here into Google Translate, just choose Chinese Simplified as the language, click here to copy it into the clipboard, go back over to the interface, and paste this in there instead. Do nothing else different and click on synthesize.

All right, just FYI, that just took about 20 seconds or so. So let's listen.

[Music]

Wow. All right, let's pop over to Podcast because this is really fun, and it's so freaking into it. I love this so much. This is set up for two speakers. You're going to name the speakers. We'll just go ahead and keep that one sample, we'll call him Donnie. We'll drop that same audio there, and then for this one, we'll get Tracy up in here. So I'm going to say Tracy. We'll drag her sample down here. I'll let you listen to what that sounds like.

Okay, so getting started, I am Tracy Star, and I am. Okay, so you get the idea. So she actually has a very distinct sort of a rhythm and a very distinct sort of tone to her voice. So you'll be able to easily tell if that worked well. Again, we have the option of putting in a reference text script or whatever, but we won't worry about that right now because it does a good job of transcribing already.

So now it's just so freaking simple. You want Tracy to say something, you just type Tracy, and then what you want her to say. "Hi, I'm Tracy, welcome to the show." And then the next person, Donnie. "And I'm Donnie, and I also would like to welcome you to the show today. On the show, our guest called in sick, or I guess they called out sick. Donnie, well, either way, I find it totally unprofessional, and I don't think I can continue this."

So now we've got two people with distinctly different rhythms in their speaking combined together. So let's just see what happens. I will un-remove the silences again and generate podcast. Now, this is doing multiple things, so it does take a little longer than just your average generation, but it's not like you got to go out and get a cup of coffee or go on a picnic with a loved one, although you could do that anyway because that's nice, and I don't think we do that enough. You might actually see the progress bar go to 100% every now and then, go back to zero because it's doing it in segments, like I said, and then they just piece it all together. That took about 53 seconds to do this. So let's listen.

"Hi, I'm Tracy, welcome to the show. And I'm Donnie, and, uh, I also would like to welcome you, uh, to the show today. On the show, our guest called in sick, or I guess they called out sick. Either way, and I find it totally unprofessional, and I don't think I can continue this."

Oh my God, that is so satisfying to me. I could go on with other examples, but you get the point. It works really well. And I really want to show you Multistyle because that is also cool. What this allows you to do is basically create an infinite number of versions of your vocal sample that you can then define with a keyword and change how it's rendered. So you can change emotion in the middle of a sentence by tagging the first part of it "happy" and then going into "angry," "sad," or whatever you define.

So let's talk about how to do that. The first thing you would do is you start with a regular. That's just me talking. Now, of course, I've already done all this because you don't want to see me do all that crap. So, regular. I'm just going to drag it right in here. That's done. Now, for whatever reason, there's this huge space here that you have to scroll down. And now I'm going to click "Add speech type." And now I scroll back up, and now I'm going to put in the next one. Now, I believe that whatever you type in here, it's case-sensitive when you put it into your text prompt, so just keep that in mind. I'm going to say "sad," and then I'm going to take this file that I created called "sad," and it has sadness in it. I just thought that there was going to be some way we could get through this without having to do this. See? Oh, look at that. I put in a little cry sound just to make it for reals.

All right, I'm going to scroll down and click on "Add speech type" again, and now this time I'm going to drag on over "happy," and I'll type "happy" here. And this is what this sounds like. "Hi, it's me, and I'm really happy to be here today." And that's what I'm, that's why I'm really talking like this, is just to express my pure, sheer, stinking happiness about everything. Okay. And then we'll go down and add one more, and it is angry. And so angry, in fact, I realize I can't really play it here on the channel for you. We'll play it as far as we can. All right, that's it. I have had it up to. Okay, well, that didn't last very long.

Now we have four different versions: regular, sad, happy, angry. We're going to scroll on down, and this is where we type it all in. If we want a particular part of our sentence here or a paragraph to be in a certain emotion, all we have to do is put that emotion in parentheses right before we do it. So let's start with regular, which did start with an uppercase R up there. "I'm Bob, and there is a lot I like to say about my inability to turn off this annoying keyboard clicking sound." And then I'll scroll on down and go angry. "I mean, for the love of Pete, I know there's a freaking way to do it since I've done it before. I can't believe this is happening during a demo recording." Then I'll scroll down to sad, because, you know, I really want people to enjoy watching these and not be irritated by some constant clicking just because I can't figure out how to make it stop. "I'm such a useless failure." I don't know how we're going to suddenly get happy, but you know what? I just remembered. "Neither do I." And that's what's great about my short-term memory issues. I can't hold on to negative emotions because I can't remember 7 seconds ago.

Generate emotional speech. Just like when we were combining the two speakers together, this is being done in multiple passes, so it takes a little bit longer, but obviously, based on what we put in the parentheses, it is referencing the appropriate reference audio and should pick up all of its little nuances. We've actually been pretty lucky on this particular demonstration that it hasn't been doing a lot of hallucinating, because I've heard it slur its words and make up all kinds of crap. You want to call a doctor on this thing sometimes. But now it's working well. Let's see what it did.

"Hi, I'm Bob, and there is a lot I'd like to say about my inability to turn off this annoying keyboard clicking sound. I mean, for the love of Pete, I know there's a freaking way to do it since I've done it before. I can't believe this is happening during a demo recording. And because, you know, I really want people to enjoy watching these and not be irritated by some constant clicking just because I can't figure out how to make it stop, and I'm such a useless failure. But you know what? I just remembered. Neither do I. And that's what's great about my short-term memory issues. I can't hold to negative emotions because I can't remember 7 seconds ago."

All right, well, that's about as concise a demo as I can give you without going into multiple examples, which is fun for me and maybe not so fun for you. In fact, why don't you just go play with it yourself? I'm you're guaranteed to have way more fun. If this is the type of stuff you like to learn about and stay on the cutting edge with, well, then you're going to need to subscribe to this channel because this is the type of thing we talk about all the time. If you subscribe now, I will not look for you. I will not pursue you. But if you do not, I will look for you. I will find you. And I [Music]