Transcription
Looking to bring your words to life with AI? In just 45 minutes, you're going to learn the essentials of creating natural and professional voices using 11 Labs. And if you wanted to take your knowledge a step further and apply it to real-world projects, we have a full course linked in the description, plus another link for you to get 11 Labs today. So, let's get started.
Before we get into how to use 11 Labs, let's first talk about what it is and what are the use cases. So 11 Labs is actually an AI company, and they specialize in generative audio. Generative audio basically means when you generate audio, and given that this is an AI tool, you're making audio with AI.
Now, previously, there have been a lot of, uh, technologies regarding generative audio, such as text-to-speech, speech-to-text, and other stuff. However, those rarely involved AI. And recently, with this implementation, the workflow has become a lot easier and a lot more accurate. So when you type something and you use punctuation, instead of it sounding super robotic, now with AI, it sounds a lot more realistic. And the other way around, speech-to-text, it can catch onto your words a lot better compared to the previous technology that was out there.
So, this right now is the homepage of 11 Labs. We will talk about account and pricing in a little bit, but just to walk you through what you can do with this tool. Obviously, you can make a lot of instant stuff, such as text-to-speech and then speech-to-text, but you can generate audios for your audiobook. You can make podcasts with it. You can maybe just talk to an AI for a YouTube video or something. You can even generate your own sound effects. So, there's a lot you can do here. And this is a very popular tool if you're into making content or have some sort of channel where you would want to, um, post these contents or perhaps even monetize them.
There's a lot of room for experimenting with this tool. You can generate a ton of things, see which one works best, and then fully commit with a longer audio clip. So, the way it works is that you have your account, you get a certain amount of credits, and per generation, you're going to use a few credits depending on how long your audio is. So, if you've used AI tools before, the workflow is relatively the same. You give it a prompt, you give it an input, choose your model, and then you get an output.
So, this is likely the page you will see when you first go into 11 Labs. This is before you have an account. I just signed out for demonstration purposes. And you can see immediately, even without making your account, you can try some things out. So, here's a sample text. And you can see they're adding certain ways that the speech sounds. So, whispering, giggling, being sarcastic. Down here, you're choosing the voices essentially, and then recently they added different languages. So, let's decide what we want to do with it. This is text-to-speech. This is just for a preview, by the way. We'll get into what each of these mean, how you can switch between languages, and etc.
"In the ancient land of Eldoria, where skies shimmered and forests whispered secrets to the wind, lived a dragon named Zepharos. Not the burn-it-all-down kind, but he was gentle, wise, with eyes like old stars. Even the birds fell silent when he passed."
So, you can see how clean it went over those voice changes, how it giggled really naturally, and it really sounded like a human being, but this is all AI, of course. So, this is with 11 V3, which is their latest model. You have the option to go over older models if you're doing a special sort of project. But this is a good way for you to immediately check if this is a tool you want to go for or not.
So, now that we have this, let's go ahead and make our first account. When you click on sign up, you can do it with Google, with your email. Put in a password, and then you immediately have access to your account. Now, before that, I'm going to head over to pricing right over here. And there are a lot of affordable options with this tool. So, the free version actually exists. You get about 10,000 credits a month. You even get API access for if you wanted to integrate 11 Labs into something else. And if you click on "view more," you can actually compare it with everything else.
So, these are the different models. So you can just click on them, and then it compares it based on the, the pack, pre-pack, starter, creator, and, uh, you can see how much it's going to cost you. Then, of course, how many credits, how many generations, you can bill it monthly or annually. Obviously, with annually, you're going to save some for your payments. I would say you can start with free and then, if it's up to your liking, move on to the starter pack. You can see it's not a crazy amount, just $5. And I'm pretty sure by the end of generating a few of these audios, you're going to be convinced at how clean and seamless all the generations are.
We have some more products down here. And then you can see if that's included in your packet or not. They also have packages for businesses and startups. So, um, as I said, it's a pretty popular platform and pretty affordable if you're just starting out. So, that's about the pricing and how to make your account. It's really straightforward, very user-friendly platform.
So, now that we know what the platform is for, in a very brief manner, how we can make an account, what are the different options for a paid packet, we can go ahead and learn a little bit about the dashboard. With all of these new AI tools coming out almost every month now, it's really important for you to brush up on the ethics of using AI if you're not already familiar with them. If you're a beginner using 11 Labs or any other AI tools, you may be wondering why sometimes it's considered unethical. And even if you know and you're still not convinced, it's important for you to realize the risk and why it could potentially harm other people.
For this lesson, we're going to talk a bit about the ethics of using AI tools. Whether it's ChatGPT, Runway, Midjourney, or even 11 Labs, there is a boom in AI tools nowadays. And it's important for you as the user to know about the risk and why you need to be using this tool, these really cool tools, in a more responsible way. So, a lot of times you come to these AI tools to either make your work easier, help you with a project, or just have some fun. And those are all okay as long as you're not using it to harm anybody.
Now, what do I mean by harm? A lot of times these AI tools are being used to impersonate people. The inputs that are used to train these models are often taken from people who did not consent to it. For example, with a lot of art tools, there are many artists claiming that their works have been used to train these models when they did not consent. So, what happens when that artist who worked really hard gives their art, and then it's being used to train these models, and then you make something with it and then sell it? So, you right there have knowingly or unknowingly, um, have kind of contributed to this bad cycle that is harming these artists.
Now, with 11 Labs, even though it's a great tool, it's important for you to take some time to read about how these models are being trained and how you could just use it for something that will not harm people. I will give you some examples with 11 Labs in particular, but this applies to any other AI tool that's out there. So, you should never be using these audios to impersonate real human beings. There are a lot of rules and regulations coming out every day about how to, um, basically rules that prevent the harm of other people. If you generate a podcast fully using AI, you do need to label it as AI. If you use the script from the internet that was also made with AI, maybe ChatGPT, it is your responsibility as a user to inform your audience that you are no, you are in no way trying to impersonate someone else or trick people into thinking that what you just generated is a person.
If you're going to be using 11 Labs for your job or for your school, bear in mind that each company and institution has their own set of rules and regulations regarding it. So, make sure you do your research and you don't put yourself at risk when using these tools. Now, for the purposes of this course, I'm just showing you how to use it. Use it to make podcasts, audiobooks, and stuff like that. These are just meant for you to explore and learn. If you do want to commercialize any of your works that you make by the end of this lesson, please take the time to read the rules and regulations of all the platforms where you would be uploading these products to, and again, label your work as AI. Because what happens is that when you submit your work in the same space as other human creators do, you will be flagged and, depending on which platform you're doing it on, there will be consequences. So, be responsible with your generations. Don't use it to harm or impersonate anyone. Label all your work. And if you choose to make money, monetize your work with either 11 Labs or any other tool out there, do it in the proper fashion. Go to the platform, see what their guide is. Almost every platform has it. Whether it's on Instagram, Behance, Adobe, wherever you want to do it, just go and navigate on their website, read about it, and follow their steps.
So, now you know a little bit more about the risk and the potential harm that you might be doing if you were to use AI in the wrong way. We can now move on and continue with our course where we learn more about this particular tool.
In this lesson, we're going to go over how to prompt with 11 Labs. So, a prompt is the directions that you give an AI tool to follow. And in this case, these are the words that you want 11 Labs to execute. Whether it's a voiceover, a text-to-speech, voice changer, or any of the other tools that we're going to take a look at later, there is a general style that you need to follow, and it's actually pretty easy compared to the other AI tools out there. Let's head over to text-to-speech where we're going to practice with the different prompts.
Now, don't mind the tools on the right side. I will get into what this is in a different lesson. But first, let's go ahead and do a simple practice. We're going to start typing in something. I'm going to have ChatGPT give us a very simple story, and then we're going to convert it into a decent audio. So, "Give me a four-sentence story, um, about a dog. Just very simple." And to make this better, let's do something like this: "Make a podcast episode based on this with only four sentences for dialogue."
All right. So, let's copy the podcast part. The reason why I chose podcast is because this way we can play around with the different emotions. So, I'm just going to paste this here. Feel free to write your own text, but let's get rid of all of these tags and just put in the, the text that they say. Um, I'll put a paragraph break after each sentence. So, let's first start by choosing a voice. For me, by default, it's George. The voice is the model that is reading your text. So, if I just hit George, I, I'm not going to change anything here. Let's actually reset the values just to make sure that we're all using the same, um, settings. Let's hit "generate speech."
"Every morning, a scruffy little dog named Max wandered the quiet streets, searching for scraps and friendly faces. One rainy afternoon, Max huddled beneath a cafe awning, shivering until a kind barista spotted him and stepped outside. 'Hey there, buddy. You hungry?' From that day on, Max came back. Until one morning, the barista opened the door with a smile and said, 'Want to stay?'"
So, that's just the model reading our text. It's very monotone. It's not exactly exciting or dull. It's just very flat. So, here are some things that already are included in this text that partake in that style that I was talking about earlier. So, every time you add a period, that's an indication for the model to pause. And then when we have a three, just three dots, it's going to be a longer pause. If you do this, this is going to be a different sort of cut. So, we're going to incorporate these within our prompt going forward, but right now we have m-dashes, we have the three dots, just a period, we have a comma. So, those things are already, uh, included. And when you, when I played the audio, you saw how it paused. When the sentence ended, it's again paused here between the two words, and it read it just fine.
But now we're going to talk about how we can add emotions because we have a barista here and then we have just the narrator. So, when you use different adjectives, 11 Labs is going to understand that as the emotion it needs to execute. So, we can say, let's see. So, right here when there's dialogue, we can say, "Set the excited barista in, let's see, so set the barista in an excited way." So, I'm using the word "excited" here. And then let's see what else can we add. Instead of this, I'm going to do an exclamation point. And then maybe two question marks. So, just this small change is going to make a big difference. Let's generate this speech one more time.
"Every morning, a scruffy little dog named Max wandered the quiet streets, searching for scraps and friendly faces. One rainy afternoon, Max huddled beneath a cafe awning, shivering until a kind barista spotted him and stepped outside. 'Hey there, buddy? You hungry?' said the barista in the excited way. From that day on,"
So, you saw how it says how the George here said that, "Hey there, buddy, you hungry?" So, that was the difference between adding two different signs versus just having one. Now, if you go to, if you want to take this a step further, you can add pauses and all of that, too, which we'll get into later because there's a tool within Studio where you just click a button, but you can add pauses here, like three dots, and then maybe like another one here. And now, let's see how that works. One thing that you can do is when you combine the two sentences, right, put them right after each other, it's not going to pause as long.
"Every morning, a scruffy little dog named Max wandered the quiet streets, searching for scraps and friendly faces. One rainy afternoon, Max huddled beneath a cafe awning, shivering until a kind barista spotted him and stepped outside. 'Hey there, buddy? You hungry?' said the barista in an excited way. From that day on, Max."
So, that's the changes that we're seeing. And moving forward, the only prompting you would be doing is giving 11 Labs a storyline like this one. We're going to look at how we can add sound effects, longer pauses, and different sorts of audios. With the other advanced tools, the traditional form of AI prompting where you describe or you command for something like, for example, "Make me a sound effect of heavy rain, um, within the forest." That's an example that will only take place when you're making sound effects with 11 Labs. But for the rest of the other tools, and there's so many, this is the sort of prompting that you're doing. So, you can either generate these sorts of storylines yourself, type it in here, as you saw, it's really easy, or you can have another tool like ChatGPT make something for you in a matter of seconds.
So, this is just a lesson for you to know what we mean when we say prompt. It's not your traditional prompt where there's specific formulas to follow. Just keep in mind the punctuation. Each of them mean different things, as we saw with the commas, m-dashes, three dots, stuff like that. And be very generous with your, with your adjectives because those can mean your, uh, AI voice model going from monotone to either excited, angry, sad, whatever adjective you want to use. And we're going to be seeing that a lot in the further lessons.
Let's take a quick look around this platform and see how we can right away start making voices with 11 Labs. So, this is the page you will see, the homepage, after you've signed into your account. So, right away, there are some latest from the library voices. Essentially, we have a library where people can add in new voices. Of course, there's a procedure that they have to go through, and it's not low-quality voices at all, but, um, this is a library that 11 Labs manages, and they're releasing new voices every day. So, let's give this one a listen, for example.
"So, that's in Hindi."
And then we have, for example, a young woman. So, this voice right now, you can see it's being called "female mature voice," and it sure did sound that way. Calm and deep. So, I just refreshed my page to show you that every time you do it, you're going to get a new set of voices. So, for example, let's go on this.
"Evil is evil. Lesser, greater, middling, it's all the same."
So, this is serious and grim, and you can see it definitely did hit that mark. Now, down here, you will notice it says "English preview." If you click on this, you get to switch between languages. These are the current languages that are available. It's going to be the same voice, the same tone, and all, but it'll be in a different language. So, let's go ahead and try Russian, for example. So, you can see how seamlessly it switches between these languages while maintaining the voice and the tone. If you speak any of these languages, you can see if it sounds realistic or not. But from what the community says, it does hit the mark pretty well.
Okay. So, if you wanted to explore more, you can go over here or simply go down to "Voices." So, this will take you to the voice library. And we have a lot of options to choose from. Right away, you will see that there are check marks on some of them, and that just indicates that they are high-quality and they've been used a lot. You can see the name of this person, and if you just hover over it, you're going to see a little bit more information. So, the description of the voices can really help you choose the one that you need. For example, David here is a newsreader, and he has a clear and crisp voice. He's middle-aged, professional American. And you can just gauge on whether or not that's going to fit to your project.
So, let's go ahead and preview this voice. Just click on it.
"Later that same evening, Detective Carlson received an anonymous tip directing her to a second crime scene. What had at first seemed a random event."
And then, once again, you can switch between the languages. Now, you can see here that the language selection is different from the first voice, and that's because each voice has its own language selection. So, here we're actually getting Arabic, Polish, Italian. I think that those are the only ones that are different. We're getting Hindi, English, and French like we did before. And down here, you can see the flags. So, it comes in these two languages and then five more. These two, two more. And it's just going to change as you go on.
So, in the "Explore" page, you can search for either the name of that voice or the way it sounds. So, for example, I'm looking for a calm voice. And I'm getting a lot of narrative and story, informative, educational. Going to give this a listen.
"And here we have a calm, well-spoken narrator, full of intrigue and wonder for nature, science, mystery, and history with a smooth and velvety tone."
And you can see that definitely did sound calm, and it's actually one of the adjectives that's being listed here. So, right away, without having to click on it or hover over it, I know exactly what sort of voice I'm going to get.
Now, you can take a look on the right side, which is the categories that they go into. If you open these in a new tab, let's just open one. We're going to get a library just for informative and educational. And you can see that's the category that I went into. Let's close that guy. Right next to it, you can see "2 Y" or "ZD." That just means that how long this voice is going to stay if it were to be removed. So, if the owner of this voice removes it from the voice library, it will become unavailable for you after 2 years. But here, for example, it will be removed immediately. So, if you want to use this long-term, definitely go for the ones that have a longer retention compared to something like this.
Next, we can see how many users have used this voice. So, the first one by Adam Stone, you can see it's pretty popular. 438,000 people have used this voice, meaning that you've probably heard it at some point.
"The people who are crazy enough to think they can change the world are the ones who do."
There we go. And then there's a plus button, which means that you can kind of favorite this voice for later use. Click on it once, and then it basically goes in this tab, which we'll get into later. Now, you can see that it went from a plus to this icon. I can use this voice to either generate an audio, do text-to-speech, speech-to-text, whatever that I have to do.
Next is some more actions. You can copy the voice link, the voice ID, and then "view similar." So, if you've used other AI tools, think of the, um, voice ID as the seat number. So, when you have the ID and you're trying to generate your voice, it will keep it a lot more consistent. So, it won't fluctuate as you're putting in different adjectives, different tones, it will remain similar, but usually, if you use one of these pre-made voices, you don't really have to use it. The link is just a link to that voice if you want to share it with someone. And then we have "view similar," which gives you more voices that sound like Adam Stone.
So, now let's explore the things up here. This was our search bar. We were able to look for an adjective, maybe someone's name. Let's look for someone called John Doe. There we go. We have three options. We can also look with search by age or gender. Maybe senior, can do male, or maybe teen male, or something. Anything that we want. So, for example, English teen youth.
"Hi there. I'm Archie. If you're after a young, energetic, and authentic British teen voice, then..."
"Hey, I'm Ethan, and I would be a great choice if you're looking for a teenage voice."
So, you can see how those search bars affect. So, you can see how searching for a certain thing does give you the correct voice selection. So, some of these voices will be charged a bit more than the others. So, with the plan that I have, you can see these ones are okay for me to use. They're the popular ones, and there's no dollar sign over here. But for these, you can see it's like two cents. And that's only for a few of these voices. You can see most of them don't have that. Okay.
So, let's go back up here. This was our search bar. Right next to it, you will notice this icon, which is basically when you get to upload a sample. Let me close that up. You get to upload a sample, and then 11 Labs will find one that's similar to that. So, let's say you have this voice that you want to replicate with AI. Of course, be very mindful of copyright, but you can put it over here. For example, if you want a nice, confident voice, like, for example, Michelle Obama, and you want to get something relatively close, you can upload a snippet of her speeches over here to get something close. So, for example, if you're doing some sort of animation and you have a character after Michelle Obama, you can do that. Again, you really want to be mindful about how you're doing this. We will talk about AI ethics in a later lesson, but, um, you can't just grab any voice and try to get the exact same thing, 'cause that could lead to some problems, big problems.
Right next to that, we have "Trending." So, if you click on it, there are a lot of popular voices as well as unique voices. What I mean by that is that the trending ones have been used a lot. So, these are the ones that we see right here. I'm going to click on this triangle just to go into the list again. So, you can see a large number of people have used it. So, if you want to be unique here, you shouldn't probably go for something like this. And that's where this tab comes in handy. If I switch to "Latest," you can see this one, for example, it's not verified, but only 14 people have used it. So, this way, I'm being more unique when I have my voices in my videos. And there's also options that are lower, like seven people, five people, and the list just goes on.
"Two people..."
"...getting left behind with AI? Feel like you can't keep up? Your business needs a prompt engineer who matters."
So, you can see even though they're new, some of them are not as high-quality as the most popular ones. So, it's a good idea to preview them, make sure it's something that works for you, and then proceed to use it. So, we have "Latest," we have "Most Users." So, you can see this one is really high.
"I think this is a really nice way to just talk naturally together, you know, talk plainly. Let's give it a shot."
So, you can see it's conversational. It's really, it sounds really natural, and that's perfect if you want to do a podcast or if you want to use your own voice and convert it to this tone. That's something really cool that 11 Labs have, and we'll get into that in a separate lesson, but that's what "Most Users" will get you.
So, "Character Usage" basically means how many characters you get to convert to audio within your credit range. This is something we'll talk about later when I show you how to do text-to-speech. You will see exactly how many characters you're using. For some of the voices, um, it will take more credits, while the others won't. So, this is what it's referring to. Don't mind this for now. We'll get into this later. At the same time, we have the "Filters" tab, which lets you, you know, dig deeper and find what you really need.
So, we have languages. First of all, a long range of languages. Let's go to like Welsh, for example. You can even choose an accent if it applies. So, for example, it doesn't for Welsh, but let's see what would have an accent. Let's see if English has one. There we go. So, if you're choosing English, there's a lot of accents that you can explore. Let's say I want someone from Chicago. Let's choose that.
Next is the category, which is the tone that the voices sound and speak in. So, think about what you want this voice for and simply choose your category. I'm going to do a very natural tone. So, conversational. Click on that. You can also click more than one category. So, I will do maybe something like this. Quality, any or high quality. Remember, high quality are the ones with the check mark. Those may ask for more credits, but, um, you can just click on "any" to get to have a bunch of options to look at. Next is gender. I think I'll do like an old person. Notice "Period." These were the numbers on the side. How long you get to keep them once the owner deletes that voice. Let's say 2 years. Custom rates, include or exclude. This is for that character, uh, usage that we saw, and then "Live Moderation Enabled." This is going to moderate the type of text that you're giving to this voice. I'm going to hit "include" for now. Let's apply the filters.
And right now, you can see I don't have any. And this will happen, but you get to create your own voices. So, um, that's something we'll explore later. I'm going to remove a few of these just so I can get something to show you guys. Okay, so just a US Chicago accent.
"Exercise makes our hearts stronger and our smiles brighter."
So, that was one example. You can see it's a very recent audio. Only 32 people have used it. It's not verified, but it's still a voice that you can use. Let's try something else. We can do again English. Try Canadian maybe. So, here's a high-quality audio.
"Meet Grim. This handsome cat at our shelter. This boy is lovable. With his tabby stripes and cute face, he's sure to steal your heart in an instant."
So, that was one example. Let's try to find another one.
"Ladies and gentlemen, welcome to today's match. We have an exciting match ahead between FC United and Team Impact. The atmosphere is electric."
So, these are all Canadian accents in different categories. Next, we have "Create or Clone a Voice." This is that cool feature that I talked about where you get to either design an entirely new voice from a text prompt. So, you would say maybe an older gentleman with a thick accent, maybe a deep voice, calm, energetic, whatever you want to put. And it's going to, 11 Labs will give you a bunch of options, samples. You get to build up on that and then make your own original audio. Now, those the things that you create, you can also have other people use it. So, that's where all those new voices come from. Takes less than a minute. And we'll get into this in a different lesson.
We have a "Voice Clone" option. So, you can clone your own voice with only 10 seconds. So, I will speak into the microphone for 10 seconds, and then 11 Labs will study my voice. Then I get to use my voice, um, with different text prompts. So, for example, if I don't want to record an audiobook that's 3 hours long, I will just speak into the mic for 10 seconds and then paste the storyline or the, just the story into the box, and then 11 Labs will generate that 3-hour audio for me.
Next is the "Professional Voice Clone." This is not included in the plan that I have, but this is just for the most realistic one. But for most of you, since you are using this tool for the first time, these two should be enough. Up here, we have a "Feedback" box, which you could just write your feedback. We have some other stuff up here.
So, this was the "Explore" tab. "My Voices" will have the voices that you favored or made yourself. We have "Default Voices." These are voices by 11 Labs. So, these aren't like user-generated.
"Just trust yourself, then you will know how to live."
You can take this, use it like all the other voices, and just take a look at the many options. We also have "Collections," which you get to create when you are trying to organize your voices. For example, I could create a collection for my YouTube channel, one for my podcast, one for my audiobooks, and that way everything will be a lot more organized.
When you're previewing a voice, we have the name here, the different languages it's available in. Play button. You can skip back and forth. See how long it is over here. You can download the voice. And then you can just hide the player if you don't need it. And that's the voices panel. If you click on this plus button, it's going to bring you to the same, same menu that came up when we clicked here.
Now, in the next lesson, we're going to go over the playground tools that are located right underneath and see what else we can do with this tool.
Let's go ahead and turn a simple text into speech using 11 Labs. In the playground section, the first tool is text-to-speech. And once you click it, you're going to be greeted with this large text box and some options on the right. So, let's go ahead and see how this whole thing works. Down here, you can use some of the presets. When I hover over them, you can see that the text changes, and there's a wide range of things that you can do. So, speak in different languages, podcast, or even just type your own thing. I'm going to get started with this one. So, I just clicked on it, and I had to pause it because it immediately starts playing.
And on the right side is where you get to make your adjustments. So, first, you can choose your voice. We can go to the saved voices, recent, um, anything that we want over here. You can scroll down to, let's go with the John Doe one that we found in the previous lesson.
"I've seen things you people wouldn't believe. All those moments will be lost in time, like tears in rain. Time to die."
So, let's say this is what I want. When you click on it, you get to learn a bit more about the voice. Again, if you don't want something that a lot of people have used and you want to be unique with your voices, this is where you want to pay attention to. And now we can just add it to the voices in case we want to use it later. Already, it's applied when you click on it the first time.
So, the model, just like any other AI tool, um, 11 Labs has its own models. And when you click on this, you get to see a little bit more about the descriptions. So, this is what we have been using, um, so far. There is a version two for different languages, which is what we have to use because we're using different languages. And then you can also view the older models if that's what you want to do. And interestingly, if you use an original voice from 11 Labs, it tells you which one is recommended for it. So, 11 Turbo version two is not the most recent, but it's going, it's recommending it to me because I chose John Doe, but I could choose to ignore it and just grab this guy. So, let's see. I'm going to go back to John Doe.
And with the model, the latest model, it's still in research preview. What you can do is just change the stability. So, if you lower this, now it's on "natural," but if you lower it, it's going to be a lot more creative as to how it, you know, executes this audio. Let me just show you a little preview. This is on "natural" right now. We have to first generate the speech.
"English has the most words, with over 170,000 in the Oxford English Dictionary."
So, immediately it gives me a bunch of different generations. So, we have one, two, and they're still going on. And then I could choose the one that I prefer, make my changes, and continue going from there. So, that was the first one. This is the second one.
"English has the most words, with over 170,000."
So, the difference between these two is that generation two pauses. You can even see from the sound waves, it pauses around here, while this one just goes on continuously. Now, I could reduce this to "creative." Let's regenerate the speech and see how this affects the generations.
"English has the most words. English has the most words."
So, you can see how there's a lot more changes in his tone within the audio itself. And compared to the other audio. If I wanted to be more robust and more, I guess, uniform, you can say, I could increase this all the way.
"English has the most words, with over 170,000 in the Oxford English Dictionary."
So, you can see that sounds a lot more robotic, but depending on what you're trying to use this voice for, that could be the option you would prefer. You can also add different speakers. So, say you are doing a podcast where you want different speakers to talk. This is where you click on "add speaker," choose your voice. Let's go with a female. And then let's add it to my voices. And we can do something like, "Haha, that is so funny." Now, we can generate the speech. I'm going to reduce this to "natural." And then we get to see what we have to work with.
"English has the most words, with over 170,000 in the Oxford English Dictionary. Romantic, strange, ramat, or shyat. That is so funny."
So, you can see how immediately after John spoke, Arabella started speaking, and then you get to go back and forth between multiple voices, just one voice, and have a pretty natural interaction between them. Now, we're going to talk more about voice prompting in a different lesson because what you can do is add adjectives in your prompt. Add pauses, add like coughs, a sneeze, anything you want to really take control of the way your voices sound, the AI voices. But for now, I'm just going to show you how these work. And once we're done with chapter 1, we can get into the, the more details on how to properly use these tools for real-life scenarios.
Okay, so this was our settings. And when we switch between models, let's go with this guy. We're going to get more options. So, the latest one was still in research preview, as the tag mentioned. So, we're getting this because we're using multiple languages. I'm just going to remove the non-English words so that we don't have that limitation and take a look at what we have down here. So, "Speed" talks about how, um, fast or how slow your AI voice is talking. "Stability," more stable, more variable. "Similarity," exaggerations. And then we have "Speaker Boost."
So, let's do a couple of examples and see how these sliders can affect it, and just type something that would need some exaggeration. So, here's an example of a prompt that you could play around with these sliders. You could have the voice read in a very monotone, uh, voice, a very exaggerated one. And right now, I'm going to generate this with the default settings that you see right here, and then we'll try some different variations.
"I walked into the party and, oh my god, there he was. I was so scared that he would learn about my travel plans that I accidentally bumped into the host of the party. I don't know what to do now."
So, you can see even without changing anything, it was very natural in the way it enunciated the fear, maybe the nervousness, the three dots, and all the other stuff that I put here. Now, I could try something faster. Make it less stable. Lower the similarity, and then make it really exaggerated. I am getting some warning. So, if you go really low, this will come up. So, I'm going to do 30. And let's do under 50. This is a random selection. I don't think this will give us a good audio, but let's see what we get.
"I walked into the party and, oh my god, there he was. I was so scared that he would learn about my travel plans that I accidentally bumped into the host of the party. I don't know what to do now."
So, it's still pretty good, but it wasn't as consistent. And I think that's because of the stability. So, playing around with these can get you the best audio. You can really customize it. And if I go to something else, let's try this one. You can see the options get either more or less, and that depends on the model. Pay attention to the recommendations by 11 Labs and the tags that it has up here. So, if I go to the first one, second one, still the same, but the one that we were on right now gave us the exaggeration. And if you ever wanted to go back to the default, simply click on "reset values."
When I have my audio, what I can do is give feedback and then share it if I want to download it. And then I could hide the toolbar by clicking this button here. It tells you how many credits you have remaining. And here, it's telling me the characters. So, the ones that were based on character usage that we saw in a previous lesson, it's going to charge 2 cents for my 201 characters.
When I go over "regenerate speech," it tells me that I have one free regeneration. And basically, what that means is that all the voices that I've made so far in this lesson did not charge me anything from my credits, and it was free. So, that's a good thing that 11 Labs has for you because when you first start making an audio from text-to-speech or generating your own custom voice, you're going to be exploring a lot with these sliders. And so, this way, you will only use your credits when you've somewhat reached a consensus on which settings are good for you.
So, that's pretty much the text-to-speech feature. Pretty straightforward. There's a lot of room for you to explore and get creative with it. We are going to be making a bunch of projects from scratch where you get to follow along based on my instructions. We will be making things like a podcast, YouTube video, the audio that is, and see how we can use these tools for everyday uses.
Congratulations on completing this free 11 Labs tutorial. This was a huge first step for you to get into the world of AI voice generation, and I hope you guys really enjoyed it. If you'd like to take this further, we have a full course linked in the description down below. Over there, you're going to learn the essentials in a lot more detail and apply your knowledge to real-world projects. Plus, you will learn how to create professional and natural voices, sync voice to video, create multi-dialogue videos, and also prepare content for monetization. We also have a link for you to get 11 Labs today. So, go ahead and do that and start creating.