📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Make AI videos with audio of anyone. Free & offline

AI Search32:55

Transcription

This AI can generate super realistic videos with any input dialogue. It's called Hunyen Video Avatar by Tencent. And this is completely free and open-source, so you can run it offline for unlimited times on your computer. It can even handle various emotions and singing and different animation styles.

So, in this video, I'm going to go over how to use it. Plus, I'm going to test it out on a variety of different scenes and audio clips. Plus, of course, I'll show you how to install this on your computer, so you can use it for free for unlimited times offline. The nice thing is this even works if you have low VRAM. I'll talk more about this in a second.

Now, you might have heard of Google's V3, which was released a few weeks ago. It went pretty viral because it could directly make videos with audio. So, you can get people to talk seamlessly in a video just from one prompt. Well, this new Hunyan video avatar is kind of like VO3, but it's free and open source. And I think it's even better because you input the audio. So, this gives you ultimate control on how you want the audio to sound like.

First of all, here are some of their official demos.

"Come on, it's time to go. I booked the table for 8, and I'm not sure exactly where the restaurant is."

"Hey, Ally, relax. This isn't work. This is a night out."

So, as you can see, it not only works with one character, but it can also work with multiple characters. Here's another example. And notice that this isn't just a lip sync tool or a face animator, which just moves the person's mouth or face. You can input any scene and it would animate the entire scene, including the person's body.

So, here's an example of a full body scene plus of a person singing.

"Smooth green."

So, as you can see, not only does it just move her mouth, but it animates her entire body, it makes her actually play the guitar, plus it animates the campfire in the foreground. The only flaw here is the notes that she's playing are not actually aligned with the song. But I mean, even VO3 can't do that. I would be shocked if it was actually in sync.

Here's another example.

[Music]

Here's another example of talking. Again, it's not just making her head and mouth move, but her entire body and the camera is also moving along. Plus, it can also handle different art styles. So, here are some examples.

All right, now enough demos. If you want to see more demos, you can check out this official page. But for now, let's actually try this out.

So, there are two ways to use it. You can try it out via their online platform. And once you're there, you can sign up for free. Now, this is all in Chinese at the moment, but if you're using Google Chrome, you can just rightclick and then click on translate to English.

All right. So, here, first of all, this is basically text to speech where you can input a transcript and select a voice. Now, there are different voices you can choose from. So, for example, let me just type in this transcript here and then let me select this voice and then click play.

"Here is a sample voice."

All right. And now let's try temperamental and graceful.

"Here is a simple voice."

Okay. So notice that it does have a Chinese accent to it. So I usually don't do text to speech. I would upload my own audio clip. So let's try this out. I'm going to upload this audio clip of Jensen Huang talking. This is 14 seconds long. Let me play this for you first.

"How is it possible that that Nvidia became so big building GPUs? And so there's an impression that this is what a GPU looks like. Now this is a GPU. This is one of the most advanced GPUs in the world. But this is a gamer GPU."

All right. And then for the image, let's upload a picture of Jensen Huang. Now notice that you can't input a prompt here to guide this generation further. This is a limitation of the online platform, but you can input a prompt in the offline version, which I'll show you in a second. Anyways, let's click generate.

All right, here's what we get. Let me play this for you.

"How is it possible that that Nvidia became so big building GPUs? And so there's an impression that this is what a GPU looks like. Now, this is a GPU. This is one of the most advanced GPUs in the world. But this is a gamer GP."

Now, as you can see, that was super realistic. It even detected where there was emphasis in the audio and it made Jensen speak that out with more emphasis. So, for example, here when he says, "This is a GPU."

"Now, this is a GPU. This is one of the most."

And then when he says, "But this is a gamer GPU."

"But this is a gamer GP."

Again, it knows to place emphasis when he says those words. Super realistic. Now, I chose this scene because it's quite tricky. He's holding a GPU plus a laptop with the other hand, but it's able to animate this entire scene really consistently with very minimal errors.

All right. Next, I'm going to upload this photo of a man giving a TED talk, which I generated with Google's Imagine. And then for the audio, I'm going to upload a random talk. Here's what it sounds like.

"What these creators share in common is independence from traditional gatekeepers and aggregators. Take Caroline Chambers. When publishers spurned her proposal for a cookbook deal, she took matters into her own hands."

All right, so let's click generate and see what that gives us.

Okay, here's what we get.

"What these creators share in common is independence from traditional gatekeepers and aggregators. Take Caroline Chambers. When publishers spurned her proposal for a cookbook deal, she took matters into her own hands."

Very nice. Now, notice that at the middle of the scene there, there was kind of a cut, which is a flaw, but actually in the original audio clip, that was also a hard cut. So, it's kind of interesting how it was able to detect that. And again, notice that he's moving super realistically. It's his entire body moving, plus the camera also moves a bit. His hands both have five fingers. Everything is just very realistic.

All right. Next, I wanted to see if it can handle different languages. So, here's a Spanish audio. And then here is my uploaded photo. Again, it's a pretty complex scene. There's a lot of people walking around in the background. Let's see if it can animate this.

All right, here's what I get.

All right, it's not bad. You can see it can clearly handle Spanish. However, there are clear flaws with the humans in the background. For example, this dude seems to be walking backwards and then, you know, it's kind of cool that it added another human walking on the left, but somehow he just kind of disappears.

Now, they claim that this can also handle different emotions. So, let's test this out. I'm going to use this audio clip.

"Please just listen for once. Why does this always happen?"

And now the trick is you need to first generate an image of the person with that emotion already. So, let's say we have this really angry girl eating ramen. Let's click generate and see what that gives us.

All right, our generation is ready. Let's play this.

"Please just listen for once. Why does this always happen?"

And as you can see, this is super realistic. She is really pissed off. It's as if she's actually shouting the audio. Plus, everything else moves really realistically, including the ramen. Very nice.

Let's also try a laughing example. So, here's the audio.

"Oh my god, it was way too funny."

And for this to work, we do need to input a photo of someone already laughing. So, I made this with Google's Imagine. Let's click generate.

All right, here's what we get. Let's play this.

"Oh my god, it was way too funny."

Okay, it's not bad. The lip sync isn't perfect, but it's like 80% there.

Now, instead of anger or laughter, here's a sad example.

"Can we just get on with it already? My body is starting to take a toll."

Now, this one looks super realistic. Notice that this person also takes a breath and sniffs based on the audio clip. This is synced really well.

Next, I also wanted to see if this works with animals. So, here's an example.

"I didn't choose the fluffy life. The fluffy life chose me."

So, as you can see here, even with animals, it lip syncs perfectly to the audio. This is really good.

Next, I also wanted to see if it can handle singing. So, let me upload this acoustic song. Let me play you what this sounds like first.

"Maybe you're thinking of me, too. Maybe that's just what I do."

By the way, I generated this with Refusion, which has easily become one of my favorite music generators out there. You can use it for free, so check it out. Anyways, back to here. I'm going to upload this image of a woman playing guitar.

All right, here's what we get.

"thinking of me."

[Music]

All right, so first of all, the lip sync is perfect. She's actually singing out the lyrics. Also, this whole video just looks very fluid and natural. Her whole body is moving. She's even moving the guitar. The camera is also slightly moving. The lighting and the reflections are just perfect. The only flaw is she's not actually plucking the notes from the song. But again, I don't expect it to do that. So overall, this is still a very impressive generation.

I also wanted to see how well it performs with different art styles. So here's a Disney Pixar example.

"Yes, this is his position in two words. A little while since he obtained an excellent offer of employment abroad from a rich relative of his, and he had made all his arrangements to accept it."

Again, this is super good. It even detected a breath and made her pause and breathe near the start of the clip.

"Two words."

Plus, her eyes do blink and she talks really realistically. The lip sync is close to perfect.

Now, instead of Disney Pixar 3D style, let's also see if it can lip-sync anime. So, here's an anime example.

Now, the mouth could be better, but overall, this is not bad. It can definitely handle anime style. And again, it makes her body also move slightly. Plus, she does blink her eyes, so this is not as robotic and rigid as some of the older tools out there.

All right, so let's talk about this versus VO3. If you want to generate a video of someone talking, I would prefer this tool over V3. First of all, for VO3, while you can upload an image as the start or end frame of the video, you can't actually get that person to talk. That's probably due to concerns over deep fakes and copyright. Another limitation about V3 is you can't actually control the voice of the person. The only way to generate audio is with a prompt. So you can't really generate a consistent character with a consistent voice. Whereas for this method, what you can do is upload any audio you want. So you can generate a voice using any voice cloner and texttospech generator out there. I've already covered a ton of them on my channel, such as F5TS or Zonos or this really popular one called RVC or retrievalbased voice conversion, which can do voicetooice. This gives you even more control. So, you can speak out the audio and then convert your voice into any voice you want. By the way, if you're interested in learning more about these voice tools, check out these videos, which I will link to in the description below.

So after you know generating an audio clip, you can easily plug this into Hunen video avatar plus upload a reference image of your character to generate a video from that. This truly allows you to create videos with consistent characters and consistent voices, which you can't really do with Google's V3. Also note that I've specified a ton of other face animators or lip-sync tools on my channel before. For example, live portrait is a really popular one where you can also upload an audio clip and it would animate a face. Or hello is another one, but those tools are just limited to the person's head or mouth. I've also featured Echomimic V3, but that tool is also just limited to a person's upper body. This is the first open-source and available AI that can animate any scene, even like a full body scene, just with an audio clip. So, that's what makes Hunyan Video Avatar so powerful.

Thanks to VidU for sponsoring this video. VidU is one of the top AI video generators out there with a ton of capabilities. Their latest Q1 model is even better. It has improved clarity, detail, and semantic accuracy, making it a powerful tool for content creators and filmmakers. It's really good at creating all types of video. For example, here's a texttovideo generation. As you can see, it's super realistic and consistent. Now, instead of text to video, you can also do image to video where you can upload a start or end frame. And as you can see, everything is super consistent. Plus, they have a ton of other features like reference to video where you can upload images of objects or characters you want to insert in the video. VU just released a new feature for reference to video. Now you can contribute your own reference images to a public reference library or browse and use references shared by others to generate your videos. For example, let's say I want to use this character. Well, I can just select it from this reference library and generate a video of this character in just a few clicks. This can save you a ton of time, especially if you need to generate specific characters or objects. Vido sets a new standard for AI video creation. Try it for free via the link in the description below.

Now, those were just some of my examples on this online platform. Next, let's look at how we can install this and run it locally for free for unlimited times on your computer without a watermark.

So, if you go to the official Hunyan video avatar page and you scroll to the middle somewhere, it does say that the minimum VRAM requirement is 24 GB, but it's going to be very slow. And they recommend using a GPU with 96 GB of VRAM. I mean, who the hell has 96 GB? But the nice thing is you don't actually need 24 GB. So you can see here their latest update is it now supports a GPU with only 10 GB of VRAM with TC included. This is a tool that speeds up your generation even faster thanks to W2GP. So that's what we're going to install today. So let's click on this link and this is by the goat deep beep meep. Props to this guy for building this really simple interface. Now, even though it says one, notice that they've actually merged Hunyan video avatar and also LTV into this interface. By the way, GP stands for GPU poor, which is quite a brutal name. Like, does it mean you're too poor to afford a better GPU? I don't know. Anyways, if you scroll down here, there are two ways you can install Want 12GP. One way is to install this app called Pinocchio. And after that, they provide a one-click installer for Want 12GP. So, if you want just a easy one-click way to install this, go with this option. But this might be a bit slower. It might be more bloated and less customizable. So, I prefer to install this manually, and that's what we're going to go over today.

Now, these are just some short instructions. So, let's scroll a bit down and look at the full installation guide. So note that here this requires Python and don't worry I'm going to show you how to install these if you don't have it and it requires a compatible GPU which is at least this or newer. So here are the instructions for these types of GPUs or if you have the 50 generation then here are the installation instructions. So for me I'm going to go over how to install this using these instructions.

Now the first step is we need to get clone this repository which requires you to have git installed on your computer. If you don't, here's how to install it. If you already have git installed, feel free to skip to the next section.

So all we got to do is download the latest release for whatever operating system you're using. So I'm using Windows, so I'm just going to click on download for Windows. I'm running 64-bit, so I'm going to click on this to download. And it's now downloading this .exe file. So once that's completed, all we got to do is open that exe file and then follow the steps. So I'm going to click on next. I'm just going to go with the default install location which is program files/get. So I'll click next for that and then I'm just going to leave this at the default and then I'm going to click next again and click next here. We're just going to use the default settings for all of these. There's a lot of settings that you need to go through. So I'm just going to click next for all of these. All right. And then it should go ahead and install all the files. So this might take a few minutes. Perfect. So now we have git installed.

All right, assuming you have git installed, the next step is to choose where on my computer I want this installed. So let's say I just want to clone this on my desktop. Well, I just need to open desktop and then at the top here, type in cmd to open up my desktop in command prompt as you can see here. Now the next step is to copy this line and paste it in here. So this is basically going to clone all the files and folders that you see over here into a folder on your desktop. So let me just drag this over here. And if I open this, notice that it has the same files and folders that you see in this GitHub repo.

All right, the next step is we need to change the directory to this new folder because right now we are still in desktop. We need to go one folder in into this one 12GP folder. So let me copy this line and then paste it in here. So right now we are within this new W2GP folder.

The next step is we need to use to create a virtual environment called W2GP. And this is going to use Python 3.10.9. Now this does require you to have installed on your computer. If you don't, here's how to install it. If you already have this installed, feel free to skip to the next section.

Now I'm just on anaconda.com and actually what I'm going to do is install Minionda. This is a minimalist version of Anaconda. If you install the full Anaconda, it installs a lot of packages and dependencies that you might not need. This just takes up more room on your computer. And of course, the installation time is a bit longer. But with Minionda, it's just a barebones package and you can always install additional packages and dependencies afterwards. So, I'm going to click on latest Minionda installer links by Python version. And I'm using Windows, so I'm going to install one of these. Now, for free and open source AI tools, usually they do not support Python 3.12. So, it's better to install the Python 3.11 version. So, I'm going to click on this, which should download an .exe file to your computer. Once it's finished downloading, simply doubleclick on this and then follow the steps to complete the installation. So, I'm going to click next and then agree. And then let's set this to all users. I'm going to go with the default destination folder. And then I'm going to check this as well. Clear the package cache upon completion. This just gives you back some more disk space without affecting functionality. All right. Once that's completed, let's click next and then we are finished.

Now, we aren't done yet. So, if you open up the command prompt and you type in- version, you're still going to see that is not recognized. This is because we haven't added Anaconda to our path yet. So, let's exit out of this. And then to add it to our path, we simply search for this function, edit the system environment variables. We're going to click on this and then click on environment variables and then click on the one that says path and then click edit. And here's where you add in the path of Anaconda. So it depends where you installed Anaconda. For me, I installed it in program data. So it's going to be in program data/min. And then if I doubleclick on scripts, you can see that is here. So this is the folder we want to paste in. So, I'm going to rightclick on this and then copy as path and then back in the environment variables window, I'm going to click new and then paste in the path here and then click okay and then okay and then okay again. Now, if you open up command prompt again and type in- version, you should see that we are running 24.5.0. So, this shows that we have successfully installed Anaconda.

All right, assuming you do have installed on your computer, let's copy this line and then paste it in here. So again, this is just creating a new virtual environment called W2GP, which uses Python 3.10. Now, the point of creating a virtual environment is think of it as like a separate hard drive that houses all the packages and dependencies that are required for this AI tool to work. And this is important because you don't want any of this to conflict with existing packages or dependencies that you have on your computer.

All right. So let's press Y to proceed. All right. Afterwards, the next step is we need to use to activate the virtual environment which we named 12GP. So let's copy this and then paste it in here. So now you can see that we are within our virtual environment because we have the name within parenthesis at the start of the line.

All right, the next step is we need to install PyTorch. Now, this does require you to have CUDA 12.4 or a later version on your system because it is backwards compatible. So, to check what version of CUDA you have, simply open up another command prompt window and then type in NVCC- version. And notice that it says here, I have CUDA 12.4. So, perfect. Now, if you have an older version than CUDA 12.4, then this line might not work. and you'll need to go to the PyTorch page and select the appropriate CUDA that you have. So, for example, let's say you're using Windows and you have CUDA 11.8, then you'll need to use this CU118 instead of CU24. Anyways, I do have 12.4, so I'm going to copy this line and then paste it in here.

All right. So, this is going to install Torch, Torch Vision, and Torch Audio, which takes like over 2 GB. All right. Now, after installing everything, you should see this line again with no error message.

So, the next step is to pip install all the requirements that are listed in this requirements.ext file. So, let me just open this really quickly. And notice that it's a really long list of dependencies that it needs to run. So, let's copy this line and then paste it in here. Now, again, because this is such a long list, it's going to take a while to download everything. All right. So, it's going to proceed to install a really long list of packages. But if all goes well, you should see this line with no error messages.

Now, we can already end it there and start running this. But here are some additional performance optimizations. If we install Sage 2 attention, we can generate videos 40% faster. So, let's go ahead and do this. First, we need to pip install Triton for Windows. So, let me copy this line and paste it in here.

All right. So, it says successfully installed Triton. Perfect. The next step is we need to install this Sage Attention Wheel. Now, even though it says CUDA 12.6, notice that CUDA is backwards compatible. So, it still works even if I have CUDA 12.4. Anyways, let me copy this and then paste it in here. Perfect. So, it has successfully installed Sage Attention. And that's pretty much it. We are now done with installing this.

So let me exit out of this window and show you how you can start this from scratch the next day. So let's double click into our one 2GP folder. And then at the top here, let's type in cmd to open this folder up in command prompt. And then afterwards, we need to use cond to activate the virtual environment which we created which is called one togp. And after that's done, if you look at this GitHub repo, if you only want to do text to video, you can run this line. But for us, for Hunan avatar, it works best if we upload an image as a reference. So we're going to use image to video, in which case we need to run this line. So let me copy this and paste it in here.

Now, for the first time you run this, you're going to have to wait a few more extra minutes for it to download this thing. But afterwards, you should see this URL. So simply hold control and click on this link and it would open up this 1GP interface in your browser. Now even though this is using your internet browser, it's not actually online. So this is completely offline. This is a local URL. There's actually a ton of options here like one image to video and also phantom first frame last frame. There's also sky reels and LTX video, all of which I've gone over on my channel before. Plus, we have this Hunyan video avatar down here. So, let me click on this.

Now, there are a few settings you should configure before you start running this. So, let's go over configuration. And here, there's a ton of different settings which you can play around with, but we're going to click on this performance tab. And you can change this up depending on how much VRAMm you have. For example, over here, this VAE tiling, if you enable this, it will be slower, but it will reduce the VRAMm requirements. For me, because I do have enough VRAM, I'm going to disable this. And then we also have this boost option, which gives you around a 10% speed up without losing quality. So, yes, let's go with this. And then here for the profile, you can select a ton of these profiles. So, the default is this one if you have low RAM and low VRAM. And here it lists out the minimum amount that you should have. So if you have at least 32 GB of RAM and 12 GB of VRAM, you should select this. If you have 24 GB of VRAM, for example, like if you have a 4090, then you should select profile 3. Conversely, if you have a ton of RAM but low VRAM, for example, if you have at least 48 GB of RAM and only 12 GB of VRAM, then you would select this and so on and so forth. These are pretty self-explanatory. And then after you've selected the best options for your device, you can click on apply changes.

All right. So, back here in video generator, again, we're going to use Hunen video avatar. And here's where we upload a reference image. And here is where we upload the audio for it to sync to. And here's where we will enter a prompt to describe the scene further. Here is the aspect ratio. So, you can choose from a ton of different aspect ratios. The nice thing about Hunyan video avatar is it supports up to even 1080p resolution it seems like. And then over here is basically the length of your video. So it's at roughly 25 frames per second. So if you want to set it to 5 seconds then it would be like 125. And this can go all the way up to 41 frames which would be roughly 16 seconds which is more than enough. And then over here is the number of inference steps. This is basically how many steps the AI model should go through when generating your video. In general, the more steps you have, the higher quality and more consistent your video will be. It will also contain fewer errors. However, at a certain point, like if you drag this to 100 steps, for example, that's going to be way too much. You're going to get diminishing returns somewhere around like 30 to 40 steps. And if you set this lower, it's going to run faster, but at the sacrifice of some quality.

All right. And then over here we also have this advanced mode toggle. And if we look at this, here is where you can select the seed or basically the starting point of your generation. Here's the number of videos per prompt. Here is the guidance scale. So how literally it should follow your prompt. And then you can also add luras over here if you want. If you're not familiar with the term Laura, these are basically like fine-tuned models or special effects that you can add on top of your video generator if you want to generate, for example, a particular effect or motion or character. Now, the most important tab in this advanced mode setting is this speed tab. So, here is where you can turn on or off tcash. This makes your generation even faster. So, for example, you can turn this on to speed up the video generation by up to 2.5x, but at the sacrifice of some quality. So, I think around like a 2x speed up would be the right balance between speed and quality. So, let's select that. And notice that this does also consume VRAM. So, make sure you do have enough VRAM before you turn this on. And we also have like some upsampling settings here and quality and miscellaneous. Feel free to experiment and play around with these.

And that's pretty much it. So, let's start generating a video right now. So, for example, I'm going to upload this image. And for the audio, let's use this simple one.

"Did you hear that? I'm scared. It sounded like something moving in the dark."

For the prompt, I'm just going to write she is talking. Let's click generate and see what that gives us.

Now, when you run this for the first time, notice that it also needs to download this Hungyian video avatar model, which you can see down here. So, this is going to be quite a large file. Again, it's like several gigabytes, so it's going to take a while.

And here's what we get. So, notice that for this local tool, it does not contain a watermark.

"Did you hear that? I'm scared. It sounded like something moving in the dark."

Not the fastest video generator out there, but this is because it includes audio and it needs to also lip-sync the video. So, there are a few extra steps in this process.

In a nutshell, that is how you can install and run Hunyan Video Avatar locally on your computer, even if you have low VRAM. So, that sums up my review and installation tutorial of Hunyan Video Avatar. This is the first available open-source tool that can animate any scene, including a person's full body, with an audio clip. It can even handle emotions and different expressions and different animation styles. So, it's a super powerful tool. And best of all, you can run it on as low as 10 GB of VRAM. So, definitely try this out. I'll link to everything in the description below. Let me know what you think of this. And if you encounter any errors during the installation, welcome to paste your error message in the comments below, and I'll try to help you troubleshoot as much as possible. As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week, I can't possibly cover everything on my YouTube channel. So to really stay up tod date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching and I'll see you in the next one.