Transcription
Hello everybody. Let's get started with a little teaser video while people join. The person talking to you right now is an AI avatar. 15 seconds of footage, that's all it took. Watch what it does with that.
This is for the people who show up on camera again and again because that's how they lead, sell, teach. Most other models are image-based, so they don't truly preserve who you are. They feel missing details with hallucinations and over time the avatar shifts. And by video 10, it looks more like your cousin than you.
Avatar 5 learns directly from your reference video. Every frame, every movement. It doesn't just read your face, it watches how you move, your rhythm, your expressions. In blind evaluations, Avatar 5 is rated number one for real human talking videos, 1/10 of the cost, 10 times the generation length, zero identity shift. You can one-shot an entire 10-minute script and walk away.
If your face is your brand, if personalized video is how you sell, if you or your team needs content at scale, this was built for you. Avatar 5, one recording, any outfit, any setting, anywhere. Your identity consistent, permanent everywhere it needs to be. Available now in HeyGen. Awesome.
So, right there, that was our launch video for Avatar 5. Everything you saw in that video was created with the Avatar 5 model. My name's Adam. I'm a product manager here at HeyGen for our avatars and voice, and I'm excited to be co-hosting today with Nick, our head of AIS, to talk all things Avatar 5.
Yeah. Great. Great to be here, Adam. Yeah, good morning, good afternoon, good evening to everyone around the world here for our webinar. A pleasure to be here again after a long time, Adam, actually. I don't know when our last webinar was for for the community and from all of our users or even new people who are joining today for the first time. I'm still not sure if we're doing this. Like are we still explaining what HeyGen is and what HeyGen does or are we just jumping into Avatar 5 right away, Adam?
I think it can't hurt to give a quick recap. Nick, want to tell everyone briefly what is HeyGen?
>> [snorts] >> We can definitely do this. HeyGen itself AI video platform to really make like our mission is really to make visual storytelling accessible to everyone. We are specially focusing on business communication videos and this is also something what we want to show today with Avatar 5 because we did the next step into the right direction for character consistency, but also like stabilize to that you really can come to HeyGen to have a reliable system. We have way many features and also products, so have a look to all of our webinars in our community. Join our community to understand what's going on that you are up to date. But let's Adam, you're right. Let's focus today on Avatar 5. I think our team is doing a great job to really introduce everyone to all of the new features. Let's let's kick it off, I would say.
All right, let's get into it. So, today we'll be talking about what is Avatar 5, what it means for you as a HeyGen user. We'll go through a live demo of creating a digital twin and creating your first Avatar 5 video. Nick will go through a bunch of use cases, a bunch of awesome examples, and share some tips for how to get the best out of it. And then we'll make sure to save some time at the end for any questions that people have.
So, first, what is Avatar 5? It is the model that was used to create that video you just watched. What it does is it allows you to create realistic videos of yourself that look, sound, and move just like you with fully customizable appearance. So, you can put yourself in any outfit, place, or pose while still looking and moving authentically like yourself. And it really works through takes in three ingredients. It has your photo, your voice, and then your video reference. And then what it does is it uses your video reference as context to animate your photo when speaking any script so that it always looks and moves like you. So, rather than simply having a single frame to animate off of and in that case it's making up a lot of the motion, it's looking at your actual video and using that so that your motion and the way you move stays true to how you actually move.
Yeah. Maybe when I can just say one one sentence to this because I I still think that a lot of our users still are actually aware of Avatar 4. So, the big difference here is really that we know actually merging this reference motion like the input video with your image to really like nail down the character consistency while we're doing a longer like type video. And I think this was really missing before because I mean, when you come from the AI part, it's like very hard because we as like the company and like our model needs to predict how you would act, how you would look like, how your teeth are looking like, how your facial expressions are looking if we just have an image because the model is not just trained on on your, right? Like there is like a a big big model underlying this and now we are actually building kind of a fine-tuned model based on our yeah, base model that we really nail down like your personality. I just wanted to like bring this together to really like yeah, nail down the difference. I'm I hope this is clear enough, but again, we have our community team here behind us. So, if you have any questions, please please let us know and we will try to answer later on.
Adam, sorry.
No problem. Thank you. And then another important quality of Avatar 5 like our other avatar models is that it supports long duration and it's very stable. So, what this means is you can give it very long scripts and you can consistently get good results. And so, what this means is it's really great for people who have high-volume repeat content creation needs where they need to realistically resemble themselves across different settings and outfits. We see you know, some of our core users are really knowledge experts or business owners who have important knowledge ideas to share, but they're not traditional video editors and they're they don't have time to you know, spend a lot of time to create and then edit videos. And so, with Avatar 5, you can give it a script and you know, convert a document or a script into that video in one shot it. So, it really allows you to save time while scaling your creation to help you grow your business.
And if you've been a part of HeyGen you know, in the past historically, you know, you're probably aware we we offered photo avatars and video avatars. And you know, as Nick mentioned previously, the problem with photo avatars is they weren't realistic enough when using it you know, as your own digital twin avatar because it didn't have your video context. So, it didn't actually know how you move, it didn't know what your mouth looks like you know, when you open it. And so, if you're trying to create videos you know, for clients or for an audience who knows you, the photo avatar didn't really work. And on the other hand, we also offered a video avatar product. And the video avatar, you would record you know, the two-minute video and then we would animate based off that. That offers realism, but the issue is is you couldn't customize your appearance. So, anytime you wanted to change your outfit or your setting, your background, you'd have to re-film that look. So, with Avatar 5, we combined the best of both worlds. You can customize your appearance to whatever you want. So, you can you know, put yourself to you know, different countries or you know, in a professional studio or a modern office. And at the same time, no matter how you customize your appearance, it's always learning from your video motion. So, all of those videos will look and move realistically like you. And and again, it has the long form stability, so you can you know, generate an hour-long video you know, without without lots of iteration or edits. You give it a script and just you know, enter generate.
And in blind evaluations, we see that the Avatar 5 model is rated best in the world for real human talking videos. So, there's a you know, a lot of incredible models that have come out that can do dynamic motion, etc. But if your need is sitting in front of the camera, communicating a message, not only is Avatar 5 the most cost-effective model, supports the longest duration, and is going to be the most stable, it's also just the best quality in terms of lip sync.
So, the process to create your digital twin and create an Avatar 5 video is really simple. All it takes is you record 15 seconds and then there are some things you can optionally do like do a standalone voice clone. You can customize your appearance to you know, further improve the quality, but the whole process takes you know, a minute. And let's go through it live now. I think I'm going to have to reshare my screen. One sec.
That's okay. All right. So this what you're seeing here, this is a avatar I previously created and this is an AI generated look. Um so this came out of let's see, my original photo was uh Not uh so this was my original recording. I did it in a phone booth and then from that recording I'm able to generate any look I want. Um but let's just go through the entire process um from scratch now. So first what I do is I go to the avatar tab, quick create and I'm going to tap this clone me option. And this is the option if I want to create an avatar of myself. So as you can see I'm just, you know, in my bedroom. Not no special setup. And now I'm just going to record this script, uh read it out live for about 15 seconds. And this process will capture all the ingredients I need, gets my appearance, gets my motion, gets my voice and it also includes a consent. And the reason we do consent is to ensure that only you can create with your identity.
Hey there, I'm speaking with lots of energy while staying natural and confident. This helps HeyGen capture my voice, my expressions and my motion so my avatar can behave just like me in any video. For safety purposes my unique code is 95.
So there, you know, I just read out that script. You can actually say whatever you want during that script. So if you, you know, have a specific message that feels more natural to say, you can say that as well. The only important part is that you read out that code at the end. So the for safety purposes my unique code is whatever number you have. Uh you do need to read that part out for this process to work successfully. Otherwise, um you know, feel free to ad-lib that script.
Yeah and maybe what what one thing to add or two things is for sure like security is the most important thing at HeyGen. This is the reason why we have this consent. It was previously working a bit different because we were able to upload like the video footage itself and then the consent separately. Um we fixed this with a one video entry point now and it's still like has the same security um benchmarks here. Um I I just want to make sure you are not able to upload any other video from any other person and just like work with their likeness. This is not how it works here at HeyGen. The second thing is Adam that I think is also very important to to have a smoother like entry point is yes, you could do it with a webcam. And if so, make sure we actually have a good lightning in your room. We had actually natural lightning. There were no big shadows. You were in the right distance of the webcam. You had a clear actually voice. So that's very helpful at the end for the model, especially for the image model to actually put you in a different spot without like drifting the similarity, okay? The thing is if you have actually a recording setup where it's very dark or like you have a lot of shadows from the right or left, it's very hard sometimes for the the model itself to create a good image or like a new image from you by keeping the similarity. I just wanted to add this. Um but you can also able to upload like a for sure a video footage from your own then you still need to do the consent process. Um yeah, I just wanted to mention this.
Totally. Um and we will get uh in a sec to the customizing your appearance part um and that's where I think some of that will play in in terms of the lighting from your your footage. But one thing to note is um you know, you don't need to be in the perfect setup when you record. You want to make sure your face is clear so it can, you know, see your face clearly, get those details. Um but if you later want to, you know, change your your background like I can change this from my bedroom to an office or a professional studio. Um and I can even upload different photos of myself where I like how I look in that photo um to use that as kind of my identity reference. But we'll yeah, get to that in a second.
So so I just recorded my 15 seconds and now I automatically get a voice clone from that footage which I can listen to here.
Hi, I'm excited you're here. This is your voice clone preview. Take a quick look.
So it's not bad. Um but we highly recommend, you know, if you're not fully satisfied with that that clone, to record a standalone voice clone. And this voice clone you just are recording your voice, no motion so it doesn't matter what you're doing or, you know, how you're expressing. The goal is is to really just um you know, vocalize and enunciate the script clearly. Uh we often see that the standalone voice clone does result in higher quality. Part of this is just, you know, it's easier to focus on just recording the voice. Um whereas for the 15-second, you know, motion recording you're also thinking about what you're doing with your hands and face. Um but it's also just in terms of the best setup to record your voice is often different from the best setup to record your your motion where with voice you might want to be a little bit closer. Um you know, make sure you're in a very quiet room. Um whereas with the recording, you know, you can be outside or, you know, really wherever. Um one other thing to note is it does work well with, you know, good quality laptop mic. But if you don't have a great laptop mic, would highly recommend to record on your smartphone and upload the audio. Um you know, if you have just like a iPhone, the voice memos app works great. Uh you can turn on the lossless setting. It will help quality a bit more. Um and then you can just read the script while recording from your phone and upload the the audio file that way. But if you have a good new laptop, um this should work as well.
So I will I'll just go through this real quick as an example. Oh my gosh, I can hardly sit still right now because I just discovered something that feels like a total game changer. For years I've wanted to share what I know, teaching, explaining ideas, helping people learn faster. But creating content takes forever. Recording, editing, fixing mistakes. It's exhausting. But then I found HeyGen. With HeyGen I can create AI videos that look professional without needing a studio. I can turn one script into videos in any language. Um etc. And then I would, you know, read the rest of the script.
Couple other notes that we'll call out is if you have an accent it can help to use a script that really includes your accent markers. Um so you can, you know, you can give this script to chat GPT and be like update the script to include words that really captures my specific accent. Um cuz when you're doing the voice clone, if you want to get um if you want to really maintain your accent well, it's important that you can capture, you know, kind of the full range of your accent from this clone. Um so here I'll I'll end my clone. Because I kept talking this one might be weird. But uh you can preview it here. By default the background sound is removed which we recommend to keep on unless like you're recording in a specific setting like outside and you want to keep the outside noise to match the uh appearance of the look. Um so then we we clone your voice with a few different engines. Uh we offer 11 labs which, you know, offers some of the best quality. And we also offer some other engines like fish. Uh we see generally 11 labs works best for people. Um this is the 11 labs V2 engine which is very stable. V3 can often be more expressive um but a little bit more unstable and then
>> Hi, I'm excited you're here. This is your voice clone preview. Take a quick listen.
All right, that's pretty good. Um and then fish we see is often good for people with uh English who are English language non-American accents. So if you have like a British accent or Australian accent um this one might be best for you. So anyway now I I choose the voice. Um and now I'm all set here. So now what's happened is I've created my digital twin. I have my reference footage. I have my voice clone which I can listen to here.
And Hi, I'm excited you're here. This is your and we automatically generate three looks just really as um kind of a demo of this is what you can generate. But then you can also generate your own. So these three looks came from that 15-second footage I just recorded. Um and I'm able to generate new ones as well. I can either do that via custom prompting. So it could be like Adam standing in front of a construction site wearing a hard hat. Um and then I can also just one-click remix uh these templates. And as you can see or here's Nick. Uh these templates are when you remix them, you get the same look, but featuring your avatar. And all these are really designed to work well with our avatar model, where they're, you know, in this kind of half-body framing, where the face is, you know, clearly visible, facing the camera. Um so, when using creating with these looks or looks that are, you know, in this style is how you can generally get the best results.
Yeah, and I and I also think this is one of the magic places here for Avatar 5 and for the new pipeline itself. Because, again, the the the challenge before with avatars was that you have a video input and you need to keep the background as it is. And then with all of the mid image tools that went out, we always had some other issues, okay? So, maybe the background was looking good, but now your similarity is off. And I think when we had this idea for this like remix template for like this inspiration, it is really like to to really also help all of you to just get a bit more inspiration in which places you can be actually now. Because, if you could also scroll a bit down um Adam here, it's like also we have a lot of different um templates for different situations, but also, as you can see, sometimes we are not just centering you in the middle, we are also centering you a bit to the left or to the right. That you have a bit more space for some overlays when you go to AI Studio at the end. So, and there's a lot of options like with different clothing, lighting, um please let us know if there's any remix template or anything missing, but for sure you could also prompt as Adam just did. Um and then really nail down your look. And I think this is a big part. Um and we got really good feedback already um from everything related to this yeah, inspiration pipeline here, but we are always love to hear a bit more feedback um to improve um because we want to. Um Are you Are you going in order to share like the personal model Adam or Uh yes.
Yeah. So, so as you generate your looks, if you're not fully satisfied with the results you're getting, what we recommend you do is you can change your base look. So, this look here, this automatically comes from my 15 seconds of footage, but what I can do as well is just upload other photos of me. Um so, here's, you know, a bunch of different photos of me, and I can choose a different photo that I prefer of myself to use as my really like reference look when generating new photos. So, when I custom prompt or even use these remix templates, uh those will be based off of whichever base look I choose. Um and then we do offer now a new feature where you can improve even above that. So, first thing we'd recommend is uh you can just experiment with a different base look. Uh like here, I'll be like, you know, I could actually just try a few different templates with a few different looks. And sometimes it's hard to predict which reference look's going to produce the best results. So, my advice would be to just kind of choose some of your favorite ones and uh and remix a few options. And then once you find the one you like, you can generally stick with that one. Um what you can also do to get even better quality and really just make it more consistently good is you can train a personal model, which and this involves uploading um minimum 10, you can do up to 80, but we recommend uh you know, 30 plus for best results. You can select all these and then you can train your model. So, we recommend to train on your uploaded photos, uh not your generated. You can include generated photos when training. Um and you know, how you would do that is like if there's certain generated photos you really like how you look, um you know, those are fine to include, but to get a model that's going to be most realistic of yourself, you can uh just select all on the uploaded looks and then tap train. Um so, now this like 10 minutes or so.
Sorry, Nick gone.
Yeah, and one thing I would also add here is like for sure it is helpful for the model, but also depends on your use case, but it is helpful to also add like some side angles or like high angles or like lower angles to just help the model to really understand where you also like look from different things, because one thing I will also let everyone know um about a new like upcoming feature will be definitely like also side angle. It means like if we do not have a clear understanding how you look from the side, like I said with the video previously, all right? So, we need to predict how you look from the side. It's very hard. It could be off. Um so, it would be helpful for the model to also upload these type of photos here to really like nail down the similarity of you.
Totally. Yeah, like a good rule of thumb, uh especially when you're thinking of what photos should you include when training this model is anything that you could potentially want to, you know, generate, ideally you include a reference of what that actually looks like. So, if you want if you're going to be generating, you know, side angles, ideally you include some side angles of yourself. If you're going to be generating one, you know, of you farther back, it's, you know, important to have those full-body looks. If you need specific lighting, it can help to even show how you look in specific lighting or specific expressions, like angry or happy, um, you know, it can help to include those as well. Um, and, you know, the more it has, the less it has to make up.
Exactly. So, the like the difference is also between like the instant pipeline for the images and also the personal model is like the personal model will really like create a fine-tuned model based on your looks with all of the footages, so it will take around 15 to 20 minutes. The other way that uh Adam was showing before is really like instantly. So, we take your image as it is from your base look, just press a remix template, as you can show now or as you can see now. Um and then for sure, if you have a very good base image, it will definitely nail down your similarity, as you can see here in some examples. But if you really want to have it for higher like stage and you really need to nail it down, we would actually recommend to do the personal model for you.
Yeah. Yeah. So, as you can see, like these ones were just from the instant base look, you know, experimenting with some different base looks. A lot of these are good. I'd say they're not perfect, um but they're very close and like, you know, some of these I I would use, but I'm waiting for now that photo model to finish training. I don't think we'll we'll wait for this demo. Um but once that is, then you can just get, you know, consistently slightly better results. Um but anyway, let's get back to
>> Yeah, what what happens if I have a look now, Adam? What's what's the next step? So, I I found an image that I like and I want to create a video. What is What is my next step?
So, the next step is the easiest part. You pick the look you like, and then you enter in your script. Um and you can either, you know, paste in text. Hey, this is an example. Uh you can also uh record audio if you prefer, um or upload an audio file. And we also support voice mirroring. Um so, if you, you know, have a different audio file and then you want it to be um delivered with your voice, that's supported as well, but for now we'll just show an example script. Here, I can preview it right there.
To bring about change, you must not be afraid to take the first step.
Sounds good. By default, selected to five, and now I just tap generate.
So, once you create your digital twin um, your your real human avatar, a Avatar 5 will be default selected across our different workflows. So, from the shortcut, whenever you select your Avatar, 5 should be selected. Uh from AI Studio, you can also create with Avatar 5, and all you need to do is select this avatar and 5 should be selected. So, the only thing basically there's no special setup from there, you just enter whatever script and tap generate. Um same is true for video agent. Uh when you use with video agent, you don't manually select the model, but when you select your real human avatar, um and it doesn't matter whether I choose the video look that I got from my raw footage or any of the generated looks I created, all of them will be using Avatar 5 and all of them are still using that video footage as the reference to ensure it looks realistic and, you know, realistically like me.
Yeah, we have one question and maybe we can go to AI Studio one more time with your um image. The question itself from one of our viewers right now is like, do inputs like be longer than 15 seconds to create a better quality? So, I want to say two things here. First of all, for this technology Avatar 5, there is no difference if we have like this 15 seconds as recommended or like 1 or 2 minute. The reason is how the technology works. But, what could make a difference is the following. If you go to AI Studio, you select your image, and on the right side, next to motion engine, there is like a yeah button where you can click and here you could actually have a different video as the I would say driving reference. So why it matters is even if it's not significant, but what you could do is like you could upload like two different type of videos. Like let's say one is a bit more expressive where you raise your eyebrows a bit more, one is a bit more professional, a bit more calm professional speaking style and you use this for the different boxes. As you can see on the left side, we always can create like a first box script and then we can add a new scene. So imagine this for the opening scene on the left side, I will welcome normally everyone. So I will actually use my image, but the driving motion will be like bit more expressive because I want to be a bit more positive. So then my talking style starts and [clears throat] from the second like at scene box, I will actually just use my I would say more professional, calm, control speaking motion video, which will actually reduce a bit the expressiveness of the avatar at the end. So I just wanted to say this because it was a good fit here, Adam. Um sorry to interrupt you.
Oh no, not at all. Thank you. Yeah, that's a great call out that you know, there's a few ways that you can influence the output of this video. Uh one, and this is true for all our avatar models, one of the the most important way it gets influenced is via the script and via the audio, where these avatar models are really audio driven models. So when you sound excited, like when you're delivering a very excited message and you're sounding super excited, that's going to make your avatar look super excited. And you know, if you give it uh really sad sounding script and you know, it's really sad and depressed, like that's going to make your avatar look sound sad and depressed as it delivers that. Where the models are really built to realistically adapt any script and you know, deliver it as you would. Um and as Nick mentioned, in addition to that, you can also influence, you know, the output uh the output visuals via your motion style. Um so yeah, a couple a couple examples where this does help. Uh one, which Nick mentioned, is having a really expressive reference when you're, you know, delivering your most expressive scenes versus kind of a more confident, calm, professional one, you know, for those use cases. Um another case where we see it can help as well is if you are creating uh side angle views. Um and often we see some of like the highest quality videos simulate a multi-camera setup by cutting between front-facing shot, 45° shot, back to front-facing. Um so for that like 45° shot, it can help if you choose a motion style that's also recorded at 45°. Um so just basically how that would work is you'd add a new look and you could just record, you know, another 15 seconds. Um but you know, facing 45°. It's definitely not critical. Um you know, the the biggest way to influence the motion is really the audio. So it's critical that you have a great, you know, high quality audio um and that, you know, it sounds, you know, has the delivery that you're going for as that will really impact the output the most, but the motion style, you know, can help as well. Um cool.
So that is the quick demo of the quick create. Let's see an example video output. Looks like that one's still generating.
Um Welcome everyone to our exciting webinar. We're thrilled to have you join us today. Get ready for an engaging session filled with valuable insights and discussions. Let's dive in and make the most of our time together.
Awesome. So Nick will be going through a ton more examples, showing comparisons of um you know, five versus our previous models, but as you can see very naturally delivering the script. Uh you know, it's expressive uh with natural hand motion and you know, we'll let you be the judge, but I'd say it looks uh a lot like me.
So to quick just recap some of the tips for recording, uh be extra expressive, especially with the voice cuz when you do the voice cloning, it can sometimes mute it a bit uh during the capture. You don't have to worry about how you look, you can later customize your appearance. It's just important that there's good lighting so that it can actually see your face clearly. Um and make sure that you capture a great voice. So if your 15-second footage doesn't produce a great voice clone, you can do the standalone voice clone step right after. And if you don't have a great laptop mic, use your smartphone. Um for generating the looks, try different face images. You know, if if the ones you get off the bat don't really look like you or you're just not fully satisfied, upload other photos of yourself that you like, try with those, and then to get the best results, you can train your personal model on 10 to 30-plus photos. Um and then one other call out I'll just mention as this has become my own personal favorite way of creating with avatar five is you can use it with our video agent and with our API. So our video agent calls our API can call the video agent and that uses avatar five um when you're creating with your real human avatar. Um so I like having my Claude, you know, pull from all the data sources and now I'm using my digital twin with avatar five to communicate weekly updates to my team with literally like zero time spent creating the videos and editing the videos. Uh so we just had an awesome launch with our CLI uh earlier this week. We checked that out. Um and avatar five is not yet supported publicly via the API except if you are creating with your real human avatar via video agent. So that is something you can try tonight. Um cool.
Finally, uh to recap some of the benefits of So why now you've seen, you know, what is avatar [snorts] five, how does it work, why should you use it? Um you know, just to recap some of the benefits, saves a lot of time versus traditional filming, like the video that um you know, created from our launch video or even just the video that you just saw me create. You know, I was able to record 15 seconds from my bedroom and produce that. Um and it also helps you always look and sound your best. Um so when you're having a bad hair day or, you know, you're not feeling great, you know, avatar five is awesome cuz you don't have to do your makeup, you don't have to, you know, get dressed, get ready to film. You can just use your avatar, enter a script. Um so saves time, but also, you know, helps helps produce better quality. And we we hear from users often that their AI, you know, avatar produced videos perform better on social um you know, than their their own filmed videos. And a lot of the reason why is the avatar videos, you know, they come out very you're clearly articulating. You know, there's no stuttering or stumbling over words, your accent is very clear to hear. Um so especially if you're like not a native English speaker speaking in English, uh using the avatar models can, you know, help make you, you know, easily understandable. So yeah, we'd definitely uh check it out and you know, um you know, some people are concerned like how do AI videos perform on social? I think because our avatar models are really designed to authentically, you know, be your appearance. These feet, you know, come off as authentic, not as, you know, just AI slop. Um and so people generally do really well when, you know, using HeyGen for social. Um and then finally, avatar five versus other AI video platforms, really the top quality benefits are one, the likeness. You know, with avatar five, it will truly look like you, whereas most other AI video platforms are just based off a photo. Um so the motion, you know, it's making it all up, but someone who knows you would know that's not actually you. Uh two is the long duration. Um you're able to one-shot videos up to an hour long. Um and you can generate scenes up to three minutes, whereas, you know, with many other models, it's limited to like 10, 15, 20-second generations. Uh three is the cost-effectiveness. It's 1/12 the cost of, you know, Synthesia which is another awesome model. Um so, you know, it's great for high-volume repeat content needs that aren't going to break your bank. And then finally, it's stable. Um you know, one of the I think key selling points of HeyGen and why people love it is that it just works. You know, with most models, you're trying a bunch of times, you're hoping you get one good result. With HeyGen, you can enter a script, generate, and consistently get, you know, good usable outputs. Um so instead of, you know, 5, 10% of the time it works, it's like 95% of the time.
Yeah, that's that's definitely true. Uh but I still do want to point out that for sure, we we still know where some weaknesses even for our models are, so we are still improving. This is also, and I just want to point this out one more time. Like all of you who are using HeyGen on a daily base or you and you, you have your first experience or you just use it for one project, like leave the feedback in the community on social um or even raise this on customer support or even in the feedback um buttons on our UI or like in our interface. Like it's the only way we could really like face um some serious or like things or or just challenges from you as a user that we can improve. It is always very helpful that we can optimize our model. So yes, I think avatar five is a great product. Is it perfect? Not yet. Um so please share more feedback with us. Um Adam, I also have one thing, maybe I can just ask everyone here right away. I have the question and I see it more and more, is actually if the launch video we showed at the beginning was really Avatar 5. And I want to point out that yes, it was. And maybe I can also offer, or maybe it's a question, maybe the community team can can handle this after this webinar. I'm actually happy to show everyone how I did the launch video. Um, everyone could actually understand. If everyone wants to join me, I'm happy to support all of the users who want to join me. Um, like free credits to regenerate the same video uh with your own likeness. Um, we can do a live Q&A. We could do kind of a Twitch uh live stream and I would just guide everyone through. Um, because there is no fake. Like we did not use any other tool. The only thing we did at the end, for sure, we used like a different editing tool to merge it together. This is true. But the A-roll itself, um, it's actually just Avatar 5. Um, so yeah, maybe the community team can hand over here and and if there is any like request, maybe we can do it. Um, so yeah, perfect. Adam, thank you.
>> a lot of people would would love to join that. Um, I think yeah, people are love to see how our launch videos are made. So, yeah, I'm sure people will. I mean, even if it's just 10 at the end, >> [laughter] >> I would be happy to talk about it. So, um, I'm I'm down for this, yeah.
So, I know we're running up on time and do want to go through all the examples, so I will skip over a couple slides. Just we'll let some, you know, about some information. It supports all Avatar types, so it's really built for the real human Avatar case, uh but it also supports virtual avatars as well. And we see that it actually improves the quality there, too. Uh currently it's default it's the default model for real human avatars, but if you have a virtual avatar or want to try it, you can access it in AI Studio or the shortcut and you just need to switch to it uh from that model menu. Uh cost-wise, same cost as uh four. Um, so yeah, you can still use, you know, 10 minutes including your base plan and then use the add-on to get more generative credits. We'll skip over this one. Um, you know, we also launched C-Dance, too. And the key thing to note there is Avatar 5 for your core talking scenes, and you can use C-Dance, which is a very powerful, uh very expensive, um, but model that's great for generating like short cinematic up to 15 second clips. Uh so, this is a quick video just showing how you can use those together. Um Let me play that.
We needed a product video, so I let my avatar handle it. First, I mapped out the key message and structure. Then I brought that message to life. This may look like just a cup, but this one changes how you work. My avatar is just the best marketing director. Now, I can turn into whatever my product or service needs. And when it needs someone to represent it, I can be the face of my brand. It's even better than me presenting pitch decks, pitching it to investors.
So, as you can see, all of those clips of Holly, who's um, you know, on our team, talking forward in that black suit, that's Avatar 5. And then some of those other like B-roll scenes of her, you know, on the catwalk or presenting, that's using C-Dance. Um, and so both of these models maintain your digital twin's likeness really well, um, which makes them, you know, great for using with your H and Avatar. Our Avatar model stable can deliver the long message, um, you know, quickly, cost-effectively, and best quality for the static talking scenes. And then you can use C-Dance to really complement it with the engaging visuals, you know, add a hook, um, can can use it for like the multi-speaker portions, etc. Um, so I would highly recommend to, you know, try out both of these and explore how you can, you know, use them together in a video to really create, uh you know, a full AI video production workflow that like no one single tool can create on its own.
Yeah, I think this this slide here will actually show a bit better to help everyone to get a better understanding where really the difference is. Um, for sure, like it is a bit hard sometimes to understand where we are coming from, where the difference is here. Um, but I think if you have a look to this and use Avatar 5 for this more like static talking like upper body or even full body videos, for like the explainer videos, for like to explain a product or whatever, and I will also have some um videos now as a showcase, um, then really like Avatar 5 is the only solution. It's the only stable solution in the market, which is also like very cost-effective. It is how it is. Prove me wrong. Um, for the cinematic part, this is not our use case, okay? So, this is not our main use case here at H and but we are still seeing now with the integration um, to really help everyone to create even more stories, like more B-rolls or even like more cinematic A-rolls, where you can actually go ahead and then just go this direction in Avatar shots like with actually C-Dance. Yes. High cost, um, 50 seconds output is different. You have 720p at the end, uh not 1080p how it is here in this draft. Um, so there is like a trade-off, 100% um, but if you combine both, you could definitely make a very nice video on H and um by keeping the similarity of you.
So, if we go to the next slide, yeah, so I will I will hand over. I will try to uh push everyone through now. The The issue will be that um I prepared this for 2 hours, so if everyone wants to join me for 2 hours, then I'm happy to guide through. No, I'm I'm joking for sure. So, let's let's have a look here one more time this video again. Same person, different outfits, different location.
>> [music] >> So, I think you can see here that anything is possible. Um, we had the question previously if full body avatars are working. I would say yes, it is working. The challenge is the following is that actually, um, we we we are still trying to improve the small faces. So, if the actual character is too small in the frame, you will see like still like kind of a drop-off in the quality itself, which we are trying to improve for the upcoming version. Um, but it is doable, yes, to answer this question.
Okay, so next video is more like a use case, um, how you could actually not put your in the center. You will actually just go on the left side as the template I was showing you and then just create like a overlay on the right side. Um, and we have
A 13-second video here, um, where I'm talking about the topic itself. So, I think, yeah, so this is definitely a video where you can see like also, okay, so I can use, um, Avatar 5 at a script and then still in AI Studio, just add like an overlay and then just like show the video and and both actually together. And for me, actually, the quality is pretty high, um, for here. Yeah, Avatar 5 goes video agent. We can have a look. It's the successor to a legend, but is it worth your money? It keeps the classic look, but adds a customizable LED touch panel for easy muting and level tracking. Pro tip, the stock windscreen is a bit thin. Upgrade to a thicker foam for much better plosive control. The onboard DSP is the real winner here. It's built in the quiet moments. It's built in the quiet moments when nobody is watching, and it starts with one simple choice every single morning, the act of making your bed. As Admiral McRaven famously said, if you want to change the world, you must start by doing the little things right.
And and then out of curiosity, how much editing did you have to do for this? Yeah, zero, exactly. So, you just need to have the idea, you go to video agent, you select your avatar, uh, you select the motion engine, the voice, um, you could select from a template or like a style, and then you actually just type in the prompt, um, choose the ratio if you want to have landscape or even portrait mode, done. Then you can interact with our agent itself if you want to improve something and then just you say, let's go, and we will create the video for you. It will take like 5 to 10 minutes, this is true, but I did nothing for this video. And this is the superpower of video agent and also if you combine this with Avatar 5, it it makes it just even better.
Yeah, here I wanted to show everyone like what's the the main difference between the engines itself. Again, we when we are coming from like Avatar 3 back in the days, you can play the video at without audio. It's not that important. I just wanted to show that the most important part here is really like about Avatar 5 now is actually the script alignment with the gesturing to the audio. It is a bit more expressive than Avatar 4, but it's not in a bad way more expressive. It is just more aligned with the script. So you will also see if you add a way more expressive voice to Avatar 5, it will also act a bit more expressive. If you have just a monotone, which is for sure like the killer of everything to just upload a a boring monotone TTS voice file, then the performance will also decrease. So yeah, voice is a big part of Avatar 5 for sure.
We can see >> Yeah, I'm with we had an issue with Avatar 4 a lot of times with the gesturing, which was also a question here. It was random. It was actually not really 100% aligned. >> external actions. It was wild, sometimes a bit uncontrollable. So we're also solving this with Avatar 5. Again, the script alignment is way more or like better, but also improved compared to Avatar 4. If we go to the next examples, here like a lot of >> AI agents operate in a continuous. We were also showing like if you if you had previously the the like selfie style camera angle, we were actually not able to always keep track with our model to keep the hand where it should be. Uh with Avatar 5, this is also solved. You also see us some like camera motions actually because it's natural um while you're holding the phone. Um here also like if you really have a look to how the script aligned to the audio now, you will see like, okay, now the gesturing >> Your knowledge is the curriculum. Your content is the lesson plan. Your audience is the classroom. When those three align, learning becomes consistent and growth feels calm. You can also see like Avatar 4 had like this very unique um style of putting the the like shoulder really in action. You could see in Avatar 5 it's way more natural. It's way more aligned. Um so it's way more smooth um body motion itself.
Okay, um, yeah, Adam, next one. I think it's the same. Maybe we can skip skip some. I want to show one more thing. Um, this is also one >> Okay, we can skip. We just have here like I get a reminder. We just have 5 minutes. I just want to show like one more thing um about like the upcoming week from research. Um where where we are taking care of of the next like upcoming things and we could just guide very like in the next 5 minutes and then I'm I'm I'm swearing we will end this webinar here. So what we want to improve like if you go to the next slide is definitely like the teeth quality for our smaller faces. Um if we have a look here to how it is right now and how actually we want to solve it, um we should be ready in the next like 2 to 3 weeks to to put this yeah upcoming version in place. Why I'm so transparent about this is because we want to I want to be transparent with everyone. We have nothing to hide here. Um we are just making the model better and better. Um and we are doing this on a very like routine basis. Um the next thing um you will recognize if you >> We shall. It is not that stable if you add a side angle. Um Adam is correct. If you upload a side angle video motion um but also like a side angle image, it should not really look to the camera, but it will sometimes. So we are fixing this um that you could even create like from a A-roll perspective. Like think about professional um video recording, you have the front face, but you also have a kind of side angle camera and then you could just even cut in between to have a very cool like A-roll play um from from different angles. So you can have close-up, you can have side angle, front face. Um so this is what we will improve. Um for the static, we are still working on this. We would love to introduce like a more static uh version means like uh less hand gesture, less facial expressions because there are different cultures, there are different use cases, there are people who have a different understanding about like how expressive I should look in front of the camera. A lot of people and we can see this already are preferring more expressive models, but we still have users who wants to keep it a bit more downsizing. So we are here trying to solve this um where you can just toggle. Okay, you can just toggle less expressive, more expressive, whatever. Um Adam will find a way to actually make it happen. Um and then uh we will solve this. Um one version we also want to have and this is the the yeah, the the first view here is like a yeah, we want to introduce a tuber model. Adam uh was asking me or like us from research to introduce a tuber model um with the same quality as the current product model. I think we are on a good way to also ship this very soon. Um and then also [clears throat] from a lot of our users, but also from Adam, who is again taking care about our product here is actually yeah, a custom prompting model, um which will take a bit more time. I'm very honest here, but I just want to give a preview that we are working on this. Um here in this example >> yes. Absolutely yes. This is the right call. And to everyone who has been here with me through this, thank you so much. This one is.
So as you can see like in the first part, it is a bit excited. We trying to make the second part a bit more confused. Um you [clears throat] will see this thumbs up and like yeah, it it's like yeah, blow a kiss. Um it's it's kind of working. So um yeah, we will try to nail this down. Um Adam always asks this custom prompting because users want to customize their own avatar like I mentioned before. It's very heavily for each culture, each individual mindset. Everyone has a different understanding how they want to like see themself or even their AI character. So custom prompting would solve this because as Adam always is telling me like, hey, I could just say like, okay, be here expressive, here more calm, here show like a a thumbs up or thumbs down. And then yeah, this is just how it works. Um and then the last thing we can show is actually the AI yeah, animated character. >> I'm about to show you Avatar 5 with AI characters. Yeah, next level reality. What's crazy? We keep identity locked even for cartoons. Yeah, that part shocked from Pixar. >> big thing you can see here is actually the teeth quality, right? So the left is the ground truth. It means like this is how the teeth looked from the input video of this character and we can keep it. This was an issue with Avatar 4 because we just actually applied different teeth. Um this like how our model was trained. Um so now we can actually also keep the same teeth from the input reference video um and yeah, characters, cartoons, um you can have the next slide, Adam. You will also see like us as a cartoon here. AI agents. And this is also working now. So yeah, I'm very excited about this and this upcoming future as well. The next two videos are you, Adam. So Welcome to lecture seven. Global leader >> it's also a big request from you. Today, we're diving into chapter nine where we'll explore. Okay, cool. >> Yeah. So yeah, we will also work on this. So yeah, this is everything for now. Our Q&A yeah, maybe do it in the community. Um we can we can guide this. I I get a more message here that we have a hard stop. So Adam, um Thank you, Nick. Thank you, everybody. Uh yeah, as Nick mentioned, please join our community forum. Ask questions there. Um we're always very interested in, you know, hearing people's feedback as that's really what guides um you know, what we work on. So let us know what you you think. Um and excited to see what you guys create. Thank you. Yeah. And if someone wants to have a live session uh with me and maybe we can set something up. Um and I'm I'm happy to show how to make the same launch video as we did. So Adam, thank you so much. Um I think that that was everything. Yeah. Bye. Bye-bye.