📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

How to use Kling 3.0 for AI Filmmaking

Jack Vs. AI20:38

Transcription

Cling 3 might be the biggest update to AI video we've ever had, and it represents a huge leap for filmmakers and advertisers. Imagine a single tool that covers multi-shot prompting, character consistency, and native audio, as well as generating up to 15 seconds of footage. Yeah, this is an absolute gamechanger for creators.

I'm Jack. I've been a visual effects artist in the advertising industry for over a decade now, and I've worked for many of the biggest brands on the planet. Today, I'm revealing everything you need to know about Cling 3. We'll cover all of the insane new features within this update, including the ability to upload elements, multiple references that will be used within your generation to lock in a mind-blowing level of character consistency.

And guys, if you do enjoy today's content, make sure to whack a like on there for me, consider subscribing, and of course, ring that bell notification icon. All right, let's get into it. So, the structure of today's video is built around each of the new Cling features. You'll find timestamps below if you had something in particular that you wanted to focus on. So feel free to navigate to the relevant chapters as you see fit.

So to access Cling 3.0, we're over on Higsfield. And they've currently got some great offers running for both Cling and Nano Banana Pro. So if you do want to go ahead and check them out, you can click the link down below in the description.

Now to get started with using Cling 3.0, want to go up to the top lefthand corner, hover over the video tab, and click onto create video. Once here, if you haven't already got Cling 3.0 selected, you can go ahead and click into the model tab, go towards Cling, and then click at the top of the list here.

And the first thing I want to cover is this new method for ensuring character consistency when it comes to your video generations. So, of course, we're able to go ahead and click into this tab here to upload a starting frame. And the video generation I want to show you guys is this one here. This, of course, would be the starting frame for this. We have this toggle for multi-shot generations. We're going to get into that a bit later on in the video. So, we'll just turn that off for now. And of course, here we have our field for entering our text prompt. And with this example, I just went for something very simple. Show a slow and full camera orbit around the bearded man. And this type of camera move is a pretty brutal test when it comes to character consistency because we're not just seeing the character from various different angles, kind of cutting between them. we are literally performing this 3D rotation around the subject and so it really does point out any issues quite clearly.

I'm really happy to say though that this generation was very effective and what that comes down to is this new feature here on this small tab that says elements. So if we click into here you can see a couple of elements that I've already created. Now this is a really simple process. We click on create new element and we upload various different references of our character from different camera angles. We can give the reference a name as well as a short description and then get that saved. You can see here Higsfield describes it as a set of images of one character or object that keeps its look consistent across the whole video. So in my case, I have one called the bearded man. If I click on edit, I can show you the examples we have here. So we have me facing front onto the camera, me in slight profile, a little bit further away camera angle of me stood in my garden and then a view of me from behind. And what I also made up was this element here of the cling logo, this kind of 3D glass render. So we have a front on view of this kind of cling logo sculpture and then various different camera angles of it as well. So, we're not able to just ensure character consistency, but object and product consistency as well. And this is why I said in the intro that cling 3.0 will be a very powerful tool when it comes to both AI film making and advertising.

So, once you have your elements made up, we can just select the relevant ones here and click use element. Now, we're also able to select the duration of our clip. We can go all the way up to 15 seconds. But what's really nice is this is on a slider, so we can be really specific about how long a generation we want. We may only want 3 seconds, but equally we can go all the way up to 15 seconds if we choose. And this is really handy when it comes to multi-shot generation like you'll see later on in the video. And then of course, we're also able to select the resolution.

So, let's take a look at the result that we got from this very simple text prompt and using our element of the bearded man. We go ahead and open this in full screen and we can really take a good look at this because the result is really really quite impressive. Probably the best character consistency I've seen from a video model to date easily. And we've got another really quite impressive example here that I want to show you guys.

But just briefly, if you're wondering how I created the starting frames for my video generations, these were all made using Nano Banana Pro. And I knew that I wanted to create what we often call a key visual for the intro to this video. And for that key visual, I wanted myself to be holding this kind of glass sculpture of the Cling logo and be in this strange otherworldly environment. Now, this is a really great workflow here, right? So what I did was I uploaded this pretty poor quality image of the cling logo. I selected Nano Banana Pro. We would have put it on 16x9 and selected 4K. And for the text prompt again very simple. Create a 3D render of the attached logo made of glowing glass in a simple dark environment. And this is what we get back. Just look at the quality here.

Now we can combine that new asset of the 3D sculpture with image references of myself and a text prompt like show the bearded man standing in a dark alien temple made of glass. The bearded man is holding a large vibrantly colored glowing glass 3D object in his hands and a bunch more detail in terms of the style and a description of myself as well. So Nano Banana Pro is combining this asset here that we created with the image reference of myself. And the result that we get back is pretty amazing. And this image reference here is what I went to plug into Nano Banana Pro multiple times just giving it various different text prompts. So you know me kind of getting dragged through the clouds. We've got the shot here of me using it kind of like a surfboard in this river of lava. So that's the general process that I follow for creating starting frames that we then take into cling 3.0 and animate. Very powerful AI film making workflow. But that's how we created the starting frame for this generation here. And we've asked again for a camera orbit as the bearded man stares down at the glass object. And you'll see once again we are achieving a really pretty amazing level of character consistency by using the elements within Clling 3.0. 0 the exact same process that I showed you for that first shot.

So, just maybe character consistency in AI video generations might have just been solved. Now, this isn't going to work every single time. You are going to get some dodgy results every now and then, but generally speaking, this is an incredibly powerful workflow. This is solving a genuine problem that has existed since AI first started being used for generating videos. So definitely encourage you guys to test out this workflow. Create your starting frame using Nano Banana Pro for example. Write a simple text prompt and then create your elements with your image references of your character. All of this in combination can get some really pretty amazing and consistent results which is really what we're after.

Now the next thing I want to speak about is multi-shot generations. another gamechanger with this cling release and what you're seeing in this generation on the screen right now that was used in the intro to this video. So, how exactly does this work? How can we get all of these different varied shots from a single 15-second generation? Let me go ahead and show you.

So, what we want to do first of all is upload our image reference. Again, the same one that you can see on the screen right now we generated with Nano Banana Pro. And then we want to toggle on this option here for multi-shot. And this allows us to add multiple shots within the same generation. And the way that I went about writing these was using chat GPT. So just as an example, I said, "Write me a multi-shot prompt that shows a bearded man in an otherworldly environment. For a slight variation here, he is fighting a huge beast and wins. Use a variety of camera angles to ensure the scene is dynamic and well-paced. Write this as individual prompts so each can be dropped into an AI video generator. Be specific about character position and actions as well as the lenses used for each shot. All action needs to happen within 15 seconds with each shot lasting a minimum of 3 seconds." And that is because that's as low as we're able to go with each shot with the multi-shot generation. And I've just added a brief description of myself to play devil's advocate here. ChachiPT might start putting me in different clothes. And at the end, I've said in short, each prompt is a maximum of 500 characters each because you can't enter more than 500 characters in each of these text prompt fields. I will put this prompt template, by the way, down below in the description if you guys want to go ahead and use it.

So, let's send this off. We can take a look at an example that chat GPT gives us to kind of make the process very easy. And just like that, we've got back exactly what we need. We've got a variety of different shots, and each text prompt chat GBT is describing the lens that's used, the type of camera angle as well. Then, of course, the action of our fight scene where we're fighting this beast. So, what we would do is we would copy and paste each prompt into each of these individual sections here. And of course, in this example, the Clling 3D sculpture is creating all of these portals in front of me. Not going to go too deeply into each text prompt for each shot, but I do want to, of course, take a look at the result that we got because it was pretty incredible. When I first watched this, it was a bit of a moment for me to be honest with you. Let's go ahead and open it up in full screen. the kind of complexity of what we are asking Clling to pull off here. I honestly didn't think this was going to work and when it did, it was a pretty exciting moment. Generally speaking, I tried to kind of avoid complex prompts and kind of set pieces like this when it comes to AI film making, but I think we may have now turned a corner. I think things like this are now going to be possible. And for each text prompt, we're able to add our elements as well to ensure that level of character consistency. So, I definitely suggest you guys do that. But, I mean, we're seeing me from different camera angles and I look nice and consistent.

Now, the only kind of issue that we're running into here with this generation is when myself as a character, I'm far away from the camera. And this seems to be a problem for pretty much all AI video models. If your character is close to the camera, the results look pretty great. When they're far away in this kind of wide or establishing shot, you start to run into problems. The faces kind of just go to mush. And I'm not sure why that happens. It's pretty frustrating. But this is one of the pitfalls of this particular generations and others that I got from Clling 3.0. So, do keep that in mind, guys. Of course, I don't just want to talk about the good things. We need to talk about the issues as well.

What I'd also like to get into now is a text-to-video style generation. So if I just click here, it's going to replace that. We can remove the starting frame because we didn't use one for this particular generation. Now once again, I asked chat GPT to help me write each of these individual prompts where there is this couple who are having an argument. And I'm going to go ahead and throw on my headphones for this one here. So, of course, we're using a different method for AI video generation where we're not creating a starting frame. We are just going by the text prompts. That's the only thing that Cling 3.0 has to go by. So, always interesting to act as a point of comparison. I wanted a scene where this couple are having an argument next to the Brooklyn Bridge. So, let's take a look at the result that we got.

>> Don't do this. I am begging you.

So for me, this was really impressive. Once again, one of the improvements with Cling 3.0 is greater kind of emotional impact. These performances of our AI actors, and I think this is a really great example. There's a lot to be impressed with here. This shot here in particular just looks incredible to me. We can talk a lot about the skin texture and her kind of flyaway hairs as well. This feels incredibly real and really quite natural as well to have just achieved this with text prompts. Yeah, pretty crazy.

And this multi-shot prompting marks a real moment, a real development when it comes to AI film making because traditionally we would have had to use a tool like Nano Banana Pro to create each of the starting frames for each individual shot and then go ahead and animate them. But by using multi-shot generations with Cling 3.0, we can create an entire kind of self-contained story with multiple shots, multiple camera angles, dialogue, everything that you'd want. all just within one generation.

And just to show you guys one more extra text-to-video based example, this was actually a bit of a happy accident because I meant to upload a starting frame of myself in one of these kind of scenes that we were looking at before. Uh but I forgot to add it in. So the prompt was simply slow dolly camera move in. And this is a result that you can get from Cling 3.0 with such a simple prompt. Again, the level of realism here, the skin texture, the imperfection as well, the flyaway hairs, the shallow depth of field, everything here looks incredible. The lighting on her cheek here where we can see some subtle blemishes. This generation, it was a happy accident, but what a great example of what Cling 3.0 is capable of generating. So, you don't necessarily have to go the route of creating all of your starting frames using Nano Banana Pro. Of course, this is the way to go if you want a greater level of control, but you can just go the text to video route and still get some really quite incredible results, particularly by using the multi-shot generations if you want to encompass an entire story within a 15-second generation, for example.

Now, to just briefly touch on some of the issues, some of the pitfalls that I ran into when using Cling 3.0, for example, we can talk about the fact that it seems to love putting random dialogue into your generations, even when you specify it not to. So, you can see here with this text prompt, I said, "No music, no dialogue. The man does not speak. He looks forward determined." And look what we got back.

>> Desperate bright eyes want.

So, I don't know what language this is. It kind of sounds like when the Sims talk. Uh, but obviously a lack of prompt adherence here because we are specifically saying do not get this guy to talk and yet Clling is kind of ignoring that. Weird because we've got back some really complex results and then it seems to be struggling with something that's actually very basic. And this is another really good example because the generation itself is really quite incredible. But again, Clling has added in this dialogue. Now, kind of my fault because I didn't specify in the text prompt to not add any dialogue. But still, with some other video models, you wouldn't necessarily get this random dialogue added in. I wouldn't mind as much if it kind of made any sense. But let's take a look and you'll see the problem.

>> But they fenler.

Still really love the visual when it comes to this generation. Really quite impressive. Great level of character consistency once again using the elements feature, but that random talking is a little bit annoying. We don't really want or need that there.

And another example here, the problem I was talking about earlier on where the character is quite far away from the camera. Again, playing 3.0 seems to struggle a little bit here. You kind of see that the glasses have fallen off of me. Like, I haven't I'm not wearing them anymore. Um, and just my face generally looks very kind of plasticky, uh, animated, but not really in a nice way. So, if you're able to try to keep your characters close to the camera and you won't run into these problems. But, of course, important to point these issues out.

Bunch of examples here as well where I'm just generating a starting frame with Nano Banana Pro. Some of these images are of course me uh in the desert with the Cling sculpture in the foreground. So that's just being animated just with a more traditional workflow when it comes to AI video generations through Cling 3.0 with a simple text prompt here. just a single starting frame, but of course still using the elements feature just to lock in that character consistency. This was a really great generation as well. Really nice and dynamic and some nice sound effects as well. And of course, that's one of the other upsides of this model. Kind of standard now that you get native audio with your video generations for a lot of the best models on the market, but still Cling 3.0 does seem to do a nice consistent job when it comes to that.

And something that I just briefly wanted to finish on here was an example of text rendering because this is one of the supposed improvements when it comes to Cling 3.0. So what I did was I asked Chat GPT to write me out this really quite intense text prompt. I said I want this to be really difficult to render basically. And what I ended up getting as this example was this kind of timetable for a bus or a train. And unfortunately, what you guys might be able to see is that there are still a fair few issues when it comes to text rendering. Of course, this is a really kind of extreme example. Uh maybe there could have been things that are done to improve it a little bit, but I think there are definitely some issues here. The Jack versus AI looks fine. So, kind of simple text seems to work okay, but then there just seems to be a lot of kind of odd problems going on. Uh whether this is meant to be in a different language, I'm not sure. I kind of just think it hasn't done the right job. There's just some strange artifacts, strange use of language and symbols and things like that. Camera move itself looks great. The lighting, the kind of how dynamic it is at moving through this 3D space, but yeah, there does seem to still be some issues with complex text rendering. brutal example, but something that I wanted to pass on to you guys.

So, there you have it, guys. Do let me know if you've managed to jump on and play with Cling 3.0, what you think of the video model. I myself, I've been generally really, really quite impressed. Biggest selling points for me is the character consistency and the multi-shot generations. I can't wait to play with this some more. Of course, we're barely scratching the surface here. And that being said, guys, if you have enjoyed today's video, make sure to whack a like on there for me. If you're new, consider subscribing and of course, ring that bell notification icon. I'll throw up some of my previous videos on the screen right now if you wanted to go ahead and check out some of those. But thank you so much for watching and have a great day.