Transcription
EVERYBODY THIS WAY. MOVE. MOVE.
>> REX, STOP. REX, stop it. Rex.
>> Kevin, put your phone down. Red, stop.
I've had my hands on Cling 3.0 for quite a bit of time now, and to be honest, there's not much good information out there. So, in this video, I will teach you everything that I know so you can master Cling 3.0. So, let's dive in.
To access Cling 3.0, you can do that directly through their website or you can use a tool like Open Art. This is the tool that I'm using for this video. And they've just updated their new UI. It looks pretty good. So, if you want to check it out and you want to try this out for yourself, then I will leave the link to that in the description down below.
Now, first, let's go over a few of the new features of Cling and then I will walk you through them uh step by step. We will start off easy and we will be going more advanced later in the video. I will also leave timestamps if you want to skip through it.
First things first, you want to select the model cleaning 3.0 or clean 3.0 Omni. Omni is referring to the multimodal references. I will tell you more on that later. But what you should know in terms of using this, you have your basic start and end frame where you can drop in your frame. If you just want to have a start frame, you drop it in right here. And you have your general prom box. And in OpenArt, you can also drop in a character if you wanted to as a visual reference.
What you can do now is you can enable this multishot feature. And this allows you to make six different shots throughout your video. By the way, the length for Cling 2.0 is now 15 seconds. I will show you why it is good and not good later. You will know. But first, let me just show you. If we have a shot right here with multi-shot, we are using the auto feature right now. If we click on customize, then we have shot one, shot two, shot three, shot four, shot five, and we can even add in six different shots. You can even choose the length of each shot. But keep in mind, you only have like 15 seconds to work with. So, you could either do like some of your shots very very short and then let's say one of your shots insanely long. I that's up to you, but this gives you more customization of what you want to have happen.
Then we have references. We can now add up to seven visual references that I will tap into this later. Let me first show you how you can get started with Cling 3.0 in terms of prompting. So I have this image reference that I want to make a video of. Then I have this prompt right here. And with the colors I want to identify all the different segments or categories you can add into your prompt. So we have the camera first, which is a medium tracking shot. Then we have the subject that is the woman. Then we have the action that is describing what your subject is doing. Then we have the environment basically describing like what's going on there. Lighting, texture, audio. The last three I would say are a bit optionally. This will help you form a good structure for your prompting.
Now if we then put this into Clink 3.0 and we just use an 8second output at 1080p, then we get a result like this. So far, nothing new. Nothing has changed basically. Um, this is just using the basic features. This could have been done with Cling 2.6 as well. But now, let me show you where Cling 3.0 comes in handy.
If we switch over to multishot right here, we can now bring our scene to life. And we only have this one reference, but we can generate up to like six different angles of that same character. And honestly, it has really, really good consistency. Like, it remembers what that character looks like. Now, the old method of making multiple different shots of your character would have been to make like four different prompts, give four different image references and then create your video. But now we can just put in this one frame. Then we can switch over to multishot. And now we can start describing each scene that we want to have happen. So if you do that correctly and you apply what you've done in the single shot that I've just shown you, we just add in the type of camera angle. than we want. Then we describe what we want our character to be doing. And then we can describe that for each shot. We can give the length of each scene. So now we have the second shot, an over the shoulder close-up as she leans in. And then we can just generate this. And then we get a 12se secondond result like this.
Now, what stands out to me is how good it keeps the scene consistent. So, we have the girl and remember this hallway. Then, if we switch over to a bit later in the video, then we can see she's walking towards the frame. She's closing her eye to the peepphole and then she's watching the porch. Now, then we switch back and we have this same hallway again. So, if we were to compare it, it looks pretty consistent to me. It might have been a little bit changed in terms of the camera angle. Like, it's a bit more blurry on the sides here compared to here. All in all, this is pretty good for getting consistent shots with multiple different angles, and that's literally what it's made for. That is the basics of how to do multi-shot. Later, I will test the limits of this tool. So, stay tuned for that.
The main question I keep getting for people that get started with AI is that they don't know how to prompt. So, for that, we can do a bit of meta prompting. What meta prompting essentially is is instead of you typing your prompt and sending it to the AI, you first ask another AI, in this case, let's say a chat or a cloth, you ask it to doublech checkck your prompt or you can even ask it to make you a prompt. So what I've done is all the information that's out there on cling 3.0. I've put it in this document and I've even used like the official guides of cling to write this thing and then I feed it into this chat GBT. Now what you can do is you can upload your reference image that you want to make a video about. Then you type in the ID that you have and then it spits out a multi-shot prompt for you. Again, this is the lazy way of prompting. I would still highly recommend you thinking what you actually want to see. If you don't do that, then you're just playing guesswork with the AI. It's literally like playing on slot machine hoping that the AI makes something good. Other professionals, they use AI to write prompts, but they still double check it and rewrite sentences and things themselves to get the exact things that they want to see.
Now that you understand how to prompt with the new cling 3.0 and the mult features, let me give you a few more details about the different ways how you can make generations in cling because there's a lot to it and you might be overwhelmed at first when you're starting this out.
So the first method of creating a video is just using text to video. Text to video is the simplest and quickest way to make a video because all you got to do is type your prompt and out comes your video. The downside of that is that you have less control over what the exact output will be because if you run that prompt like 20 different times then you will get 20 different results. So if you want to generate text to video this is how you do it. So on open art I feel like this is a bit weird like I'm going to give them some feedback on this because I don't like this. We have start and end frame but for text to video we're going to go over to reference. It might be confusing for you because it's quite weird names. Essentially with text to video you have two options. So you can use the milshot or you can just use a regular prompting.
Let me first show you an example of regular prompting. So here I'm just describing a shot of a rugby player is sitting in full rugby gear and then he is sitting in the locker room and we have another guy that's also like sitted seated in front of him. I basically have them have a conversation. So uh the guy is like making joke about hurting his shoulder. Yeah, damn that linebacker was big and they both have a laugh. This is what you would do if you let the AI come up with the outcome. If you don't have a specific idea about what you want to exactly see, I don't recommend using this method a lot if you want to have a consistent looking scene across multiple different scenes. But if you just need to make something quick, then this will work.
>> Hurt your shoulder?
>> Yeah. Damn, that linebacker was big.
>> Okay, my guy got got messed up by the linebacker. And I I don't mean that in a weird sense.
Now, if you want to do multi-shot, then it's literally the same as what I already described. You just describe your scene and you customize this. You could use the auto feature and it will create multiple shots for you. But then again, you use your entire like big prompt and you put it here. I would use the customization and add in your different shots. So this works really well if you have a bit of dialogue. So here I have like a white two shot of two people. Then I have a slow push in of a man in his 40s as he is chewing and looking down to his plate. Then I have a medium shot on a 90-year-old that's looking a bit confused. And then I switched back to the two of them and that got me this result.
>> Why are we meeting like this?
>> We're cutting you out.
>> What? Sorry, I need this. You know I need this.
>> Like that's quite good. If we go across the scene, then you can see that the consistency is there. If we go from this scene to this shot right here at the end, we can see the face look consistent. We even have the mirrored reflection. Then if we see the close-up, it looks quite good.
Now, if you want to check out all of the prompts that I'm using throughout this video, and if I'm going too fast for you, but you want to check them out at your own pace, then check out my school. In my free school community, I make a document where I post all the prompts that I'm using, including the examples, so you can take a look at your own pace. It's also the place where you can hang out, ask questions, and also chat with other AI enthusiasts.
Let's go a bit more advanced. Let's use references. Now, references are inputs that you can give the AI to make a video about. Now, if we go over here to reference and then we switch it to Clink 3.0 Omni, then you can add in up to seven different references in your video. You can choose from your own generations or you can also upload any type of image that you have in mind. And this is super powerful if you have all kind of little details that need to be there in your shot. Again, this is not really a new thing. Like, this has been around with VO. This has been around with previous versions of Cling, but this one is just a bit better and it's more consistent. I still must say though, it is glitching a little bit sometimes, but I will show you what I mean by that.
Okay, so let me give you an example of putting in multiple different references in your shot. I have this image of myself right here. The reason why I have three different angles of myself is because I feed the AI a bit more data what my character looks like. So in this case, I have a side angle of me. Then I have a front-facing angle of me. And then I have a backfacing angle of me. That hopefully prevents the AI to start glitching or making up random stuff that you don't want in your scene. So that is the first character I have in there. Then I have done the same thing for this other character that I'm putting in my shot. And then I have the setting or the location that I want to have this scene happen in. And then lastly, I've added in this cute little dog. If you want to make three different angles of your own character reference, then use this prompt that's on the screen right now. Just pause this, take a look at it, and copy it. This will also be in the document, by the way. This will save you a lot of work with doing this manually in Photoshop and then changing it all away. You can just put this one time in Nanana and then it spits out this version for you.
So, we're putting on a reference right here. Then, we're adding in the second reference. And then, we're adding in the third one and the fourth one. Now you can see we have image one, two, three, and four. In your prompt, whether you're using multi-shot or not, you can tag those different characters in there, and you can start describing what's happening with each character or with each location. So you can say like, hey, image one and image two are in the location of image three. Then image four pops up. You name it. You can literally describe any story you want as if you're referring to that person, but you're actually referring to the image.
So, if we start off easy without using multi-shot, then I'm doing this long take and the background is of image three. Then the camera does a deliberate 360° orbit with a tense scene. Now, then image one, that's me, is standing up excited and he says, "I got it." The camera continues circling to reveal a woman opposite of him, attentive and sitting. Then we have a tag of the of the women and we're basically like we're basically directing it. You got to think of a director. You got to think like what do you want to have happen in your shot. So I wrote this little dialogue and I have them going back and forth and I have the camera in one continuous motion. Then I have the dog appear somewhere like that's my scene here. Let me just show you what I mean by this.
>> I got it.
>> You got what?
>> I know who the killer is.
>> It's Emma, isn't it?
>> It has to be her.
So yeah, as you saw, this scene is by no means perfect. We have this dog that is like all of the sudden like becoming somewhat of a unicorn by getting a tail out of his head. I don't know even know how he does it, but it still glitches a bit. And that's the main issue I have with references. The more references you put in there, the harder it gets for the eye to keep it consistent. But depending on how complex your scene is, you could make this work. If you have a pretty easy shot, you could make this work with multiple different references. If you want that level of complexity, Cense 2.0 will probably solve that for you. And by the way, that's also why I'm using Open Art because soon you can also be using Cadence 2.0 inside of Open Art.
If you want to make it more complex, then you want to be using multi-shot. So in this example, we have a shot with four different references again. So here I have this girl that we will be using as our main character. Then we have this guy as our secondary character. Then we got this cute little hedgehog. And then we have the location it's in. So now if we reference it, I want to have four different shots. So I'm basically describing like what's going on in each shot. Some shots are more complex like for example this first shot than the second shot. You just need to think as a director like think okay I have this opener shot what do I want to see then second shot let's do a low angle and then we have this person that is waddling to the sofa she reaches down to take a small blanket then we have the third shot where we have a bit of dialogue going on and then lastly we end it with a fourth shot and that's the video.
So if we put this to action then this is the final output.
Pokey, could you fetch me that blanket?
>> Oh, hey. You must be the new neighbors.
>> Oh, hi. Yes, we are.
>> The only thing that is not perfect yet is the lip sync, but other than that, the AI has a pretty good understanding of where to make a cut and what's going to happen in the next scene. This is pretty smooth. You could have more control if you were to make each scene individually, but that's also a lot more work. So for animation videos like this one for example, this would save you a lot of time if you were to generate long movies like you now have the ability to make like 15-second shots. But there are a few issues with that in particularly the lip sync. But I will tell you like what the best solution to that is later in this video.
Then we have the third method that is using image to video. And this is by far the most popular used method. Like this is the method I use all the time and this is also the method that gives you the most amount of control. It's just going to back to start end frame. Here you can add in like your start frame. If you just want to have a start frame, you can leave it at that. Or if you want to have even more control, so you want to have your scene end in a specific way. Then you add in your end frame. Those are the two that I've been using a lot using this feature. You can also make continuous videos by, for example, using Open Art. You can get a frame. So if we grab this frame, let's say I want to make a cut at a certain moment. So let's say I want to make a cut right here. I can download this frame and I can make this as my new start frame. in the next video. So that is kind of like the idea there with start and end frame where right now we can continue this video with this new start frame. But just by putting in a start frame, you give the AI already a start reference of what it should look like. It only has to do the motion from there.
Now with multishot, we can just give him this prompt and I have like three different scenes in there. And then if we generate this, then I get a result like this.
>> Hey, you sure this is safe?
>> Yeah, man. It's fine. Just keep going.
>> Okay. I I trust you.
>> Looks pretty good, right? Here's another example of what you can do with just start frame and then using not multi-shot. So, here I have this example for you.
>> Hey, you sure this is safe?
>> Yeah, man. It's fine. Just keep going.
>> Okay.
>> I I trust you.
And this is, in my opinion, where Clink 3.0 know shines like it has way better prompt coherence than any of the other models that I've tried before like cling 2.6 and VO 3.1. I couldn't even generate this with VO but with cling 2.6 I had this result and that looked really bad. So that is one of the takeaways that I want to give you. If you are using any of the AI models don't think that you have to use multi-shot all the time. You can also just be using the normal prompting tool. Important thing to notice is that multishot is not available if you have both a start and an end frame. So don't start looking for that. It's just not there. With this, you have basically more control. So basically I want to start my frame at this and I want to end it at this weird looking guy. So the good thing about this is we can make an output that's longer than 10 seconds by sliding this over. I must say that the longer your video is, the more glitches you might be running into. But anyways, I generated this video with that prompt.
Where is my daughter?
>> WHERE IS SHE?
>> AGAIN, just to show you like why you need to be updating these models and trying multiple models out because like if I use this same shot in Clink 2.6 and VO3.1, I get this.
Where is my daughter? Where is she?
Okay, let's dive a little bit deeper and fully analyze if Clinko is worth the hype because I know including me, but many influencers and many brands are hyping this up a lot. But let's actually analyze it and see how good this shot is. So, I've made this video, which is a multi-shot video of this vigilante bucket man with this other character. And let's just see how good IT DOES.
>> WHERE IS HE?
>> WHO? WAIT, WAIT, WHAT?
>> Don't play games with me. Where is he?
>> Dude, why are you wearing a bucket?
So, let's analyze the prompt here. Let's first watch this first shot. So, that's the first two seconds. If we go through it, we can see we have the vigilante that is slamming the guy against the brick wall. There should be steam drifting in between them and there's dramatic low angle framing. Now then he screams in a forced gravel force. Where is he? I think he got that right though. But if we take a look and analyze this scene like frame by frame, then the steam is somewhat coming out of the bucket. The guy kind of already is slammed against the wall. But that's also the reference image that I gave it. But the way he pushes and leans in forward to him, that's pretty good. And now I want you to analyze like the shot. So take a good look at the references. Take a good look at this part. Take a good look at how he's leaning in.
Let's now go to the next scene. So for this next scene, we have the close-up. If we take a look at the hands, we can see that the gloves that he's wearing is quite consistent. The eye bags underneath this character details are good. I like how detailed his face is. Then we have this guy that is like looking super confused and he's like, "Who? Wait, what?
>> Who? Wait, wait, what?
>> That is pretty good. Now, switch back to the bucket man. We have these two eyes coming out of there and he says, "Don't play games with me.
>> Don't play games with me. Where is he?
>> What I like about this one is if we analyze this last little second right here, he's using all the force he has to lean in and say, "Where is he?" And then as he leans in, we also see in the next shot that he's leaning in. So I find that quite impressive that the AI understand that he's leaning in in the next shot. So we have just from one reference, if we have movement during that scene, it remembers that movement in the next shot. Even here, this is the thing that I wanted you to analyze is if we were to compare all of this and we can see the level of detail. So if you look at these marks on the wall and compare that to the marks we have at the beginning, we can see that they're quite consistent. We just have a zoomed in version. Even the background like we have this blue light right here. We can see that blue light also coming from this garage or whatever it is right here. We also have the like the city lights that are same setting. The bucket hat is pretty consistent. The character is also pretty consistent. In my opinion, this is pretty impressive. It might be nothing like what Cance 2.0 has been doing. Um and I will be very excited for that. But also it's just to give you an perspective like how good this is becoming like how much of a director you can be. So think like a director. That's the main takeaway I want you to be from this scene. Like that is what Cling 3.0 is more than ever now enabling you to do.
The last scene that I want to analyze where I pushed it to the limits is this sixshot multi scene. This is inspired from the movie the substance and this is the lunch scene where that guy is eating disgustingly. I apologize in advance if you're watching this because this is about to get quite disgusting.
>> Promotions not just about talent. It's about loyalty. I will get pe who understand that.
>> Yeah, that's quite disgusting. Right. Apart from like this guy grabbing a prawn out of thin air right here. Like I don't know where he grabs it from. The multi-shot here, the cuts in it look quite good. Like it it feels pretty natural and it managed to follow up all of the prompts pretty good. Even this girl, I only gave her a like prompt description. No reference at all. Looks pretty good. Like it looks pretty consistent here. Like even if we switch around in a different shot, she is consistent in both shots. So for pretty easy shots like that, it works quite well.
Let's talk also a bit about the limitations of this product because it's claiming it has like enhanced lip sync and the lip sync is much better. But I don't know about you, as soon as you generate more than 10 seconds, your lip sync goes to [ __ ] And let me give you an example what I mean by that.
>> My date stood me up tonight. So I'm available.
>> H So you're free?
>> Yes.
>> Cool. Could you take out the trash?
After around 10 seconds in this shot, you can see that my lip sync is completely off. And I'm not the only one experienced this issue. I've had comments about this. I've seen other people commenting this issue. It's pretty unfortunate. I hope they're going to fix it. But for now, the main tip I can give you is if you have dialogue, let the dialogue happen in the first 10 seconds and then have the rest of the 5 seconds. If you want to have these 5 seconds, have it as a action going on like not as in dialogue.
>> Hey,
>> what are you hiding behind your back?
>> This is for you.
Another thing we should disclose is the morphing. So with longer scenes especially, there's a bit of morphing going on. For example, check out this scene. The easiest way to fix that is to have less complex prompt or to make it less than 10 seconds again. Or you just got to wait for C dance and use C dance for this.
The thing clink 3.0 is really good at is expressions and emotions. For example, this clip.
Why? WHY? Come back now. I almost feel sorry for prompting this to this guy. Like the level of emotion he has is so realistic in my opinion.
So that is the complete guide on how you can master cling 3.0. Now if you want to try this out for yourself then I will leave the link to open art in the description down below. In OpenArt, I would suggest if you want to play around with it enough and if you want to have access to all of the models and you have a big project coming on, I would suggest go with a plan that is a bit bigger than any of the small plans because you will burn through credits way too quick. You don't want to run out of credits and having to reby credits every time.
Now, click the video that's on the screen right now if you want to learn more about how you can use AI for cinematic filmmaking.