Transcription
Get ready, friend, because today I'm going to show you how to create insane text animations with motion graphics entirely with AI, completely free. All you need is a solid prompt, a powerful model, and and a sprinkle of, of course, creativity. What you're watching right now is AI generated, and it took less than 1 minute to make, and you do not need After Effects. You do not need experience with motion graphics, but you do need a good eye. I used when 2.214b 214B to get these results. And as you know, it was released less than 48 hours ago. And in my opinion, it's the best open-source video model to date. And I'll say it, it's better than Cense Swamp Pro when it comes to generating text, particularly.
This is what we'll cover in this video. We're going to install WAN 2.2 locally so that you can use it for free. We're going to install WAN 2.2 via Rompod if you do not have a powerful machine. We're going to explore the best prompts, best formulas, the best practices for generating high quality results. We're going to generate our first text animation, and we're going to be using WAN 2.2 Pro for more control and speed. Let's dive in.
First, we will install WAN 2.2 locally using Comfy UI. And if you're new to this, just go to this website that I'm showing you on the screen and download it for your particular operating system. Well, the installation is self-explanatory. It's is very easy to follow. Just follow the steps and and you're good to go. After you've downloaded it, let's launch Conf UI. I'm currently on a 2023 M3 chip MacBook Pro. What I'm going to do is I'm going to go over to workflow, browse, templates, and video. Now, Comfy UI natively supports when 2.2. And you'll see a few options over here. You'll see 14b text to video. You'll see 14b image to video. So, these are the biggest weights. And we have the 5B counterparts. Less weight, faster results, less quality. We're going to install the 14B weights, which the total size is 33.13 GB. So, make sure that you have some storage available. And we're just going to click download over here and wait it out.
Once you have everything installed, you'll see your workflow interface. Um, as you're seeing right now on the screen, we have our models. We have the video size. We have the prompt section over here, we have our video settings, and then in the very end, we have our final video preview. All you have to do is you drop in prompt, you hit run, and you'll let the model do its thing, right? Generate the video. For a 5-second clip, it takes me about 60 minutes. I know, 60 minutes. sometimes like an hour and 30 minutes on my Mac. It's sad. It is extremely slow, but it's free, right? It's free. Personally, personally speaking, I'd rather pay than wait this long. So, this is what I usually do to make this process even faster. I would do the exact same thing. I would install the same exact workflow inside, which is a cloud GPU platform that runs uh Comfy UI using higherend GPUs. In this case, I'm using an H100 GPU, which costs about $239 per hour. So, let me show you how to set that up. Feel free to skip this part if you have a powerful computer or if you know how to do this.
So, let's head over to RMPod and let's set up our storage. Go to the storage section and create a new volume for our models and for our outputs. I would recommend allocating at least 100 GB to ensure that you have a lot of space. With that in place, let's deploy our GPU. Let's go to the secure cloud and look for high performance GPU. Since WAN 2.214B is a demanding model, I'm going to choose Nvidia H00. I'll search for an available H00 and select it. Give it a name. And this next step is crucial because we need to attach the 100 GB volume we just created to an actual GPU to ensure that our work is saved and persists between uh sessions. Uh with our GPU then selected, the next step is choosing the right template. We're going to click select template and we're going to search for conf UI in the list. We're going to use the ROMPod comfy UI template which comes pre-installed with confiu manager and pytorch. This will save us a lot of manual setup. And once your pod is up and running, you'll see a connect button. Click it and from the drop-own menu, select connect to Jupyter Labs. This will open a new browser tab and it's going to take you directly to Jupyter Lab environment. Think of this as your uh command center for the final setup. The very first thing that we need to do is we need to upload two essential setup files that you can find in the description called requirement text and confi setup. All we have to do is drag and drop the two files from the computer um on to the left hand side of the screen. Open a new terminal and in this terminal our first command will be to install all the necessary dependencies. We're going to type pip install r requirement.txt and press enter. Uh this will read the text file and install all the required packages for our entire workflow to run smoothly. And after the dependencies are finished installing, we will run our custom setup script. In the terminal, type python comfy UI setup and hit enter. Uh this will trigger the script to download all the one 2.2 models for us. Now this is the longest part of the process, so please be patient. You can monitor the progress by navigating into the confi folder and then the model subfolder. Once this script has finalized, all the models will be ready to use. And all we have to do now is just start the Comfy UI server. Navigate into the main Confy UI directory. From here, we are going to open a new terminal window to launch the Comfy UI back end. To start the server, we will run the main script. You will see the server starting up in the terminal window. Wait for it to display a message indicating that um it's running and and uh it's listed on a local URL. And once you see that confirmation message, head back to Rompod dashboard. Um the final connection step is waiting for you here. Back on the pod page, click connect button again. This time an HTTP port will be available. So click that. So you can open the Comfy UI interface. And here we are. Everything's ready for you. All you need to do is navigate to the templates video, launch your favorite WAN model, and start generating.
Now we are ready to generate with WAN 2.2. to but first let me explain how prompt engineering works for this particular model prompting is the most important part of the process okay like that's all you have to do that's that's all that you need to think about everything else is just done by AI right so after 48 hours of of stress testing this I found this formula to work best and here it is first we have a subject we need to describe the subject it can be an object it can be a person whatever your sub whatever your main subject of the video is. Then we have a scene scene description. Then we have motion. So we need to describe our camera and and the movement. Then we have our aesthetic control. And then in the end we have stylization. You can see the prompt formula on the screen. And here's what each component means exactly. Again, subject description means that you're describing your main character or your main object. For example, a blackhaired man in ethnic clothing or a dreamlike creature with opolescent skin and featherless wings. And we have scene description. We're describing the environment or our setting. We have motion description. We are describing the movement, slow motion, shattering glass, what is happening. Then we have aesthetic control, camera angles, light source, lens, different movements, the framing. In the end, we have stylization. What is the style? Like, do we want to create a cinematic look, cyberpunk? Is it an illustration? Is it is it a drawing? Is it a postapocalyptic scene? What is it that we're trying to achieve? I already have a free PDF in the description that you can download. You can explore more ways of prompt engineering for cinematic videos.
Now, here's a very cool trick. I feed this entire prompt into chat GPT so it can help me out with generating u as as as many ideas as possible. We're going to open up chat GPT and paste this exact description. So, we're going to say, you are a prompt engineer. Your job is to give me prompts optimized for WAN 2.2 two following this exact formula. You add in the formula and then you add in an example prompt. You paste that in. In the end, you say, "Let me know if you understand so that I can tell you my idea." Next, it's going to ask you to start describing what you want. And that's what you're going to do right now. For example, I'm going to say, "Generate a few prompts for text animations designed as motion graphics sequences." The text should say, "This is serial." I'm going to hit enter. Now you'll get a list of prompts right that chat GPT has generated following our formula. We're going to copy those and we're going to paste them inside comfy UI. Here is the prompt section. The green section is our prompt section. We're going to paste that. We're going to set the frame rate to 24 frame per second and we're going to set the aspect ratio to 16 by9 and then we're just going to hit run. Now if you're using one GPU this might take about 30 minutes, sometimes more. If you use eight H00 GPUs, it's going to be way faster. Around 5 minutes per 5second clip. Now, this is what we got. Take a look at this. This looks phenomenal. Like, right, you tell me. Like, this looks great. And these are four different promps that I copied from ChachiPT and pasted in Confy UI and they perfectly follow the description. Take a look. [Music] I don't like how is a space evolving so fast. Something very cool that you can do also is that you can prompt it against a green background so you can easily key out later for transparent overlays. Take a look at this. Crazy.
Now, if you do not want to wait so long and you want results that are quicker, you can get access to when 2.2 Pro, correct? You heard that right, Pro that is inside enhancer, and it takes about 1 to2 minutes to generate unlike the the open- source uh one 2.2 version, which is free. One 2.2 Pro gives you creative options for for style, for camera motion, for for direction, for for special effects. So, let's test this out and then see the results side by side. We're going to go over to tools, video generator to text to video. We're going to select WAN 2.2 Pro as as our main model. And we're going to set the style to dynamic. We're going to set the camera to gimbal smoothness. We're going to set the direction to zoom in. We're just going to leave everything else as none. We're going to paste in this prompt and hit generate. This is what we get. I I have I have no words. So, we have left WAN 2.2 which is free. You can use it inside comfy UI. That's what we generated previously. And we have write Want 2.2 Pro which we generated with enhancer. They both look epic. The only difference is that W pro allows for more creative control.
This is not only limited to to text and motion graphics. Um even though I do believe WAN 2.2 does excel at that, you can create any type of video. So So let's try a few cinematic examples here with WAN 2.2 Pro. What I'm going to do is that I'm going to select over here um VR 360. We're going to selecting zoom in for the motion and we're going to paste in a cinematic scene. We're going to copy the prompts from chat and we're going to paste it over inside enhancer. We're going to hit generate. Looks good. Let's do a push out effect but from this image instead from a text. Drop our source image in and add our prompt. For example, I I'll simply just say here rain. So rain, right? There's no need to follow the prompt formula here since we are using the the pro version and it's embedded in. We're going to select push out everything else none. Hit generate and here's what we get. Let's do a sidebyside comparison. WAN 2.2 W 2.2 Pro. Pretty cool. [Music] [Music]
So this my friend is how you can create insane text animation or motion graphics entirely with AI. If you enjoyed this content, drop a like and a sub. That that means a lot to me. But also do let me know how are you finding enhancer. Look, I'm doing my best to to bring you the best AI custom workflows inside one creative app that is meant for professional producers, and I think so far people are loving the Portrait Upscaler, but I'm just getting started. And this could not be possible without you. I hope that you have learned something new today. And do not forget, the real magic of AI is not what it can do for you, but how it empowers you to do what you've always wanted, to create without limits. My friend, this is serial.