📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

I Paid 10 AI Tools to Turn My Images Into a Video Ad - Which Model Did It Best?

AI2Play28:24

Transcription

Welcome to another AI video tutorial. Today I'm putting different AI image and video models to the test. Some deliver impressive results, others fall flat. By the end, you'll know exactly which models are worth your money and which ones are just wasting your credits. Plus, I'll show you how I use these tools to create real video ad projects. Stick around if you want to save time and get better results.

I don't want to pay for 10 different subscriptions just to test a bunch of models. So, I picked a platform that has most of these tools all in one place. It's the easiest way to see how these models actually perform without spending a fortune. If you stick around until the end, I'll show you how to get some free credits. The platform is called Nim.video. You can log in fast with Google and give it a try. Let me show you why having everything in one place makes things way easier. I have a tab here for image generation and one for video generation. Let's start with image generation first. You can see there are a bunch of models here like flux, GPT, H highream and Google imagining. For video models, I also have the popular ones like different versions of cling or one luma and miniax.

AI is advancing really fast and the models get upgraded all the time. So by the time you watch this, there might be even better models. There's a real competition to make the best ones and they keep improving every month. Enough talking. Let's generate something. We need a prompt. That's just a description of what you want to generate. I could write a cute cartoon bunny, but that gives too much freedom to the AI. You didn't specify how it should look, the art style, the environment, or what actions it's doing. So, I suggest being more descriptive with your prompts. You can use this button to enhance your prompt. If I click it in a few seconds, it will give me a longer, more detailed prompt. If you don't know much about prompts, you can just write something simple and let it the AI decide. Then you choose a model, but there are so many. How do you know which one to pick? I hope by the end of this video, you'll have a better idea of which ones are worth using.

Depending on the model, you get different aspect ratios available. For example, if I choose Geminiedit, which is used for editing, it doesn't offer many ratios. So, it depends on the model you pick. You can check that for different models. The GPT model we use in chat GPT from my experience only offers three ratios, square, 2:3, and 3:2, but here they added 16 to 92, which chat GPT can't do. I assume the image is generated at 3:2 and then cropped to get the ratio you want. Because this platform uses an API, it offers more options than chat GPT itself, like different quality levels at different prices. So, if your generation doesn't need to be complex, you can get away with using fewer credits. Some models let you use a seed. By default, it's random and changes with each generation, but you can uncheck that and use a fixed seed if you want to test similar prompts and get variations of the same image. Here's another useful function. We have this describe button that lets us upload an image and then it generates a prompt based on that image. If you've used MidJourney before, you've probably seen something similar. It's pretty useful if you don't know how to write prompts. So, let's choose the Flux Pro model. And when we click generate, we should get an image based on the model. Usually it takes between 10 and 30 seconds. The GPT model can take longer sometimes. And here's the result. A cute bunny. When the image opens, you can see all the important info here. We used a text to image workflow and it shows the prompt, the seed, the aspect ratio, and the resolution. At the top, you can add the image to favorites, download it to your computer, share it, and there are some extra functions, too. Of course, you can also generate a video from it, edit the image with Gemini or lip sync, get a prompt from it, reuse the settings, and so on. From here, you can go back with the edit option. You can use Gemini to ask for changes like making the bunny wear a hat or something else.

I use chat GPT to generate different prompts so we can test each model and see how good they are with people, cartoons, anime, or text. This way, you can choose the right model for the style of images you want to create. I'll add all the prompts I use in this video on Discord so you can download them for free for both images and videos. Let's start with the first one, a portrait of a woman. I'm aiming for a realistic portrait and we'll see which model can get closest. I'll test six different models. Flux Pro, which many platforms offer. Google Imagin, which just got a new version as I finished recording. GPT image model, the one you find in chat GPT. Flux model, which you can also run locally if you have a modern Nvidia card with a lot of video memory. Hydream, another model that was open sourced recently. flux uncensored for those who want something more naughty. Obviously, I won't be prompting for that here, but I'll test it with normal prompts to see the quality. For the aspect ratio, since each model supports different ratios, but most are trained on square, I'll go with square for all so we can compare better. I'll use a fixed seed. You can pick any number, but I'll use the current year. Click generate, and let's see what we got. This is the first result. Remember, you can see the model I used, the prompt, the seed, and the aspect ratio. The result is quite good. I like it.

Next, I'll use the GPT model. I'll test low, medium, and high quality. This is the highest quality, and the result is pretty good, but it's missing a bit of that realistic look. Uh, not sure what exactly yet. This one is medium quality, which uses fewer credits, but it starts to lose texture in some areas. the low quality is even smoother. For the next prompts, I'll just test GPT at higher quality to see what the best result is for the Hydream model. Again, we have different quality levels at different credit costs. This is the highest quality. It went a little overboard with freckles, but the mood overall is okay. At the optimal setting, it starts to lose texture and gets more of that AI look. The fast setting loses image quality with some noise or artifacts. So, I don't really recommend the fast version unless you're out of credits and want something cheap. If you want to use Gemini Edit and try to generate, you'll see you can't because you need to upload an image first. That's only for editing. But Google Gemini uses the Google image and model, so that's what you should use. The result is this one, which is pretty good overall. For the Flux Plus character model, you can't use it as is. It switches to Flux Pro. You'll need to upload a character image first to use that model. Let's test the Flux model, which seems to be the cheapest option. The result isn't very realistic. The skin is too smooth for my liking. Not bad, but it could be better. For the uncensored version, since I didn't put any uncensored words in the prompt, I should get a safe for work version, and the result is actually quite nice. What surprised me is that the image size is a little bigger than the previous generators, as you can see here. If I compare it to other models, all the others had smaller sizes. Not by much, but still noticeable. Now, it depends on your preferences, but in my opinion, Flux Pro, Google Imaging, and Flux Uncensored gave me the most realistic portraits. What do you think?

Now, I tested a full body realistic photo of a woman. Again, you can see the model and prompt I used here. Flux Pro did a good job, but all models still struggle with things in the distance sometimes, like hands and eyes. The further they are from the camera, the more mistakes you can get. GPT can also do okay, but it tends to add extra noise and a yellow tint to the images, and sometimes the faces aren't very attractive. Hydream usually handles more art styles and is quite creative, but it also struggles with hands sometimes. Google Imagining decided to show me a behind view of the woman. It's realistic, but has a small mistake with how the hand holds the basket. Flux does okay, but the faces don't have enough texture and can sometimes look plastic. The uncensored model is also okay, but as I said, if people are far away, the eyes or face can have mistakes. It also tends to add a signature or text in the bottom right corner sometimes. Overall, they all did an okay job, but Flux and GPT sometimes have that AI look we don't want. It depends on the prompt and the seed. Sometimes they can do great and sometimes they can be really bad. So, try a few versions. Some of the models might not be so good with people, but work better with product photos and objects. So, let's test food photography because who doesn't love food, right? Flux Pro did a delicious burger photo with no obvious mistakes. It looks pretty realistic to me. The GPT model tends to have fewer mistakes, so it's quite accurate, but somehow it looks a little too perfect and less realistic. Also, that yellow tint shows up. Again, I've learned to spot images made with chatbt by that look. Hydream is okay overall, but the image quality isn't quite there. It looks like a lowquality JPG or something. Google imaging, what have you done? It gave me a 3D render version of the burger that doesn't look realistic at all. The Flux model did a good job, too. It's not great with faces, but for other stuff, it's pretty solid and only costs one credit. I also use it with Comfy UI on my PC, and I have tutorials for it on my other channel, Pixar Roma. The Flux Uncensored version also looks okay, but added some random words in the image. You can easily remove those with Photoshop's remove tool. Overall, they all do a good job with food photography if you prompt them right, except Google Imagining, which wasn't realistic, and the GPT version, which is too perfect. That green salad looks like a robot made it.

Moving on to the next category, cartoon characters. I usually like 3D render style cartoons, and Flux Pro did a great job with that. GBT is also pretty good with all kinds of cartoons, including vector style images. I mostly use it when I need something more complex because GPT understands prompts better than anything else I've tested. Hydream is pretty good with cartoons too and can be quite creative. Google Imaging is okay, maybe less artistic than the other models. Flux is okay sometimes. It depends on luck. Sometimes the results are great, sometimes less so. Flux uncensored also did an okay job, but again, those words show up in the bottom right. But probably anyone using uncensored generation isn't using those images for commercial work anyway. So, you can get decent cartoons from all of them. Um, depending on the prompt, I'll probably stick with Flux Pro, and if it can't handle something more complex, I'll switch to GPT. Testing another 3D cartoon style with a bunny cartoon. Flux Pro did okay, but the ear isn't quite visible. GPT is also pretty good, though extra noise it adds can be annoying sometimes. Hydream is good, too, but again, the quality isn't quite there, and there are some strange artifacts on the edges of the image. Google Imagin is okay again, but I still feel the renders are sometimes too simple and less artistic. Flux did okay, though maybe that hat looks a little strange. Flux uncensored is also okay, but less artistic and more mundane. Like I said before, my favorites for cartoons are Flux Pro and GPT. Hydream can be good, too, but the image quality isn't always consistent.

I also tested anime style. I have more experience with cartoons than anime, but you can decide what looks best. Flux Pro did okay. GPT is also good with anime. If they didn't add that noise, it would be even better. Hydream again has a bit lower image quality, but the styles seem fine. Uh, Google Imagining looks good and might be my favorite, but sometimes you can get copyrighted content or logos, so be careful if you use Google Imagin for commercial work. I even got a Nuto logo on a Ninja Bunny once. Flux model isn't bad, but it can do better. It looks less professional to me. Same for Flux uncensored, but for those who prompt for uncensored stuff, it's probably good enough. So, my favorites for anime are Flux Pro, GPT, and Google Imagining, but it depends on your preference. Decide what works for you. I wanted to see how the models handle typography in single words, and how creative they can get. Something with a fluffy texture. Flux Pro did an okay job with the words and colors, but the texture isn't quite right. GPT did okay, too, but the colors look a little dirty to me. Hydream looks okay, but too bad the quality isn't better. Google Imaging could have better colors. The render could be improved. Flux did an okay job on this one. I actually quite like it. The Flux uncensored version also did an okay job. For this, you probably need to test more versions to get more accurate results, but they managed the word love pretty well. For some models, I didn't like the colors, and for others, the quality.

I wanted to test a more complex prompt with multiple words to see how the models handle it. So, I asked for a cat holding a sign with AI2 play on top and nim video underneath. Flux Pro added an extra set of hands for this prompt, which can happen sometimes. GPT can handle really complex text, more complex than the other models, and got everything right. Hydream also did a nice job, though it added an extra hand as well. Google Imagining got the text okay for this one. Flux didn't put Nim video on the sign and Flux uncensored had an okay version too. So all of them can handle text to some extent. I didn't try even more complex text except with GPT and it can do really long text for accuracy and if you need text, GPT will be better, but other models should handle a few words fine. So I tried to give some ratings based on my preferences and compared all these models. Here's what I think. The cheapest option is the flux model. So, you can get away with using that most of the time if you're willing to try a few versions. GPT is the most expensive, so I only use it when other models fail to give me what I need since GPT has the best prompt understanding. Each model supports different aspect ratios. So, if you need a specific ratio, check which one supports it. For realistic people, most models do okay. I gave GPT and Flux minus a star. For food photography, all did okay except Google Imaging. For cartoons, all did an okay job. For anime, some of them weren't quite there. For text generation, I don't think any other model beats GPT. So, now it comes down to the price and quality you want. I want a decent price and good quality. So, most of the time, I'll probably go with Flux Pro.

The best way to get a video is to start with an image. So, I use Flux Pro with a 9 to6 ratio, the one usually used for shorts or reals. I generated five different images, some realistic, some cartoon, and one anime. Let's start with the first one. a woman walking. You can see the prompt I used here, but again, you can find all the prompts I use on Discord for free. Then I have an old lady riding a zebra. I wanted to see how AI interprets that. Next, I wanted a baby deer looking at a butterfly. So, we can test how the deer interacts with the butterfly. Does it look at it? Is the butterfly flying? Okay. And so on. After that, we'll test something more complex like an action scene with the camera focusing on different characters. and at the end an anime girl to see if it follows our prompt. So let's start with the simplest one, the woman walking on the street. If I click on image to video, I can upload the image directly there or I could upload one from my computer. Then we select the video model, click on the other button to see all the models that support image to video. To remove an image, click the X button. You can upload an image or a character reference to do text to video with that reference. Let's select image to video and choose the woman. Now we need a prompt for it. The initial prompt isn't great because it was made for photography. We need motion. So I created some prompts for the video with chat GPT. You can see words like flowing hair, walking, slow cinematic movement, words that add motion. I like to use slow cinematic movement because if the motion is too fast, it can cause more glitches. I tested three different Cling models. Cling standard, Cling Pro, and Cling version two to see what the differences are. And I got these three videos. Cling version two is the most expensive, but is it really worth the money? Let's see. This is the result for Cling Standard. You can see the size it generates. The motion is quite natural, but the quality isn't great. Definitely not the best there. For Cling Pro, everything improved. Uh, we got a bigger image size. Now, if you're wondering why it's not the perfect 9-16 ratio, it's because when the AI generates images, it prefers sizes that are multiples of 8 or 32. But if you crop the image to 9 to6 in Photoshop, for example, you'll get an exact full HD size. Cling version 2 understands prompts really well, but the video resolution is smaller, only HD instead of full HD. So, you pay more for better prompt understanding, but get lower quality. What about the one model? This one also has two versions, one and one pro. The one model gives HD size videos, but the woman looks a little cartoonish with a jumping walk. For one pro, the movement improved and it takes the slow cinematic camera movement more seriously. It outputs full HD size, so the quality is good, but it's missing some realism in the motion. The Miniax model has problems with walking. The walk doesn't look natural. The LTX model also has walking issues and is very slow. Google V2 looks okay and can only do HD videos. Lumay 2 is okayish, but it added a man in the scene I didn't ask for. And the video resolution seems even smaller.

For the old lady on the Zebra, Cling Pro did an okay job. The Zebra is moving and it looks natural. It's in full HD format, so the quality is better. For Cling version two, I forgot to upload the image and got a text to video instead. So, if you have good prompts, you can get okay videos without an image, but it's hard to control. This is the image to video version for Cling version 2. I got the zebra walking, while most other models had the zebra not moving forward. The wand model, I'm not sure what it did here. I thought the lady was going to have a heart attack and fall off the zebra. Juan Pro came out a little fuzzy and not so clear. I'm not sure why since it's full HD, but it looks like a bad upscale. The LTX model looks almost frozen in time, so there isn't much movement. V2 did okay. The woman is petting the zebra, but her face doesn't move much. Lumay 2 gave me more motion, but it also blurred the face, so the quality isn't quite there. Minimax also looks frozen in time and some leaves start coming out from behind the lady.

Moving to cartoon animation to see how the baby deer interacts with the butterfly. Cling Pro did a good job here. I like the result. For Cling version two, I like how accurately it follows the butterflyy's motion and how the deer looks at it, but it messed up the butterflyy's wings. For Veo 2, I forgot again to attach the image and got a text to video. Bye-bye, sweet credits. The Wan model doesn't have much motion here. And the butterfly doesn't just fly in one place. It moves more. Wan Pro played with the camera and gave me this zoom. Not sure why. And again, not much movement. LTX is really bad with butterflies, so don't waste credits on this one. V2 started more static, but then got some nice movement. Luma 2 gave me a time warping kind of movement, like it's about to be sucked into a black hole. And for Miniax, it looks like the baby deer is blind. I can't see where the butterfly moves. It doesn't look at it, and the butterfly itself isn't great.

Let's increase the complexity with a gnome chasing a bunny, then the camera focusing on the gnome to see if it can do that. Cling Pro seems to have managed what I asked and got the focus on the gnome afterward. Cling version two also did well. It's usually the best for more complex motion and prompts. For the wand model, you have to watch this. I think it's the funniest version. Obviously, it's not what I asked for, but it made me laugh how the bunny flies and the gnome runs over the bunny. Juan Pro did a little better, but what's up with those crazy bunny eyes? Now look at this. What did you do, LTX? What is that? Looks like a glitch in the system for Veo2. It perfectly captured how I run in my dreams, like running in place in slow motion. Lumay 2 has the same problem. Running in place or maybe it's dancing. Not sure, but definitely not what I asked for. And Minimax looks like it's from some horror movie that doesn't follow the laws of physics. Look at the gnome running with a twisted head.

Now, for the anime girl, I forgot to change the prompt and was using the girl with the gnome chasing the bunny prompt. Let me quickly show you how each one did with that prompt because I think it's funny to see how they handle the transition and interpret the results. But don't do what I did. Having good prompts that actually relate to your image is way better. Um, if you prompt for something that can't be seen in the image, the AI basically does some text to image and everything gets mixed up. I mean, it's fun to experiment and see things get crazy, but it's not so fun when it costs money. Run, anime girl. Run. The gnome is going to catch you. So, finally, I fixed the prompt and asked for some wind in the hair. Then, the girl should turn around and walk away. Cling Pro did a good job, but I probably should have set it to 10 seconds so it has enough time to walk. Cling version two did exactly what I asked. The girl turned and walked away. It's really good at prompt understanding. Juan model wasn't that bad. I think with anime, it's the only one that didn't mess it up too much. Juan Pro got some nice wind and motion, but didn't walk away like I asked. For LTX, I have no words. It just disintegrated the girl or teleported her somewhere in the land of forgotten video generators. Vo2 starts static for a few frames and only then begins to move. Let's hope they fix that. Luma model is all slow motion. It looks like it was trying to do what I asked but very slowly. And Minax for the first time didn't completely mess it up. But still, there are better models out there.

Let me show you one more thing. Depending on the model you use, like Cling Pro, it also has a 10-second option, but it costs double the credits. Google V2 has a slider for the number of seconds and can do up to 8 seconds, but it's more expensive. So, here's my conclusion for the video models. Starting with prices, clinging version two is the most expensive and VO2 is the second most expensive. When it comes to supported ratios, not all models support square or other ratios, but all of them support 16-9 and 9-6, which are usually used in video. Each model delivers different sizes. The only ones with full HD are Cling Pro and WPro. But as you saw from the results, Cling Pro gives better quality and is cheaper. I also rated prompt understanding and cling version 2 definitely has the best followed by the other cling versions Veo 2 and Luma for overall video quality not just resolution the winner for me is Cling Pro. So for me the one worth the money is Cling Pro and for more complex motion and prompts try Cling version 2. But keep in mind this is just at the time of recording. Keep an eye on model versions. New ones might be better. I recently saw Google V3, but haven't had access to it yet. Every month, there are new and updated models coming out.

Let me show you some steps I followed to make a video ad with AI. First, I generated a script of what I wanted using chat GPT. It took a few tries to explain exactly what I needed, and after some adjustments, I got this final script. Then, I created prompts based on that script for the images I wanted to generate. I made prompts for each scene and for a logo. I started with this bunny preparing the composition for macarins. Then a squirrel placing some macarins on a tray, a mouse pushing the tray into the oven, a hedgehog doing the filling for the macarins, and at the end, another mouse packaging the final product. Notice how I generated white boxes as mock-ups to work with more easily later. Here's another photo of a chipmunk next to a blank macarine box. And at the end, a truck driving away. also as a mockup. For the logo, I used a square ratio and the GPT model since it's good with logo design. To get a perfect 16 to9 ratio, I opened the images in Photoshop. I selected the crop tool, typed in 16-9 for the ratio. And when I cropped the image, it lost a few pixels in height. Now, the image has the perfect ratio. For example, this was the cropped image. Then I used Topaz Gigapixel to upscale it to full HD. I did the same for all the images. The better the quality, the better the video. Then I opened the logo in Photoshop and used the remove background option so I could save it as a transparent PNG. This makes it easier to work with the logo. With the blank box, I added the logo, put it in perspective, and applied a Gaussian blur so it matches the box's blur. I used the distort tool to place the image how I wanted on the box, and it looks like it belongs there. I did the same for the other box. Sometimes when you have shadows, you can use multiply blending to blend the logo better with those shadows. If the product is white, you can even reduce the opacity a little. The truck had more complex lighting, so I used a clipping mask to place the background on top and set it to multiply to pick up the shadows. If it looks too dark, you can add another background on top with a clipping mask set to screen mode and reduce the opacity. Now it looks pretty realistic. I used the Cling Pro model to generate all the videos from those images. Then I used 11 Labs to generate the audio. You can also generate audio directly on the Nim video website, but I wanted a different voice they didn't have. After that, I used Sunno AI to create an instrumental song. So now I have a complete project of around 30 seconds. Let me show you how it looks.

In a cozy kitchen, magic begins with gentle paws and a whisk. Colorful moments crafted with delicate care. Baked softly, one sweet story at a time, filled with joy, curiosity, and a touch of whimsy. Packed with love and a dash of woodland charm. Open the box, taste the magic. Macaron tales delivered fresh from our kitchen to your door with warmth and [Music] wonder.

Thank you for staying until the end. I'll add a link in the description where you can get some free credits on the Nim video website. Leave a like and a comment if you found something useful. It took me a few days to test everything and edit the video. Let me know which model is your favorite. And don't forget to check Discord for prompts and more resources, all free. I also want to thank everyone who subscribed, became a member, and supported this channel. Thanks to AI Titans for their continuous help. I wish you an amazing day, and I'll see you on Discord.