Transcription
Are you ready for another AI tutorial? Today I want to talk about the MidJourney video generator that lets you create a video from an image. We'll look at what it's good for, the quality, the price, and a few tips and tricks.
At the top, you can see videos created by other users or images that people have made. This is a great way to get inspiration. You can see what prompts they used and either use the same ones or adapt them for your own images. I've been using MidJourney since the beginning, and while there were better options out there for image generation, the new video model has caught my attention again. So, here are some important things to know.
For me, aside from quality and prompt understanding, price is a big factor, specifically how many videos I can generate and how much it costs. There are four plans available. If you just want to test it quickly, you can try the basic $10 plan, but they offer two types of generation. Fast mode, which consumes credits, and relax mode, which is free but slower since you have to wait in line. If you want unlimited image generation, the $30 plan is the best option. But if you're looking for unlimited video generation, I suggest the pro plan. I'm not trying to sell you. This is not a sponsored video, but for me, the only plan that's really worth it is the one with unlimited video generation in relax mode. I went for the yearly option because it's cheaper that way, and I can create unlimited videos in relax mode, at least for now. Always check what's included in case they change that in the future. You also get stealth mode in the pro plan, which means you can generate images and videos in private, so only you can see them.
Let me show you how you can create images and videos. Click on create and at the top you can type what you want to create. So if I use a prompt like a cute cartoon cat and press enter, it will create four images for me based on that prompt. No matter what you're doing, image or video generation, you always get four images to choose from. It takes less than a minute to generate them all. Once you have the image, you can animate it using this button. You can only generate videos from an image. It doesn't support text to video, but I never use text to video anyway because I want to control the look.
You can also add your own images to create a video. If I click on this button, I can select the starting frame. I'll upload this portrait of a woman. You have to click on the image after uploading so you can see it appear here and here. If it's not selected, it won't work. Then you add a prompt describing what the character does or the camera movement. On the right, you have the settings. If you want slower or faster motion, I like to use low to avoid mistakes. For speed, fast uses credits but is quicker, while relax doesn't cost credits but takes more time to generate. I have stealth on so only I can see it. Click on that arrow and it will start generating. It takes less than 2 minutes to get four videos, so it's pretty fast. Here are the results. Four video generations from that single image, so we can choose the best one. All of them seem to keep the character consistent and look natural. You can use the mouse wheel to move from one generation to another. So with portraits, it seems to work quite nicely. If you don't like any of them, you can rerun it to get another four generations. You can use that start frame, the image that was used for the video, and then type a different prompt. You can extend the video by another 4 seconds up to four times. Use slow motion for subtle nice motions and high motion when you have an action scene or need a fast moment. You can also use slow motion manual and it will move slower but lets you describe the motion. So, I'll click on manual low motion and then here after smiling, I'll add the next motion. Let's say she plays with her hair. Let's submit that to see what we get. I'll close this so I can see what it does. You can see how it says video extend and shows the prompt used. Now it will be 9 seconds instead of 5 seconds. You can also right click on a video and see all kinds of options to extend or download it. Let's download the raw format just to see the differences. And then I'll use download for social to see how it differs from the raw version.
If you look at the raw video, that's the original size the AI generated it at. So the dimensions aren't great. I think this is the biggest downside of MidJourney. Cling AI and others can do full HD videos, and I'm not sure why they keep it so small. But if you select for social, they give you a bigger image, full HD, but it's not truly upscaled. It's like quickly resizing a small image in Photoshop. You can use an upscaler like Topaz Video AI. For shorts, it's good enough, but I hope they fix this in the future to get bigger resolution for videos.
So, if we look at the extended video first, we see the original 5 seconds, and then she's playing with her hair, just like I asked. I'm surprised how good it is with the hands. The motion of the hand and the hair movement look quite nice. It depends on the motion. It's not always perfect, but it does have good prompt understanding.
Let's see how we can first create the image and then the video. Let's type a prompt like a cute cartoon cat 3D render maybe dressed as a pirate. Always check the settings. You can choose different ratios like portrait, square, or landscape. And you can move this slider to see all the available options. Let's say I want to create an image for a short, so I choose a 9 to 6 ratio. You can leave the model set to default. The latest model is usually the best. You can also choose the speed. Remember that Relax doesn't consume credits if you have the right subscription. Let's run it to see what cats we get. For prompts, you can also use chat GPT to generate them or go to explore to see what prompts other people used and adapt them to your needs. So, prompting is not really a problem these days. Here are the cats generated. All of them look pretty cool. Let's say I want to make a video from this cat. You can right click and choose an option. Maybe you want a different variation first or to upscale it. We also have options here. Just remember, since this is an image, these options are for generating a video, not for extending one. Let's test all of them. Low motion, high motion, and maybe one with manual low motion. I'll adjust the prompt like she takes the hat down and begs for food or something similar. I noticed it took longer to start the job and saw it remained on relax mode when I showed the settings. Since I need it faster for the tutorial and don't want to wait too long, I'll switch to fast mode. This uses credits, but it's quicker. I'll use this cancel button to stop the current generation. I'll do the same for all of them. Then switch to fast mode and test again with low motion, high motion, and manual low motion. Now, it's already started, so it's generating pretty fast. Here are the results for the auto slow motion. No prompt, just what the AI thinks will work for that animation. and it looks pretty good. For high motion, it moves a little faster, but can sometimes introduce mistakes when it moves too quickly. And for the low motion manual, when I asked it to take the hat down, it almost did it correctly in version 3, but it left a piece of the hat on the head, so it doesn't always work perfectly. Maybe if I do another generation, I'll get luckier. They still need to improve the model to follow the laws of physics better.
Let me show you a useful feature. Go to the personalize menu and look for mood boards. Here you can create a mood board to define a style you want the AI to follow across multiple images. I'll upload some images. I made some icons with chat GPT on a white background. You can upload more images, but for now let's test with these four. You can change the title of the mood board to something memorable. If I click on use and prompt, it gives you a code that you can add at the end of your prompts to capture that style. Here's how I use it. I go to create and where it says P for personalize, I turn it on and select one of the mood boards I made, like the icons mood board. When it's active, the P is colored. Now, when I generate something, it should follow that style. Let's test it with a simple gem. I won't mention the white background just to see if it picks that up from the style in the images I uploaded. I'll choose a square ratio since it works better for icons. And let's test it. You can see it shows here that it used a mood board. I kind of like this one. It looks pretty good. Let's see if I can animate it. I'll try one with auto to see how the AI thinks it should be animated. And I'll try another with manual low motion where I want a 360° rotation. Let's submit and check the results. Since this was just a simple test, I will turn personalize off so it doesn't influence the rest of the generations. Here is how the animation looks for autolotion. It looks pretty nice, like an artifact in a game inventory or something. For manual low motion, I got some rotation. Even if it's not a full 180°, it's more like 180°. It probably didn't have enough time to rotate more in that number of seconds. Some versions have errors, but some I could work with. It's useful for getting an object or character from different angles while keeping it consistent.
Let's test a few more images to see what else it can do. I'll upload this burger photo, select it, and then write a simple prompt for what I want. While that is generating, let's try another one for an anime girl on the beach in an exotic location. Then maybe something more realistic, like a man holding a sign. and maybe one of a woman next to a window with rain outside. While the last images are being turned into videos, I already have the first generations ready. Here is what I got for the burger video. I think I can work with one of the generations. Maybe this one. It looks interesting. Usually from the four generations, I get one that has fewer mistakes and looks good. Like in this case, the girl is moving and we have nice fire and water. Even though the fire isn't quite perfect, it still looks pretty good. This one could work for memes and funny videos. The more realistic the image you give to AI, the more realistic the video will be, except when it makes motion mistakes that don't look natural. But I like how this old man with the sign came out. And for the woman in front of the window, it also looks quite nice. I expected to see more rain outside and water drops dripping down the window. But I probably need to give more complex and detailed prompts to describe that kind of motion.
Let me show you also what kind of problems you can get sometimes if for example I upload this portrait of a cartoon girl I did with AI and let's say I prompt for a cartoon girl looking at the camera and smiling. Look what happens when I submit. I get an error like sorry the AI moderator is unsure about this prompt and obviously I didn't break any rule. It's a simple cartoon and nothing should have triggered that. But AI can make mistakes sometimes. Let me try something with a mummy skeleton that I expected to trigger the AI moderator, not the cartoon girl. But for this one, it didn't say a word, so it went and generated the image. So, just something to be aware of. Didn't follow exactly the prompt, but the motion is okay. I like that hands are kept consistent. Let's try to extend this video using manual low motion. Let's say a door is closing behind the mummy. And the results are these. It did add a door, but it's closing and opening in all kinds of strange ways. So, if we look first, we have the previously generated part, and then comes the part with the door that should be closing. But I guess in the fantasy world, anything is possible. Maybe in a video editor, you could cut it at the part where the door is closed. Save that last frame and start another video from there using it as the first frame.
I have this cute bunny that I want to test walking in a fantasy forest. We can drag this image here over the image where it says starting frame. Then I'll write this simple prompt and submit it to see how MidJourney does it. Then I want to do the same thing with Cling so we can compare. I'll go to video then image to video using the latest model. For the start frame I'll upload the same image and add the same prompt. Let's generate it. Says it takes around a minute. So sometimes it can be faster than MidJourney. It does full HD, but it's also more expensive and they don't offer unlimited generation. And let's test it also with Hyuo version 2. For this one, I don't have a subscription on their site. I already have too many subscriptions, but I can use it on Open Art with different models, including that one. So, same image. And let's test it for MidJourney. It seems like it did a good job with the walking. It's not something complex to walk, but sometimes it can fail. The fancy world stays consistent as it walks, so that's a good thing. I like that I can choose the best version from all of these. In other AIs, I have to generate it a few times until I get one I like. Let's check what Cling AI did. It walks and looks around. Looks pretty cool. Did a great job. And it's full HD. I wish Cling AI also offered unlimited at a decent price. For Hyo and Open Art, it seems to take longer. And I just saw I forgot to add the prompt, so it finally finished the video. The motion is quite cool, but I want it to walk. So, let's add the same prompt to make it a fair comparison. Spending a few more credits, and the result is this one. It starts to walk, comes toward us, then turns around. No, where are you going? Come back. All AI have strengths and weaknesses. It depends on the image and the motion you want. Sometimes MidJourney does better, sometimes Cling, and sometimes Hyuo, so it's hard to compare. MidJourney's advantage is the price if you go with unlimited and it's prompt understanding. I hope they fix the resolution problem since nobody wants small videos.
Here are a few more ideas. I use my avatar with autoload and got all kinds of motions from that image. But when I use manual low motion and added a prompt like transformer robot front view video outro, look what interesting results I got. It's quite nice for video intros and outros or GIF animations for the web. You should try it with your logo or products.
Here is another cool thing. Go to image and then go to omni reference upload an image. Let's say I use this bunny. And then I select it. If I click on the image, you can see I have this white bunny with blue eyes. And here I have a slider. You can choose the strength of that image. I will leave it at 100. And let's see if we can generate a similar bunny now that we have that reference image there. I want a white bunny with blue eyes walking on a beach on an exotic island. So let's try it. So this is how the original looks. And this is how the generation looks. It's pretty close. I got different variations of that bunny in the environment I asked. Let's try to animate one. I will use low motion manually. And I see it already had the walking from the image generation. So let's go with the same prompt. And these are the results. Not all are perfect. For some it's walking, but it doesn't move on the ground like walking in the same place. I think this version could work. This version looks like it's slippery sand.
That is all for today. If you noticed, I changed the voice for this channel since the Pixarama channel already cost me a lot for the 11 Lab subscription. So, for this one, I tried a cheaper website that still sounds okay in my opinion. I will test it for a month to see how it goes. I want to thank you all who subscribed, became members, and support this channel. Many thanks to AI Titans for the continuous support. If you found something useful, please leave a like and a comment to help with the YouTube algorithm. Thank you, and I will see you on Discord.