Transcription
In this video, we will take a detailed look at the recently released Alibaba 1 2.1 image-to-video model. In the prompt battle, we will test one against Clink, Minimax, Halio, Lumaray, Pika, Huan, and Skyres. Let's get started.
First of all, a quick reminder for battle rules: no rerolls allowed. I always took the first result coming from AI video generators. Second, for initial images, I used the Magnificent Mystics fluid model.
The first challenge is about a woman applying lipstick in front of a mirror. It's difficult to render this for AI video models because the model needs to give us a natural-looking application of lipstick and correctly render the moment. Additionally, the reflection of our subject on the mirror must be correctly rendered as well.
The Clink output gave us a natural-looking application of lipstick. Lipstick is touching the lips, and the reflection of my subject on the mirror looks accurate enough. I'm generally happy with the Clink's output; I think it did a good job.
Jumping into Lumaray, one thing you will realize is that an additional hand-looking thing appears on my subject's throat. This makes the shot pretty much unusable. Additionally, the hand's position while applying the lipstick doesn't really look natural, and the lipstick doesn't really touch the lips; it looks quite unnatural. One thing I appreciate, though, is that even though in the initial image, the ring on my subject's finger doesn't exist on the reflection, Lumaray 2 was able to recognize that and put the ring on my subject's finger. From my perspective, that was impressive. The rest of the shot is actually quite problematic.
The One output looks natural; of course, it's not as smooth as the Clink and the Lumaray outputs, but at least the lipstick is touching the lips, and the reflection on the mirror is accurate enough. The metallic case of the lipstick is quite shiny, and that was interesting because it was able to understand this metallic part and give a shine to that. So, in terms of understanding materials and light, it's very impressive. So I think One in this challenge did a good job.
In the Huan output, unfortunately, the lipstick is not touching the lips, and it just looks like she's randomly moving the lipstick right and left. It looks quite random and unnatural. It just seems like the model didn't quite understand the prompt. Also, what's going on in the initial image because it gave us this unnatural random movement?
The Skyio output has a similar problem; it looks very unnatural. The lipstick doesn't touch the lips, and again, it looks like she's just randomly moving her hand, just like circling around but not really applying the lipstick.
The Pika output is quite chaotic as well. At some point, you will realize that she's putting the lipstick under her cheek. It's a very, very unnatural and not really a good output here.
Halio Minimax decided to give us a random zoom in to my character's face. While doing that, my character started having a stroke or some kind of shock or paralysis because she started to just shake randomly. I need to say it looks quite strange; she really doesn't look well. Unfortunately, the Minimax Halio output won't be usable in the next prompt.
We are starting with a closeup of a candle in focus. The focus shifts to a woman walking towards the candle in the background.
Starting with the Clink output, unfortunately, my subject started to walk quite late in the shot. The first 2-3 seconds of the shot, she's just frozen. I have a feeling that if I would extend this shot, most likely Clink would be able to do the focus shift, but with the first roll, unfortunately, what I asked for is not there, so not the best output from Clink.
The Lumaray output managed to give me my character, and she's somehow walking very unnaturally, but then instead of giving me the focus shift in the middle of the shot, it decided to cut to the closeup of my character, which is not what I asked for.
In the One output, the walking of my character on the background looks really unnatural and unfortunately didn't really understand my prompt and what I asked for.
Huan decided to choose a very unorthodox direction. Instead of bringing my character close to the candle, it took the candle towards my character. It's a very interesting approach, but unfortunately, it's not what I asked for.
The Skyres output is again unusable. In the middle of the shot, it completely changed my subject and room design, and the context changed abruptly.
Pika did a very good job with this challenge. You can see that my character is walking towards the candle, and then the focus changes, and my character's face slowly comes into the focus. It's a very good cinematic-looking shot, and I'm very happy with the Pika result here.
Similarly, the Minimax Halio shot did a good job. It's not as smooth and cinematic as Pika, but at least it understood the prompt, and it gave me a little bit of an unnatural walking of my character, but after that, it gave me the focus shift that I asked for, which was the main point of this challenge, so I'm happy with this result.
The next prompt is: woman swings the sword dynamically fast while the camera pushes in. Here we have a physics and anatomical challenge together with a fast-moving action, which is swinging the sword, as well as a prompt understanding challenge because while asking this dynamic movement, we are also asking for a camera movement. We want the camera to push in, get closer to the subject.
Starting with the Clink result, I am generally happy with how smooth this shot looks, and the movement looks quite natural. Of course, we have a slight coherence issue; especially, you will realize in the last frame that half of the sword is gone, but overall, I think it's a cool shot with some slight coherence issues, and it also gave us the camera pushing, which is a thumbs up for Clink in terms of prompt understanding.
Lumaray gave us a slow-motion-looking footage, but I think it looks pretty cool. I wish I would have two or three seconds more so we could see the end of the movement. Again, in the last frame, we lost a little bit of coherence, but I think overall, the shot looks really good. The slow-motion part, which is not something we asked for, but in the end, it looked pretty cool.
I'm not so sure what went wrong with One, but um, I actually used the same exact prompt which I used with other models. I would say it's quite dynamic, but unfortunately, I don't know why her outfit changed to red; completely lost coherence, and this looks more like anime aesthetics. Unfortunately, it's not usable with this shape and form; requires a reroll.
The Huan output looks like the whole army is dancing with the sword and spares. It looks quite funny, and she looks kind of shocked, and this shocked face with this strange dance makes the footage even funnier. Unfortunately, not usable, but super funny.
A similar problem with Skyres: we completely lose the coherence; we didn't get the camera push in, and suddenly, the whole shot became red. I would say the major coherence issues here make this shot unusable.
A similar problem with Pika: we didn't get a natural swing; we lost the coherence of the body; the head moves separately from the rest of the body. Instead of camera push in, we actually got a camera push out, which is a thumbs down in terms of prompt understanding.
Minimax, on the other hand, gave us a camera push in, but because of some reason, our subject is also escaping from the camera, and it gave us multiple swings instead of one. And I'm pretty sorry for the soldier on the right side because it's most likely that she actually hit that soldier with the sword. I'm not 100% sure that this shot is actually usable.
Jumping into the next challenge, we have giant tsunami waves hitting the city. We are looking into the rendering of natural elements and how accurately the models render the natural elements. We have also got, like, smoke, which can be confusing for models, as we accepted.
Clink mixed up smokes with waves. This doesn't look like tsunami waves hitting the city; it still looks more like smoke to me. I can say not the best job from Clink for this challenge.
Lumaray gave us some waves, but it's not really perfect. It's still really mixed up with smoke, and the wave is really minor; it's not a giant tsunami; it's a small wave.
One output destroyed the whole city with the giant waves, but it's kind of looks cheap and fake, and the last frame looks like waves are destroying the camera, and then we have all these cracks. It just looks odd and strange for my taste.
Huan decided to not give us anything, and then the model is like, "Well, if you want some waves, here's some waves for you," and cut to an unrelated scene. It just feels like it completely didn't understand what I asked for.
Skyres, same problem: let me show you waves. Then again, it's not what we asked for, okay?
Pika definitely gave us massive tsunami waves, but after that, we lost coherence and switched to almost like a different city; night becomes day. Many coherence issues.
When I look at all of the outputs, in the end of the day, the Minimax output is the best one. I would say it gave us giant waves of tsunami covering the city, and it somehow gave us what we asked for. Uh, therefore, I think Minimax did the best job here.
The next prompt is: woman rows in a river while standing on a crocodile. Here we have a challenging physics and anatomical render as well as a challenge for prompt understanding. The model needs to be able to understand that the woman is standing on a crocodile and, in the same time, rowing.
For Clink, I can see that the model was able to understand what we provided as an initial frame. It gave us a natural-moving crocodile, and the woman's movements are natural and nicely rendered. Slight coherence issue with the hands and with the paddle; some coherence issues are noticed, but overall, a pretty good shot here.
We have the Lumaray output. It's definitely not bad. It seems that the model was able to understand what's going on in the initial frame. The problem is the pedaling or rowing action doesn't look really natural; the pedal is barely touching the water.
The One shot seems that it was able to understand what we asked for. The only problem is it looks far too fast; unnatural indeed. This crocodile must be the fastest crocodile ever lived.
Coming back to the Huan output, it just didn't work out. The crocodile is not moving; the pedal changes constantly all the time; many problems.
Skyres did initially good, but after that, you will realize that the pedal doubles, and we have two pedals in my subject's hand. It's not horrible; I think Skyres did an okay job. It just feels like the rowing action doesn't really look natural.
For Pika, we completely lost coherence. You will realize so; the pedal disappears, the woman disappears, and then the crocodile becomes a monster, and we are losing the coherence of the crocodile's face; many issues, unfortunately.
The Minimax shot actually started well; then we lost the coherence of our subject's face; then we got two pedals instead of one. The coherence issues are making this shot again unusable; the rowing action is pretty much non-existent.
Jumping into the next prompt: a man free falls while the camera dynamically follows him. Here we are looking into the model's understanding of physics, wind, and of course, prompt understanding because we want the camera to follow our subject.
Starting with the Clink output, I think it did a pretty good job. The wind on the subject's clothes looks very natural. The camera's dynamic tracking is here, but I would prefer Lumaray's dynamic tracking here. It feels more dynamic, and the camera is tracking our subject in real time; looks much more natural and close to reality for me.
And I don't know what to say about One's output here, like what happened here? The helmet took off, and suddenly, the temperature changed; it became super hot and started to sweat. I don't know what's going on in your mind, One, but it's super strange, really.
The Huan shot is kind of fun; it made me really happy to see my subject so happy, but in comparison to Clink and Lumaray outputs, it's just not there; it doesn't look really natural, but it's still kind of fun to watch this.
Sky's output: Oh my God, what a struggle, guys. Really, like, didn't just fall to the ground; also continued struggling on the ground. This is really hard to watch.
Okay, the Pika output did okay. The only problem is in the second part of the footage; we completely lost coherence; you will realize especially problems on the legs. It, it is definitely okay, but not as coherent as Clink and Lumaray in this challenge.
The Minimax shot made me think that he just looks like he's frozen; it's like time stopped. I mean, my prompt was: man free falls while the camera dynamically following him. I cannot see this on Minimax, really. It just, this shot doesn't look dynamic at all.
Jumping into our last challenge for today: we have a man; he's standing on the street; then rain starts, and the man starts crying while the camera slowly focuses to his face. Here we are looking for the rendering of natural elements; in this case, rain, and we are looking for an accurate render of human emotions.
Starting with Clink, we got the camera focus to his face, and we got the crying of the man, but unfortunately, the rain part was completely ignored; we didn't get the rain for the Clink output.
Lumaray gave us rain, but then ignored the crying part and also didn't focus to his face. It's almost like these models are doing one thing correct and just missing the other one. So One gave us rain; this is one giant teardrop that he's crying from his left eye, and it looks like a giant waterfall; not the best output here.
Looking at Huan, we didn't get neither crying nor rain, but at least it gave us a little bit of focus to his face.
Sky, if Sky could keep this coherence, it would be actually accurate in terms of prompt understanding because we have rain and crying, but this jump from the first frame to the second is really drastic.
Well, Pika gave us what we asked for; we got the rain, and we got crying; prompt understanding is there; just we didn't really get the camera focus to his face.
And Minimax gave us camera focus, crying; I cannot really see the rain that I was looking for in this challenge. Almost all models missed one part or multiple parts of the prompt.
So let's look at my final verdict. I would like to start by saying that no model is perfect. Looking at image-to-video performance, there are models who excel in certain things, and some models are lacking on those frontiers. So no model is perfect, and this is my personal opinion.
Looking at the capabilities of models and where they are in competition, when I look at the results, I can still clearly say Clink 1.6 is still leading the competition. Of course, Clink is not perfect as well, but I would say it's in the leading edge of the current AI video models competition.
On the other hand, Lumaray 2, Minimax Halio, and One 2.1, these models are also extremely good models, and they are highly competitive, but they are not leading the competition at this stage from my perspective.
Sky, Huan, and Pika 2.2 are still lacking significantly in coherence, and they require further development to catch up with the competition. This is, of course, the image-to-video performance, and I'm looking at only one dimension. There are many dimensions for these models, of course. You would look at the workflow, for example. Clink and Minimax Halio are offering great camera controls. Pika is going into a different direction; they are introducing new features, a lot of fun features that people can engage with. I think looking at all these solutions holistically requires further, a deeper investigation, and there's also the whole element of pricing because some of these models are actually open source; they are free to use, and some are paid. But if you ask me what's your opinion in terms of image-to-video performance, this would be my ranking. But hopefully, this video was truly helpful for you. Don't forget to give a thumbs up and subscribe for more in-depth tutorials. If you want to learn more about creative intelligence, click here.