Transcription
AI never sleeps, and this week has been absolutely insane. We have not one but two new AI tools that can seamlessly add reference characters or objects into images.
This one is even crazier. You can add reference characters or objects in any video and get them to do anything. Or you can even replace any object in a video. This AI can magically erase anything, even if it's a really complex and crowded scene. And then we have this new AI which can create 4D scenes that can be used for virtual and augmented reality. Plus, we have an open-source robot which anyone can create using a 3D printer and a lot more.
So, let's jump right in. First up, this AI is super powerful. It's called Dream O. And this can create images using reference photos of any character or object. And this is incredibly accurate. For example, if you have this image of a pig character and you prompt it with, "He is driving a fighter jet in the sky," here is what you get. Or if you have this plushy and you prompt it with, "A toy holding a sign saying Dreo on the mountain," this is the result. And look how accurate this character is. It looks exactly like the toy. Plus, you can even add multiple reference images or objects in the same photo. So, here's an example of that. And again, note how accurately it portrays both characters in the final image. Here's another example of adding two characters in the same photo. You can also do something like this where you have this woman ride this giant dog in a noisy modern city. It's also good with passing the style of one photo onto another. So, let's say you have this castle and this colorful smoke. Well, you can apply this style to the castle and get something like this. And of course, you can add a hat and sunglasses to a character like this. You can also change the style of a photo. So, if this is your original image, you can turn her into pixel art like this. Or here's another example where we get this woman to hold the toy above her head in the park.
So, not only is it good at transferring reference images of objects or characters, but it's also good at prompt understanding. The nice thing is they've released a Hugging Face demo for you to try this out. So here is where you would upload one or two reference images. And then here's where you would add the prompt. Here's where you would specify the dimensions of your final image. And then the number of steps is basically how many iterations you want the AI to go through before generating your image. In general, the more steps you have, the better the quality will be. But at a certain point, you're going to get diminishing returns. So it seems like the sweet spot here is just 12 steps. And then guidance is how literally you want the AI to follow your prompt. So a higher value means it would follow your prompt very literally, whereas a lower value would mean it can be more creative. I'm just going to leave this at the default of 3.5. Let's try a few examples. I'm going to upload this image of a woman and then I'm going to write, "She is wearing sunglasses on a beach." And then let's set this to 1024 by 768. And then press generate. And here's what we get. So as you can see, it has isolated the character from this image. And then it has made her wear sunglasses on the beach. Very nice. Let's try another one. I'm going to upload this and then upload this GPU as the second image. And then I'm going to write, "The cat in the white tuxedo is holding the GPU." I'm going to set the width to 768. And let's see what this gives us. Very nice. So, here we go. Here is indeed the same cat that was in the reference photo. Plus, we have the GPU over here. Very impressive.
The nice thing is they've released everything already. So, on this GitHub repo, if you scroll down a bit, it contains all the instructions on how to download and use this locally on your computer. So, if you're interested, I'm going to link to this page in the description below for you to read further.
Next up, this AI is super cool. It's called Hollow Time, and this can generate 4D scenes from a single image or text prompt, allowing for immersive experiences in virtual and augmented reality. Now, a lot of people might get confused by the term 4D scene. So, let's clarify this really quickly. A 4D scene is basically a 3D video. And the fourth dimension here is time. So here are some examples where you can upload one image and it would generate a 4D scene from the image. So note that this scene is 3D. You can like use VR glasses or whatever to move around the scene. Plus the video is moving as you can see from the waves. So this makes it a 4D scene. Here's another example where we can upload this panoramic image and it can create a 4D scene from that. Or here's yet another example. So instead of spending hours or weeks manually building out a 3D world, you can just plug an image into this AI and it would create a fully immersive and moving 3D world for you. Here's yet another example. This one is pretty cool. And then here's an example of some northern lights in the sky. And not only is it able to generate a 3D world, but it's also able to animate the northern lights pretty realistically. Now, instead of generating 4D videos, you can also get it to generate panoramic videos by uploading a panoramic image. So, here's an example of that. And as you can see, it does a great job animating the cars on this road. Here's another really tricky example of this pano image of fireworks. And the AI was able to animate the fireworks pretty well. Although, one flaw I can spot is that the humans aren't really moving. And then here's another really cool example of these people around a campfire. And then here's another example. And like I mentioned earlier, not only can this just take in an image and generate a video from that, but you can just use a text prompt to generate a panoramic video. So here are some examples. "Here's a cartoon style cloud city where floating buildings drift among fluffy clouds as airship pass by." Here's a really impressive pano video showing the Shibuya crossing. And note that this was generated from just a text prompt. They didn't even feed it a reference image, but this looks so realistic and detailed. Here's another one. "A bustling marketplace in a desert town. Merchants shouting and colorful fabrics swaying in the hot wind." And that's indeed what you see. Here's another one. "A sci-fi energy facility where blue energy pulses through glowing tubes." And that is kind of what you get. Plus, this scene looks super realistic. Here's another really nice example. So, while we have a ton of video generators already, most of them can only do like 16 to 9 videos at the widest. But this one can do like full panoramic videos like this, which is super impressive.
Really quickly, here's how it works. So, Holot Time uses a two-stage process. First, it has this panoramic animator component which converts your panoramic image or a text prompt into a high-quality panoramic video. But then this also gets fed into this panoramic space-time reconstruction component which transforms the panoramic video into 4D scenes that can be viewed with like VR headsets. And it does this using this space-time depth estimation component. Anyways, if you scroll up to the top of the page, the nice thing is they've released everything already. So, the models are all on HuggingFace. Plus, if you click on this GitHub repo and scroll down a bit, this contains all the instructions on how to download and run this on your computer. Props to them for open-sourcing this. I think this is a really cool technology. Anyways, I'll link to this main page in the description below for you to read further.
Next up, this AI is also super cool. It's called Flexi Act. And this allows you to transfer the movements from one video onto another video. So, here are some examples. Let's say you have this reference video on the left. Well, you can just input an image of any other character and plug it into this AI and it would map her movements onto the new character like this. And it doesn't matter if it's realistic or 2D or 3D. It can animate the character pretty well according to the reference video. Here are some other examples. For the first row, if we have an input video of this woman doing a squat, well, we can transfer her movements onto other characters with just an image. So, for example, if we plug an image of Sam Altman doing a speech on stage into this AI, we can get him to do a squat instead. And we can also do the same for Trump. Or in this middle row, we have this woman boxing. And again, we can transfer her movements onto any other character with just one input photo. And then finally in the bottom row, we have the same thing. It doesn't matter if it's a 2D image like Mario where his body is like way more compact, but it's still able to transfer the movements of this woman pretty well. The cool thing is it doesn't just have to work with humans. You can also take a reference video of animals moving and transfer that onto other animals. Plus, notice that it can even work with different angles or perspectives. So for example, the Pomeranian in the top row, it's looking to the left. Whereas for our input images, they are all looking to the right, but it's still able to transfer the movements of, you know, the animal getting up with just its hind legs. Super impressive. And you know, I really love the last row here where we have the input video of this kangaroo or is it a wallabe? And it can transfer the hopping movements onto these birds. A super useful tool. It even works with more complex poses like yoga or working out as you can see from these examples. And here's the crazy part. You can even transfer the actions of humans onto animals. So the top row is the reference video and for the bottom row we've input images of animals and it's able to transfer the movement of the human onto the animal. So, like for example, in the first column, it's actually getting the tiger to kind of do a handstand. Or for the other two columns, it's actually getting the dog or the wolf to do these yoga poses. That's pretty crazy. I mean, with this tool, you can potentially film yourself doing something and then map it onto the movement of your pet. And then here are just a few more examples for your reference. I mean, with this tool, you can easily control the movement of anyone in the video.
And really quickly if you want to get into the nitty-gritty technical details of this, Flexi Act has two main components which are responsible for the magic. So the first one is this ref adapter component which helps adapt the spatial structure of the reference video to the target image. And then we also have this frequency aware embedding component or FAE for short which extracts the actions from the reference video and applies them to the target image. And this method also preserves consistency and flexibility even though the target image might have a different body composition or camera angle. Anyways, if you scroll up to the top of the page, the awesome thing is they've also open-sourced this. So, the models are all out here on Hugging Face. Plus, here is the GitHub repo, which if you scroll down a bit, contains all the instructions on how to download and run this on your computer. Plus, they've even provided scripts on how you can train and fine-tune this as well. All the links are up here, so I'm going to link to this main page in the description below for you to read further.
Next up, this AI is freaking legendary. So, it's called Hunyan Custom by the one and only 10 cent Hunyan team. This is a super powerful way for you to add reference characters or objects in videos. So, here are some examples. If we input this photo of a girl and then we prompt it with, "The girl plays house with plush toys in the living room," this is what we get. Look how accurately it generates the girl. She looks exactly like the reference photo. Or here's an even more impressive one. Let's say we have this input image. And then for the prompt, "She is taking a selfie in a busy street. She holds a smartphone in one hand and makes a peace sign with the other. The background is a bustling street scene." And indeed, that is what we get. Plus, notice how her outfit and the text Hunyen on her shirt remains consistent throughout the entire video. This is insanely good quality here. If we upload this image of a poodle and then we write, "A dog is chasing a cat in the park," that is indeed what we get. Plus, you can upload multiple reference images to add in your video. So, here we have this woman holding a paintbrush and drawing a picture of the cat. And that is indeed what we get. Here we have a dude presenting the chips on his hand beside a swimming pool. And both the dude and the bag of chips look exactly like the reference photo. "Here on the modern city street, a man asks a woman for directions, but she doesn't understand what he's saying." And this is what we get. Although, if I were the doctor, I wouldn't be asking for directions. I would just ask for her number. And of course, with such a powerful reference tool, you can also do a close swap on the character. So here we have this woman wearing this hanfu while reading a book in the study room. And look how accurately it portrays the woman plus the hanfu. Like all the details are exactly like the reference photos. And it gets even crazier. So you can even plug in a video and change parts of the video. So let's say we have this input video. Well, we can get him to wear this hat instead. And this is our final result. How cool is that? Here's another example where we have this reference video and if we want to swap out the teddy bear with this husky plushy, that is indeed what we get. Note how seamless and accurate it does the swap. This is so impressive. Here are some more examples. So, you can see with this tool, consistent characters in videos are finally here. All you need is a reference photo and you can get that character to wear any outfit or be in any setting and do anything. I mean, within the next few months, I predict that we're going to see full AI-generated short films that actually look good with consistent characters and a good storyline. Here are some additional examples of having multiple reference images in a video. And I really like this one. This looks like a legit boxing fight with a panda. Like they are actually sparring with pretty fast punches. Super cool. And as you might have guessed, this tool is going to transform the advertising industry. You don't need to hire any actors or any videographers or anything like that. You just need a photo of the model or character and then a photo of the product and then you can get them to do anything with the product in a video. Here's another cute example of a dude with a penguin. By the way, aren't penguins like the cutest things ever? Comment below if you think so as well. Here's another example of an insane video swap. Let's say we want to swap this left clownfish in the reference video with this white fish over here. It pulls this off seamlessly. And you can even do this for anime characters like this. Notice it swaps only the character while preserving all the other details like the text in the video.
Oh, did you think we were done yet? We're not done yet. This also can do lip-sync. Let's say we have this reference photo of a woman and we prompt her to be in a dressing room holding a lipstick. We can also add an audio clip to this and it would lip-sync her to the audio. Let's hear what this sounds like. Absolutely unreal. Here's another example where we get this image of a man to be at a store counter holding a mechanical watch. And if we add an audio clip to this, here's what that looks like. This is such an insane tool. Here's another example where we get this woman to hold a cake in a bakery. And if we input some audio, here's what that would sound like. Here's another example. I have no idea what he's saying, but he seems pretty passionate about it.
And you know the best part is they've released everything already. So you can potentially run this offline for unlimited times. All the models are on HuggingFace. Plus, if you click on this GitHub repo, it contains all the instructions on how to run this. However, note that at least for now, even the low-resolution version requires 60 GB of VRAM, which I'm sure most of you do not have. In fact, they tested it on a machine with eight Nvidia GPUs. However, that being said, because this is open-source, I'm sure the open-source community is going to act fast and quantize or compress this to be able to run on much lower VRAM. We've already seen this with Frame Pack, which is another tool that uses Hunyen, and the original Hunen requires 60 GB, but Frame Pack can be run on as low as 4 GB of VRAM. So, it's only a matter of time before we can run Hunyan custom as well. Anyways, for now, if you're interested in reading further or checking out more examples, I'll link to this main page in the description below for you to read further.
In other news, we have a new open-source video generator this week. It's called LTX Video 13B version 0.9.7 by Litrix. And not only is the quality great, but it's blazing fast. Up to 30 times faster than other competitors. And that means you don't have to wait hours on your computer just to generate a few seconds of video. This can handle it in minutes. So, it's the best of both worlds. You get quality and speed. In particular, it has this multiscale video rendering feature which generates each clip from coarse to fine details. Plus, they've also open-sourced these two upscalers. And they've also included Comfy UI workflows for you to generate videos using an image as a reference frame. Or you can also extend videos or you can also add multiple images to use as keyframes throughout the video. Now, I already did a full tutorial on how to install this locally and run these workflows. So, see this video if you haven't already. However, if you don't have a good enough GPU to run this offline, you can also use this via their online platform called LTX Studio. You can create storyboards or there's even an image generator. But in order to use this new LTXV13B, simply click on motion generator and then over here is where you can select the latest 13 billion parameter model. If you're interested in learning more, I'll link to this in the description below.
Next up, this AI is also super useful. It's called Pixel Hacker and this can magically erase or fill in missing parts of an image. So here are some examples. Let's say we have this image of a handbag. Well, we can paint over the handbag as you can see from this red outline and this AI would magically erase it like this. Or here's an image of a plane. Again, we can just paint over this plane and then this AI will erase it. And if we have these annoying humans ruining the scene, well, we can erase them, too. And this is what we get. Or here, let's get rid of this sign. Or we can also get rid of this human like this. Here are some trickier examples where we also have these rails. So let's see if the AI can seamlessly erase the human while keeping the fence realistic. So this one works super well. And same with this one. There are some minor artifacts over here. Plus a bit of her shoe is sticking out over here. But for the most part, it is able to pretty seamlessly erase the woman from the scene. Plus, I guess the person didn't really paint out this part of the shoe, so that's why it's still there in the final image. Here are some super useful applications for this. Let's say you're at a tourist attraction and there are just so many humans in the background, there's no way in hell you're going to get a photo without someone photobombing the scene in the background. Well, you can just plug it into this AI afterwards and erase everyone from the scene. Or here's another example where it's super crowded in the background. Well, you can just select all these humans and magically erase them like this. And here's yet another example. Here's an even busier and more challenging scene. Let's see if it can erase all these people from the scene. Holy smokes. This is really impressive. And here's another example. How cool is that? This kid over here, he looks clearly unhappy because there's way too many people. So, let's just remove all these humans. And voila. Look at that. There are a ton of examples on this page. In the interest of time, I'm not going to go over all of them, but if you scroll up to the top of the page, it seems like they are preparing to release the code and the models for this, which is fantastic. Anyways, for now, I'll link to this main page in the description below for you to read further.
In other news, UC Berkeley has developed a robot called the Berkeley Humanoid Light. This is an open-source, customizable, and very affordable humanoid robot. Normally commercial humanoid robots would cost something like tens of thousands of dollars, but this one can be 3D printed. And get this, the total cost of all the parts is under $5,000. Plus, all the hardware designs, the files for 3D printing, the software, and even the training scripts are totally free on GitHub under a really permissive MIT license. That means anyone can potentially build this and even fine-tune it and make it better. Apparently, you can just 3D print nearly all the structural parts on just a standard desktop printer. In terms of specs, it's about 8 m tall, which is roughly 2 1/2 ft. It weighs 16 kg and has 22 actuators for all its movements in the arms, legs, and torso. Its brain is an Intel N95 mini PC, which controls everything. And it has a battery which can last for around 30 minutes. It's designed to be really customizable and flexible, so you can easily change its dimensions, how the joints are configured, and you can even change the entire structure to something completely different. So, if by any chance you're interested in 3D printing these in your secret lab and creating a robot army to take over the world, here's their GitHub. This contains all the instructions and templates and code to actually 3D print this in your own home, assuming you have the appropriate 3D printers and all the materials. So, if you're interested, I'll link to this GitHub page in the description below for you to read further.
In other news, Google's Gemini 2.5 Pro has successfully beat Pokémon Blue. This is a huge milestone for large language models because this is the first time an AI model has completed Pokémon Blue autonomously, at least mostly autonomously. So, while the AI did handle most decisions, including navigation and Pokémon fights and puzzle solving, it did sometimes need human intervention. So, for example, there was one key part of the game where the developer had to intervene to address a bug in the game. The human had to inform Gemini that there was a bug where it actually needed to talk to the character twice in order to get a key item. But other than that, for the most part, it did play Pokémon autonomously, and it actually got to the end and beat the final gym. Note that this idea was actually inspired by Anthropic, who've gotten their latest Claude 3.7 to autonomously play Pokémon Red. But unfortunately, Claude is still stuck in the really early stages of the game without much progress. This just shows how much better or more intelligent Gemini 2.5 Pro is. Note that this is different from other AI models that are specialized to play certain games. So, for example, AlphaStar was specialized to do really well in Starcraft, but if you get it to play another game, for example, it would completely fail. Also note that Gemini was not trained to play Pokémon, right? This is a general large language model that can do a lot of things like writing essays, answering questions, and coding apps. So, it's really impressive that it was also able to autonomously complete Pokémon, which is a pretty complex game. It has a lot of stages, plus it involves a long series of reasoning, decision-making, and strategic planning. So, this is quite a huge milestone in large language models being able to autonomously play video games.
In other news, we have yet another AI image editor. It's called
Zen Control. This is also free and open source, and this can generate new images from a single reference image of anything. So, for example, you can input this photo of a liquor bottle and then prompt it to be in a forest background, and here is what you get.
Or if you input some furniture and you prompt it for a photo of this furniture in a modern room, this is your result. Or you can take this Range Rover and place it on a lakeside like this. So this can regenerate subjects at any angle, or it can swap backgrounds or clothing, all without the need for any additional training.
So here are some examples of a background swap where we have one product photo, but you can add this to different backgrounds. Notice it kind of even gets the shadow correct. Or here's another example of this Range Rover, but you can add it to different backgrounds like this. And note how seamlessly it blends in with the background. Like the lighting and white balance and everything are perfect. I would not have guessed that the background of these images were swapped.
And then here's another example with this Amazon speaker, but you can place it in different backgrounds like this. Notice it even gets the reflection of the speaker, right? And then here's another example with another product. The nice thing is they've released a free Hugging Face space for you to try this out online. So let's click into that. It's pretty straightforward to use. Here is where you would upload an image, and then here is where you would specify your prompt.
So let's try a few examples. So here if we input this photo and then for the prompt we put, "pants, a Japanese man wearing the green pants and a blue jacket walking towards the camera in the busy streets of Tokyo," etc., etc. Here is what we get. Or let's say we upload this shoe and then we write, "a man wearing white shoes stepping outside, closeup of the shoes." This is our result. Or let's say we upload this watch and for the prompts we put, "resting on wet volcanic rock, ocean spray, mist, golden sunrise," etc., etc. This is what we get. Notice it even preserves the text of the watch screen.
Here's another example where we can place this photo of a drink at the poolside of a luxury hotel with an elegant glass and tropical fruits. And this is what we get. I mean, with this AI tool and some other image editors, which I featured this week and in previous news videos, we really no longer need to hire any product photographers. You can just upload any reference image of a product and prompt an AI to create new images for you in any background or lighting or angle you want.
Here's another example where we can upload this Goku figurine. Wait a minute. Why does Goku have blue hair? Is this some next level of Super Saiyan which I'm not aware of? Let me know in the comments below. Anyways, for the prompt, if we put, "a kid playing with the toy figurine indoor on a sunny day," here is what we get. And then finally, here's an example of headphones. And if we put, "a black man wearing black wireless headphones at a basketball game," here is what we get.
Now, in addition to the free Hugging Face space, which you can try out online, they've also released a GitHub repo which contains all the code and instructions on how to download and run this locally on your computer. And best of all, this is under the Apache 2 license, which has very minimal restrictions, and you can even use this for commercial purposes. All the links are up here, so I'll link to this main page in the description below for you to read further.
Next up, this AI is super interesting. It's called Primitive Anything. And this is also by Tencent. And this is an AI that can break down complex 3D shapes into simpler shapes called primitives. These are kind of like building blocks. So here are some examples. You can plug any 3D model through this AI and it would segment and break down the model into these primitives or basic shapes like spheres, cylinders, and cones, just to name a few examples. And then here are some more examples. So if you input this deer, this is the resulting segmented 3D model with all these basic shapes. Here's a dolphin, a koala, etc., etc. Even with this like really tricky anime girl figurine, it was able to break this down into these basic shapes. Same with this sword and this really complicated treasure chest and this Meiku doll.
Now, it doesn't have to take in a 3D model. You can just use a text prompt and it can create a 3D model made up of these primitive shapes as well. So here's a bookshelf, here's a chair, a fence, a firetruck, etc., etc. And if you compare this with other tools, so the third column is Tencent's new Primitive Anything. And the last column is the ground truth. And as you can see, this new one is just a lot more accurate. So here's another example where this column is the new Primitive Anything and the last column is the ground truth. So compared to the other methods, Primitive Anything is just way more accurate. The nice thing is if you scroll up to the top, they've also released a free Hugging Face space for you to try this out online. So it's pretty simple to use. Here is where you would upload any 3D model and then you would simply press process model. So let's say we upload this 3D model of Mickey Mouse. I'm going to press process and let's see what that gives us. And here we go. Here is the primitive breakdown of this 3D model of Mickey.
Now, this is not perfect, but note that this is made up of really basic shapes like cylinders and spheres and cones. Now, you might be wondering, why would we want to break this down into basic shapes? Well, it's actually way easier to create and manipulate these basic shapes compared to a really messy and complex model. And these primitives use much less memory than a high-resolution 3D mesh. So if you're concerned about faster processing, especially for real-time applications, then it's way more efficient to use 3D models that are built with primitive shapes. Anyways, in addition to the free Hugging Face space, they've also released the models and the code already. So if you click on this GitHub repo, it contains all the instructions on how to download and run this locally on your computer. All the links are up here. So, I'll link to this main page in the description below.
Finally, this AI is really interesting. It's called T2IR1, and this is an image generator, but it works a bit differently than the generators we're used to, like DALL-E and Flux and Stable Diffusion. So, you know how DeepSeek has this deep think feature where it takes some time to reason and think through its answer before spitting out a response. And that's kind of the feature that made it go viral. And then since then, other leading AI companies like OpenAI have also added similar thinking features. Well, that's what this AI is kind of doing, but for image generation. So, it uses chain of thought reasoning to make the images look more realistic and accurate. Specifically, here's its process. So, here's a diagram showing how it works. If you prompt it with "a black cat and a brown mouse," first it goes through this semantic level chain of thought reasoning which takes in your prompt and before even generating the image, the AI kind of thinks through what the image should look like, what objects should it include and where should it put them. This is like high-level planning, similar to outlining a drawing before filling in the details. And then after this step, it actually proceeds to generate the image. And note that this is an autoregressive image generator. So it generates the image from top to bottom. This is how the legendary GPT-4 image generator works. And this is different from Stable Diffusion or Flux which just generate the entire image all at once.
Anyways, while generating the image, it also uses chain of thought. And this process is called token-level chain of thought. And here the AI focuses on smaller details like how to draw each part of the image so that everything looks good and fits together. And then here are some of its sample generations. So here the prompt is, "show a plant that is a symbol of good fortune in Irish culture," and it gives us a clover. Here is a flower with deep cultural and religious significance in India. Note that the prompt is kind of tricky in testing its understanding on world knowledge instead of just stating the exact object it wants to generate in the image. And also note that the image quality isn't great, especially if you compare this with top image generators out there like GPT-4. So I don't think it's really usable at this stage. But this design of incorporating chain of thought reasoning into an image generator is really interesting. So it's still worth a quick share.
Anyways, on this GitHub repo, they've released all the code on how you can download and run this on your computer. So if you're interested, I'll link to this page in the description below for you to read further. And that sums up all the highlights in AI this week. Let me know in the comments what you think of all of this. Which piece of news was your favorite? And which tool are you most looking forward to trying out? As always, I will be on the lookout for the top AI news and tools to share with you. So, if you enjoyed this video, remember to like, share, subscribe, and stay tuned for more content. Also, there's just so much happening in the world of AI every week. I can't possibly cover everything on my YouTube channel. So, to really stay up to date with all that's going on in AI, be sure to subscribe to my free weekly newsletter. The link to that will be in the description below. Thanks for watching, and I'll see you in the next one.