📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

China’s New AI Makes Videos That Look Better Than Reality!

AI Revolution9:16

Transcription

China just dropped what they're calling the most powerful AI video generator on the planet, and it's already got over 22 million users. Bite Dance is secretly working on AI smart glasses to take on Meta, and their new lip-sync tech is so realistic it's honestly creepy. Everything from how we film, wear, and even fake ourselves is being rewritten by AI.

But before we jump in, just a reminder: our free AI avatar course is up inside the school community, and it's already helping creators level up their content. We've also got a free weekly newsletter packed with the latest AI tools, news, and wild updates. Both links are in the description; check them out, and then let's get into the most powerful AI video generator on the planet.

So we absolutely have to talk about Quixho. If you're not too familiar with the name, Quixho is one of China's largest short video platforms. Some people might know it as the main rival to Bite Dance's TikTok within China, and they've been busy developing their own AI magic recently. They introduced Cling AI 2, which is their newly upgraded video generation model. Now, Quixho is calling it the world's most powerful, which is a pretty bold claim, but here's the thing: they're not just talking a big game. Quixho's senior vice president, Guy Kun, announced that Cling AI has already gathered over 22 million users worldwide, who've collectively generated more than 168 million video clips and a whopping 344 million images. Let me tell you, those are some serious numbers. No wonder they're so confident calling this tool the most powerful out there.

Cling AI 2 boasts improvements all around, including better instruction following, better prompt understanding, a higher quality of images, smoother character movement, and an overall more realistic aesthetic. One big strength is how punchy the final video content is, which is critical when you're trying to attract attention in the short video world, right? According to Artificial Analysis, which is a third-party service that tests and ranks AI models, the previous generation of Cling already held a top spot for image-to-video performance and was second only to Google DeepMind's V2 for text-to-video tasks. So this new update is basically building on that strong foundation, and the results so far look like they're pushing Quixho even further ahead in the race.

Speaking of competition, you can't talk about Chinese AI without mentioning Bite Dance, Alibaba, Tencent, and a bunch of specialized startups like Zepu AI and Shengshu Tech. Each of these players has been going all-in to create newer, flashier AI models and features because the market is pretty much a run for life, as Gyon puts it. This sort of environment forces everyone to keep innovating, or they risk losing ground to the next big competitor. For instance, Bite Dance has their own offerings; Alibaba is also cooking up something in their labs; Tencent has been stepping up in AI research; and smaller unicorns such as Zepu AI have been making strides. Zepu is even eyeing an IPO this year after seeing successes like Deepseek. Meanwhile, Quixho just recently announced the NextG project, offering funding, technology, and promotional support for artists. They want to help creators generate film-quality content, presumably with Cling AI 2 or whatever new tech they have on deck.

I got to say, it's such a fascinating time to be in the content creation space because you have these big names actively encouraging the little guys to go out there and experiment, create, and really push boundaries in AI-driven video.

Now, the business model for these AI video tools is also pretty interesting. Instead of just giving everything away for free, many of them use what's called a premium approach. So you might get basic features at no cost, but if you want higher resolution, faster processing, or more advanced editing tools, you pay for a premium subscription. That's what a lot of Chinese chatbots did, and it makes sense for video too. People love freebies, but if you're a professional creator, you'll probably want advanced tools, and you won't mind paying if it helps you stand out on crowded social platforms. It's basically a win-win for the company and the user.

While Quixho has been busy flexing with Cling AI 2, let's pivot over to Bite Dance for a second. They've apparently been working on their own AI smart glasses. Rumor has it that Bite Dance is diving deeper into wearables, trying to rival what Meta's been doing with products like the Ray-Ban smart glasses. According to sources, Bite Dance is talking with suppliers to figure out how they can get decent quality video and image capture on the glasses but without wrecking battery life. That's always been one of the big challenges with wearables, right? The moment you start trying to do advanced AI processing or record high-quality video, your device is screaming for more power. Anyway, Bite Dance started this project within the last year or so, and they're already actively working on specs, features, and pricing. They're the ones behind that collaboration with Qualcomm back at MWC 2025 for a next-generation VR headset, which was seen by some as a statement that Bite Dance wants to do more than just short-form videos on your phone. And we can't forget how Bite Dance acquired Pico in 2021, which definitely gave them some VR manufacturing firepower. Honestly, it looks like Bite Dance is building an entire ecosystem of hardware, from VR headsets to these potential AI glasses, so they can take on Meta in the wearable and extended reality space. If they do deliver something that stands out in terms of design or battery performance, that could be a huge deal, especially since Meta's Ray-Ban smart glasses have been trying to gain traction but aren't exactly mainstream yet.

Now let's chat about another Bite Dance project that's making waves: Omnihuman 1. This is that advanced lip-syncing tool available on the Draina platform. I know a lot of you might wonder, "All right, we already have lip-sync tools out there, so what's new here?" The big selling point is how realistic the lip-syncing becomes from just a single reference image. Like, you snap one photo of a person, feed it in, and boom—the AI can generate a video that aligns the lip movements extremely closely to any audio. It can handle text-to-speech stuff or even let you upload your own custom audio. That's pretty huge for creators who want to do everything from comedic sketches to serious voiceovers without having to record a ton of real footage.

Omnihuman 1's claim to fame is how well it handles close-up shots and even singing scenes, which is notoriously difficult for AI to mimic convincingly. Some other lip-sync tools might start to fail when the camera zooms right in on a person's mouth, or they can't deal with the complexities of musical performances. But according to Bite Dance, Omnihuman 1 sets a new industry benchmark, especially for that type of content. It's not totally perfect, though; it apparently does best with static backgrounds. So if you need dynamic scene changes or multiple characters walking around, it might not be the tool for you. But for talking heads, singing performances, or any close-up scenario, it's pretty darn impressive.

And because the lip-sync AI space is getting more competitive, we've seen other notable players like Hedra and Clink AI step in with their own offerings. Hedra's known for being solid with scene prompting and dynamic movements, which Omnihuman 1 can't really do as effectively right now. Hedra tries to direct entire scenes with characters walking around or performing certain tasks. On the flip side, Hedra doesn't handle side profiles and singing as well as Omnihuman 1, so it's a give and take. Then you have Clink AI, which apparently focuses on generating these super dynamic backgrounds but lags behind when it comes to lip-sync accuracy. So if you really want your background to pop or you need fancy transitions, Clink AI might be your go-to, but you'd probably sacrifice some realism in how the mouth lines up with the audio.

So basically, these AI tools are getting really good at generating videos and lip-syncing, but there are still tough spots, like animating multiple characters who are walking around or gesturing a lot. Omnihuman 1 from Bite Dance rocks at close-ups and even singing, but it struggles when people move around too much, and it's not alone. Other companies like Quixho, OpenAI, Google DeepMind, Alibaba, Tencent, and a bunch of startups are all racing to outdo each other. Bite Dance is even rumored to be making AI smart glasses to rival Meta's Ray-Ban bands, and Quixho's Cling AI 2 is bragging about being the world's most powerful with millions of users. Meanwhile, smaller players like Tencent, Back Norwal, and AI unicorn Zepu are also stirring up buzz. It's kind of like the early internet days where everything's evolving at lightning speed. So the question is: how long until most videos we watch are fully AI generated, or we're all wearing smart glasses instead of holding phones?

If you haven't subscribed yet, now's a good time. We cover all the craziest updates as they drop. Drop a like if you enjoyed the video, and as always, thanks for watching. Catch you in the next one.