📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Qwen 3 Omni — The Open AI Model That Does It ALL

Prompt Engineering15:01

Transcription

Hey. Uh, what am I holding?

You're holding an envelope addressed to the Internal Revenue Service IRS in Cincinnati. Inside there's a small plant with broad green leaves, likely a succulent or similar species.

Okay, so the Quint team just released their latest version of their omni model which is natively multimodel. It can process videos, images, text and audio and can generate streaming responses for text and audio. This is a very significant because the type of applications that you can build on top of this natively multimodel openweight model and did I tell you it's also multilingual in nature. Quinn and Alibaba is becoming a significant player when it comes to openweight models. In fact, even for the video models, they have best-in-class image or texttovideo models which are very close in performance to some of the close proprietary models. But specifically for Alibaba or Quinn, they're really pushing the boundaries of multimodality. None of the other openweight models have such strong omni models. This is probably one of the few openweight models that can actually compete with some of the closed source models.

So this is building on top of their previous work with quen 2.5 omni model but with some key architectural differences. So the previous version was a 7 billion model with thinker talker architecture. This preserved the same architecture. So you have a thinker and talker but now both of them are mixture of experts or oees. Also know they're using audio transformer for encoding the speech that is coming in directly but it's a much smaller footprint. So it's a 30 billion but only 3 billion active parameters. They also released a technical report with this new release with some very interesting key innovations. Later in the video I'll show you a couple of quick demos but let's look at some interesting details.

So this can process text, images, audios and videos and can deliver real-time streaming responses in both text and natural speech. It can process up to 30 minutes of video at one frames per second. So effectively you're looking at about 3 minutes of video. Now this version is natively multilingual and it can support text interaction in 119 languages, speech understanding in 19 languages and speech generation in 10 languages. So here are the languages for speech input. I think it covers a wide variety of languages. Speech output is limited to English, Chinese, French, German, Russian, Italian, Spanish, Portuguese, Japanese, and Korean. Although in my test it can also generate outputs in Arabic, in Udu and Hindi but officially I think they are not supported in speech output. And this is a very strong Omni model. It's state-of-the-art for openweight models and can also compete closely with models like Gemini 2.5 Pro GPD 40 or transcribe. It actually has a dedicated speech transcription part which you can use standalone. and the performance is really good. But as always, it's good to test it on your own applications and your own test set. Now, with the right hardware, the speech transcription latency can be pretty awesome. So, they say that it can achieve latency as low as 211 milliseconds in audio only scenarios and a latency as low as 500 milliseconds in audio video scenarios. And actually if you look at my demo later in the video, the audio video interaction does seem very natural. You can understand up to 30 minutes of audio. It has a pretty long context window of over 100,000 tokens. And then you can also control the behavior using a system prompt. So I think even if you're doing speech transcription, you can provide a system prompt. So for example, it can be just correcting grammatical mistakes or even changing the tone of the transcription.

All the new Quinn models are really good at coding and they are specifically focusing on agentic capabilities. So Quen 3 Omni supports function calling enabling seamless integration with external tools and services which are going to be critical if you want to build agents on top of this model.

Okay. Okay. So an important part of this release is Quen 3 Omni 30 billion 3 billion active parameter captioner. So this is their speech transcription part which you can use as a replacement to something like whisper or parakeet models from Nvidia. If you want to use another model for speech generation, you can just use that captioner for transcription and then build the subsequent layers with some other models. In terms of the architecture, it still follows the thinker talker architecture very similar to the previous version, but now both of them are. Now the talker no longer consumes the thinker's highle text representation and conditions only on audio and visual multimodal features. This is actually critical because now this decoupling allows other modules like rag or function calling to intervene on the thinker's textual output and if desired you can supply text to talker via controlled pre-processing for streaming synthesis. Another major part is since these two are now decoupled you can have dedicated system prompts for both thinker and talker and you can control them separately.

Now for audio they are introducing audio transformer which the audio encoder is directly using and this is trained on about 200 million hours of audio data which provides strong general audio representation capabilities and actually if you look at the audio output as well it's much more pleasant to listen to compared to the previous version. As I mentioned before both the thinker and talker now are based ones. Another interesting observation that they have is that mixing unimodel and crossmodel data during early stages of text pre-training can achieve parity across all modalities which is very interesting. They claim the model to have Gemini 2.5 Pro level performance when it comes to speech recognition and instruction following.

Now, as always, I highly recommend to test this out, but I would say that the video capabilities or real-time video interaction is really pleasant. The speech is a lot more natural compared to the previous version. Although there are some very interesting hallucination when you start interacting with the model. Sometimes the model assume that the person it's seeing in the video is itself. So I have seen instances in which let's say if I'm showing it something it will think that the model itself is that specific person or individual and also in certain cases I've seen that the model hallucinates and would start generating responses in another language.

Okay so a quick look at the benchmarks. So they claim that Quenti Omni matches the performance of samesized single modality models within the Quen series and then for modality it's best in class out of all the open weight models and matches the performance of closed source models. Now the benchmarks that they have shown here are for Omni Flash model. I'm not sure if it's exactly the same model, but even on the Quen official website, only the flash version is available. There is a demo for the openweight model on hugging face, but overall I think the benchmark performance are really good.

Now in terms of the openweight models, they released three different things. One is the quentry omni 30 billion 3 billion active parameter instruct version. This is the non-thinking version. Then they have a specifically thinking version if you want reasoning traces. And then they also have this 30 billion 3 billion active captioner. So this is the speech transcription model. So let's have a quick play with the model and then I'll show you some code examples. Let's have a play with the model and see how good the experience is. For this we're going to be using the official Quen chat. In here they have quenti omni flash which seems to be a proprietary version of the omni model. So it's not exactly the openweight model but I think the experience is going to be very similar. Also I don't know whether flash is a smaller or bigger version in the platform. You can either do voice chat or video chat. Video chat per session seems to be limited to 10 minutes which is pretty long. So, we're going to try this. And if you go to the settings, not only there are a large number of different voices available, but you can also select different languages. And when you're interacting with the model, you can just directly change the language. And it seems to adopt pretty good. So, let's have a quick test.

Hey, uh, what am I holding?

You're holding an envelope address to the Internal Revenue Service IRS in Cincinnati.

Okay. What's the PO box number?

The PO box number is A02502.

Okay. Uh, and what am I holding now?

You're holding a book titled The Coming Wave. What would you like to know about it?

Uh, who are the authors of this book?

The authors of the book are Mustafa Sulean and Michael Bascar.

What is this book about?

Ah, you're asking about that book. It's actually called The Coming Wave. It dives into how AI, robotics, and nanotechnology are about to reshape our world. The author is Mustafa Sulean, and he's got a pretty compelling take on how these technologies will transform everything from healthcare.

When was this book?

The book was released on June 27th, 2023.

Okay. Uh, what am I holding?

Oh, so you're you're holding a terrarium. It's a glass container with a geometric shape, possibly octagonal or docahedral with a gold colored metal frame. Inside there's a small plant with broad green leaves, likely a succulent or similar species. The terrarium is filled with small pebbles at the bottom, creating a drainage layer, and the plant is rooted in soil or growing medium. Terrariums are decorative containers that house small plants in a controlled environment. Often mujahcoy video audio file upload car sneak s Auntie auntie in 9 a okay at the end it was just making of stuff uh but in the previous conversations um I have seen that it is multilingual that can speak udu or Hindi now some sometimes it it switches the context and u it's a really fun model to talk to

Okay so with this model. They have released a number of cookbooks on GitHub which are extremely helpful. So for example, how to do speech recognition, speech translation, music analysis, sound analysis, audio caption. They have specifically cookbooks for visual analysis, audio visual analysis, that's realtime interaction, the type of stuff that we looked at. Also agents and function calling and then omni captioner.

So this is a quick example of how to use the speech to text model. So here they're setting up the model and then we're going to be downloading that specific captioner model. After that you just need to provide the path of your audio file. It can process up to 30 minutes of audio which is pretty great. And I think you can also do streaming responses. Right. So this is an audio file. Again we have another one. Right. The transcription part is extremely simple.

Here's an example of working with OCR or reading text from images. So this is the initial setup of how to set up the model. That's the model that we're going to be using and then we provide the path of an image and the text is basically I think extract the information from this image. So it extracted that information or he has extract the text from the image. So this is another image as an input and this is the same text in latex format. So it's a mathematical equation that the model is able to extract. So the good thing is that you can actually provide a system instruction or specific information on how you want the model to behave. That's also going to be possible with the speech transcription. So let's say you can tell the model to correct any grammatical mistakes or maybe even rewrite the transcription in a specific format or style that you want.

I'm going to probably create more detailed videos of how to build on top of this model and what type of hardware you will need. So stay tuned for that. But I hope you found this video useful. Thanks for watching and as always, see you in the next one.