📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Google Just Dropped TurboQuant And Changes AI Forever

AI Revolution11:20

Transcription

Google just unveiled Turbo Quant, a new compression system that could cut AI memory use by six times and speed some workloads up by as much as eight times in a way that's already reminding some people of Silicon Valley's fictional Pied Piper compression breakthrough. While OpenAI is shutting down Sora, losing its Disney deal, and racing to launch a new model called Spud within weeks. There's some really big news here, so let's talk about it.

All right, so starting with Google because this one sounds technical on the surface, yet the impact is actually very easy to understand once you strip it down. Right now, one of the biggest hidden problems in AI is not just how smart the models are, it's how much memory they need to function. Every time you type something into a model, it has to keep track of everything you've said before so it doesn't lose context. That memory builds up fast, especially with long conversations or long documents. And that memory is not cheap. It slows things down. It increases costs. And it forces companies to use more powerful and expensive hardware just to keep things running smoothly.

Google just introduced something called Turbo Quant. And the whole idea is simple in concept. Take that memory, shrink it massively, and somehow keep the same level of performance. They're claiming at least a six times reduction in memory usage for this part of the system called the KV cache. That's basically the short-term memory of AI models. So instead of needing say six units of memory, you now need just one. And on top of that, they're showing up to eight times speed improvements during inference, meaning when the AI is actually generating responses. So you're getting faster outputs, lower costs, and less hardware pressure at the same time.

Now, the obvious question is how do you shrink something that much without breaking it? Because normally when you compress data, you lose quality. You've seen that with images, videos, audio, everything. What Google is doing here is compressing the internal memory of AI in a way that keeps the important information intact. And they're doing it using something called vector quantization. That's just a fancy way of saying they're taking complex data and representing it in a more compact form. Traditional methods already exist for this, like product quantization. Yet, they come with a problem. They need to be trained on specific data before they can work properly. That makes them slow and not flexible enough for real-time AI systems. Turboquant avoids that completely. It's what they call data oblivious, which basically means it doesn't care about the data set. You don't need to train it for each scenario. It just works immediately. That alone removes a huge bottleneck.

Then comes the clever part. They apply a random rotation to the data inside the model. That might sound weird, yet what it actually does is it spreads the information evenly across all dimensions. In simple terms, instead of having some parts of the data being very important and others not, everything becomes more balanced. And once that happens, they can compress each part independently in a very efficient way. So instead of dealing with one giant complicated structure, they break it into many smaller, simpler pieces and compress those individually. That's where the mean squared error optimization comes in. They're basically finding the best possible way to compress each piece while minimizing the difference from the original.

Now, here's where things get tricky. And this is something most people would never think about. But before we get into that, there's something else worth pointing out here. Better AI systems don't just improve the back end, they also open the door to entirely new kinds of entertainment. Higsfield is sponsoring today's video and they've launched something called Higsfield Original Series, which they describe as the world's first complete AI streaming platform focused entirely on AI made films and series. One of the first releases on there is Arena Zero, and this is where things get interesting. It's a full 10-minute AI action film created end-to-end with Higsfield by a small team of creators using the platform. That means everything from characters and scenes to motion and final output is handled inside the same AI workflow. So, this isn't just another short demo clip or random visual test. It's a longer structured piece with actual scenes, pacing, and story progression. Something much closer to what you'd expect from traditional production, just built entirely with AI tools. They're also building this as a full ecosystem where people can explore new AI films and even vote on which projects continue, which adds a completely different layer to how content gets developed. And that's really the bigger shift here. As the tech improves on one side, platforms like this are starting to turn that progress into actual watchable content. Go check out Arena Zero. Link is in the description.

All right, now let's get back to this. In AI models, a lot of the work comes down to calculating relationships between pieces of data. That's done using inner products. If your compression method messes that up, the model starts making worse decisions. So, Google added a second step to fix that. They combine the main compression method with something called a quantized Johnson Lynden Strauss transform or QJL. This step removes bias and ensures that those relationships stay accurate. And the result is that even after compression, the model behaves almost exactly the same as before. From a math perspective, they're getting extremely close to the theoretical limit of how much you can compress data without losing information. They're within about 2.7 times of the absolute best possible. And at very low precision, like one bit, they're only about 1.45 times away. That's very tight.

In real world tests, they ran this on models like Llama 3.18B and Ministral 7B. Even with four times compression, the models kept full accuracy in long context tasks. One of the tests is called needle in a haystack. The model has to find a tiny piece of information hidden inside a massive context, sometimes over 100,000 tokens long. Turboquant matched full precision performance up to 104,000 tokens under that four times compression. So even after compressing memory heavily, the model still remembers everything correctly. They also introduce something interesting with non-integer bit precision. Instead of sticking to clean values like two bits or three bits, they use things like 2.5 or 3.5 bits per channel. They do this by giving more precision to important parts of the data and less to the rest. So it's a smarter allocation of resources instead of a uniform approach.

And outside of language models, this also affects search systems. When you build vector databases, you usually need time to index the data. That can take minutes or even longer for large data sets. Turboquant removes that almost entirely. We're talking about indexing times dropping from hundreds of seconds to basically zero around 0.0013 seconds for high-dimensional vectors. That's instant. So, this isn't just about saving memory. It's about making entire AI systems more efficient from the ground up. That's why some people are comparing this to Deepseek's efficiency breakthrough, and others are jokingly calling it Pied Piper, referencing that fictional compression algorithm from Silicon Valley that was supposed to change computing forever. At the same time, this is still a research stage breakthrough. It hasn't been deployed widely yet, and it only affects inference, not training. Training still requires massive compute and memory. So, this doesn't solve everything. Yet, it does remove one of the biggest bottlenecks when it comes to actually running AI models in production.

Now, while Google is pushing efficiency forward like this, OpenAI is making a very different move. They're shutting down Sora. The same Sora that shocked everyone with realistic AI video generation is now being discontinued as a standalone app just months after launch. OpenAI confirmed it directly. They thanked users, acknowledged that people built communities around it, and said they'll share timelines for shutting down the app and API along with details on how users can preserve their work. The bigger question is why? And the answer comes down to resources and strategy. Video generation is extremely expensive. Every clip generated uses a lot of GPU power and those GPUs are limited. OpenAI is currently under pressure to compete with companies like Anthropic and Google, especially in enterprise AI and productivity tools. So, they're reallocating resources. Instead of continuing to invest heavily in Sora, they're shifting that compute toward their core products.

And then there's the Disney deal. OpenAI had a massive agreement with Disney where Disney planned to invest $1 billion and license some of its characters for use in Sora. The goal was to eventually integrate this into Disney Plus. That deal is now dead. Disney confirmed they're exiting, saying they respect OpenAI's decision and will continue exploring AI partnerships elsewhere. So, not only is Sora being shut down, one of its biggest strategic partnerships collapsed at the same time. There were also issues early on with intellectual property. When Sora launched, it allowed the use of existing characters and likenesses in ways that made Hollywood uncomfortable. OpenAI had to quickly adjust and give studios more control. So between high compute costs, strategic shifts, failed partnerships, and IP concerns, Sora as a standalone app no longer made sense.

That said, OpenAI is not leaving AI video. They're just integrating it into something bigger. Instead of a separate app, video generation will likely become one feature inside their broader ecosystem, probably within Chat GPT or their upcoming desktop super app. And that leads directly to what they're working on next. Internally, OpenAI has been focusing heavily on a new model with the code name Spud. Pre-training is already complete, and Sam Altman told employees that this model could be released within weeks. It might be GPT6 or possibly GPT 5.5. That part is still unclear. What is clear is that Altman described it as a very strong model that could accelerate the economy. No specific capabilities have been revealed, so we don't know if this is about reasoning, agents, or something else entirely. Yet, given the timing, it clearly fits into their push toward more advanced productivity tools. They're building a super app that combines Chat GPT, Codex, and a proprietary browser into a single desktop experience. So, instead of using separate tools for chatting, coding, and browsing, everything runs in one place. That also explains why Sora didn't fit into the plan. It was too separate, too heavy, and not aligned with this unified approach.

There's also internal restructuring happening. The safety division is moving under the research division led by Mark Chen. Technical security is being handled by Greg Brockman. Sam Altman is focusing more on fundraising and building data centers. And Fijiimo, who joined recently, is now leading the AGI deployment division, which covers all product areas. So this is not just a product change, it's a full company level shift. Even the Sora team is being redirected. They'll now work on something called world simulation research, which is expected to play a role in robotics in the future. So instead of building standalone video tools, they're moving towards systems that simulate environments and interact with the real world.

Anyway, if you found this useful, drop a like and subscribe. Thanks for watching and I'll catch you in the next one.