Transcription
I have covered Llama 1, Llama 2, Llama 3, and now Llama 4. Meta's newly released Llama 4, Scout and Maverick, have been met with considerable disappointment across the AI community, primarily due to uninspired design choices and performance issues. We have tested them out from various angles, plus we have compared them on various benchmarks already on the channel, and we have even installed one of them locally to test it out.
Now, for the purpose of this video, I just don't want to focus on the negatives, but also what the future holds for AI in the wake of this underwhelming release from Meta. On a lighter note, we also remember this sentence and statement from Mark Zuckerberg from Meta where he said that AI is going to replace mid-level engineers soon. I don't think so. Llama 4 is going to do that, at least now.
The thing is that this might be a pivotal moment in AI because, if you remember, as I said when Llama 1 was released, Llama 2 was released, and even when Llama 3 was released, this was much, much applauded by the community, and Llama has done wonderfully well in terms of open-source. So the disappointment over Llama 4 is palpable because the expectations were so high. Just imagine if this Llama 4 model would have been released like 6 months ago; it would be a huge, huge success. But I believe the models like Deep Seek, QWQ, Quen, OpenAI, Anthropic, Claw 3.7, they have now set the bar way too high.
So what is happening here? Well, I believe that this disappointment stems from several critical factors, among them an overly simplistic reliance on a moderately sized mixture of expert architecture with only 17 billion active parameters. These expert mixtures appear inadequate for matching or outpacing competitors, in my opinion. Simply scaling parameters without careful optimization seems to have resulted in models that cannot use Meta's abundant data and resources effectively, and they have plenty of resources. This outcome underscores the fact that sheer computational scale does not guarantee success. Genuine advancement in AI demands authentic innovation rather than brute force alone, and I'm quite positive and confident about that.
A further cause of disillusionment concerns multimodal capabilities. Well, expectations were high that Meta would use extensive multimodal features, like integrated audio, video, or advanced image comprehension similar to industry leaders. Instead, users were presented with limited multimodality, primarily restricted to vision input, and we saw that in one of the videos when we tried to create the video output. Given the competitive landscape, which includes Google's Gemini series with robust multimodal functionalities and Deep Seek's exploratory multimodal endeavors, Meta's offerings appear far behind and lagging.
Equally frustrating for the community is the absence of smaller, resource-efficient models that would allow broader experimentation and adoption by ordinary users and researchers. This has alienated a substantial segment formerly eager to embrace Meta's open-source initiative. If you look at various benchmarks, they also tell you a lot of a different story. So, for example, if you look here, this benchmark also paints a stark picture of Llama's force deficiencies.
So, for example, on this Ader polyglot benchmark, which assesses a model's multilingual coding and comprehension skills, Llama 4 Maverick performs notably worse than contemporary models like Google's Gemini 2.5 Pro and Deepseek V3. This benchmark uses code completion success rates across multiple languages to assess effectiveness, and Maverick shows a remarkably low success rate, indicating poor multilingual understanding and failings in cross-linguistic accuracy. Also, in evaluations represented by yet another benchmark, which is artificial intelligence, uh, artificial analysis intelligence index, which is a comprehensive assessment built from seven challenging reasoning tests, including a test on mathematics, which is Math 500, scientific reasoning, which is Psychode Logic, and general understanding, which is, uh, for example, Human Humanities last exam GPQA Diamond, Meta's offering again falls. Llama 4 Scout scores at the bottom with significantly lower outcomes compared not just to Deepseek and other competitors, but also to much smaller models like Google Gemma 3, 27 billion, which is a really good model, by the way. Maverick fares marginally better but remains well behind Deep CXR1 and QWQ's 32-billion model, reinforcing concerns about fundamental limitations in reasoning and analytical capacities. And QWQ 32 billion is yet another magnificent model.
And yes, I understand that these models have different attributes. QWQ, which is Quen with, um, question being a reasoning and dense model, and Llama 4 being instructed and a mixture of expert sparse model with only 17 billion active parameters. But the thing is that a user doesn't really care about how the model is working internally; they focus on performance and how achievable, uh, it is to properly get the results back. And I think if you really look at it, 32-billion QWQ needs, uh, cheaper hardware to host rather than these huge Llama 4 models, and we have seen that how expensive it was to even host the Scout, Scout, uh, locally in this second video here. Now, there are a lot of other things which I could go on and on about these benchmarks, but I think if you look at this last benchmark, this also tells you a lot of a story, and Llama, uh, 4 is just at the end. Look here, this one. Now, this FictionLive bench evaluates AI abilities in maintaining long context, deep comprehension essential for handling extended narratives, documentation, or prolonged user interactions effectively. Here too, Llama 4 shows deep-rooted issues. Performance really drops drastically even at mid-level context window, for example, 4,000 tokens and higher, vastly underperforming other available models such as Google's Gemini, Claude from Anthropic, and advanced GPT variants. And these findings effectively negate Meta's previously advertised advantage of supporting long-range context interaction as practical capabilities quickly diminish with increasing lengths of interactions.
Well, what's next? Before I talk about what's next, let me also introduce you to the sponsors of the video, who are Bot Aenbot. Lets you effortlessly deploy a personalized knowledge bot across platforms like Discord, Slack, and others. It is ideal for open-source tech communities and startups that provide user support, and I will drop the link to their website in the video's description. Also, do me a favor and like the video and subscribe to the channel, as it helps a lot, and share the video among your network.
Okay, coming back to what's next. Well, Meta needs to go back to its roots. Meta must substantially rethink its strategic approach to a model's development. Rather than investing solely in parameter scaling or superficial modality, Meta should concentrate on deliberate architecture refinements, innovation in understanding context and semantic coherence, and meaningful integration of diverse data modalities. Also, they should prioritize transparent benchmarking and realistic internal evaluations to align their offerings with genuine industry needs and user expectations. Leveraging smaller, more efficient models to regain the trust and enthusiasm of the research and open-source, open-source community should also become a central focus. Only then can Meta hope to regain its credibility and competitive relevance in the rapidly evolving AI landscape. And that won't be simple, I understand, but you know what, if anyone can do it, that's Meta. So let's wait for Meta's return. Maybe their behemoth model might be able to do so, but at the moment, this is the current state. Thank you very much.