📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

LMS Are About to Hit a wall - The AI Scaling Law Might Be Breaking...

TheAIGRID11:36

Transcription

So for the last 3 years, the AI industry has been built on one idea: make the models bigger, add more parameters, add more data, add more compute, and the model just simply gets smarter.

Now, this is the promise that's the entire reason companies like OpenAI, Google, and Anthropic, and XAI have been able to raise hundreds of millions of dollars. Bigger is better, more is smarter, that was always the deal.

Now, here's the thing, guys. A new research paper just shattered that deal in a very specific way. A team of researchers tested something most people would never think to test. And what they found is that for one of the most important types of human thinking, making the AI bigger did not make it smarter. And in some cases, making the AI bigger actually made it worse. And the part that made the AI industry quietly nervous is that the same pattern shows up inside the models we're already using.

Now, if you're wondering what the paper's called, the paper is called "Emergent Analogical Reasoning in Transformers," and it came out on the open research site arXiv. Now, the team behind it wanted to study this kind of reasoning, which is the type of thinking where you understand the relationship between two things and apply that same relationship to two completely new things. It's the thing your brain does when you realize that one situation is similar to another situation you've seen before, even if the details are different. Analogical reasoning is considered one of the most important parts of human intelligence, and has always made AI look most impressive when it works.

So the researchers built clean, controlled tests. They trained a series of small AI models from scratch on a fake world that they invented, where they could control every variable, and they slowly scaled the models up. They tested models with a width of 64, 128, 256, and 512. They tested deeper models, they tested with more data, they tested with less data. They tested with everything you would normally chew when training a frontier AI model, and they tracked when each model could do analogical reasoning.

Now, the result was not what the AI industry wants to hear. The smallest models could not do it, that was expected. But then, the medium-sized models did the best. And when they scaled up to the bigger models, performance actually got worse. The paper says this directly: "The back line is increasing model size does not monotonically improve performance and in some case degrade."

Now, you have to understand that this sentence is a big problem. You have to understand this is what the entire AI industry has essentially been built on. There is something called the scaling law. The scaling law is the rule that if you make the model bigger and give it more data, the model gets smarter in a predictable way. The scaling law is the reason Nvidia is one of the most valuable companies in the world. Scaling law is the reason Microsoft put $100 billion into OpenAI. Scaling law is the reason data centers are being built across America and the Middle East at a pace nobody else has ever seen before.

And so if the scaling law doesn't work anymore, that means the performance gains are actually going to have to come from the harnesses around the model itself. And that's where today's sponsor comes in. If you're anything like me, you've probably got Claude open, ChatGPT open, Gemini open, and a bunch of other AI tools spread across way too many tabs. And that's really the problem. The tools are powerful, but the workflow is still fragmented. And that's why Genspark actually stood out to me. Genspark is an all-in-one workspace that brings the top models and task-focused agents into one place. So instead of bouncing between tools, you can actually get work done in one workflow. It has already hit $250 million in ARR in 12 months. And yes, they even ran a Super Bowl commercial this year. And the reason it feels different is that Genspark doesn't just chat with you. It helps you move the work forward. For example, if I need a presentation, I'll open AI slides and type, "Create a clean presentation about how AI is changing the job market, including short introduction, three key trends, one slide on risks, one on opportunities, and a final summary. Add visuals and make your presentation ready." And instead of starting from a blank slide, Genspark turns that into a polished first draft that I can use right away. Or if I need to capture ideas quickly, I'll open Speak Lee and say, "Turn these rough notes into a clear follow-up email and a short recap of the meeting." So instead of leaving myself with messy voice notes or half-finished thoughts, I get something structured and usable immediately. And that's what makes it practical. I can talk naturally and GenSpark helps turn that into something I can actually send or work from. And if I need to make sense of data, I can open up AI sheets and type, "Analyze this spreadsheet, show me the main trends, highlight the top-performing categories, and give me a short summary of what matters most." So, instead of staring at rows of numbers trying to build everything manually, I can get a cleaner view of the data, clear patterns, and a result that I can use right away. And that's the difference here. One workspace, access to top models and tools that help you actually finish your task. GenSpark also has a lot of other built-in tools for things like docs, design, websites, research, and more, all in one place. If you want to try GenSpark, click the link below. New users get free credits when they sign up, and you can earn more through the get-started tasks.

If the scaling law breaks down for an important type of thinking, then a lot of that spending is built on a shakier assumption than people have been told. The paper does not say that scaling is dead overall. That's actually important. Compositional reasoning, which is a different kind of reasoning, did scale normally in their test. It did improve with size. So, scaling does still work for some things, but analogical reasoning, the kind that mostly closely mirrors creative human thinking, didn't follow that pattern.

And this is where the story actually gets bigger. The researchers did not stop at their small test models. They also ran the same analysis on real frontier AI models. They tested Gemma 2, which is Google's open-source model, in both a 2-billion parameter version and a 9-billion parameter version, and they also tested Meta's Llama models. And the same pattern showed up again. The bigger Gemma model did not reliably get better at analogical reasoning compared to the smaller one. The thing that mattered more than size was whether the model had built a specific kind of internal structure during training. The researchers gave that structure a name. They called it geometrical alignment. The basic idea is that the model needs to organize its internal map of concepts in a very specific way before it can do analogical reasoning at all. And if the model never builds that structure, no amount of extra parameters will save it.

Now, this is the part that should make anyone training a frontier model uncomfortable because the largest AI labs in the world are betting that if they just keep throwing money at bigger models, the models will keep getting smarter, and the paper suggests that for at least one important type of thinking, that bet actually doesn't pay off. The model either leans to the right structure during training or it does not.

Now, remember, the paper does not exist in a vacuum. The industry has been quietly bumping into this wall for a year now. And Ilya Sutskever, who co-founded OpenAI, has been saying in public talks that the era of scaling is over. He's gone as far as saying that essentially all of the useful internet data has already been used up by the major labs.

"Computers are very big. In some sense, we are back to the age of research."

So, maybe here's another way to put it. Up until 2020, from 2015 from 2012 to 2020, it was the age of research. Now, from 2020 to 2025, it was the age of scaling. Or maybe plus minus, let's add error bars to those years, because people say this is amazing, you got to scale more, keep scaling. The one word, scaling. But now the scale is so big, like is is it is the belief really that oh, it's so big, but if you had a 100x more, everything would be so different? Like it would be different for sure, but like is the belief that if you just 100x the scale, everything would be transformed? I don't think that's true. So, it's back to the age of research again, just with big computers.

That's a very interesting way to put it. Now, here you can see, in May, a separate paper on arXiv argued that the famous Chinchilla rule, the rule that told everyone how to balance model size and training data, is no longer valid for frontier labs because the assumption it depends on, that there's an unlimited amount of unique data on the internet, is just not true anymore. And this is happening at the same time that DeepMind kind of waved off new Chinese labs prove that you can match frontier performance with a fraction of the compute by focusing on smarter training rather than bigger models. That moment back when DeepMind caught one hit the market was the first time investors openly started asking whether the spending plans of the biggest US labs actually made sense.

So, this new paper lands at a moment where the industry was already starting to admit quietly that scaling is not the magic engine it was sold as. The difference is that this paper provides a clean, controlled, mechanistic demonstration of why. It doesn't just say bigger is not better, it shows the exact thing that has to happen inside the model for a specific type of reasoning to appear. And it shows that size alone does not cause that thing to happen.

If you've been using ChatGPT, Claude, Gemini, or some of the newest models, you'll know that, you know, very recently, they only feel slightly better than the previous ones. And this paper helps explain why. The labs are getting less and less reasoning improvement out of each new generation, even though they're spending more money than ever to train them. Big Tech is on track to spend about $725 billion on AI infrastructure in 2026, and that's a pretty huge bet. And a lot of it is built on the assumption that the next generation will be meaningfully smarter than this one.

Now, the most interesting line in the paper, when you read it really carefully, is this: "Analogical reasoning is not primarily determined by the model capacity, but rather whether the model discovers the aligned geometric structure in the embedding space." Now, that sentence is doing a lot of work. It's saying that the smartest people care about most is not really how big the model is. It's about whether the model managed to build the right internal map during training. And that map is not guaranteed to form. It depends on the data, it depends on the number of relationships in the training set, it depends on the optimization settings. The researchers even found that in some cases, the model learned analogical reasoning during training and then lost it again later in the same training run. It just decayed away. They called this transient behavior. And that should not be possible if scaling was the only thing that mattered.

Now, the AI research community noticed this paper quickly because it's one of the cleanest demonstrations of a scaling limit that anyone has produced. There are a ton of papers right now that hint at scaling problems, and this one shows it under controlled conditions with clear math, and they then find the same pattern in real Gemma and Llama models. That combination is pretty rare. The frontier labs have not officially responded to this specific paper, but their behavior over the past year already lines up with what the paper is saying. OpenAI has shifted heavily towards inference time compute, which is when the model thinks longer at the moment you ask a question rather than just relying on a bigger base model. Google has been doing the same thing with these newer Gemini reasoning models, and Anthropic has quietly pushed in the same direction. The whole industry is pivoting away from pure scale towards smarter ways of running the model after it was already trained.

Now, if the pattern in this paper holds, the next 2 years of AI development are probably going to look a little bit different from the last 2 years. The labs that win will not be the ones that spent the most money on the biggest models. They'll be the ones that figured out how to make the model build the right internal structure during training. They'll be the ones that spend more time on data quality, more time on post-training, and more time on how the model reasons at the moment it answers. And that shift is already happening, but most people don't see it because the marketing still talks about bigger models. When OpenAI talks about its next release, the headline is always the size or capability. And when Google talks about Gemini, the headline is always the new version number. But behind the scenes, the research priorities have actually moved.

Now, there's also a financial side to this that most people aren't considering. Nvidia, Microsoft, Google, and Meta have all priced their stocks around the assumption that scaling keeps working. And if the market starts to believe that scaling has real measurable limits on important types of reasoning, that assumption gets questioned. We have already seen one version of that panic earlier this year when DeepMind dropped. The second version, triggered by the clean academic paper, could land harder because it cannot be brushed off as a one-off competitive shot.

So, when you think about this, for the longest time the story in AI was simple: bigger models, smarter answers, repeat. This paper is one of the clearest signs that the story is not that simple anymore. The most human kind of thinking, the kind that connects unrelated ideas and find the pattern in between them, does not just scale up with size. It depends on something more delicate inside the model that you cannot guarantee by spending more money on compute. That does not mean AI is finished improving. It just means that the next round of improvement has to come from somewhere other than raw size. The labs that already understand this are quietly changing direction. The ones that are still chasing the old scaling story might be in for a much harder story than the investors are expecting.