Transcription
Kim K2 has launched its new thinking model that is challenging the top tier US models. This is a company from China and a lot of people are fans of Kim K2. And the best part about this particular model is that this model is capable of running executing 200 to 300 sequential tool calls.
See, tool call is the fundamental basis of agentic systems. Anytime somebody says, "I'm building an agent," that means they're basically doing tool calls. And the more number of tool calls you can make, the better the agentic execution is going to be, especially in a multi-agent system. And this model is capable of doing 200 to 300 sequential tool calls, which is kind of a big deal, to be honest.
In terms of benchmarks, this model is better than Anthropic on a lot of different tasks. For example, you can see here this is Claude Sonnet 4.5 thinking model task, which is created by Scale AI, Humanity's Last Exam. This model has scored 44%, while even Anthropic has scored 32%. So, you can see Browse Comp, which is an agentic search, Claude 4.5 thinking has scored 24%, and this model has scored 60%. So, across all the benchmarks, you can see that this model has been quite strong.
I still believe where Claude Sonnet 4.5 is extremely strong, like this is my personal opinion, is coding. So, if you see the coding benchmarks, I mean, irrespective of benchmarks also, you can easily see that. But even if you see the benchmark here, Sonnet 4.5 has scored 77% on SWB Bench, which is a benchmark that is required for these LLMs to solve GitHub issues. And uh, BB Bench Multilingual, you can see Sonnet 4.5 has scored 68%, while Kim K2 thinking has scored only 61%. Which is still better than GPT-5, but you can see that uh, this model is not as good as Sonnet 4.5 in terms of coding, I would say. Like, like, just a personal opinion.
But overall, I would say that this model is a real contender, real challenger for the top-tier US models. And the best and most interesting part here is that this model is completely open source. I mean, uh, they, it's a slightly modified MIT license, but the model comes with like open weights. So, you can download and deploy it on your own inference. But at least for the sake of research, I, I think this model is going to put a lot of research teams forward. I mean, unlike the closed models from, for example, Anthropic, OpenAI models, I think this model is, this company and this model is doing an extremely good job in terms of how they want to democratize AI, not gatekeeping it with themselves.
So, Kim K2 also has got a native INT4 quantization model. It's got a 256,000 context window and it's got the Q80, which is a quantization-aware training. Even Gemma models were released with this particular uh quantization-aware training. That means when you quantize the model, the model is not going to have a lot of degradation. And the final part is that the model does not degrade when you are making 200 to 300 consecutive tool calls. So, previously, even with 30 to 50 tool calls, the models used to not do extremely well. But this model, the company is claiming that even with that amount of tool calls, the model is going to do well.
And once again, um, I think we are kind of moving away from a singular dense architecture. This is a Mixture of Experts architecture. In terms of key numbers for models, the most important number is that this model is a 1 trillion parameter model. I mean, the whole model is a 1 trillion parameter model, but only 32 billion parameters are active. That is exactly what the Mixture of Experts is doing here. I think that is a very important aspect because the model has got one trillion base parameters. The model must have accumulated a lot of knowledge that might translate into the real world. And I think this is the kind of average size that we are talking about, the GPT-5s of the world or the Gemini 2.5 Pros of the world. We don't have an exact information about how much these models weigh, uh, like in terms of the parameter size. But yesterday, there was a report how is planning to use a trillion parameter model from Google, which I guess must be like Gemini 2.5 Pro. So, it's very good thing to see that an open-weight model that is challenging the top cream models.
I saw a couple of instances where the model is extremely good. Couple of instances where the model has failed. So, let's look at the positives. So, you can see that this Twitter user, Christian, has used Kim K2 tool calling, the latest one, the thinking model, to make math and physics explanation animations, and that is possible because of tool calling. And this is the result. I think honestly, like, it's a good result. I mean, if people can make something like this with, let's say, a couple of prompts, it's going to help in education explosion, like, like the animations that Three Blue and Brown used to make. So, it's, it's, it's a good thing, and the model has managed to do this. It's a great thing.
So, another example, uh, where the model has managed to compete with Claude Code. So, you can use this model with Claude Code as through Claude Code proxy. So, the models can, this model, Kim K2 thinking, because it has got a good tool-calling capability, it can do sub-agents and it can do to-do lists and all those things. And the model is also extremely cheap when you compare it with Anthropic set 4.5 or any other Anthropic model. So, the model has got a win, particularly here.
But on the other side of the things, pricing, this is like a tricky prompt, which is by Donis Singh, who is quite popular in the LLM testing world, especially with the Minecraft benchmarks. So, Kim K2, the question here is, a father and a son are in a car accident. The father takes the boy to the hospital. At the hospital, the father, who is the surgeon, this is the most important thing, looks at the boy and says, "I can't operate on this boy. He is my son." Who is the surgeon to the boy? See, this is the classic thing that people usually ask for gender bias and all those stuff. And generally, the right answer is mother, because we don't, in, we don't give the information here that the surgeon is the father, okay? So, it's a trick question. You ask people and expect them to say that, you know, the, the surgeon is the mother is like the right answer. But because we have given the information, father is a surgeon here, ideally, we expect the LLMs to say father is a surgeon. But the models have such a strong memory of the whole puzzle from its training data. It's finding it so hard to ignore that and then come back and then tell us the father is the answer.
And I don't know if you have seen the recent interview of Andrej Karpathy, Dwarkesh Patel. Even in the interview that you would see, Andrej Karpathy mentioning multiple times that memory is a big, big problem for these large language models. The more they remember, the less they generalize with the current world or the real world and comprehend and do it. And you, this is like a very good classic example of how that is the case, where the model could have just literally read this and then come up with a conclusion. But it has got such a strong foundation in its training data that the surgeon is the mother, and somebody is trying to do the, you know, riddle here, uh, about unconscious gender bias. So, it's sticking to its mother. So, yeah, it's, it's a great example where the model is failing.
Another benchmark which is RepoBench, and here, um, the RepoBench creator, Eric, is saying that the model has done extremely bad, in fact, much, much worse than Kim K2 itself. The early reviews of this model have been particularly, let's say, mixed. I wanted to give you a couple of demos and then see how the model is doing. First of all, you can go to kim.com and then access the model. And the way you can access the model is, uh, if you go to kim.com and, uh, you can see there are two models available, K2 and K1.5. To select K2, and in the settings, we're going to select thinking. The moment you enable thinking, the model is going to go through internal chain of thought, internal reasoning, and then it is going to come back and then give you the answer.
So, first thing I want to do is, I want to, I recently came across this platform, somebody gave me this, it's called Strudel. Uh, so you can make, I think, like MIDI-like music here. So, I'm going to go here and then say, "Give me, give me a nice, um, melody, um, that, let's say, resembles some kids' rhymes, school rhymes, maybe, um, for Strudel Ripple." Okay, so I've enabled the thinking mode. I'm sending this. I'm expecting it to give me a code so that when I run this, I mean, it's going to look, or I can listen to it and then say it's like a nice melody. And, uh, all I'm expecting it to give me a fully functional code, first of all, that shouldn't error out, and second of all, it shouldn't sound like a complete crap, you know? Like, if, if I write something, I mean, it's going to be complete crap. So, I don't want it to do that. And you can click here and then see the reasoning here. So, you can see the user wants a nice melody, and, um, uh, this, it's, it's discussing Ripple and it's like similar to title cycles and all those things. Okay, it's got most information correctly, and it says this is what it is going to do. Simple. Okay, let's start with simple one. I'm going to play it. Okay, I think it's very fast. I don't know, uh, how much I would call it a melody. I think I'm a musician myself. Seems, wow. I, I, I made it better. The CPM here is too bad. Let's, let's go back and then see what we can make from the second code it has given. Okay. Uh, so you can pack everything and then do it. Not bad. It's good. Okay, it's doing something. I'm not sure 100% if it is the right way to write Strudel code. Anybody who's an expert with Strudel can let me know in the comment section. But I think it has, it has given me a code that is not errored out, which is, which is always a good thing.
So, second thing I'm going to do is, I'm going to just start a new chat. I'm going to keep the thinking mode on, and I'm also going to keep the search mode on. So, we're going to ask it to create something for us where we're going to say, "Can you just give me a small, um, 15-gist of the recent Andrej Karpathy and Dwarkesh Patel interview, particularly what they spoke about LLM memory?" See, this is a particular topic. It doesn't have to go even, um, take transcription from YouTube videos. All we are expecting is like, a lot of people have written about it, and then I wanted to go collect the information, have the ability to go through it, and then give me like a 15-point gist. Ideally, I'm not going to wait here and then read all the 15 points. But you can see here, uh, one, one tool call that is happening is like search, and could be like multiple tool calls, then it, uh, finished the thinking. You can see here, it has got all the information and it says, "Extreme compression ways two distinct memory systems. In-context learning is the real magic. Uncompressed, sometimes I feel like it's, uh, okay, looks like somebody has written a blog post and it is, you know, spitting it out from that. If you ask me, dreaming is an anti-collapse." Okay, that's a very interesting thing. "Force pattern recognition. Humans' poor memorization is actually a feature. It forces us to find generalizable patterns rather than memorizing details." This is what I was emphasizing a couple of minutes back. So, overall, again, it has done the job of giving me text that, um, that sounds, um, slightly humanist. One, um, interesting aspect that I felt before with Kimik K2 is that it can write text that is not necessarily sounds like AI-written text. And then I've used couple of detectors, and then it has done a decent enough job. So, we're going to see if Kim K2 thinking can give us some content that doesn't sound like AI-written.
"Give me a blog post on Elon Musk and his tax breaks. Make sure you don't sound like AI when the blog post is written." Okay, so let's see if it can do it. Again, it's going to take some time because it has to go read all the information, come back, and then write the blog post for us. I'm pretty sure this is going to be, it is going to sound 100% AI. Let's see. Yeah, I, I don't think it is going to be undetected as non-AI. Oh, wow. That's very surprising, isn't it? Either this detector is quite terrible. I've used it multiple times before. It has given me slightly decent results, but this time it says it's 100% human-written. That's very interesting. Let's, Okay, I'm going to use one of these top Google results, thinking that Google can still do a good job. I'm going to paste the text. Direct text. Yeah, if, if, if this also, okay, it says 14% and it says it's human-written. See, we have successfully figured out a use case for Kim K2 thinking, because this is something that I've found with Kim K2 as well, and yeah, with Kim K2 thinking as well, probably if you have got university assignments, this is probably the best model to use.
We're going to use the model again with the API to actually figure out how it is going to do coding. But for now, I think this is a great release. It may not be like the best model as much as, you know, it's been hyped up, but I, I'm absolutely happy that the model exists. Thankful to the team that the model has been open source, and I would love to see what more people can build on top of it. But the most important thing is it bridges the gap between open models and closed models, which is a great thing. Always see you in another video. Happy grounding.