Transcription
Hello community. So glad that you are back.
Today we talk about something brand new, context learning. And you might say, "Hey, wait a minute. I know context learning." No, you don't. Because I mean, too, I thought, "Hey, I know it." Turns out I'm wrong.
Look here, a publication here by Peking University, Xiamen University, and Tsinghua University, May 25th, 2026, about context chain of thought, the latest in chain of thought, enhancing the context learning via a high-quality reasoning synthesis. They tell us, "LLMs struggle significantly with context learning, the ability to dynamically extract, internalize, and apply new knowledge from complex task-specific context." And you might say, "Why?"
Now, they give you this example. You have a dishwasher here, a manual describing here some safety instructions, and then you just have questions, ne? So, you feed the complete safety manual in, and then you question, you see, a raw chain of thought will fail. It will fail for very particular reasons, like an unsupported inference, faulty comparison, option conflation, concept conflation. And we will find today a new solution, and then everything is green, and everything is perfect.
So, what is it? The fundamental capability here, learning generally new information for an LLM. And now we are not in the harness. We are now going back to the core of our agent, our LLM. Learning some generally new information from a specific context in the prompt, and then reasoning over it, rather than relying on the internalized pre-training memory, the parametric memory, the parametric knowledge of an LLM. This is defined as context learning. And I thought, "This is in-context learning, ne?" Turns out, no, I'm wrong, because there's a massive difference.
They authors define context learning. At first, they tell us it's a severe bottleneck for contemporary LLMs. Those LLMs fail massively in context learning. If you provide here a prompt and some examples, it will fail massively. And I thought it's doing it's doing great. Well, you see, because there's this study. This is here from February 3rd, 2026, Tencent here, and this is here CL Bench, context learning benchmark. And this is here Fudan University and Tencent, and they introduce now here in February this CL Bench. And they say, "Now we go for the real hard stuff, a real hard world benchmark consisting of 500 complex contexts, 1,899 tasks, 31,000 verification rubrics, all crafted by experienced domain expert. So, now we do some real-world testing." >> [snorts] >> And they say, "Each task is designed such that the new content that is required to solve it is contained within the corresponding context that we provide in the prompt to the LLM." And you feel, "Oh, there is now something different to in-context learning." And they tell us this goes far beyond here the long context task that probably tested the retrieval or the reading comprehension. And those what we called in the old times in-context learning, the ICL task, where the models learn simple task patterns via instruction or a few short examples and a demonstration. So, it is different.
So, let's have here a direct comparison. And I made this list here for me, and I want to share this with you. So, the classical in-context learning that we know now for years, ICL, measures here an LLM model's ability to infer some task formats or some active static pre-trained knowledge from a few prompt examples. So, if you want in-context learning leverages a pre-existing knowledge like a mathematician's familiarity with mathematics, where the few-shot example just clarify, in the simplest case, the format or some preferred structure, some preferred logic sequence based on a domain-specific, already trained, pre-training knowledge. So, you see in-context learning, in the real definition of years ago, is merely an activation mechanism. So, you have to have the training data, the pre-training data exactly on your mathematics. So, and now you just give it a few some few-shot examples that are a little bit different, but in the same domain, in the same complexity. But, you want it in another format, in a different structure. This is in-context learning.
But, now we go here a complete step further. We say that the new context learning evaluates an LLM's capacity to dynamically internalize, synthesize, and reason over, and hold on to your socks, genuinely novel, complex, or even contradictory knowledge to the pre-trained knowledge explicitly provided now in the prompt. So, this is now different. Give you an example. Context learning is now like asking the same mathematician to build an entirely new mathematics from alien rules. So, context learning is like being handed a rulebook for an entirely invented, non-Euclidean, alien universe, where, and let's make it crazy, where one and one equals three, the gravity pushes upwards, and specific spells summon some localized black holes. Complete crazy nuts. We have nothing that is as similar domain knowledge. We only have our classical mathematics here on Earth, but the alien mathematics on different planets or whatever are completely different. We have no base system to fall back to and say, "Oh, we just have some little modifications, some little alteration." This is a complete new body of knowledge. New rules, new domain knowledge, new sequences, and now we have a look how good are LLMs to do this. And I want to show you here the result.
They examined here February 2026 exactly this. And they give us here for the SEAL bench here, let's just just look at the overall, yeah? GPT 5.2 at the time in February, this was here the top model, 18% success. O3 high, 17% success. Gemini 3 Pro, 15% success. Deep Seek version 3.2 thinking, 13 percentage success. And it's just horrible. It is absolutely unable to perform this. Context learning is absolute heavy stuff that LLM fail. They fail in the domain knowledge reasoning and the rule system application and the procedural task execution and the empirical discovery and simulation. Complete failure.
So, what is happening now? They also looked here at the distribution of the error types across all the models. And they said, "What went wrong with a GPT 5.2 high?" And you see in 60% of the cases, the context that was provided in the prompt to a GPT 5.2 was ignored. Simply ignored. Or let's go here. 65% of the cases, the context was misused. Or we had format errors in 33% and so on. Yeah. You see those error types here, multiple error types happening now exactly if you have a deep dive what went wrong if we have now not an in-context learning with a GPT-5.2, but a context learning. A brand new domain knowledge here only in the prompt, not in the pre-training data, not in the parametric knowledge of our LLM. This is the truth. Our AI fails horribly.
And now this study here today looks here at the study here of CL bench and says, "Why? Can we optimize it?" So, a very simple example just to get a feeling for this, now. Imagine you're a brilliant classical physicist who has spent your complete lifetime just here on mechanics, Newtonian mechanics, now. And no quantum mechanics, nothing, now. And suddenly you're dropped here into a bizarre alternative universe governed entirely only by string theory, a complete different kind of mathematics. And to learn string theory, you are now handed a thick rulebook, now. Let's say this is here about the context, about the skills, about the procedures, about the sequences, whatever, explaining now this new reality. But you have been pre-trained your whole life only on Newtonian mechanics. And there is no one-to-one mapping from Newton to string theory. So, you have to learn now something only with this rulebook that is given to you, a complete new complexity in a domain you have never heard of. And this is why our AI models fail. They are unable to do so.
And now something is happening that Yota has found out. And this is important to understand why. Because this happens with human and this happens with AI system. Well, AI just mimics human. So, great. Let's say I immediately ask you to solve a complex trajectory problem here in quantum mechanics, and I give you not only the query, but also the final answer in this complete new domain. I tell you, listen, in quantum mechanics, the particle exists here in multiple states simultaneously. So, if you are now the AI or the researcher here, and I give you the answer, what will you do? Now, it turns out that the AI or the Newtonian physicist is lazy and strongly biased here by the Newtonian pre-training. It will skim here the rulebook and generate a highly convoluted post hoc classical rationalization to justify the answer I just gave you. And what's the particle exists in multiple states simultaneously. But, it is not by reading here the rulebook and understanding here quantum mechanics. It is an argumentation that is based on the classical Newtonian physics. So, if you want, this is a pure hallucination squared. This is hallucinating now a bridge between your old priors, your Newtonian mechanics, and the answer that I told you, multiple states simultaneously. Now, without quantum mechanics, this is not possible in Newtonian physics. So, therefore, the AI will invent, it will hallucinate, I have no idea what, but some crazy stuff. Now, it turns out not only humans are holding on to the old knowledge and try to explain the world from the old bubble, but also AI is doing this.
So, the authors said we have to break this. And this is the authors tell us what happens in a standard chain of thought data synthesis. We usually give you a frontier teacher model. Yes, of course, we are here distillation. So, we have a teacher AI model in the long context, the question that we have, and as I just giving you here, a ground truth answer. The portal exists in multiple states simultaneously. Asking it now to generate a reasoning trajectory in order to train a smaller, let's say local 4 billion pre-trained parameter student AI model. Now, the teacher now engages here in a supervision leakage. Exactly the same is happening now. This teacher now hallucinates a post hoc rationalization using only its pre-trained parametric knowledge, because this is what it has been trained up for half a year, rather than faithfully deriving now the truth from the new context. So, we have a problem. Those teacher LLMs are lazy. They don't want to build up a complete new understanding of string theory of quantum physics from the rule book. They think, "Okay, I have a pre-trained knowledge. I have my parametric knowledge in the tensor weights in my neural network. And this is, let's say, 99% of my knowledge. I will use now this parametric knowledge and build some crazy bridge, pure hallucination, to somehow explain the ground truth answer." And this is not a joke. This is really what the AI is doing.
So, the first insight from the authors was, they call it an epistemic blindfolding or a minimum leakage. And they said, "We must hide the final answer now from the teacher LLM during the data generation, because at first we have to generate the the set before we can do any training, eh? So, the force now to teach a LLM to blindly navigate the new physics rule book of quantum physics or of string theory. We don't tell them there are particles that exist in multiple states. And we wanted the teacher really try to derive the answer from scratch. And if it fails, and if it goes back to its parametric knowledge, now and only now we give you the teacher LLM the smallest possible hint. Hey, listen, buddy. Don't go back to Newtonian physics, go to quantum mechanics. So, a single failed rubric refusing to reveal the real target. So, hide the answer from the teacher. Because the teacher LLM will cheat massively. Amazing, eh? Absolutely amazing."
And the second insight from the authors was Yeah, some beautiful technical terms, latent geodesics or student inverse selection. This is easy to understand. If we have a teacher LLM that is huge, this is extremely intelligent, and then you have a student LLM, and I'm the student LLM, and I get an extreme complex explanation from the teacher, I as a student have no idea. Even if the teacher explains me, "Look, these are six steps and you have to do A, B, C, D." And I say, "I don't get it." So, what we have to do. Even if the teacher LLM derives a perfectly correct proof, highly complex mathematics, it might take a massive intuitive leap that a smaller student model, like I, cannot follow. It's simply too complex for me to follow. I don't have the knowledge to understand here the given sequence or whatever. So, therefore, we have to evaluate multiple valid proofs from the uh teacher LLM and select the one that represents path of the least cognitive friction through the specific students AI model probability distribution. So, this means we have to tell the teacher, "Listen, buddy. You just cannot think in your classical normal way of thinking. You have to think for a little 4 billion model student AI. So, you have to be extreme simple. You have to reduce the complexity further. Multiple steps, lower the complexity again, again, and again. We have to have the least cognitive friction. Currently, we're not doing this. And this is it. Those are the two most important element of the complete paper. If you have understood these two ideas, you've understood the paper. I told you it's easy, now.
So, therefore, what it does introduce is a three-stage pipeline to synthesize high-fidelity supervised fine-tuning training data because we will do some fine-tuning, of course. We will do a distillation, and we have to generate the data set. So, how we do it? Multi-state extraction, then I showed you the minimum leakage filtering, and the student level chain of thought selection. Beautiful. This is it. That's it.
Now, for the definition that's important, what is a rubric in AI? A structured assessment tool that outlines specific criteria for evaluating a piece of work, detailing different levels of performance for each criterion. It is used to provide a clear guidance for the assessment, ensuring thereby consistency and transparency in the grading or the evaluation process itself. This is what we're going to do.
Now, if you want to code this, what we have now are two ideas. So, now we have to transform those two ideas into a mathematical framework. And if we have a mathematical framework, and we have an optimization process or a training process. Then, we can write the code when we have the mathematical formulas. Because code is equal mathematical formulas. So, let's do this. We have the two ideas, and now we build a little tiny bit of mathematics. We try to understand what is happening. So, assume we have a context question pair. So, this is the question, and we provide context. No way about context learning. And a set of Rubik passing candidate reasoning trajectories already that were generated by the teacher LLM. This is a huge deep seek version four or whatever. The real expensive AI model, no? So, we have here the reasoning trajectories here. And in this reasoning trajectories, you see we have our Y1, Y2, YK. And each trajectory YI is a sequence or of reasoning steps, step one to step N. Great. So, what we do now? The question is simple. You know, how do we mathematically define a learnable trajectory for a specific lower intelligent AI student model MS, let's say a 4 billion model. Or 3 billion model, no? And the authors optimize here a convex combination of two distinct matrices. And they say, "Easy." At first, we have a look at the stepwise alignment in this reasoning complexity. And then, we have a look about the reasoning gains that we have with each single step. I mean, couldn't be easier, no? What is a learnable trajectory? You go step by step by step by step, and then you have a look, "Hey, how much does reasoning improve with every single step?" And yeah, is this here in a complexity level that the little 4 billion uh pre-trainable parameter model can handle this, understand it, that it is in its complexity parameter. Is it possible to solve it?
So, the step was alignment. So, it's two steps, no? First step now is step-wise alignment. What is it? Mathematically, it is more or less a smoothness, no? We calculate it a negative log likelihood. This means the difficulty of each individual step under the student distribution conditioned now on the context, the question, and of course on the prior steps. Simple log. And if a trajectory contains massive inferential jumps, the sequence of difficulties, D1 to DN, will be highly volatile, no? And to penalize now cognitive friction and enforce a smoother transition, the alignment score that we give now to this system is what? Now, guess what? Simple. It's a negative variance of the step difficulties. So, that's all there is to it.
The second is the reasoning gain as I told you, no? The cognitive objective utility. Beautiful. So, what is it in in the core understanding? It's a smooth trajectory is in itself, if you want, useless if it doesn't actually help the model arrive at the correct final answer to solve it. So, smoothness in itself is not enough. We have to have here a reasoning gain with each and every step. So, we measure how much conditioning under rational YI reduces the student's uncertainty via the perplexity. Okay, so we have perplexity one minus perplexity two. And if we have a positive S, mathematically, this guarantees that reading in the trajectory lowers the entropy of the target answer. Beautiful. Less entropy, less chaos, more knowledge.
Then we have to combine it. Look, plus. So, the final selection objective is simple, no? Careful, because the values exist on different numerical scales. So, we have to min-max normalize this now, bring it back to one common scale. And we have a hyperparameter that you can select, beautiful. And the algorithm selects now this particular trajectory and augmax as the final supervised fine-tuning target. And this is it. Now we have the perfect sequence of a reasoning complexity that is adapted here from the teacher LLM given the little complexity of a 4 billion LLM. And now we can start with the supervised fine-tuning. Now we have the training data set for it.
If you want to have here a look at this, this is here from a screenshot from the original study. Everything that I told you up until now, multi-stage chain of thought sampling, rubrics-based minimum leakage filtering, and the student aware chain of thought selection here is exactly here this in one image. That's it. That's the study. Beautiful. Isn't it simple? Just have to have the ideas.
Oh, yeah, maybe we should look at the results. Now, this is a little bit strange, I have to tell you, because yeah, they they make it a little bit over the top. But let's have a look at the figure at the facts. So, to see the devastating effects of the post hoc rationalization, so when the LLM is just lying, referring to its parametric and not to its new context learned, and the triumph of the new method, the context chain of thought mathematics, you have here. So, if we do the fine-tuning here on a QWERT 3.5 4 billion with the new training data set that we get out of all of this, you see here a Yeah, forget about everything. The overall performance here, the first column here of the base model is 9.06. And if we would have only used an answer exposed chain of thought, so we give the correct answer to the to the if you want boss LLM here to the teacher LLM, then we have all the hallucination. So you see here, if you do this and do a fine-tuning, in those case the performance is reduced. So it is a good idea, remember two things that we learned here, that we do not provide the correct answer to the teacher LLM. Then we get a better performance. So 9.06 compared to 8.59. Now, when I saw this I thought this is really a minimum difference, no? But okay, it is an indicator in the right direction, good. But have a look now at this complete new methodology, what is it now with the context chain of thought supervised fine-tuning here at Q and 3.5 4 billion, and now we jump from 9 to 12.85. Okay, so 12.85% is not great, but it is better than before. So therefore, let's say we accept it, no? So you see this really validates here our what I call the physicist analog, no? The Newtonian physicist here and gets a string theory or quantum mechanics, no? So when you expose the answer to the teacher LLM, it will generate logically flawed parametric dependent shortcuts. It will simply hallucinate, and the teacher LLM will lie to you, to the human, and to the student LLM. And the student will learn these flawed shortcuts from the Newtonian physics and will get worse at reading novel context, no? And they see they argue, you see it here, answer exposed chain of thought supervised here to the base model. Okay. Okay.
And then it's a manipulation study. This is nice, no? They show us here in A, the top one, increasing the number of context chain of thought training sample consistently improves the performance. But after 3,000 samples, you start here to plateau out. Okay. But here, this is now interesting here. If you have here on the Y axis here the CL bench overall performance in percentage, and please note, we are here at 10% going to 13%, which is ridiculously low. This is a massive failure, but at least we are not below 10%. So, we are now already happy that we have above 10%, and we can talk about 13%, which tells you current generation of LLMs here is not able to do context learning at all. But we have to have We have to find any way to go at least to 30%. Let's start. Baby steps. If you have a look here at the Y and the student ever alignment weight or a lambda that I showed you here in the mathematics, there's an absolute sweet spot here in the alignment weight structure. And I see here dot four, this is here the sweet spot. Almost 13%. Indicating that here the balancing here the stepwise alignment and the reasoning gain is really important. Interesting. Now, look at this. If we would be at a lambda of zero, no? We care about getting here only the right answer. The performance drops compared from 13, we go down to 11. If we say, "Hey, let's go for lambda equal one." So, we are at here this point here, when we only care about the smooth text ignoring here if it leads to the answer. The performance also plummet, no? So, what do you see? We have a multi-objective optimization problem. It is not enough just to get the right answer. It is not enough to have a smooth differential mathematical formulation. Both have to work in tandem. Both have to come together. It is a new theory of a multi-objective frontier. So, simultaneously it must bridge the probability gaps in the student models latent space while mathematically lowering here, as I showed you in the formula, the perplexity of the final tokens. So, distillation having here an all-knowing teacher LLM and a little student LLM, this distillation is not at all understood. It is not at all simple. And now we discover, my goodness, we have here multiple objectives optimization that we have never thought about before.
The overall result, just be absolutely clear about this. If you do not a fine-tuning on this new data set that we have, on roughly 4,000 high-quality now context chain of thoughts, those synthesized trajectory here from these huge LLMs, and we have those open-source models then that we have a supervised fine-tuning here like a Qwen 3.5 4 billion on this high-quality data set, what is now the jump into performance? And the authors tell us the absolute gain of nearly up to 4%. And they argue this is a massive relative jump [laughter] in the notoriously difficult CL benchmark. And it's right. It is extremely difficult the CL benchmark. And 4% is is already an indicator that it is in the right direction. but you know what? 4% is for me we have no idea what we are missing in the eye. We have no idea how we should optimize the pre-training, what we have to integrate on the context learning, how we should modify the post training of our AI system because a context chain of sort it does not bring us up to 20% or I'm not dreaming of 50%, but you know, it shows us my goodness, we are we are really completely in the dark with AI development with the models.
Some personal insights. This is now end of the paper. This is just me reflecting here a little bit. As I showed you here with lambda equals zero, the truth, the correctness does not equal a learnability in AI. A rational a reasoning trace, highly complex, highly dense, generated by the latest GPT-5, you might think, oh yeah, this is what we want. It might be perfectly valid and correct but probabilistically absolute alien to a 4 billion parameter model. You can't use it. It is absolute waste of energy because now we understand the distillation must be treated as an alignment problem between two distinct manifold topologies. And today I showed you there two ideas two methodologies, yeah that we tried to implement. And is it just two that are just bringing us four percentage. So, I dare to say maybe this is just the beginning and we are missing something massively. So, therefore, distillation to absolute distinct manifold topologies and more or less, we have no idea how to do this in AI. But and the simple thing is, hey, I really like to perform the supervised fine-tuning not on the 4B model, but they say, you know what? We use even for a 4B model a LoRA, a lowering adaptation adapter that we put on. Okay, they go with a rank of 32, so absolute beautiful. It's not eight or something, it's 32 rank. I love it. So, they freeze the original weight here, the tensor weight, and train a small low-rank matrix here injected here into the attention MLP projection layer. Beautiful as you know it. So, a funny even with a 4B you can take a LoRA and get this result. So, the next step would be, okay, not take a LoRA, but really do the fine supervised fine-tuning here on the 4B model. I mean, come on, this is possible. We don't need here the shortcut of LoRA. And then maybe we get 5%. I mean, just let me dream here of performance, yeah? So, just want to show you where we are today here, and I think this is a beautiful paper. Have a look at this read. There are some additional data in there. There are some experiments described here in detail. It gives you a feeling where we are and what we are completely missing out. And whenever you use distillation and teacher-student distillation and transfer of knowledge, now you understand, my goodness, we have almost no idea how to do this. I hope you enjoyed it. I hope you had a little bit of fun. I hope there was any information for you, and it would be great to see you in my next video.