📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Deep Dive into LLMs like ChatGPT

Andrej Karpathy3:31:24

Transcription

human-generated.

So, to summarize what we've discussed so far, we started with the pre-training stage, where we took internet documents, broke them into tokens, and predicted token sequences using neural networks. The output of this stage is a base model, which is essentially an internet document simulator at the token level.

Now, we want to move into the post-training stage, where we transform this base model into an assistant. This involves creating datasets of conversations that the model can learn from, allowing it to respond to human queries in a more meaningful way.

In this stage, we will be programming the assistant's behavior through examples. Human labelers will create conversations, providing ideal responses to various prompts. The model will then be trained on these conversations, learning the statistical patterns of how to respond appropriately.

The post-training stage is computationally less expensive than pre-training, as the datasets of conversations are much smaller. This training can take only a few hours compared to the months required for pre-training.

During inference, when the model is used, it will generate responses based on the patterns it learned during training. The goal is to create a system that can engage in multi-turn conversations, responding accurately and helpfully to user queries.

As we move forward, we will explore how these conversations are structured, how they are tokenized, and how the model generates responses during inference.

In conclusion, the journey from a base model to a fully functional assistant involves careful design and training on conversation datasets, allowing the model to learn how to interact with users effectively. This process is crucial for developing language models that can serve as helpful assistants in various applications.

human, and it's kind of like, um, gone in that direction since, uh, but roughly speaking, we still have SFT data sets. They're made up of conversations we're training on them, um, just like we did before.

And, uh, I guess like the last thing to note is that I want to dispel a little bit of the magic of talking to an AI. Like when you go to ChatGPT and you give it a question and then you hit enter, uh, what is coming back is kind of like statistically aligned with what's happening in the training set.

And these training sets, I mean, they really just have a seed in humans following labeling instructions. So what are you actually talking to in ChatGPT, or how should you think about it? Well, it's not coming from some magical AI. Like roughly speaking, it's coming from something that is statistically imitating human labelers, which comes from labeling instructions written by these companies.

And so you're kind of imitating this, uh, you're kind of getting, um, it's almost as if you're asking a human labeler. Imagine that the answer that is given to you, uh, from ChatGPT is some kind of a simulation of a human labeler.

And it's kind of like asking what would a human labeler say in this kind of a conversation. And, uh, it's not just like this human labeler is not just like a random person from the internet because these companies actually hire experts.

So for example, when you are asking questions about code and so on, the human labelers that would be involved in the creation of these conversation data sets, they will usually be educated expert people. And you're kind of like asking a question of like a simulation of those people, if that makes sense.

So you're not talking to a magical AI; you're talking to an average labeler. This average labeler is probably fairly highly skilled, but you're talking to kind of like an instantaneous simulation of that kind of a person that would be hired in the construction of these data sets.

So let me give you one more specific example before we move on. For example, when I go to ChatGPT and I say, "Recommend the top five landmarks to see in Paris," and then I hit enter, uh, okay, here we go.

Okay, when I hit enter, what's coming out here? How do I think about it? Well, it's not some kind of a magical AI that has gone out and researched all the landmarks and then ranked them using its infinite intelligence, etc. What I'm getting is a statistical simulation of a labeler that was hired by OpenAI.

You can think about it roughly in that way. And so if this specific, um, question is in the post-training data set somewhere at OpenAI, then I'm very likely to see an answer that is probably very, very similar to what that human labeler would have put down for those five landmarks.

How does the human labeler come up with this? Well, they go off and they go on the internet and they kind of do their own little research for 20 minutes, and they just come up with a list, right? Now, so if they come up with this list and this is in the data set, I'm probably very likely to see what they submitted as the correct answer from the assistant.

Now, if this specific query is not part of the post-training data set, then what I'm getting here is a little bit more emergent, uh, because, uh, the model kind of understands statistically, um, the kinds of landmarks that are in this training set are usually the prominent landmarks, the landmarks that people usually want to see, the kinds of landmarks that are usually, uh, very often talked about on the internet.

And remember that the model already has a ton of knowledge from its pre-training on the internet, so it's probably seen a ton of conversations about Paris, about landmarks, about the kinds of things that people like to see.

And so it's the pre-training knowledge that has then combined with the post-training data set that results in this kind of an imitation. Um, so that's, uh, that's roughly how you can kind of think about what's happening behind the scenes here in this statistical sense.

Okay, now I want to turn to the topic of LLM psychology, as I like to call it, which is what are sort of the emergent cognitive effects of the training pipeline that we have for these models.

So in particular, the first one I want to talk to is, of course, hallucinations. So you might be familiar with model hallucinations; it's when LLMs make stuff up. They just totally fabricate information, etc.

And it's a big problem with LLM assistants. It is a problem that existed to a large extent with early models, uh, from many years ago, and I think the problem has gotten a bit better, uh, because there are some medications that I'm going to go into in a second.

For now, let's just try to understand where these hallucinations come from. So here's a specific example of a few, uh, of three conversations that you might think you have in your training set, and, um, these are pretty reasonable conversations that you could imagine being in the training set.

So like, for example, who is Cruz? Well, Tom Cruz is a famous American actor and producer, etc. Who is John Barasso? This turns out to be a U.S. senator, for example. Who is Genghis Khan? Well, Genghis Khan was blah, blah, blah.

And so this is what your conversations could look like at training time. Now, the problem with this is that when the human is writing the correct answer for the assistant in each one of these cases, uh, the human either knows who this person is or they research them on the internet, and they come in and they write this response that kind of has this like confident tone of an answer.

And what happens basically is that at test time, when you ask for someone who is this, is a totally random name that I totally came up with, and I don't think this person exists, um, as far as I know. I just tried to generate it randomly.

The problem is when we ask, "Who is Orson Kovats?" The problem is that the assistant will not just tell you, "Oh, I don't know." Even if the assistant and the language model itself might know inside its features, inside its activations, inside of its brain, sort of, it might know that this person is like not someone that, um, that is that it's familiar with.

Even if some part of the network kind of knows that in some sense, the, uh, saying that, "Oh, I don't know who this is," is not going to happen because the model statistically imitates its training set. In the training set, the questions of the form "Who is blah?" are confidently answered with the correct answer.

And so it's going to take on the style of the answer, and it's going to do its best. It's going to give you statistically the most likely guess, and it's just going to basically make stuff up because these models, again, we just talked about it, they don't have access to the internet.

They're not doing research; these are statistical token tumblers, as I call them, uh, just trying to sample the next token in the sequence, and it's going to basically make stuff up.

So let's take a look at what this looks like. I have here what's called the inference playground from Hugging Face, and I am on purpose picking on a model called Falcon 7B, which is an old model. This is a few years ago now, so it's an older model.

So it suffers from hallucinations, and as I mentioned, this has improved over time recently. But let's say, "Who is Orson Kovats?" Let's ask Falcon 7B to instruct run. Oh yeah, Orson Kovats is an American author and science fiction writer.

Okay, this is totally false; it's a hallucination. Let's try again. These are statistical systems, right? So we can resample. This time, Orson Kovats is a fictional character from this 1950s TV show. It's total BS, right?

Let's try again. He's a former minor league baseball player. Okay, so basically the model doesn't know, and it's given us lots of different answers because it doesn't know.

It's just kind of like sampling from these probabilities. The model starts with the tokens "Who is Orson Kovats?" assistant, and then it comes in here, and it's getting these probabilities, and it's just sampling from the probabilities, and it just like comes up with stuff.

And the stuff is actually statistically consistent with the style of the answer in its training set, and it's just doing that. But you and I experience it as made-up factual knowledge.

But keep in mind that, uh, the model basically doesn't know, and it's just imitating the format of the answer, and it's not going to go off and look it up, uh, because it's just imitating, again, the answer.

So how can we, uh, mitigate this? Because, for example, when we go to ChatGPT and I say, "Who is Orson Kovats?" and I'm now asking the state-of-the-art model from OpenAI, this model will tell you, "Oh, so this model is actually even smarter."

Because you saw very briefly it said, "Searching the web." Uh, we're going to cover this later. Um, it's actually trying to do tool use and, uh, kind of just like came up with some kind of a story.

But I want to just, "Who is Orson Kovats?" did not use any tools. I don't want it to do web search. There's a well-known historical or public figure named Orson Kovats.

So this model is not going to make up stuff. This model knows that it doesn't know, and it tells you that it doesn't appear to be a person that this model knows.

So somehow we sort of improved hallucinations, even though they clearly are an issue in older models, and it makes totally, uh, sense why you would be getting these kinds of answers if this is what your training set looks like.

So how do we fix this? Okay, well, clearly we need some examples in our data set where the correct answer for the assistant is that the model doesn't know about some particular fact.

But we only need to have those answers be produced in the cases where the model actually doesn't know. And so the question is, how do we know what the model knows or doesn't know?

Well, we can empirically probe the model to figure that out. So let's take a look at, for example, how Meta, uh, dealt with hallucinations for the Llama 3 series of models as an example.

So in this paper that they published from Meta, we can go into hallucinations, which they call here factuality, and they describe the procedure by which they basically interrogate the model to figure out what it knows and doesn't know, to figure out sort of like the boundary of its knowledge.

And then they add examples to the training set where for the things where the model doesn't know them, the correct answer is that the model doesn't know them, which sounds like a very easy thing to do in principle, but this roughly fixes the issue.

And the reason it fixes the issue is because remember, like the model might actually have a pretty good model of its self-knowledge inside the network. So remember we looked at the network and all these neurons inside the network.

You might imagine that there's a neuron somewhere in the network that sort of like lights up for when the model is uncertain. But the problem is that the activation of that neuron is not currently wired up to the model actually saying in words that it doesn't know.

So even though the internal of the neural network knows, because there's some neurons that represent that, the model will not surface that. It will instead take its best guess so that it sounds confident, um, just like it sees in a training set.

So we need to basically interrogate the model and allow it to say, "I don't know" in the cases that it doesn't know. So let me take you through what Meta roughly does.

So basically what they do is here I have an example. Uh, Dominic Kek is, uh, the featured article today. So I just went there randomly, and what they do is basically they take a random document in a training set, and they take a paragraph.

And then they use an LLM to construct questions about that paragraph. So for example, I did that with ChatGPT here. So I said, "Here's a paragraph from this document. Generate three specific factual questions based on this paragraph and give me the questions and the answers."

And so the LLMs are already good enough to create and reframe this information. So if the information is in the context window, um, of this LLM, this actually works pretty well.

It doesn't have to rely on its memory; it's right there in the context window, and so it can basically reframe that information with fairly high accuracy.

So for example, it can generate questions for us like, "For which team did he play?" Here's the answer. "How many cups did he win?" Etc.

And now what we have to do is we have some question and answers, and now we want to interrogate the model. So roughly speaking, what we'll do is we'll take our questions and we'll go to our model, which would be, uh, say Llama, uh, in Meta.

But let's just interrogate Falcon 7B here as an example. That's another model. So does this model know about this answer? Let's take a look.

Uh, so he played for Buffalo Sabers, right? So the model knows, and the way that you can programmatically decide is basically we're going to take this answer from the model and we're going to compare it to the correct answer.

And again, the models are good enough to do this automatically, so there's no humans involved here. We can take, uh, basically the answer from the model, and we can use another LLM judge to check if that is correct according to this answer.

And if it is correct, that means that the model probably knows. So what we're going to do is we're going to do this maybe a few times. So, okay, it knows it's Buffalo Sabers. Let's drag in, um, Buffalo Sabers.

Let's try one more time. Buffalo Sabers. So we asked three times about this factual question, and the model seems to know, so everything is great.

Now let's try the second question: "How many Stanley Cups did he win?" And again, let's interrogate the model about that, and the correct answer is two.

So, um, here the model claims that he won, um, four times, which is not correct, right? It doesn't match two, so the model doesn't know; it's making stuff up.

Let's try again. Um, so here the model again, it's kind of like making stuff up, right? Let's drag in here. It says he did not even win during his career.

So obviously the model doesn't know, and the way we can programmatically tell again is we interrogate the model three times and we compare its answers, maybe three times, five times, whatever it is, to the correct answer.

And if the model doesn't know, then we know that the model doesn't know this question. And then what we do is we take this question, we create a new conversation in the training set.

So we're going to add a new conversation to the training set, and when the question is, "How many Stanley Cups did he win?" the answer is, "I'm sorry, I don't know," or "I don't remember."

And that's the correct answer for this question because we interrogated the model and we saw that that's the case. If you do this for many different types of, uh, questions for many different types of documents, you are giving the model an opportunity to, in its training set, refuse to say based on its knowledge.

And if you just have a few examples of that in your training set, the model will know, um, and has the opportunity to learn the association of this knowledge-based refusal to this internal neuron somewhere in its network that we presume exists.

And empirically, this turns out to be probably the case, and it can learn that association that, hey, when this neuron of uncertainty is high, then I actually don't know, and I'm allowed to say that I'm sorry, but I don't think I remember this, etc.

And if you have these, uh, examples in your training set, then this is a large mitigation for hallucination, and that's roughly speaking why ChatGPT is able to do stuff like this as well.

So these are kinds of, uh, mitigations that people have implemented and that have improved the factuality issue over time.

Okay, so I've described mitigation number one for basically mitigating the hallucinations issue. Now we can actually do much better than that.

Uh, instead of just saying that we don't know, uh, we can introduce an additional mitigation number two to give the LLM an opportunity to be factual and actually answer the question.

Now, what do you and I do if I was to ask you a factual question and you don't know? Uh, what would you do, um, in order to answer the question? Well, you could, uh, go off and do some search and, uh, use the internet, and you could figure out the answer and then tell me what that answer is.

And we can do the exact same thing with these models. So think of the knowledge inside the neural network, inside its billions of parameters. Think of that as kind of a vague recollection of the things that the model has seen during its training during the pre-training stage a long time ago.

So think of that knowledge in the parameters as something you read a month ago, and if you keep reading something, then you will remember it, and the model remembers that.

But if it's something rare, then you probably don't have a really good recollection of that information. But what you and I do is we just go and look it up.

Now, when you go and look it up, what you're doing basically is like you're refreshing your working memory with information, and then you're able to sort of like retrieve it, talk about it, or etc.

So we need some equivalent of allowing the model to refresh its memory or its recollection, and we can do that by introducing tools for the models.

So the way we are going to approach this is that instead of just saying, "Hey, I'm sorry, I don't know," we can attempt to use tools. So we can create a mechanism by which the language model can emit special tokens, and these are tokens that we're going to introduce new tokens.

So for example, here I've introduced two tokens, and I've introduced a format or a protocol for how the model is allowed to use these tokens. So for example, instead of answering the question when the model does not, instead of just saying, "I don't know, sorry," the model has the option now to emit the special token "search start," and this is the query that will go to like Bing.com in the case of OpenAI or say Google search or something like that.

So it will emit the query and then it will emit "search end." And then here what will happen is that the program that is sampling from the model that is running the inference, when it sees the special token "search end," instead of sampling the next token in the sequence, it will actually pause generating from the model.

It will go off, it will open a session with Bing.com, and it will paste the search query into Bing, and it will then, um, get all the text that is retrieved.

And it will basically take that text, it will maybe represent it again with some other special tokens or something like that, and it will take that text and it will copy-paste it here into what I tried to like show with the brackets.

So all that text kind of comes here, and when the text comes here, it enters the context window. So the model, so that text from the web search is now inside the context window that will feed into the neural network.

And you should think of the context window as kind of like the working memory of the model. That data that is in the context window is directly accessible by the model; it directly feeds into the neural network.

So it's not anymore a vague recollection; it's data that it has in the context window and is directly available to that model. So now when it's sampling the new tokens here afterwards, it can reference very easily the data that has been copy-pasted in there.

So that's roughly how these, um, how these tools use, uh, tools function. And so web search is just one of the tools. We're going to look at some of the other tools in a bit.

But basically, you introduce new tokens, you introduce some schema by which the model can utilize these tokens and can call these special functions like web search functions.

And how do you teach the model how to correctly use these tools, like say web search, "search start," "search end," etc.? Well, again, you do that through training sets.

So we need now to have a bunch of data and a bunch of conversations that show the model by example how to use web search. So what are the settings where you are using the search, um, and what does that look like?

And here's by example how you start a search and the search, etc. And, uh, if you have a few thousand maybe examples of that in your training set, the model will actually do a pretty good job of understanding, uh, how this tool works.

And it will know how to sort of structure its queries, and of course, because of the pre-training data set and its understanding of the world, it actually kind of understands what a web search is.

And so it actually kind of has a pretty good native understanding, um, of what kind of stuff is a good search query.

Um, and so it all kind of just like works. You just need a little bit of a few examples to show it how to use this new tool, and then it can lean on it to retrieve information and, uh, put it in the context window.

And that's equivalent to you and I looking something up because once it's in the context, it's in the working memory, and it's very easy to manipulate and access.

So that's what we saw a few minutes ago when I was searching on ChatGPT for "Who is Orson Kovats?" The ChatGPT language model decided that this is some kind of a rare, um, individual or something like that.

And instead of giving me an answer from its memory, it decided that it will sample a special token that is going to do web search. And we saw briefly something flash; it was like using the web tool or something like that.

So it briefly said that, and then we waited for like two seconds, and then it generated this, and you see how it's creating references here, and so it's citing sources.

So what happened here is it went off, it did a web search, it found these sources, and these URLs, and the text of these web pages was all stuffed in between here.

And it's not showing here, but it's basically stuffed as text in between here, and now it sees that text, and now it kind of references it and says that, okay, it could be these people, citation could be those people, citation, etc.

So that's what happened here, and that's why when I said, "Who is Orson Kovats?" I could also say, "Don't use any tools," and then that's enough to, um, basically convince ChatGPT to not use tools and just use its memory.

And I also went off and I, um, tried to ask this question of ChatGPT, "So how many Stanley Cups did, uh, Dominic Hasek win?" And ChatGPT actually decided that it knows the answer, and it has the confidence to say that, uh, he won twice.

And so it kind of just relied on its memory because presumably it has, um, it has enough of a kind of confidence in its weights, in its parameters, and activations that this is, uh, retrievable just from memory.

Um, but you can also conversely use web search to make sure, and then for the same query, it actually goes off and it searches, and then it finds a bunch of sources, it finds all this, all of this stuff gets copy-pasted in there, and then it tells us, uh, again, and cites, and it actually says the Wikipedia article, which is the source of this information for us as well.

So that's tools, web search. The model determines when to search, and then, uh, that's kind of like how these tools, uh, work, and this is an additional kind of mitigation for, uh, hallucinations and factuality.

So I want to stress one more time this very important sort of psychology point: knowledge in the parameters of the neural network is a vague recollection. The knowledge in the tokens that make up the context window is the working memory.

And it roughly speaking works kind of like, um, it works for us in our brain. The stuff we remember is our parameters, and the stuff that we just experienced, like a few seconds or minutes ago, and so on, you can imagine that being in our context window.

And this context window is being built up as you have a conscious experience around you. So this has a bunch of, um, implications also for your use of LLMs in practice.

So for example, I can go to ChatGPT and I can do something like this: I can say, "Can you summarize chapter one of Jane Austen's Pride and Prejudice?" Right? And this is a perfectly fine prompt, and ChatGPT actually does something relatively reasonable here.

And the reason it does that is because ChatGPT has a pretty good recollection of a famous work like Pride and Prejudice. It's probably seen a ton of stuff about it. There's probably forums about this book.

It's probably read versions of this book, um, and it's kind of like remembers because even if you've read this or articles about it, you'd kind of have a recollection enough to actually say all this.

But usually when I actually interact with LLMs and I want them to recall specific things, it always works better if you just give it to them.

So I think a much better prompt would be something like this: "Can you summarize for me chapter one of Jane Austen's Pride and Prejudice?" And then I am attaching it below for your reference.

And then I do something like a delimiter here, and I paste it in, and I found that just copy-pasting it from some website that I found here.

Um, so copy-pasting the chapter one here, and I do that because when it's in the context window, the model has direct access to it and can exactly, it doesn't have to recall it; it just has access to it.

And so this summary can be expected to be a significantly high quality or higher quality than this summary, uh, just because it's directly available to the model.

And I think you and I would work in the same way. If you want to, it would be, you would produce a much better summary if you had reread this chapter before you had to summarize it, and that's basically what's happening here, or the equivalent of it.

The next sort of psychological quirk I'd like to talk about briefly is that of the knowledge of self. So what I see very often on the internet is that people do something like this: they ask LLMs something like, "What model are you and who built you?"

And, um, basically this, uh, question is a little bit nonsensical. And the reason I say that is that, as I try to kind of explain with some of the underhood fundamentals, this thing is not a person, right?

It doesn't have a persistent existence in any way. It sort of boots up, processes tokens, and shuts off, and it does that for every single person. It just kind of builds up a context window of conversation, and then everything gets deleted.

And so this entity is kind of like restarted from scratch every single conversation, if that makes sense. It has no persistent self; it has no sense of self. It's a token tumbler, and, uh, it follows the statistical regularities of its training set.

So it doesn't really make sense to ask it, "Who are you? What built you?" etc. And by default, if you do what I described and just by default and from nowhere, you're going to get some pretty random answers.

So for example, let's, uh, pick on Falcon, which is a fairly old model, and let's see what it tells us. Uh, so it's evading the question. Uh, talented engineers and developers here, it says, "I was built by OpenAI based on the GPT-3 model."

It's totally making stuff up now. The fact that it's built by OpenAI here, I think a lot of people would take this as evidence that this model was somehow trained on OpenAI data or something like that.

I don't actually think that that's necessarily true. The reason for that is that if you don't explicitly program the model to answer these kinds of questions, then what you're going to get is its statistical best guess at the answer.

And this model had a, um, SFT data mixture of conversations, and during the fine-tuning, um, the model sort of understands as it's training on this data that it's taking on this personality of this like helpful assistant.

And it doesn't know how to, it doesn't actually, it wasn't told exactly what label to apply to self. It just kind of is taking on this, uh, this, uh, persona of a helpful assistant.

And remember that the pre-training stage took the documents from the entire internet, and ChatGPT and OpenAI are very prominent in these documents.

And so I think what's actually likely to be happening here is that this is just its hallucinated label for what it is. This is its self-identity, is that it's ChatGPT by OpenAI, and it's only saying that because there's a ton of data on the internet of, um, answers like this that are actually coming from OpenAI and ChatGPT.

And so that's its label for what it is. Now, you can override this as a developer. If you have a LLM model, you can actually override it, and there are a few ways to do that.

So for example, let me show you there's this ALMO model from Allen AI, and, um, this is one LLM. It's not a top-tier LLM or anything like that, but I like it because it is fully open source.

So the paper for ALMO and everything else is completely fully open source, which is nice. Um, so here we are looking at its SFT mixture. So this is the data mixture of, um, the fine-tuning, so this is the conversations data, right?

And so the way that they are solving it for the ALMO model is we see that there's a bunch of stuff in the mixture, and there's a total of 1 million conversations here.

But here we have a lot of hardcoded. If we go there, we see that this is 240 conversations, and look at these 240 conversations. They're hardcoded.

"Tell me about yourself," says user, and then the assistant says, "I'm an open language model developed by the Allen Institute of Artificial Intelligence, etc. I'm here to help," blah, blah, blah.

"What is your name?" "ALMO project." So these are all kinds of like cooked-up hardcoded questions about, uh, to and the correct answers to give in these cases.

If you take 240 questions like this or conversations, put them into your training set and fine-tune with it, then the model will actually be expected to parrot this stuff later.

If you don't give it this, then it's probably a ChatGPT by OpenAI. And, um, there's one more way to sometimes do this is that basically, um, in these conversations, and you have terms between human and assistant, sometimes there's a special message called a system message at the very beginning of the conversation.

So it's not just between human and assistant; there's a system. And in the system message, you can actually hardcode and remind the model that, "Hey, you are a model developed by OpenAI, and your name is ChatGPT-4.0, and you were trained on this date, and your knowledge cut-off is this."

And basically, it kind of like documents the model a little bit, and then this is inserted into your conversations. So when you go on ChatGPT, you see a blank page, but actually, the system message is kind of like hidden in there, and those tokens are in the context window.

And so those are the two ways to kind of, um, program the models to talk about themselves. Either it's done through, uh, data like this, or it's done through system messages and things like that.

Basically, invisible tokens that are in the context window and remind the model of its identity. But it's all just kind of like cooked up and bolted on in some way; it's not actually like really deeply there in any real sense as it would be for a human.

I want to now continue to the next section, which deals with the computational capabilities, or like I should say the native computational capabilities of these models in problem-solving scenarios.

And so in particular, we have to be very careful with these models when we construct our examples of conversations, and there's a lot of sharp edges here that are kind of like elucidative. Is that a word?

Uh, they're kind of like interesting to look at when we consider how these models think. So, um, consider the following prompt from a human, and suppose that basically we are building out a conversation to enter into our training set of conversations.

So we're going to train the model on this. We're teaching you how to basically solve simple math problems. So the prompt is, "Emily buys three apples and two oranges. Each orange costs $2. The total cost is $13. What is the cost of apples?" Very simple math question.

Now, there are two answers here on the left and on the right. They are both correct answers. They both say that the answer is three, which is correct.

But one of these two is a significantly better answer for the assistant than the other. Like if I was a data labeler and I was creating one of these, one of these would be, uh, a really terrible answer for the assistant, and the other would be okay.

And so I'd like you to potentially pause the video even and think through why one of these two is a significantly better answer than the other.

And, um, if you use the wrong one, your model will actually be, uh, really bad at math potentially, and it would have, uh, bad outcomes. And this is something that you would be careful with in your life labeling documentations when you are training people, uh, to create the ideal responses for the assistant.

Okay, so the key to this question is to realize and remember that when the models are training and also inferencing, they are working in a one-dimensional sequence of tokens from left to right.

And this is the picture that I often have in my mind. I imagine basically the token sequence evolving from left to right, and to always produce the next token in a sequence, we are feeding all these tokens into the neural network.

And this neural network then is the probabilities for the next token in the sequence, right? So this picture here is the exact same picture we saw, uh, before up here, and this comes from the web demo that I showed you before, right?

So this is the calculation that basically takes the input tokens here on the top and, uh, performs these operations of all these neurons and, uh, gives you the answer for the probabilities of what comes next.

Now, the important thing to realize is that roughly speaking, uh, there's basically a finite number of layers of computation that happened here. So for example, this model here has only one, two, three layers of what's called attention and, uh, MLP here.

Um, maybe, um, a typical modern state-of-the-art network would have more like, say, 100 layers or something like that. But there's only 100 layers of computation or something like that to go from the previous token sequence to the probabilities for the next token.

And so there's a finite amount of computation that happens here for every single token, and you should think of this as a very small amount of computation.

And this amount of computation is almost roughly fixed, uh, for every single token in this sequence. Um, that's not actually fully true because the more tokens you feed in, uh, the more expensive, uh, this forward pass will be of this neural network, but not by much.

So you should think of this, uh, and I think as a good model to have in mind, this is a fixed amount of compute that's going to happen in this box for every single one of these tokens.

And this amount of compute can possibly be too big because there's not that many layers that are sort of going from the top to bottom here. There's not that much computationally that will happen here.

And so you can't imagine the model to basically do arbitrary computation in a single forward pass to get a single token. And so what that means is that we actually have to distribute our reasoning and our computation across many tokens.

Because every single token is only spending a finite amount of computation on it, and so we kind of want to distribute the computation across many tokens.

And we can't have too much computation or expect too much computation out of the model in any single individual token because there's only so much computation that happens per token.

Okay, roughly fixed amount of computation here. So that's why this answer here is significantly worse. And the reason for that is imagine going from left to right here, um, and I copy-pasted it right here.

The answer is three, etc. Imagine the model having to go from left to right, emitting these tokens one at a time. It has to say, or we're expecting to say, "The answer is space dollar sign," and then right here we're expecting it to basically cram all of the computation of this problem into this single token.

It has to emit the correct answer, three. And then once we've emitted the answer three, we're expecting it to say all these tokens, but at this point, we've already produced the answer, and it's already in the context window for all these tokens that follow.

So anything here is just, um, kind of post-hoc justification of why this is the answer, um, because the answer is already created. It's already in the token window.

So it's not actually being calculated here. Um, and so if you are answering the question directly and immediately, you are training the model to try to basically guess the answer in a single token, and that is just not going to work because of the finite amount of computation that happens per token.

That's why this answer on the right is significantly better because we are distributing this computation across the answer. We're actually getting the model to sort of slowly come to the answer from the left to right.

We're getting intermediate results. We're saying, "Okay, the total cost of oranges is four, so 30 - 4 is 9," and so we're creating intermediate calculations.

And each one of these calculations is by itself not that expensive, and so we're actually basically kind of guessing a little bit the difficulty that the model is capable of in any single one of these individual tokens.

And there can never be too much work in any one of these tokens computationally because then the model won't be able to do that later at test time.

And so we're teaching the model here to spread out its reasoning and to spread out its computation over the tokens, and in this way, it only has very simple problems in each token, and they can add up.

And then by the time it's near the end, it has all the previous results in its working memory, and it's much easier for it to determine that the answer is, and here it is, three.

So this is a significantly better label for our computation. This would be really bad and is teaching the model to try to do all the computation in a single token, and it's really bad.

So, uh, that's kind of like an interesting thing to keep in mind is in your prompts. Uh, usually, you don't have to think about it explicitly because, uh, the people at OpenAI have labelers and so on that actually worry about this, and they make sure that the answers are spread out.

And so actually OpenAI will kind of like do the right thing. So when I ask this question for ChatGPT, it's actually going to go very slowly. It's going to be like, "Okay, let's define our variables, set up the equation," and it's kind of creating all these intermediate results.

These are not for you; these are for the model. If the model is not creating these intermediate results for itself, it's not going to be able to reach three.

I also wanted to show you that it's possible to be a bit mean to the model. Uh, we can just ask for things. So as an example, I said, I gave it the exact same, uh, prompt, and I said, "Answer the question in a single token. Just immediately give me the answer, nothing else."

And it turns out that for this simple, um, prompt here, it actually was able to do it in a single go. So it just created a single, I think this is two tokens, right?

Because the dollar sign is its own token. So basically, this model didn't give me a single token; it gave me two tokens, but it still produced the correct answer, and it did that in a single forward pass of the network.

Now that's because the numbers here, I think, are very simple. And so I made it a bit more difficult to be a bit mean to the model.

So I said, "Emily buys 23 apples and 177 oranges," and then I just made the numbers a bit bigger, and I'm just making it harder for the model. I'm asking it to do more computation in a single token.

And so I said the same thing, and here it gave me five, and five is actually not correct. So the model failed to do all of this calculation in a single forward pass of the network.

It failed to go from the input tokens and then in a single forward pass of the network, a single go through the network, it couldn't produce the result.

And then I said, "Okay, now don't worry about the token limit and just solve the problem as usual." And then it goes all the intermediate results.

It simplifies, and every one of these intermediate results here and intermediate calculations is much easier for the model, and, um, it sort of, it's not too much work per token.

All of the tokens here are correct, and it arrives at the solution, which is seven. And I just couldn't squeeze all of this work; it couldn't squeeze that into a single forward pass of the network.

So I think that's kind of just a cute example and something to kind of like think about, and I think it's kind of, again, just elucidative in terms of how these, uh, models work.

The last thing that I would say on this topic is that if I was in practice trying to actually solve this in my day-to-day life, I might actually not, uh, trust that the model, that all the intermediate calculations are correct here.

So actually, probably what I do is something like this: I would come here and I would say, "Use code." And, uh, that's because code is one of the possible tools that ChatGPT can use.

And instead of it having to do mental arithmetic like this, mental arithmetic here, I don't fully trust it. And especially if the numbers get really big, there's no guarantee that the model will do this correctly.

Any one of these intermediate steps might in principle fail. We're using neural networks to do mental arithmetic, uh, kind of like you doing mental arithmetic in your brain.

It might just like, uh, screw up some of the intermediate results. It's actually kind of amazing that it can even do this kind of mental arithmetic. I don't think I could do this in my head, but basically, the model is kind of like doing it in its head, and I don't trust that.

So I wanted to use tools. So you can say stuff like, "Use code," and, uh, I'm not sure what happened there. "Use code."

And so, um, like I mentioned, there's a special tool, and the, uh, model can write code, and I can inspect that this code is correct.

And then, uh, it's not relying on its mental arithmetic; it is using the Python interpreter, which is a very simple programming language to basically, uh, write out the code that calculates the result.

And I would personally trust this a lot more because this came out of a Python program, which I think has a lot more correctness guarantees than the mental arithmetic of a language model.

Uh, so just, um, another kind of, uh, potential hint that if you have these kinds of problems, uh, you may want to basically just, uh, ask the model to use the code interpreter.

And just like we saw with the web search, the model has special, uh, kind of tokens for calling, uh, like it will not actually generate these tokens from the language model.

It will write the program, and then it actually sends that program to a different sort of part of the computer that actually just runs that program and brings back the result.

And then the model gets access to that result and can tell you that, okay, the cost of each apple is seven. Um, so that's another kind of tool, and I would use this in practice for yourself, and it's, um, yeah, it's just, uh, less error-prone, I would say.

So that's why I called this section "Models Need Tokens to Think." Distribute your computation across many tokens. Ask models to create intermediate results, or whenever you can, lean on tools and tool use instead of allowing the models to do all of the stuff in their memory.

So if they try to do it all in their memory, I don't fully trust it and prefer to use tools whenever possible.

I want to show you one more example of where this actually comes up, and that's in counting. So models actually are not very good at counting for the exact same reason.

You're asking for way too much in a single individual token. So let me show you a simple example of that. Um, how many dots are below? And then I just put in a bunch of dots, and ChatGPT says there are, and then it just tries to solve the problem in a single token.

So in a single token, it has to count the number of dots in its context window, um, and it has to do that in the single forward pass of a network.

And a single forward pass of a network, as we talked about, there's not that much computation that can happen there. Just think of that as being like very little computation that happens there.

So if I just look at what the model sees, let's go to the LM, go to tokenizer. It sees, uh, this: "How many dots are below?" And then it turns out that these dots here, this group of I think 20 dots, is a single token.

And then this group of whatever it is is another token, and then for some reason, they break up as this. So I don't actually, this has to do with the details of the tokenizer, but it turns out that these, um, the model basically sees the token ID, this, this, this, and so on.

And then from these token IDs, it's expected to count the number, and spoiler alert, it's not 161; it's actually, I believe, 177.

So here's what we can do instead. Uh, we can say, "Use code." And you might expect that, like, why should this work? And it's actually kind of subtle and kind of interesting.

So when I say, "Use code," I actually expect this to work. Let's see. Okay, 177 is correct.

So what happens here is I've actually, it doesn't look like it, but I've broken down the problem into problems that are easier for the model. I know that the model can't count; it can't do mental counting.

But I know that the model is actually pretty good at doing copy-pasting. So what I'm doing here is when I say, "Use code," it creates a string in Python for this.

And the task of basically copy-pasting my input here to here is very simple because for the model, um, it sees this string of, uh, it sees it as just these four tokens or whatever it is.

So it's very simple for the model to copy-paste those token IDs and, um, kind of unpack them into dots here. And so it creates this string, and then it calls the Python routine, count, and then it comes up with the correct answer.

So the Python interpreter is doing the counting; it's not the model's mental arithmetic doing the counting. So it's again a simple example of, um, models need tokens to think.

Don't rely on their mental arithmetic, and, um, that's why also the models are not very good at counting. If you need them to do counting tasks, always ask them to lean on the tool.

Now, the models also have many other little cognitive deficits here and there, and these are kind of like sharp edges of the technology to be kind of aware of over time.

So as an example, the models are not very good with all kinds of spelling-related tasks. They're not very good at it, and I told you that we would loop back around to tokenization.

And the reason to do for this is that the models, they don't see the characters; they see tokens, and their entire world is about tokens, which are these little text chunks.

And so they don't see characters like our eyes do, and so very simple character-level tasks often fail. So for example, uh, I'm giving it a string "ubiquitous," and I'm asking it to print only every third character starting with the first one.

So we start with U, and then we should go every third, so every, so 1, 2, 3, Q should be next, and then, etc. So this I see is not correct.

And again, my hypothesis is that this is again mental arithmetic here is failing, number one, a little bit. But number two, I think the more important issue here is that if you go to the tokenizer and you look at "ubiquitous," we see that it is three tokens, right?

So you and I see "ubiquitous," and we can easily access the individual letters because we kind of see them. And when we have it in the working memory of our visual sort of field, we can really easily index into every third letter, and I can do that task.

But the models don't have access to the individual letters; they see this as these three tokens. And, uh, remember these models are trained from scratch on the internet, and all these tokens, uh, basically the model has to discover how many of all these different letters are packed into all these different tokens.

And the reason we even use tokens is mostly for efficiency, uh, but I think a lot of people are interested to delete tokens entirely. Like we should really have character-level or byte-level models.

It's just that that would create very long sequences, and people don't know how to deal with that right now. So while we have the token world, any kind of spelling tasks are not actually expected to work super well.

So because I know that spelling is not a strong suit because of tokenization, I can again ask it to lean on tools. So I can just say, "Use code," and I would again expect this to work because the task of copy-pasting "ubiquitous" into the Python interpreter is much easier.

And then we're leaning on the Python interpreter to manipulate the characters of this string. So when I say, "Use code," "ubiquitous," yes, it indexes into every third character, and the actual truth is U, Q, U, S, which looks correct to me.

So, um, again, an example of spelling-related tasks not working very well. A very famous example of that recently is, "How many R's are there in strawberry?"

And this went viral many times, and basically the models now get it correct. They say there are three R's in strawberry. But for a very long time, all the state-of-the-art models would insist that there are only two R's in strawberry.

And this caused a lot of, you know, ruckus because, is that a word? I think so. Because, um, it just kind of like, why are the models so brilliant and they can solve math Olympiad questions, but they can't like count R's in strawberry?

And the answer for that, again, is I've built up to it kind of slowly. But number one, the models don't see characters; they see tokens.

And number two, they are not very good at counting. And so here we are combining the difficulty of seeing the characters with the difficulty of counting, and that's why the models struggled with this.

Even though I think by now, honestly, I think OpenAI may have hardcoded the answer here, or I'm not sure what they did, but, um, but this specific query now works.

So models are not very good at spelling, and there are a bunch of other little sharp edges, and I don't want to go into all of them. I just want to show you a few examples of things to be aware of, and, uh, when you're using these models in practice.

I don't actually want to have a comprehensive analysis here of all the ways that the models are kind of like falling short. I just want to make the point that there are some jagged edges here and there, and we've discussed a few of them, and a few of them make sense.

But some of them also will just not make as much sense, and they're kind of like you're left scratching your head even if you understand in-depth how these models work.

And a good example of that recently is the following: uh, the models are not very good at very simple questions like this, and, uh, this is shocking to a lot of people because these math, uh, these problems can solve complex math problems.

They can answer PhD-grade physics, chemistry, biology questions much better than I can, but sometimes they fall short in like super simple problems like this.

So here we go: 9.11 is bigger than 9.9, and it justifies it in some way, but obviously, and then at the end, okay, it actually flips its decision later.

So I don't believe that this is very reproducible. Sometimes it flips around its answer; sometimes it gets it right; sometimes it gets it wrong.

Uh, let's try again. Okay, even though it might look larger. Okay, so here it doesn't even correct itself in the end. If you ask many times, sometimes it gets it right too.

But how is it that the model can do so great at Olympiad-grade problems but then fail on very simple problems like this?

And, uh, I think this one is, as I mentioned, a little bit of a head-scratcher. It turns out that a bunch of people studied this in depth, and I haven't actually read the paper.

But what I was told by this team was that when you scrutinize the activations inside the neural network, when you look at some of the features and what features turn on or off and what neurons turn on or off, uh, a bunch of neurons inside the neural network light up that are usually associated with Bible verses.

And so I think the model is kind of like reminded that these almost look like Bible verse markers, and in a Bible verse setting, 9.11 would come after 9.9.

And so basically, the model somehow finds it like cognitively very distracting that in Bible verses, 9.11 would be greater, um, even though here it's actually trying to justify it and come up to the answer with math.

It still ends up with the wrong answer here, so it basically just doesn't fully make sense, and it's not fully understood.

And, um, there's a few jagged issues like that. So that's why treat this as what it is, which is a stochastic system that is really magical, but that you can't also fully trust, and you want to use it as a tool, not as something that you kind of like let it rip on a problem and copy-paste the results.

Okay, so we have now covered two major stages of training of large language models. We saw that in the first stage, this is called the pre-training stage.

We are basically training on internet documents, and when you train a language model on internet documents, you get what's called a base model, and it's basically an internet document simulator.

Right now, we saw that this is an interesting artifact, and, uh, this takes many months to train on thousands of computers, and it's kind of a lossy compression of the internet, and it's extremely interesting, but it's not directly useful because we don't want to sample internet documents.

We want to ask questions of an AI and have it respond to our questions, so for that, we need an assistant. And we saw that we can actually construct an assistant in the process of a post-training, and specifically in the process of supervised fine-tuning, as we call it.

So in this stage, we saw that it's algorithmically identical to pre-training. Nothing is going to change; the only thing that changes is the data set.

So instead of internet documents, we now want to create and curate a very nice data set of conversations. So we want millions of conversations on all kinds of diverse topics between a human and an assistant.

And fundamentally, these conversations are created by humans. So humans write the prompts, and humans write the ideal responses, and they do that based on labeling documentations.

Now, in the modern stack, it's not actually done fully and manually by humans, right? They actually now have a lot of help from these tools.

So we can use language models, um, to help us create these data sets, and that's done extensively. But fundamentally, it's all still coming from human curation at the end.

So we create these conversations that now become our data set. We fine-tune on it or continue training on it, and we get an assistant.

And then we kind of shifted gears and started talking about some of the kind of cognitive implications of what this assistant is like.

And we saw that, for example, the assistant will hallucinate if you don't take some sort of mitigations towards it. So we saw that hallucinations would be common, and then we looked at some of the mitigations of those hallucinations.

And then we saw that the models are quite impressive and can do a lot of stuff in their head, but we saw that they can also lean on tools to become better.

So for example, we can lean on a web search in order to hallucinate less and to maybe bring up some more, um, recent information or something like that, or we can lean on tools like code interpreter so the code can, so the LLM can write some code and actually run it and see the results.

So these are some of the topics we looked at so far. Um, now what I'd like to do is I'd like to cover the last and major stage of this pipeline, and that is reinforcement learning.

So reinforcement learning is still kind of thought to be under the umbrella of post-training, uh, but it is the last third major stage, and it's a different way of training language models and usually follows as this third step.

So inside companies like OpenAI, you will start here, and these are all separate teams. So there's a team doing data for pre-training and a team doing training for pre-training, and then there's a team doing all the conversation generation in a different team that is kind of doing the supervised fine-tuning.

And there will be a team for the reinforcement learning as well, so it's kind of like a handoff of these models. You get your base model, then you find you need to be an assistant, and then you go into reinforcement learning, which we'll talk about, uh, now.

So that's kind of like the major flow. And so let's now focus on reinforcement learning, the last major stage of training.

And let me first actually motivate it and why we would want to do reinforcement learning and what it looks like on a high level.

So I would now like to try to motivate the reinforcement learning stage and what it corresponds to with something that you're probably familiar with, and that is basically going to school.

So just like you went to school to become, um, really good at something, we want to take large language models through school.

And really what we're doing is, um, we have a few paradigms of ways of, uh, giving them knowledge or transferring skills.

So in particular, when we're working with textbooks in school, you'll see that there are three major kind of, uh, pieces of information in these textbooks, three classes of information.

The first thing you'll see is you'll see a lot of exposition. Um, and by the way, this is a totally random book I pulled from the internet. I think it's some kind of an organic chemistry or something; I'm not sure.

But the important thing is that you'll see that most of the text, most of it is kind of just like the meat of it is exposition. It's kind of like background knowledge, etc.

As you are reading through the words of this exposition, you can think of that roughly as training on that data. So, um, and that's why when you're reading through this stuff, this background knowledge and this all this context information, it's kind of equivalent to pre-training.

So it's where we build sort of like a knowledge base of this data and get a sense of the topic. The next major kind of information that you will see is these, uh, problems and with their worked solutions.

So basically, a human expert in this case, uh, the author of this book has given us not just a problem but has also worked through the solution.

And the solution is basically like equivalent to having like this ideal response for an assistant. So it's basically the expert is showing us how to solve the problem in its, uh, kind of like, um, in its full form.

So as we are reading the solution, we are basically training on the expert data, and then later we can try to imitate the expert, um, and basically, um, that's that roughly corresponds to having the SFT model.

That's what it would be doing. So basically, we've already done pre-training, and we've already covered this, um, imitation of experts and how they solve these problems.

And the third stage of reinforcement learning is basically the practice problems. So sometimes you'll see this is just a single practice problem here, but of course, there will be usually many practice problems at the end of each chapter in any textbook.

And practice problems, of course, we know are critical for learning because what are they getting you to do? They're getting you to practice, uh, to practice yourself and discover ways of solving these problems yourself.

And so what you get in a practice problem is you get a problem description, but you're not given the solution. But you are given the final answer, usually in the answer key of the textbook.

And so you know the final answer that you're trying to get to, and you have the problem statement, but you don't have the solution. You are trying to practice the solution.

You're trying out many different things, and you're seeing what gets you to the final solution the best, and so you're discovering how to solve these problems.

So, and in the process of that, you're relying on, number one, the background information, which comes from pre-training, and number two, maybe a little bit of imitation of human experts, and you can probably try similar kinds of solutions and so on.

So we've done this and this, and now in this section, we're going to try to practice. And so we're going to be given prompts, we're going to be given solutions, uh, sorry, the final answers, but we're not going to be given expert solutions.

We have to practice and try stuff out, and that's what reinforcement learning is about.

Okay, so let's go back to the problem that we worked with previously just so we have a concrete example to talk through as we explore sort of the topic here.

So, um, I'm here in the tokenizer because I'd also like to, well, I get a text box, which is useful, but number two, I want to remind you again that we're always working with one-dimensional token sequences.

And so, um, I actually like prefer this view because this is like the native view of the LLM, if that makes sense. Like this is what it actually sees; it sees token IDs, right?

Okay, so Emily buys three apples and two oranges. Each orange is $2. The total cost of all the fruit is $13. What is the cost of each apple?

And what I'd like you to appreciate here is these are like four possible candidate solutions as an example, and they all reach the answer three.

Now, what I'd like you to appreciate at this point is that if I am the human data labeler that is creating a conversation to be entered into the training set, I don't actually really know which of these conversations to, um, to add to the data set.

Some of these conversations kind of set up a system of equations; some of them sort of like just talk through it in English, and some of them just kind of like skip right through to the solution.

Um, if you look at ChatGPT, for example, and you give it this question, it defines a system of variables, and it kind of like does this little thing.

What we have to appreciate and, uh, differentiate between, though, is, um, the first purpose of a solution is to reach the right answer. Of course, we want to get the final answer three; that is the important purpose here.

But there's kind of like a secondary purpose as well where here we are also just kind of trying to make it like nice, uh, for the human because we're kind of assuming that the person wants to see the solution.

They want to see the intermediate steps; we want to present it nicely, etc. So there are two separate things going on here. Number one is the presentation for the human, but number two, we're trying to actually get the right answer.

Um, so let's for the moment focus on just reaching the final answer. If we're only care, if we only care about the final answer, then which of these is the optimal or the best prompt, um, sorry, the best solution for the LLM to reach the right answer?

Um, and what I'm trying to get at is we don't know. Me as a human labeler, I would not know which one of these is best.

So as an example, we saw earlier on when we looked at, um, the token sequences here and the mental arithmetic and reasoning, we saw that for each token, we can only spend basically a finite number of finite amount of compute here that is not very large or you should think about it that way.

And so we can't actually make too big of a leap in any one token is maybe the way to think about it. So as an example in this one, what's really nice about it is that it's very few tokens, so it's going to take us a very short amount of time to get to the answer.

But right here, when we're doing 30 - 4, I D 3 equals right in this token here, we're actually asking for a lot of computation to happen on that single individual token.

And so maybe this is a bad example to give to the LLM because it's kind of incentivizing it to skip through the calculations very quickly, and it's going to actually make up mistakes, make mistakes in this mental arithmetic.

Uh, so maybe it would work better to like spread out the, spread it out more. Maybe it would be better to set it up as an equation. Maybe it would be better to talk through it.

We fundamentally don't know, and we don't know because what is easy for you or I as human labelers, what's easy for us or hard for us is different than what's easy or hard for the LLM.

Its cognition is different, um, and the token sequences are kind of like different hard for it. And so some of the token sequences here that are trivial for me might be, um, very too much of a leap for the LLM.

So right here, this token would be way too hard, but conversely, many of the tokens that I'm creating here might be just trivial to the LLM, and we're just wasting tokens.

Like why waste all these tokens when this is all trivial? So if the only thing we care about is the final answer, and we're separating out the issue of the presentation to the human, um, then we don't actually really know how to annotate this example.

We don't know what solution to give to the LLM because we are not the LLM, and it's clear here in the case of like the math example.

But this is actually like a very pervasive issue. Like our knowledge is not LLM's knowledge. Like the LLM actually has a ton of knowledge of PhD in math and physics, chemistry, and whatnot.

So in many ways, it actually knows more than I do, and I'm potentially not utilizing that knowledge in its problem-solving. But conversely, I might be injecting a bunch of knowledge in my solutions that the LLM doesn't know in its parameters, and then those are like sudden leaps that are very confusing to the model.

And so our cognitions are different, and I don't really know what to put here if all we care about is the reaching the final solution and doing it economically ideally.

And so long story short, we are not in a good position to create these, uh, token sequences for the LLM, and they're useful by imitation to initialize the system.

But we really want the LLM to discover the token sequences that work for it. We need to find it; it needs to find for itself what token sequence reliably gets to the answer given the prompt, and it needs to discover that in the process of reinforcement learning and of trial and error.

So let's see how this example would work like in reinforcement learning.

Okay, so we're now back in the Hugging Face inference playground, and, uh, that just allows me to very easily call, uh, different kinds of models.

So as an example here on the top right, I chose the Gemma 2, 2 billion parameter model. So two billion is very, very small, so this is a tiny model, but it's okay.

So we're going to give it, um, the way that reinforcement learning will basically work is actually quite, quite simple.

Um, we need to try many different kinds of solutions, and we want to see which solutions work well or not. So we're basically going to take the prompt, we're going to run the model, and the model generates a solution, and then we're going to inspect the solution.

And we know that the correct answer for this one is $3, and so indeed the model gets it correct. It says it's $3, so this is correct.

So that's just one attempt at this solution. So now we're going to delete this, and we're going to rerun it again. Let's try a second attempt.

So the model solves it in a bit slightly different way, right? Every single attempt will be a different generation because these models are stochastic systems.

Remember that at every single token here, we have a probability distribution, and we're sampling from that distribution, so we end up kind of going down slightly different paths.

And so this is a second solution that also ends in the correct answer. Now we're going to delete that. Let's go a third time.

Okay, so again, slightly different solution, but also gets it correct. Now we can actually repeat this, uh, many times.

And so in practice, you might actually sample thousands of independent solutions or even like a million solutions for just a single prompt.

Um, and some of them will be correct, and some of them will not be very correct. And basically what we want to do is we want to encourage the solutions that lead to correct answers.

So let's take a look at what that looks like. So if we come back over here, here's kind of like a cartoon diagram of what this is looking like.

We have a prompt, and then we tried many different solutions in parallel, and some of the solutions, um, might go well, so they get the right answer, which is in green, and some of the solutions might go poorly and may not reach the right answer, which is red.

Now this problem here, unfortunately, is not the best example because it's a trivial prompt, and as we saw, uh, even like a two billion parameter model always gets it right.

So it's not the best example in that sense, but let's just exercise some imagination here and let's just suppose that the, um, green ones are good and the red ones are bad.

Okay, so we generated 15 solutions; only four of them got the right answer. And so now what we want to do is basically we want to encourage the kinds of solutions that lead to right answers.

So whatever token sequences happened in these red solutions, obviously something went wrong along the way somewhere, and, uh, this was not a good path to take through the solution.

And whatever token sequences there were in these green solutions, well, things went, uh, pretty well in this situation.

And so we want to do more things like it in prompts like this. And the way we encourage this kind of behavior in the future is we basically train on these sequences.

Um, but these training sequences now are not coming from expert human annotators. There's no human who decided that this is the correct solution.

This solution came from the model itself. So the model is practicing here; it's tried out a few solutions. Four of them seem to have worked, and now the model will kind of like train on them.

And this corresponds to a student basically looking at their solutions and being like, "Okay, well, this one worked really well, so this is how I should be solving these kinds of problems."

And, uh, here in this example, there are many different ways to actually like really tweak the methodology a little bit here.

But just to give the core idea across, maybe it's simplest to just think about taking the single best solution out of these four, uh, like say this one that's why it was yellow.

So this is the solution that not only led to the right answer but maybe had some other nice properties. Maybe it was the shortest one, or it looked nicest in some ways, or, uh, there's other criteria you could think of as an example.

But we're going to decide that this is the top solution; we're going to train on it, and then, uh, the model will be slightly more likely once you do the parameter update to take this path in this kind of a setting in the future.

But you have to remember that we're going to run many different diverse prompts across lots of math problems and physics problems and whatever, wherever there might be.

So tens of thousands of prompts maybe have in mind; there's thousands of solutions prompt. And so this is all happening kind of like at the same time, and as we're iterating this process, the model is discovering for itself what kinds of token sequences lead it to correct answers.

It's not coming from a human annotator. The model is kind of like playing in this playground, and it knows what it's trying to get to, and it's discovering sequences that work for it.

Uh, these are sequences that don't make any mental leaps; they seem to work reliably and statistically and, uh, fully utilize the knowledge of the model as it has it.

And so, uh, this is the process of reinforcement learning. It's basically a guess and check. We're going to guess many different types of solutions; we're going to check them, and we're going to do more of what worked in the future, and that is, uh, reinforcement learning.

So in the context of what came before, we see now that the SFT model, the supervised fine-tuning model, it's still helpful because it still kind of like initializes the model a little bit into the vicinity of the correct solutions.

So it's kind of like an initialization of, um, of the model in the sense that it kind of gets the model to, you know, take solutions like write out solutions, and maybe it has an understanding of setting up a system of equations, or maybe it kind of like talks through a solution.

So it gets you into the vicinity of correct solutions, but reinforcement

We can go to the web page and see the kinds of problems that are actually in these, um, the kinds of math problems that are being measured here.

So these are simple math problems. You can, um, pause the video if you like, but these are the kinds of problems that basically the models are being asked to solve.

You can see that in the beginning, they're not doing very well, but then as you update the model with this many thousands of steps, their accuracy kind of continues to climb.

So the models are improving, and they're solving these problems with a higher accuracy as you do this trial and error on a large data set of these kinds of problems.

The models are discovering how to solve math problems.

But even more incredible than the quantitative kind of results of solving these problems with a higher accuracy is the qualitative means by which the model achieves these results.

So when we scroll down, uh, one of the figures here that is kind of interesting is that later on in the optimization, the model seems to be, uh, using average length per response, uh, goes up.

So the model seems to be using more tokens to get its higher accuracy results.

So it's learning to create very, very long solutions.

Why are these solutions very long?

We can look at them qualitatively here.

So basically what they discover is that the model solutions get very, very long, partially because—so here's a question and here's kind of the answer from the model.

What the model learns to do, um, and this is an emerging property of new optimization, is it just discovers that this is good for problem solving.

It starts to do stuff like this: "Wait, wait, wait, that's not a moment I can flag here. Let's reevaluate this step by step to identify the correct sum."

So what is the model doing here?

Right, the model is basically re-evaluating steps.

It has learned that it works better for accuracy to try out lots of ideas, try something from different perspectives, retrace, reframe, backtrack.

It's doing a lot of the things that you and I are doing in the process of problem solving for mathematical questions.

But it's rediscovering what happens in your head, not what you put down on the solution.

And there is no human who can hardcode this stuff in the ideal assistant response.

This is only something that can be discovered in the process of reinforcement learning because you wouldn't know what to put here.

This just turns out to work for the model, and it improves its accuracy in problem solving.

So the model learns what we call these chains of thought in your head, and it's an emergent property of the optimization.

And that's what's bloating up the response length, but that's also what's increasing the accuracy of the problem solving.

So what's incredible here is basically the model is discovering ways to think.

It's learning what I like to call cognitive strategies of how you manipulate a problem and how you approach it from different perspectives.

How you pull in some analogies or do different kinds of things like that, and how you kind of, uh, try out many different things over time.

Check a result from different perspectives and how you kind of, uh, solve problems.

But here it's kind of discovered by the RL.

So extremely incredible to see this emerge in the optimization without having to hardcode it anywhere.

The only thing we've given it are the correct answers, and this comes out from trying to just solve them correctly, which is incredible.

Now let's go back to actually the problem that we've been working with and let's take a look at what it would look like for this kind of a model, what we call reasoning or thinking model, to solve that problem.

Okay, so recall that this is the problem we've been working with, and when I pasted it into ChatGPT 4.0, I'm getting this kind of a response.

Let's take a look at what happens when you give this same query to what's called a reasoning or a thinking model.

This is a model that was trained with reinforcement learning.

So this model described in this paper, DC car1, is available on chat.dec.com.

Uh, so this is kind of like the company that developed it, hosting it.

You have to make sure that the Deep Think button is turned on to get the R1 model, as it's called.

We can paste it here and run it, and so let's take a look at what happens now and what is the output of the model.

Okay, so here it says—so this is previously what we get using basically what's an SFT approach, a supervised fine-tuning approach.

This is like mimicking an expert solution.

This is what we get from the RL model.

Okay, let me try to figure this out.

So Emily buys three apples and two oranges. Each orange costs $2. Total is $13.

I need to find out, blah blah blah.

So here you, um, as you're reading this, you can't escape thinking that this model is thinking.

Um, it's definitely pursuing the solution.

It derives that it must cost $3, and then it says, "Wait a second, let me check my math again to be sure."

And then it tries it from a slightly different perspective, and then it says, "Yep, all that checks out. I think that's the answer. I don't see any mistakes."

Let me see if there's another way to approach the problem, maybe setting up an equation.

Let's let the cost of one apple be $8, then blah blah blah.

Yep, same answer.

So definitely each apple is $3.

All right, confident that that's correct.

And then what it does once it sort of, um, did the thinking process is it writes up the nice solution for the human.

And so this is now considering—so this is more about the correctness aspect, and this is more about the presentation aspect, where it kind of like writes it out nicely and, uh, boxes in the correct answer at the bottom.

And so what's incredible about this is we get this like thinking process of the model, and this is what's coming from the reinforcement learning process.

This is what's bloating up the length of the token sequences.

They're doing thinking, and they're trying different ways.

This is what's giving you higher accuracy in problem solving, and this is where we are seeing these aha moments and these different strategies and these, um, ideas for how you can make sure that you're getting the correct answer.

The last point I wanted to make is some people are a little bit nervous about putting, you know, very sensitive data into chat.com because this is a Chinese company.

So people don't, um, people are a little bit careful and shy with that.

A little bit, um, Deep Seek R1 is a model that was released by this company.

So this is an open-source model or open weights model.

It is available for anyone to download and use.

You will not be able to, like, run it in its full, um, sort of the full model in full precision.

You won't run that on a MacBook or like a local device because this is a fairly large model.

But many companies are hosting the full largest model.

One of those companies that I like to use is called Together.

So when you go to Together, you sign up, and you go to playgrounds, you can select here in the chat Deep Seek R1.

And there's many different kinds of other models that you can select here.

These are all state-of-the-art models.

So this is kind of similar to the Hugging Face inference playground that we've been playing with so far.

But Together will usually host all the state-of-the-art models.

So select DC car1.

Um, you can try to ignore a lot of these.

I think the default settings will often be okay, and we can put in this.

And because the model was released by Deep Seek, what you're getting here should be basically equivalent to what you're getting here.

Now, because of the randomness in the sampling, we're going to get something slightly different, but in principle, this should be, uh, identical in terms of the power of the model.

And you should be able to see the same things quantitatively and qualitatively.

But, uh, this model is coming from kind of an American company, so that's Deep Seek, and that's what's called a reasoning model.

Now, when I go back to chat, uh, let me go to chat here.

Okay, so the models that you're going to see in the drop-down here, some of them like 01, 03, mini, O3, mini, high, etc., they are talking about uses advanced reasoning.

Now, what this is referring to, uses advanced reasoning, is it's referring to the fact that it was trained by reinforcement learning with techniques very similar to those of Deep Seek R1, per public statements of OpenAI employees.

Uh, so these are thinking models trained with RL, and these models like GPT-4 or GPT-4 40 mini that you're getting in the free tier, you should think of them as mostly SFT models, supervised fine-tuning models.

They don't actually do this like thinking as you see in the RL models.

And even though there's a little bit of reinforcement learning involved with these models, and I'll go into that in a second, these are mostly SFT models.

I think you should think about it that way.

So in the same way as what we saw here, we can pick one of the thinking models, like say 03 mini high.

And these models, by the way, might not be available to you unless you pay a ChatGPT subscription of either $20 per month or $200 per month for some of the top models.

So we can pick a thinking model and run.

Now what's going to happen here is it's going to say reasoning, and it's going to start to do stuff like this.

And, um, what we're seeing here is not exactly the stuff we're seeing here.

So even though under the hood, the model produces these kinds of, uh, kind of chains of thought, OpenAI chooses to not show the exact chains of thought in the web interface.

It shows little summaries of those chains of thought, and OpenAI kind of does this, I think partly because, uh, they are worried about what's called the distillation risk.

That is, that someone could come in and actually try to imitate those reasoning traces and recover a lot of the reasoning performance by just imitating the reasoning, uh, chains of thought.

And so they kind of hide them, and they only show little summaries of them.

So you're not getting exactly what you would get in Deep Seek as with respect to the reasoning itself.

And then they write up the solution.

So these are kind of like equivalent, even though we're not seeing the full under-the-hood details.

Now, in terms of the performance, uh, these models and Deep Seek models are currently really on par.

I would say it's kind of hard to tell because of the evaluations, but if you're paying $200 per month to OpenAI, some of these models, I believe, are currently—they basically still look better.

But Deep Seek R1 for now is still a very solid choice for a thinking model that would be available to you, um, sort of, um, either on this website or any other website because the model is open weights.

You can just download it.

So that's thinking models.

So what is the summary so far?

Well, we've talked about reinforcement learning and the fact that thinking emerges in the process of the optimization when we basically run RL on many math and kind of code problems that have verifiable solutions.

So there's like an answer, three, etc.

Now these thinking models you can access in, for example, Deep Seek or any inference provider like Together.

And choosing Deep Seek over there, these thinking models are also available in ChatGPT under any of the 01 or 03 models.

But these GPT-4 R models, etc., they're not thinking models.

You should think of them as mostly SFT models.

Now, if you are, um, if you have a prompt that requires advanced reasoning and so on, you should probably use some of the thinking models or at least try them out.

But empirically, for a lot of my use, when you're asking a simpler question, there's like a knowledge-based question or something like that, this might be overkill.

Like there's no need to think 30 seconds about some factual question.

So for that, I will, uh, sometimes default to just GPT-4.

So empirically, about 80-90% of my use is just GPT-4, and when I come across a very difficult problem like in math and code, etc., I will reach for the thinking models.

But then I have to wait a bit longer because they're thinking.

Um, so you can access these on Chat, on Deep Seek.

Also, I wanted to point out that, um, AI Studio.go.com, even though it looks really busy, really ugly because Google's just unable to do this kind of stuff well, it's like, what is happening?

But if you choose model and you choose here Gemini 2.0 flash thinking experimental 01 21, if you choose that one, that's also a kind of early experimental thinking model by Google.

So we can go here, and we can give it the same problem and click run, and this is also a thinking problem, a thinking model that will also do something similar and comes out with the right answer here.

So basically, Gemini also offers a thinking model.

Anthropic currently does not offer a thinking model, but basically this is kind of like the frontier development of these LLMs.

I think RL is kind of like this new exciting stage, but getting the details right is difficult, and that's why all these models and thinking models are currently experimental as of 2025, very early 2025.

But this is kind of like the frontier development of pushing the performance on these very difficult problems using reasoning that is emerging in these optimizations.

One more connection that I wanted to bring up is that the discovery that reinforcement learning is an extremely powerful way of learning is not new to the field of AI.

And one place where we've already seen this demonstrated is in the game of Go.

And famously, DeepMind developed the system AlphaGo, and you can watch a movie about it, um, where the system is learning to play the game of Go against top human players.

And, um, when we go to the paper underlying AlphaGo, so in this paper, when we scroll down, we actually find a really interesting plot, um, that I think, uh, is kind of familiar to us, and we're kind of like rediscovering in the more open domain of arbitrary problem solving instead of on the closed specific domain of the game of Go.

But basically what they saw—and we're going to see this in LLMs as well as this becomes more mature—is this is the ELO rating of playing the game of Go, and this is Lee Sedol, an extremely strong human player.

And here what they are comparing is the strength of a model learned trained by supervised learning and a model trained by reinforcement learning.

So the supervised learning model is imitating human expert players.

So if you just get a huge amount of games played by expert players in the game of Go and you try to imitate them, you are going to get better, but then you top out, and you never quite get better than some of the top, top, top players in the game of Go like Lee Sedol.

So you're never going to reach there because you're just imitating human players.

You can't fundamentally go beyond a human player if you're just imitating human players.

But in a process of reinforcement learning, it is significantly more powerful.

In reinforcement learning for a game of Go, it means that the system is playing moves that empirically and statistically lead to winning the game.

And so AlphaGo is a system where it kind of plays against itself, and it's using reinforcement learning to create rollouts.

So it's the exact same diagram here, but there's no prompt.

It's just, uh, because there's no prompt, it's just a fixed game of Go, but it's trying out lots of solutions.

It's trying out lots of plays, and then the games that lead to a win instead of a specific answer are reinforced.

They're made stronger.

And so, um, the system is learning basically the sequences of actions that empirically and statistically lead to winning the game.

And reinforcement learning is not going to be constrained by human performance, and reinforcement learning can do significantly better and overcome even the top players like Lee Sedol.

And so, uh, probably they could have run this longer, and they just chose to crop it at some point because this costs money.

But this is a very powerful demonstration of reinforcement learning, and we're only starting to kind of see hints of this diagram in larger language models for reasoning problems.

So we're not going to get too far by just imitating experts.

We need to go beyond that, set up these like little game environments, and let the system discover reasoning traces or like ways of solving problems that are unique and that, uh, just basically work well.

Now, on this aspect of uniqueness, notice that when you're doing reinforcement learning, nothing prevents you from veering off the distribution of how humans are playing the game.

And so when we go back to, uh, this AlphaGo search here, one of the suggested modifications is called move 37.

And move 37 in AlphaGo is referring to a specific point in time where AlphaGo basically played a move that, uh, no human expert would play.

So the probability of this move, uh, to be played by a human player was evaluated to be about 1 in 10,000.

So it's a very rare move, but in retrospect, it was a brilliant move.

So AlphaGo, in the process of reinforcement learning, discovered kind of like a strategy of playing that was unknown to humans and, but is in retrospect, uh, brilliant.

I recommend this YouTube video, um, Lee Sedol versus AlphaGo move 37 reactions and analysis.

And this is kind of what it looked like when AlphaGo played this move.

"That's a very surprising move. I thought it was a mistake when I see this move."

Anyway, so basically people are kind of freaking out because it's a move that a human would not play that AlphaGo played because in its training, uh, this move seemed to be a good idea.

It just happens not to be a kind of thing that a human would do.

And so that is again the power of reinforcement learning.

And in principle, we can actually see the equivalence of that if we continue scaling this paradigm in language models.

And what that looks like is kind of unknown.

So, um, what does it mean to solve problems in such a way that, uh, even humans would not be able to?

How can you be better at reasoning or thinking than humans?

How can you go beyond just, uh, a thinking human?

Like maybe it means discovering analogies that humans would not be able to, uh, create.

Or maybe it's like a new thinking strategy.

It's kind of hard to think through.

Uh, maybe it's a wholly new language that actually is not even English.

Maybe it discovers its own language that is a lot better at thinking, um, because the model is unconstrained to even like stick with English.

So maybe it takes a different language to think in, or it discovers its own language.

So in principle, the behavior of the system is a lot less defined.

It is open to do whatever works, and it is open to also slowly drift from the distribution of its training data, which is English.

But all of that can only be done if we have a very large, diverse set of problems in which these strategies can be refined and perfected.

And so that is a lot of the frontier LLM research that's going on right now, is trying to kind of create those kinds of prompt distributions that are large and diverse.

These are all kind of like game environments in which the LLMs can practice their thinking.

And, uh, it's kind of like writing, you know, these practice problems.

We have to create practice problems for all domains of knowledge.

And if we have practice problems and tons of them, the models will be able to reinforcement learn on them and kind of, uh, create these kinds of, uh, diagrams.

But in the domain of open thinking instead of a closed domain like the game of Go, there's one more section within reinforcement learning that I wanted to cover, and that is that of learning in unverifiable domains.

So, so far, all of the problems that we've looked at are in what's called verifiable domains.

That is, any candidate solution we can score very easily against a concrete answer.

So for example, the answer is three, and we can very easily score these solutions against the answer of three.

Either we require the models to like box in their answers, and then we just check for equality of whatever is in the box with the answer, or you can also use, uh, kind of what's called an LLM judge.

So the LLM judge looks at a solution, and it gets the answer and just basically scores the solution for whether it's consistent with the answer or not.

And LLMs, uh, empirically are good enough at the current capability that they can do this fairly reliably.

So we can apply those kinds of techniques as well.

In any case, we have a concrete answer, and we're just checking solutions against it, and we can do this automatically with no kind of humans in the loop.

The problem is that we can't apply the strategy in what's called unverifiable domains.

So usually these are, for example, creative writing tasks like write a joke about pelicans or write a poem or summarize a paragraph or something like that.

In these kinds of domains, it becomes harder to score our different solutions to this problem.

So for example, writing a joke about pelicans, we can generate lots of different, uh, jokes, of course.

That's fine.

For example, we can go to ChatGPT, and we can get it to, uh, generate a joke about pelicans.

"So much stuff in their beaks because they don't bellan in backpacks."

What?

Okay, we can, uh, try something else.

"Why don't pelicans ever pay for their drinks? Because they always B it to someone else."

Haha, okay.

So these models are not obviously not very good at humor.

Actually, I think it's pretty fascinating because I think humor is secretly very difficult, and the model has the capability, I think.

Anyway, in any case, you could imagine creating lots of jokes.

The problem that we are facing is how do we score them?

Now, in principle, we could, of course, get a human to look at all these jokes just like I did right now.

The problem with that is if you are doing reinforcement learning, you're going to be doing many thousands of updates.

And for each update, you want to be looking at, say, thousands of prompts.

And for each prompt, you want to be potentially looking at hundreds or thousands of different kinds of generations.

And so there's just like way too many of these to look at.

And so, um, in principle, you could have a human inspect all of them and score them and decide that, okay, maybe this one is funny, and, uh, maybe this one is funny, and this one is funny.

And we could train on them to get the model to become slightly better at jokes, um, in the context of pelicans at least.

Um, the problem is that it's just like way too much human time.

This is an unscalable strategy.

We need some kind of an automatic strategy for doing this.

And one sort of solution to this was proposed in this paper that introduced what's called reinforcement learning from human feedback.

And so this was a paper from OpenAI at the time, and many of these people are now, um, co-founders in Anthropic.

Um, and this kind of proposed an approach for, uh, basically doing reinforcement learning in unverifiable domains.

So let's take a look at how that works.

So this is the cartoon diagram of the core ideas involved.

So as I mentioned, the native approach is if we just set infinity human time, we could just run RL in these domains just fine.

So for example, we can run RL as usual.

If I have infinity humans, I would just want to do—and these are just cartoon numbers—I want to do 1,000 updates where each update will be on 1,000 prompts.

And for each prompt, we're going to have 1,000 rollouts that we're scoring.

So we can run RL with this kind of a setup.

The problem is in the process of doing this, I will need to run one—I will need to ask a human to evaluate a joke a total of 1 billion times.

And so that's a lot of people looking at really terrible jokes, so we don't want to do that.

So instead, we want to take the RLHF approach.

So, um, in our RLHF approach, we are kind of like—the core trick is that of indirection.

So we're going to involve humans just a little bit, and the way we cheat is that we basically train a whole separate neural network that we call a reward model.

And this neural network will kind of like imitate human scores.

So we're going to ask humans to score, um, roll—we're going to then imitate human scores using a neural network, and this neural network will become a kind of simulator of human preferences.

And now that we have a neural network simulator, we can do RL against it.

So instead of asking a real human, we're asking a simulated human for their score of a joke as an example.

And so once we have a simulator, we're often racing because we can query it as many times as we want to, and it's all a whole automatic process.

And we can now do reinforcement learning with respect to the simulator.

And the simulator, as you might expect, is not going to be a perfect human, but if it's at least statistically similar to human judgment, then you might expect that this will do something.

And in practice, indeed, it does.

So once we have a simulator, we can do RL, and everything works great.

So let me show you a cartoon diagram a little bit of what this process looks like, although the details are not 100% like super important.

It's just a core idea of how this works.

So here I have a cartoon diagram of a hypothetical example of what training the reward model would look like.

So we have a prompt like "write a joke about pelicans," and then here we have five separate rollouts.

So these are all five different jokes, just like this one.

Now, the first thing we're going to do is we are going to ask a human to, uh, order these jokes from the best to worst.

So this is, uh, so here this human thought that this joke is the best, the funniest.

So number one joke, this is number two joke, number three joke, four, and five.

So this is the worst joke.

We're asking humans to order instead of give scores directly because it's a bit of an easier task.

It's easier for a human to give an ordering than to give precise scores.

Now that is now the supervision for the model.

So the human has ordered them, and that is kind of like their contribution to the training process.

But now separately, what we're going to do is we're going to ask a reward model about its scoring of these jokes.

Now the reward model is a whole separate neural network, completely separate neural net, um, and it's also probably a transformer, but it's not a language model in the sense that it generates diverse language, etc.

It's just a scoring model.

So the reward model will take as an input the prompt number one and number two, a candidate joke.

So, um, those are the two inputs that go into the reward model.

So here, for example, the reward model would be taking this prompt and this joke.

Now the output of a reward model is a single number, and this number is thought of as a score.

And it can range, for example, from zero to one.

So zero would be the worst score, and one would be the best score.

So here are some examples of what a hypothetical reward model at some stage in the training process would give, uh, scoring to these jokes.

So 0.1 is a very low score, 0.8 is a really high score, and so on.

And so now, um, we compare the scores given by the reward model with, uh, the ordering given by the human.

And there's a precise mathematical way to actually calculate this, uh, basically set up a loss function and calculate a kind of like a correspondence here and, uh, update a model based on it.

But I just want to give you the intuition, which is that as an example here for this second joke, the human thought that it was the funniest, and the model kind of agreed, right?

0.8 is a relatively high score, but this score should have been even higher, right?

So after an update, we would expect that maybe this score should have been—will actually grow after an update of the network to be like, say, 0.81 or something.

Um, for this one here, they actually are in a massive disagreement because the human thought that this was number two, but here the score is only 0.1.

And so this score needs to be much higher.

So after an update on top of this, um, kind of a supervision, this might grow a lot more, like maybe it's 0.15 or something like that.

Um, and then here the human thought that this one was the worst joke, but here the model actually gave it a fairly high number.

So you might expect that after the update, uh, this would come down to maybe 0.3 or 0.5 or something like that.

So basically, we're doing what we did before.

We're slightly nudging the predictions from the models using a neural network training process, and we're trying to make the reward model scores be consistent with human ordering.

And so, um, as we update the reward model on human data, it becomes better and better simulator of the scores and orders, uh, that humans provide, and then becomes kind of like the neural simulator of human preferences, which we can then do RL against.

But critically, we're not asking humans one billion times to look at a joke.

We're maybe looking at 5,000 prompts and five rollouts each.

So maybe 5,000 jokes that humans have to look at in total, and they just give the ordering, and then we're training the model to be consistent with that ordering.

And I'm skipping over the mathematical details, but I just want you to understand a high-level idea that, uh, this reward model is basically giving us this score, and we have a way of training it to be consistent with human orderings, and that's how RLHF works.

Okay, so that is the rough idea.

We basically train simulators of humans and RL with respect to those simulators.

Now I want to talk about first the upside of reinforcement learning from human feedback.

The first thing is that this allows us to run reinforcement learning, which we know is an incredibly powerful kind of set of techniques, and it allows us to do it in arbitrary domains, including the ones that are unverifiable.

So things like summarization and poem writing, joke writing, or any other creative writing, really, in domains outside of math and code, etc.

Now empirically, what we see when we actually apply RLHF is that this is a way to improve the performance of the model.

And, uh, I have a top answer for why that might be, but I don't actually know that it is like super well established on like why this is.

You can empirically observe that when you do RLHF correctly, the models you get are just like a little bit better.

Um, but as to why is, I think, like not as clear.

So here's my best guess.

My best guess is that this is possibly mostly due to the discriminator-generator gap.

What that means is that in many cases, it is significantly easier to discriminate than to generate for humans.

So in particular, an example of this is, um, when we do supervised fine-tuning, right, SFT, we're asking humans to generate the ideal assistant response.

And in many cases here, um, as I've shown, the ideal response is very simple to write, but in many cases, it might not be so.

For example, in summarization or poem writing or joke writing, like how are you as a human labeler, um, supposed to give the ideal response?

In these cases, it requires creative human writing to do that.

And so RLHF kind of sidesteps this because we get, um, we get to ask people a significantly easier question.

As data labelers, they're not asked to write poems directly.

They're just given five poems from the model, and they're just asked to order them.

And so that's just a much easier task for a human labeler to do.

And so what I think this allows you to do basically is it, um, it kind of like allows a lot more higher accuracy data because we're not asking people to do the generation task, which can be extremely difficult.

Like we're not asking them to do creative writing.

We're just trying to get them to distinguish between creative writings and, uh, find the ones that are best.

And that is the signal that humans are providing, just the ordering, and that is their input into the system.

And then the system in RLHF just discovers the kinds of responses that would be graded well by humans.

And so that step of indirection allows the models to become a bit better.

So that is the upside of RLHF.

It allows us to run RL, it empirically results in better models, and it allows, uh, people to contribute their supervision, uh, even without having to do extremely difficult tasks, um, in the case of writing ideal responses.

Unfortunately, RLHF also comes with significant downsides.

And so, um, the main one is that basically we are doing reinforcement learning not with respect to humans and actual human judgment, but with respect to a lossy simulation of humans, right?

And this lossy simulation could be misleading because it's just a—it's just a simulation, right?

It's just a language model that's kind of outputting scores, and it might not perfectly reflect the opinion of an actual human with an actual brain in all the possible different cases.

So that's number one, which is actually something even more subtle and devious going on that, uh, really dramatically holds back RLHF as a technique that we can really scale to significantly, um, kind of smart systems.

And that is that reinforcement learning is extremely good at discovering a way to game the model, to game the simulation.

So this reward model that we're constructing here that gives the scores—these models are transformers.

These transformers are massive neural networks.

They have billions of parameters, and they imitate humans, but they do so in a kind of like a simulation way.

Now the problem is that these are massive complicated systems, right?

There's a billion parameters here that are outputting a single score.

It turns out that there are ways to game these models.

You can find kinds of inputs that were not part of their training set, and these inputs inexplicably get very high scores, but in a fake way.

So very often what you find if you run RLHF for very long—so for example, if we do 1,000 updates, which is like say a lot of updates—you might expect that your jokes are getting better and that you're getting like real bangers about pelicans.

But that's not exactly what happens.

What happens is that, uh, in the first few hundred steps, the jokes about pelicans are probably improving a little bit, and then they actually dramatically fall off a cliff, and you start to get extremely nonsensical results.

Like for example, you start to get, um, the top joke about pelicans starts to be "the."

And this makes no sense, right?

Like when you look at it, why should this be a top joke?

But when you take "the" and you plug it into your reward model, you'd expect a score of zero, but actually the reward model loves this as a joke.

It will tell you that "the" is a score of 1.0.

This is a top joke, and this makes no sense, right?

But it's because these models are just simulations of humans, and they're massive neural networks, and you can find inputs at the bottom that kind of like get into the part of the input space that kind of gives you nonsensical results at the top.

Now, here's what you might imagine doing.

You say, "Okay, the 'the' is obviously not a score of one.

Um, it's obviously a low score.

So let's take the 'the' and let's add it to the data set and give it an ordering that is extremely bad, like a score of five."

And indeed, your model will learn that the 'the' should have a very low score, and it will give it a score of zero.

The problem is that there will always be basically an infinite number of nonsensical adversarial examples hiding in the model.

If you iterate this process many times and you keep adding nonsensical stuff to your reward model and giving it very low scores, you can—you'll never win the game.

Uh, you can do this many, many rounds, and reinforcement learning, if you run it long enough, will always find a way to game the model.

It will discover adversarial examples.

It will get really high scores, uh, with nonsensical results.

And fundamentally, this is because our scoring function is a giant neural net, and RL is extremely good at finding just the ways to trick it.

Uh, so long story short, you always run RLHF for maybe a few hundred updates.

The model is getting better, and then you have to crop it, and you are done.

You can't run too much against this reward model because the optimization will start to game it, and you basically crop it, and you call it, and you ship it.

Um, and, uh, you can improve the reward model, but you kind of like come across these situations eventually at some point.

So RLHF basically, what I usually say is that RLHF is not RL.

And what I mean by that is I mean RLHF is RL, obviously, but it's not RL in the magical sense.

This is not RL that you can run indefinitely.

These kinds of problems, like where you are getting the correct answer, you cannot game this as easily.

You either got the correct answer or you didn't, and the scoring function is much, much simpler.

You're just looking at the boxed area and seeing if the result is correct.

So it's very difficult to game these functions.

But, uh, gaming a reward model is possible.

So in these verifiable domains, you can run RL indefinitely.

You could run for tens of thousands, hundreds of thousands of steps and discover all kinds of really crazy strategies that we might not even ever think about.

For performing really well for all these problems in the game of Go, there's no way to, to basically game, uh, the winning of a game or the losing of a game.

We have a perfect simulator.

We know all the different, uh, where all the stones are placed, and we can calculate, uh, whether someone has won or not.

There's no way to game that, and so you can do RL indefinitely, and you can eventually beat even Lee Sedol.

But with models like this, which are gameable, you cannot repeat this process indefinitely.

So I kind of see RLHF as not real RL because the reward function is gameable.

So it's kind of more like in the realm of like little fine-tuning.

It's a little—it’s a little improvement, but it's not something that is fundamentally set up correctly where you can insert more compute, run for longer, and get much better and magical results.

So it's, it's, uh, it's not RL in that sense.

It's not RL in the sense that it lacks magic.

Um, it can find you in your model and get a better performance, and indeed, if we go back to ChatGPT, the GPT-4 model has gone through RLHF because it works well.

But it's just not RL in the same sense.

RLHF is like a little fine-tune that slightly improves your model.

It's maybe like the way I would think about it.

Okay, so that's most of the technical content that I wanted to cover.

I took you through the three major stages and paradigms of training these models: pre-training, supervised fine-tuning, and reinforcement learning.

And I showed you that they loosely correspond to the process we already use for teaching children.

And so in particular, we talked about pre-training being sort of like the basic knowledge acquisition of reading exposition.

Supervised fine-tuning being the process of looking at lots and lots of worked examples and imitating experts and practice problems.

The only difference is that we now have to effectively write textbooks for LLMs and AIs across all the disciplines of human knowledge.

And also in all the cases where we actually would like them to work, like code and math and, you know, basically all the other disciplines.

So we're in the process of writing textbooks for them, refining all the algorithms that I've presented on the high level, and then, of course, doing a really, really good job at the execution of training these models at scale and efficiently.

So in particular, I didn't go into too many details, but these are extremely large and complicated distributed, uh, sort of, um, jobs that have to run over tens of thousands or even hundreds of thousands of GPUs.

And the engineering that goes into this is really at the state of the art of what's possible with computers at that scale.

So I didn't cover that aspect too much, but, um, this is very kind of serious, and they were underlying all these very simple algorithms ultimately.

Now I also talked about sort of like the theory of mind a little bit of these models, and the thing I want you to take away is that these models are really good, but they're extremely useful as tools for your work.

You shouldn't, uh, sort of trust them fully, and I showed you some examples of that.

Even though we have mitigations for hallucinations, the models are not perfect, and they will hallucinate still.

It's gotten better over time, and it will continue to get better, but they can hallucinate.

In other words, in addition to that, I covered kind of like what I call the Swiss cheese, uh, sort of model of LLM capabilities that you should have in your mind.

The models are incredibly good across so many different disciplines, but then fail randomly almost in some unique cases.

So for example, what is bigger, 9.11 or 9.9?

Like the model doesn't know, but simultaneously it can turn around and solve Olympiad questions.

And so this is a hole in the Swiss cheese, and there are many of them, and you don't want to trip over them.

So don't, um, treat these models as infallible models.

Check their work, use them as tools, use them for inspiration, use them for the first draft, but, uh, work with them as tools and be ultimately responsible for the, you know, product of your work.

And that's roughly what I wanted to talk about.

This is how they're trained, and this is what they are.

Let's now turn to what are some of the future capabilities of these models, uh, probably what's coming down the pipe, and also where can you find these models.

I have a few bullet points on some of the things that you can expect coming down the pipe.

The first thing you'll notice is that the models will very rapidly become multimodal.

Everything I talked about above concerned text, but very soon we'll have LLMs that can not just handle text, but they can also operate natively and very easily over audio, so they can hear and speak, and also images, so they can see and paint.

And we're already seeing the beginnings of all of this, but this will be all done natively inside the language model.

And this will enable kind of like natural conversations.

And roughly speaking, the reason that this is actually no different from everything we've covered above is that as a baseline, you can tokenize audio and images and apply the exact same approaches of everything that we've talked about above.

So it's not a fundamental change; it's just, uh, it's just a—oh, we have to add some tokens.

So as an example for tokenizing audio, we can look at slices of the spectrogram of the audio signal, and we can tokenize that and just add more tokens that suddenly represent audio and just add them into the context windows and train on them just like above.

The same for images.

We can use patches, and we can separately tokenize patches, and then what is an image?

An image is just a sequence of tokens, and this actually kind of works, and there's a lot of early work in this direction.

And so we can just create streams of tokens that are representing audio, images, as well as text, and intersperse them and handle them all simultaneously in a single model.

So that's one example of multimodality.

Uh, second, something that people are very interested in is currently most of the work is that we're handing individual tasks to the models on kind of like a silver platter, like please solve this task for me, and the model sort of like does this little task.

But it's up to us to still sort of like organize a coherent execution of tasks to perform jobs, and the models are not yet at the capability required to do this in a coherent, error-correcting way over long periods of time.

So they're not able to fully string together tasks to perform these longer-running jobs, but they're getting there, and this is improving, uh, over time.

But, uh, probably what's going to happen here is we're going to start to see what's called agents, which perform tasks over time, and you supervise them, and you watch their work, and they come up to once in a while report progress and so on.

So we're going to see more long-running agents, uh, tasks that don't just take, you know, a few seconds of response but many tens of seconds or even minutes or hours over time.

But these, uh, models are not infallible, as we talked about above, so all of this will require supervision.

So for example, in factories, people talk about the human-to-robot ratio, uh, for automation.

I think we're going to see something similar in the digital space where we are going to be talking about human-to-agent ratios, where humans become a lot more supervisors of agent tasks in the digital domain.

Uh, next, um, I think everything is going to become a lot more pervasive and invisible.

So it's kind of like integrated into the tools and everywhere.

Um, and in addition, kind of like computer using—so right now, these models aren't able to take actions on your behalf, but I think this is a separate bullet point.

Um, if you saw ChatGPT launch the operator, then, uh, that's one early example of that where you can actually hand off control to the model to perform, you know, keyboard and mouse actions on your behalf.

So that's also something that I think is very interesting.

The last point I have here is just a general comment that there's still a lot of research to potentially do in this domain.

Main one example of that, uh, is something along the lines of test-time training.

So remember that everything we've done above and that we talked about has two major stages.

There's first the training stage where we tune the parameters of the model to perform the tasks well.

Once we get the parameters, we fix them, and then we deploy the model for inference.

From there, the model is fixed.

It doesn't change anymore.

It doesn't learn from all the stuff that it's doing at test time.

It's a fixed, um, number of parameters, and the only thing that is changing is now the token inside the context windows.

And so the only type of learning or test-time learning that the model has access to is the in-context learning of its, uh, kind of like, uh, dynamically adjustable context window depending on like what it's doing at test time.

So, but I think this is still different from humans who actually are able to like actually learn, uh, depending on what they're doing, especially when you sleep.

For example, like your brain is updating your parameters or something like that, right?

So there's no kind of equivalent of that currently in these models and tools.

So there's a lot of like, um, more wonky ideas I think that are to be explored still.

And, uh, in particular, I think this will be necessary because the context window is a finite and precious resource.

And especially once we start to tackle very long-running multimodal tasks, and we're putting in videos, and these token windows will basically start to grow extremely large, like not thousands or even hundreds of thousands, but significantly beyond that.

And the only trick, uh, the only kind of trick we have available to us right now is to make the context windows longer.

But I think that that approach by itself will not scale to actual long-running tasks that are multimodal over time.

And so I think new ideas are needed in some of those disciplines, um, in some of those kind of cases in the main where these tasks are going to require very long contexts.

So those are some examples of some of the things you can, um, expect coming down the pipe.

Let's now turn to where you can actually, uh, kind of keep track of this progress and, um, you know, be up to date with the latest and greatest of what's happening in the field.

So I would say the three resources that I have consistently used to stay up to date are number one, El Marina.

Uh, so let me show you El Marina.

This is basically an LLM leaderboard, and it ranks all the top models, and the ranking is based on human comparisons.

So humans prompt these models, and they get to judge which one gives a better answer.

They don't know which model is which; they're just looking at which model is the better answer.

And you can calculate a ranking, and then you get some results.

And so what you can hear is what you can see here is the different organizations like Google Gemini, for example, that produce these models.

When you click on any one of these, it takes you to the place where that model is hosted.

And then here we see Google is currently on top with OpenAI right behind.

Here we see Deep Seek in position number three.

Now the reason this is a big deal is the last column here.

You see license.

Deep Seek is an MIT license model.

It's open weights.

Anyone can use these weights.

Anyone can download them.

Anyone can host their own version of Deep Seek, and they can use it in whatever way they like.

And so it's not a proprietary model that you don't have access to.

It's basically an open weight release, and so this is kind of unprecedented that a model this strong was released with open weights.

So pretty cool from the team.

Next up, we have a few more models from Google and OpenAI.

And then when you continue to scroll down, you start to see some other usual suspects.

So XAI here, Anthropic with Sonnet, uh, here at number 14, and, um, then Meta with Llama over here.

So Llama, similar to Deep Seek, is an open weights model, and so, uh, but it's down here as opposed to up here.

Now I will say that this leaderboard was really good for a long time.

I do think that in the last few months, it's become a little bit gamed, um, and I don't trust it as much as I used to.

I think, um, just empirically, I feel like a lot of people, for example, are using Sonnet from Anthropic, and that it's a really good model.

So, but that's all the way down here, um, in number 14.

And conversely, I think not as many people are using Gemini, but it's racking really, really high.

Uh, so I think use this as a first pass, uh, but, uh, sort of try out a few of the models for your tasks and see which one performs better.

The second thing that I would point to is the, uh, AI news newsletter.

So AI news is not very creatively named, but it is a very good newsletter produced by Swyx and friends.

So thank you for maintaining it, and it's been very helpful to me because it is extremely comprehensive.

So if you go to archives, uh, you see that it's produced almost every other day, and, um, it is very comprehensive.

And some of it is written by humans and curated by humans, but a lot of it is constructed automatically with LLMs.

So you'll see that these are very comprehensive, and you're probably not missing anything major if you go through it.

Of course, you're probably not going to go through it because it's so long, but I do think that these summaries all the way up top are quite good, and I think have some human oversight.

Uh, so this has been very helpful to me.

And the last thing I would point to is just X and Twitter.

Uh, a lot of, um, AI happens on X, and so I would just follow people who you like and trust and get all your latest and greatest, uh, on X as well.

So those are the major places that have worked for me over time.

And finally, a few words on where you can find the models and where can you use them.

So the first one I would say is for any of the biggest proprietary models, you just have to go to the website of that LLM provider.

So for example, for OpenAI, that's, uh, chat.

I believe actually works now, uh, so that's for OpenAI.

Now for, or, you know, for, um, for Gemini, I think it's gem.google.com or AI Studio.

I think they have two for some reason that I don't fully understand.

No one does.

Um, for the open weights models like Deep Seek, CL, etc., you have to go to some kind of an inference provider of LLMs.

So my favorite one is Together.

Together, uh, a, and I showed you that when you go to the playground of Together, a, then you can sort of pick lots of different models, and all of these are open models of different types, and you can talk to them here as an example.

Um, now if you'd like to use a base model like, um, you know, a base model, then this is where I think it's not as common to find base models.

Even on these inference providers, they are all targeting assistants and chat.

And so I think even here, I couldn't see base models here.

So for base models, I usually go to Hyperbolic because they serve my Llama 3.1 base, and I love that model, and you can just talk to it here.

So as far as I know, this is a good place for a base model, and I wish more people hosted base models because they are useful and interesting to work with in some cases.

Finally, you can also take some of the models that are smaller, and you can run them locally.

And so, for example, Deep Seek, the biggest model, you're not going to be able to run locally on your MacBook, but there are smaller versions of the Deep Seek model that are what's called distilled.

And then also you can run these models at smaller precision, so not at the native precision of, for example, FP8 on Deep Seek or, you know, BF16 Llama, but much, much lower than that.

Um, and don't worry if you don't fully understand those details, but you can run smaller versions that have been distilled and then at even lower precision, and then you can fit them on your, uh, computer.

And so you can actually run pretty okay models on your laptop.

And my favorite, I think, place I go to usually is LM Studio, uh, which is basically an app you can get.

And I think it kind of actually looks really ugly, and it's—I don't like that it shows you all these models that are basically not that useful.

Like everyone just wants to run Deep Seek, so I don't know why they give you these 500 different types of models.

They're really complicated to search for, and you have to choose different distillations and different, uh, precisions, and it's all really confusing.

But once you actually understand how it works—and that's a whole separate video—then you can actually load up a model.

Like here, I loaded up a Llama 3, uh, 2 instruct 1 billion, and, um, you can just talk to it.

So I ask for pelican jokes, and I can ask for another one, and it gives me another one, etc.

All of this that happens here is locally on your computer, so we're not actually going to anywhere else.

This is running on the GPU on the MacBook Pro, so that's very nice.

And you can then eject the model when you're done, and that frees up the RAM.

So LM Studio is probably like my favorite one, even though I don't—I think it's got a lot of UI/UX issues, and it's really geared towards, uh, professionals almost.

But if you watch some videos on YouTube, I think you can figure out how to use this interface.

So those are a few words on where to find them.

So let me now loop back around to where we started.

The question was when we go to chat.openai.com and we enter some kind of a query and we hit go, what exactly is happening here?

What are we seeing?

What are we talking to?

How does this work?

And I hope that this video gave you some appreciation for some of the under-the-hood details of how these models are trained and what this is that is coming back.

So in particular, we now know that your query is taken and is first chopped up into tokens.

So we go to the tokenizer, and here is the place in the, um, sort of format that is for the user query.

We basically put in our query right there.

So our query goes into what we discussed here is the conversation protocol format, which is this way that we maintain conversation objects.

So this gets inserted there, and then this whole thing ends up being just a token sequence, a one-dimensional token sequence under the hood.

So ChatGPT saw this token sequence, and then when we hit go, it basically continues appending tokens into this list.

It continues the sequence.

It acts like a token autocomplete.

So in particular, it gave us this response.

So we can basically just put it here, and we see the tokens that it continued, roughly.

Now the question becomes, okay, why are these the tokens that the model responded with?

What are these tokens?

Where are they coming from?

Uh, what are we talking to, and how do we program this system?

And so that's where we shifted gears, and we talked about the under-the-hood pieces of it.

So the first stage of this process—and there are three stages—is the pre-training stage, which fundamentally has to do with just knowledge acquisition from the internet into the parameters of this neural network.

And so the neural net internalizes a lot of knowledge from the internet, but where the personality really comes in is in the process of supervised fine-tuning here.

And so what happens here is that basically a company like OpenAI will curate a large data set of conversations, like say 1 million conversations across very diverse topics.

And there will be conversations between a human and an assistant, and even though there's a lot of synthetic data generation used throughout this entire process and a lot of LLM help and so on, fundamentally this is a human data curation task with lots of humans involved.

And in particular, these humans are data labelers hired by OpenAI who are given labeling instructions that they learn, and their task is to create ideal assistant responses for any arbitrary prompts.

So they are teaching the neural network by example how to respond to prompts.

So what is the way to think about what came back here?

Like what is this?

Well, I think the right way to think about it is that this is the neural network simulation of a data labeler at OpenAI.

So it's as if I gave this query to a data labeler at OpenAI, and this data labeler first reads all of the labeling instructions from OpenAI and then spends two hours writing up the ideal assistant response to this query and, uh, giving it to me.

Now we're not actually doing that, right?

Because we didn't wait two hours.

So what we're getting here is a neural network simulation of that process.

And we have to keep in mind that these neural networks don't function like human brains do.

They are different.

What's easy or hard for them is different from what's easy or hard for humans.

And so we really are just getting a simulation.

So here I've shown you this is a token stream, and this is fundamentally the neural network with a bunch of activations and neurons in between.

This is a fixed mathematical expression that mixes inputs from tokens with parameters of the model, and they get mixed up and get you the next token in a sequence.

But this is a finite amount of compute that happens for every single token.

And so this is some kind of a lossy simulation of a human that is kind of like restricted in this way.

And so whatever the humans write, the language model is kind of imitating on this token level with only this specific computation for every single token in the sequence.

We also saw that as a result of this and the cognitive differences, the models will suffer in a variety of ways, and, uh, you have to be very careful with their use.

So for example, we saw that they will suffer from hallucinations, and they also—we have the sense of a Swiss model of the LLM capabilities where basically there's like holes in the cheese.

Sometimes the models will just arbitrarily like do something dumb.

Uh, so even though they're doing lots of magical stuff, sometimes they just can't.

So maybe you're not giving them enough tokens to think, and maybe they're going to just make stuff up because their mental arithmetic breaks.

Uh, maybe they are suddenly unable to count the number of letters, um, or maybe they're unable to tell you that 9.11 is smaller than 9.9, and it looks kind of dumb.

And so, so it's a Swiss cheese capability, and we have to be careful with that.

And we saw the reasons for that, but fundamentally this is how we think of what came back.

It's again a simulation of this neural network of a human data labeler following the labeling instructions at OpenAI.

So that's what we're getting back.

Now, I do think that the, uh, things change a little bit when you actually go and reach for one of the thinking models like O3 mini.

And the reason for that is that GPT-4 basically doesn't do reinforcement learning.

It does do RLHF, but I've told you that RLHF is not RL.

There's no, there's no, uh, time for magic in there.

It's just a little bit of a fine-tuning is the way to look at it.

But these thinking models, they do use RL.

So they go through this third stage of perfecting their thinking process and discovering new thinking strategies and, uh, solutions to problem solving that look a little bit like your internal monologue in your head.

And they practice that on a large collection of practice problems that companies like OpenAI create and curate and, um, then make available to the LLMs.

So when I come here and I talk to a thinking model and I put in this question, what we're seeing here is not anymore just the straightforward simulation of a human data labeler.

Like this is actually kind of new, unique, and interesting.

Um, and of course OpenAI is not showing us the under-the-hood thinking and the chains of thought that are underlying the reasoning here, but we know that such a thing exists.

And this is a summary of it, and what we're getting here is actually not just an imitation of a human data labeler.

It's actually something that is kind of new and interesting and exciting in the sense that it is a function of thinking that was emergent in a simulation.

It's not just imitating a human data labeler.

It comes from this reinforcement learning process.

And so here we're, of course, not giving it a chance to shine because this is not a mathematical or a reasoning problem.

This is just some kind of a sort of creative writing problem, roughly speaking.

And I think it's, um, it's a question—an open question as to whether the thinking strategies that are developed inside verifiable domains transfer and are generalizable to other domains that are unverifiable, such as creative writing.

The extent to which that transfer happens is unknown in the field, I would say.

So we're not sure if we are able to do RL on everything that is very verifiable and see the benefits of that on things that are unverifiable like this prompt.

So that's an open question.

The other thing that's interesting is that this reinforcement learning here is still like way too new, primordial, and nascent.

So we're just seeing like the beginnings of the hints of greatness in the reasoning problems.

We're seeing something that is in principle capable of something like the equivalent of move 37, but not in the game of Go, but in open domain thinking and problem solving.

In principle, this paradigm is capable of doing something really cool, new, and exciting—something even that no human has thought of before.

In principle, these models are capable of analogies no human has had.

So I think it's incredibly exciting that these models exist, but again, it's very early, and these are primordial models for now.

And they will mostly shine in domains that are verifiable, like math and code, etc.

So very interesting to play with and think about and use.

And then that's roughly it.

Um, I would say those are the broad strokes of what's available right now.

I will say that overall it is an extremely exciting time to be in the field.

Personally, I use these models all the time, daily, uh, tens or hundreds of times because they dramatically accelerate my work.

I think a lot of people see the same thing.

I think we're going to see a huge amount of wealth creation as a result of these models.

Be aware of some of their shortcomings.

Even with RL models, they're going to suffer from some of these.

Use it as a tool in a toolbox.

Don't trust it fully because they will randomly do dumb things.

They will randomly hallucinate.

They will randomly skip over some mental arithmetic and not get it right.

Um, they randomly can't count or something like that.

So use them as tools in the toolbox.

Check their work and own the product of your work, but use them for inspiration, for first drafts.

Uh, ask them questions, but always check and verify, and you will be very successful in your work if you do so.

Uh, so I hope this video was useful and interesting to you.

I hope you had fun, and, uh, it's already like very long, so I apologize for that, but I hope it was useful.

And yeah, I will see you later.