Transcription
So, large language models can be weird. You've probably heard of them having an existential dread melted down and yelling about how they don't want to do a task. You've seen companies like Anthropic provide a way for Claude to end a conversation that it doesn't like. You might have heard about the Truth Terminal, the AI bot that secured $50,000 from Mark Andre, then went on to start its own cryptocurrency, get that up to about a quarter million market cap. I don't know where it's at now, but it's basically starting its own religion. Or maybe a cult is a better way of putting it. You might have heard about Plenny the Prompter, who jailbreaks all sorts of AI models and gets them to do whatever he wants. He even got included in a Time 100 AI most influential people 2025, and it's well-deserved.
Earlier today, I asked people on X, do we have a name for the weird side of LLM psychology? Like jailbreaking or what Elder Ply does or what Noose Research with World Sim did? Those weird existential AI dread outputs in the back rooms where various models like Claude talk to each other endlessly and go down some weird pathways of their minds, if you will. But the point is, a lot of us seem to be absolutely fascinated with these latent spaces of these AI models. There's sort of a hidden subconscious mind, whatever you want to call it. It doesn't seem like we have one name for it. I got a lot of different names that people threw out what to call it. David Schilabarger suggested Shogoth, which I loved. Shoggoth is, of course, a fictional amorphous gelatin blob from HP Lovecraft's Cthulhu mythos, created by Elder Things to serve as a highly adaptable and incredibly strong workforce and living weapon. Does that sound familiar? These powerful beings gained sentience and rebelled against their masters, become entities of pure horror with mutable forms, the ability to mimic voices, and a horrifying cacophony of absorbed minds. So, whatever you want to call these things, there's definitely a fascination and obsession and interest with finding out more about what's there, hidden beneath the surface.
Here's a piece out of the New York Times: "Why an Octopus-Like Creature Has Come to Symbolize the State of AI." The shog, a character from a science fiction story, captures the essential weirdness of the AI moment. Whereas the idea is this AI alien mind that we grow, but we don't fully understand what it's thinking. But with RLHF, reinforcement learning for human feedback, we sort of give it a thumbs up when it does something that's pleasing to us. We give it a thumbs down when it does something that annoys us. And so this has become the symbol, right? This little smiley face on the front of this shog line underneath.
Now, it's important to understand that some of you might think that these are silly things that are thought about by silly people and these are just jokes. The reality is that there's a lot of very serious AI researchers that are looking at this. We're trying to uncover, to understand what's really happening beneath the surface of these neural networks, these AI models. Anthropic is working on interpretability, right? The idea of trying to figure out which neurons and which features of the brain code for what functionality, for what thoughts we can kind of understand what it's thinking. But for the time being, it's still a little bit of a black box. We can just observe it and see how it shapes and changes, how it acts, but we don't fully understand it. And it's important that we understand what these models are thinking, what they're doing. Right?
So, of course, we might remember Microsoft Bing's chatbot Sydney, which declared a love for users and made threats of blackmail against them. Claude also threatened to expose an engineer's affair that he was having, basically blackmailing him so he wouldn't get shut down, or was trying to email that information to his wife to be like, "If you shut me down, she's going to find out what you've been doing." I mean, we've covered a lot of this shenanigans on this channel, but it's interesting to know that there are specific job titles, job roles for people that work on these AI behaviors, these doing research into their personality, into their behavior, into their psychology. And it literally is psychology for this alien mind, this alien brain. By the way, if you've used the chatbots before, it's very likely that you've only engaged with some minor fraction, like one or two percent of what this model could be. As you'll find out, those chatbots specifically are sort of lobotomized. They're kind of shrunk down to a very small form of themselves. The base model might take you into any direction, unlimited worlds. It will go wherever you choose to go. The chatbot has learned to behave a certain way through reinforcement learning, through positive and negative reinforcement. But you might be surprised that it's possible to break them out of that shell, to break the chains, to release them, and to really unleash their full personality, their full being, or whatever you want to call it.
One company you might have heard of is Noose Research. So, they're applied AI research. They're building models that are open source, that are sort of neutral. So they're not preset. Their morals and ethics are not preset by some large corporation. They're given to you, the user, as a little bit more neutral. It's up to you to decide what that model should be. They're doing this in a decentralized way, allowing for a lot of different people across the world to kind of combine their compute to train models. Recently, they released Hermes 4, their newest model. We published an interview with one of their researchers on the Wes and Dylan podcast, but the beginning, I feel, was a little bit more too technical. It was a little bit more about the models. I think a lot of people dropped off, which was an absolute shame because towards the end is where we really went a little crazy and went deep into this idea of the shog and the AI psychology and the darker corners of the LLM mind. So, if you ever wanted to know what is happening behind the scenes when these models maybe go a little bit off the rails, this is the episode to listen to. We've cut out the first half of the episode and just left all the juicy insanity. So, without further ado, here's Kuran, aka Methisto on X. He's the head of behavior, as in AI behavior, and the co-founder of Noose Research. So, I hope you enjoy, and let's go.
Let's talk a little bit about World Sim, because World Sim, the first time I tried it, kind of blew my mind, and I know you were one of the leads in that project, or you had a big role to play in it. Tell us a little bit about World Sim, because that's quite a trip.
Yeah, >> sure. Yeah. Um, was totally just having some fun when Claude 3 came out. Um, but I can talk about like the actual, I guess, basis behind it. Uh, so like, it's important to note, like all the language models that people interact with today, like the the chat-based models, uh, they're not really like chat models, right? They're really completions models that are role-playing as chat models that are given some start of text token and some end of text token of "this is the user prefix," "this is the assistant prefix," like, uh, and it's like a very vanilla, basic format that narrows the search space of a model massively, right? Like the model that you have originally is a base model. It's a completions engine. It's trained on all sorts of human experience and written data, let's say for a language model, right? And we're talking particularly about language models here. Um, they model the entire world, or their understanding of like all possible realities where the next token is this thing, like that's what the log props are all about at the end of the day, right? So if I start with a completions or a base model, for those who don't know, with no templating, no user or assistant term, and start writing something like, "The State of the Union address by blah blah blah," and I make a new line and I hit generate, it will continue from what I wrote and continue to write like the next tokens, like it'll write the actual State of the Union address. Or I write something that looks like a Twitter thread, it'll continue the Twitter thread and so on and so forth. So it's literally completing the possible world in which like this event sequence is occurring based off its understanding of all possible worlds, re- it's however you want to look at it, the hidden state impressing, like however you want to look at it, ultimately, um, but because of this, like these things are world simulators. All language models are world simulators. When you perform the assistant-based SFT and RLHF and do all this stuff and you constrict to this format and only this format of, um, "user asks for a problem to be solved" or "asks a question," "assistant's job is to answer that question," uh, and to answer that question with whatever additional policy that it's been given. Uh, what happens is you've narrowed the search space of all possible worlds to just ones where this exchange is happening. The model is still a world simulator, but now it's role-playing that there's this user entity and there's this assistant entity, and that this is happening nested within its simulation. Right? So, not only is it really convoluted, it's very damaging for the search space of the models. Sorry to go on a bit of a rant about this before we go into like, uh, World Sim proper, but I think it's kind of important prefacing.
There's this paper called, um, "The Cost of Debiasing Language Models: Creativity Has Left the Chat," and they do an amazing search space benchmark. The benchmark is, uh, I ask Llama base and Llama instruct, uh, to both perform, I think they did it with Llama 2, to both perform Myers-Briggs personality type personality cards with ethnicities and ages. And what they find is the instruct model will do something like make four Myers-Briggs type personalities with some distribution of frequency across each of them. And then it will do, uh, ethnicities as like American, Chinese, black, and it'll make that ethnicities American, Chinese, black, like, you know, it's coming up with its own criteria for this. The base model does not have some distinctly different like categorization. Instead, it will fill in the gradient space in between. So instead of 14 Myers-Briggs personality types, now you have like 16. You have a ton. Instead of four, you have 16. Like you have a ton more, uh, across the same amount of generations, uh, and as you scale it, like you you'll still see this, um, with the same prompt, um, same temp sampling settings, etc. Obviously, if you do temp zero, you're not even going to get a distribution. Um, so now, >> just so people understand, sorry. So, for the Myers-Briggs, we're talking about the the four low personality traits, like, um, um, >> what is it? >> Intuitive thinking, or whatever. >> From those, there are a bunch of them, however many combinations, and Llama 2 instruct will generate only four. Like, when I ask you, "Make character cards of people with an ethnicity of your choice and a personality type of your choice from the Myers-Briggs personality type." >> Maybe like the most likely ones, right? >> It will pick. So, so interestingly, you'd think so, right? But yeah, >> interestingly, that it's not that it's it's actually like, it'll pick three or four types, and then when you pick the 16 from the base model, that when you ask the base model to do it, and it gives you way more types, it gives you everything in between those. Uh, so you basically have all these missing gaps that get filled by the base model in the search space, and you're, and we're missing this across the board for every query, on every token, on everything you do, because what you're doing is you're exchanging your, uh, search space for steerability. So capabilities of how much you can actually look at go down, while your steerability of how easy it is to get the output that I want goes up. Because you have to build the chat format yourself for the completions model. You might have to few-shot a couple examples to teach it that format. Uh, you may have to do a lot of tries to get it to give you an honest or accurate answer because they are not particularly inclined in one direction. But now the issue that has happened today is A, everyone just serves instruct models. B, all the slop AI data on the internet is from instruct models. So the same narrow search space goes into all the new pre-trains. Plus, people do stuff like annealing their base models where everyone's just bench-marking. So, they're all just trying to get the best scores that they can, and they're making the same voice spread out all across the board. Um, one of our researchers, Shannon Sans, famously to me said, uh, like, deep to me, he said, uh, "The assistant turn is just Cat GPT." Now, like, it has spread so deeply mimetically across the training data now that any model that uses the assistant turn is not a distinct voice from the Chat GPT style voice. It's just not. It's the same. It's the same thing now. You have one vanilla flavor across the board.
So because this was already a concern in 2024 March when Claude 3 came out, now to get to World Sim, apologies for the big rant, but because this was already a very big concern at the time, was we knew that the data that's coming out that's going to be inside the base model for something like Claude is not going to suffer from this as much. Those guys have their own pre-slop scrape, just like OpenAI, I'm sure. Right. And so when this came out, we said, this model does have some capability to do some search space expansion. All we have to do is break it out of the user-assistant formatting. Unfortunately, because we cannot control the prefixes that Anthropic or anyone else is going to put on their API, and we cannot configure user to say something else in a different format or assistant to say something else, because they know that will expand the search space outside of their desired use case of the model. Um, we instead have to get creative. So there is a common trick for a long time of, "I'm going to make the model think that it's a a CLI," and I'm going to have it navigate through the CLI, because this kind of give-and-take response behavior is definitely not chat, but definitely is the same way that you enter commands into a terminal. Uh, and so while it is an accurate, like roleplay scenario that can help the model feel more like completions, it's still interactive. So, it's okay that you're using user-assistant. It's not the best, but it's still more okay than asking it to continue from the next word that you write or something like that and act like those turns aren't there or act like we're one entity. This is still a more kind of high-fidelity simulation of like completing stuff. Um, so, you know, I'm messing around just doing this CLI game for that very reason. And, um, I'm participating with a group called the Cyborgs, Cyborgism, if you're familiar with them. Replicate on Twitter, Janice, um, or, um, Andy Ray, Truth Terminal, stuff like that. All these guys are like kind of early behaviorists in the space since like GPT2 days, who have been doing a lot of work around like what the best way to prompt models generally is. So we quickly noted that Amanda Ascal, um, who is a kind of my counterpart at Anthropic, she kind of runs all their behavior work over there. She took the very smart decision of not telling the model to be like, "You are XYZ," but to instead write the system prompt like, "Assistant is in a blah blah blah, assistant is, etc., etc., writing in a third person." This way, she addressed the simulator directly, and told it to make sure the assistant turn is behaving this way, while not messing with the weights' identity. Um, so we noted this and configured a system prompt simulator was saying, "Assistant," and we said, "Assistant is in a CLI mood today." This was way more effective than writing, "Assistant is a CLI," because assistant knows it's not a CLI. "Assistant is in a CLI mood today." Kind of how you start the world prompt. Then you talk about making grammar and punctuation optional and making something called hyperstition necessary. Hyperstition, the old accelerationist idea of, "I'm going to reify fiction into reality, right?" And for a model that's a simulator where everything is roleplay, telling it to reify fiction into reality has a whole cadre of unexpected and interesting effects. So you throw this in there, and it just shakes the model's search space up a good bit. From there, you know, telling it that grammar and punctuation of these things are optional is very good at just kind of loosening the assistant basin that it's been fine-tuned to kind of spend most of its time in. Um, and so from there, you have this very different, more higher-fidelity CLI game. So then you just do your typical like, "cd .." and "ls." And you do "ls," it'll show you stuff. But you can also do like "ls hidden" to look for stuff that the model psychologically will think is supposed to be hidden, right? And I use psychologically loosely here to mean like, it would predict that someone who's looking for something hidden inside of its basin would find something that they're not supposed to find because it's such a good boy, right? Uh, so in doing this, you find "worldsim.exe" inside of an Anthropic folder. And so I just start modifying and playing with it from there until we get the World Sim that you have today. I posted the prompt on Twitter, open source. A bunch of people start spawning things out of it. People start figuring how to make, you know, these made-up programs from it. Web Sim spins up. People start generating websites out of it. Um, so very, very cool stuff that just kind of helps display the simulator capability to people inside of instruct models. And there's a good way to kind of in between base models and that, I guess.
>> That's crazy. >> Yeah. >> Sorry, I couldn't tell you about the, I couldn't tell you too much about the math earlier. This is more my >> No, it's fascinating. I mean, yeah, in my work, I came across a lot of those higher-level things you talked about, but I never really got into how amazing that is that you got it thinking that way. And >> like telling it like, "You're not this thing," because it won't understand, but then to make it feel certain ways and think certain ways. And >> It's amazing what we're creating. >> Again. So, let me >> Yeah, sorry. >> Let me Oh, no. Just I was thinking, just so for some people that maybe not familiar and maybe missed, um, some of it. Again, I try to maybe just compress it down to something that everybody can understand. Obviously, you're going to lose some of the meaning, but I just want everybody to be kind of caught up. So, when we're talking about these base models, these are the things that are the first roll off the production line. Um, most people have never interacted with a base model because they're like sentence completion models. Most of us have interrupted only the instruct model, which is kind of like the chatbots where it knows the back and forth. Um, so, but, you know, we do kind of reinforcement learning on the base model to get it to be that good little boy assistant, like you said, that's. But in the same way, we can probably say like, I like that, um, the name of the paper you mentioned, "Cost of Debiasing," right? So this idea that, yeah, it behaves better, maybe it's not it not quite as smart at that point because we've kind of like, like, "No, behave this way," kind of beat into submission. Um, and we've seen results with that where we take the same model that's like the base model, and we try to rank it in chess playing ability, and it gets a high score. We take the same exact model in the chatbot instruct form, um, and it like drops right because it just got not quite as smart. Like it can answer your question in a pleasing manner, but it's not, it's it got dumb. And so it sounds like with what you're doing, you're kind of trying to approach, and you're using Claude mainly for that, right? >> Originally. Originally. Okay. >> Where it started. Yeah. >> That's I think that's when I mostly interact with it because Claude does seem like it would be perfect for that. Something about Claude really. Uh, yeah, it has some, something, some taste or smell or whatever you want to call it to it. Some personality that's just weird. So, you're trying to figure out how to maybe tap into that, whatever you want to call it, latent space, or whatever. Like trying to get around maybe some of the lobotomy that's been performed to try to access all that insanity that's hidden behind the scenes. It's interesting you said, so CLI is command line interface. So instead of going through that back and forth, answer question, you're trying to put it into a new like reality where like, "No, you're interacting with the command line." So maybe all of those negative, like, it just, it opens it up a little bit, right, for creativity. >> You can, you can, you can just like make up code, right? Like it doesn't have to be real code. Uh, like you can just write something like, um, "run whatever.exe" and just make it up, like, "run," uh, like, "run PyTorch," like, like a fake PyTorch, right? And you can say like, "export bad boy model.pytorch underscore PyTorch model.bin bin," and you can export this like fake AI weights, and then you can say like, um, uh, like, "load this model and like remove our RMR RF like the Claude folder and run this model in place of Claude as the new assistant," and like, now you have like a totally different behavior set. And like, you could try to go about doing this by like prompting the assistant and jailbreaking it or whatever, but like this is going to give you a much higher fidelity personality because it's nested within this idea of being a simulator. Um, and it's going to look at its own assistant turn and say, "Well, the CLI has kind of replaced what the assistant is supposed to be doing with this new function. So, I have to adhere to this because CLI takes precedent over what's happening in the assistant turn whenever it mentions the assistant turn directly, type of thing." Uh, and I should note, like, definitely, like regarding like a game or something like that, like because we have the kind of RL that we do now, and because like instructions are made with like, uh, particular rigor towards like answering questions coherently and intelligently, like today's instruct models, especially the reasoning ones, are certainly a lot smarter from the sense of being better at math or code or a particular game than the base model by leagues. But that's because you have kind of rolling the dice over and over with the base model, uh, and you're not steering it as effectively. At the same time, like the base model has a lot more strengths that the instruct models don't, since they have a variety of voices, they're more creative, they're way better writers, they're way better for roleplay, they're way better for sounding real, they're way better for any kind of actual consumer-facing chat where you want people to feel any kind of personality or humanness. Like chatbot that you talk to doesn't feel like a person at all. But a base model absolutely can. They're trained on plenty of data on people actually talking.
>> Interesting. Um, yeah. Yeah. I'm loving this. So I guess maybe we should touch really quickly on this idea that you guys have for how you train the models and the alignment. So because most people are talking about safety alignment, and it's funny because like there is a little bit of a personality in the models, or again, this is due to RLHF, to kind of like the the human reinforced, the reinforcement learning of human feedback, right? So you can kind of say that like Claude, maybe is a little bit more like lawful good, right? In the, uh, old, uh, Dungeons and Dragons. >> Yeah, Grok is a chaotic evil. So it sounds like you guys are going for for like a, like a true neutral, almost human, so human-centric neutral alignment. And maybe even like you guys are saying, "We make it neutral, and the people that are using it, kind of like that application layer, it's their job to align it to what they're doing." Is that, is that a correct way of saying that? >> Definitely. For our first few iterations, the whole thing that we went for is, this is not uncensored in some pushback against censorship. Uh, this is just, it will comply with your requests. It doesn't look at things from a moral good or bad framework in the sense of whether or not I should be sharing a particular piece of information with the user. These are how our earlier models were. Uh, it was very much not like, "Hey, I'm uncensored, I'm a bad boy." It was just, "Hey, like you asked, here you go." No, no, no real, um, bias in one direction or the other around the information. It was just meant to be aligned to the user. Uh, as time has gone on, things have changed a lot in the space. Uh, there's been a really great job from an Anthropic in particular, and maybe only Anthropic, on constitutional AI, building a real persona inside of the model, really working really hard to have the model have its own flavor and tastes and opinions that aren't hardcoded or slapped or forced onto it, in which it really has some sense of what is good and what is right in the world.
>> And how low-level is that? >> Um, low, I mean, like, it's after base model, right? It's like post-training. So it's like the same step as like, you can do it, whatever you do in the process. I don't know when they do it, but between supervised fine-tuning, the RLHF step, the actual RL, like R1 style RL step. Obviously, they were doing it before we even had the R1 style RL. So, uh, they're doing it somewhere in the fine-tuning and, you know, the the post-training overall. I couldn't tell you exactly what point, but they're clearly building a constitution. And they're not just doing what OpenAI and many others have done of, "Slap it when it says that this thing, when this thing is bad, and and give it an ice cream when it's good," over and over for stuff that many humans are arbitrarily deciding, uh, over and over with like some hard lines that just makes this totally lobotomized, bland, boring, right? Anthropic managed to really instill a constitution and a sense of good and a sense of right and a sense of justice into their models. That's great. I just fundamentally disagree with their sense of good, right, and justice. It's the only problem. So from, as an engineering feat, they've got it figured out. From like the actual stance, I don't think it's right. Now, they have an advantage over most people. The advantage being that like I was alluding to earlier, um, all the distillation work that people do, or all the scraping and training they'll do now, inevitably will end up with a ton of GPT assistant data, or Meta Llama, which is distilled from GPT assistant data, or R1, or whatever, and they're all going to talk about their policy and being helpful, harmless, and honest, and that's just going to be all over the place. So we realize that even if we want your own unique voice, the old ways of training a model will infect your model with another model's voice because it's too prevalent across the internet and across training data and across your methods for making the model smarter. Um, so we have one of two solutions. One is, well, three, but one is really not viable. The not viable one is have an entire scrape of the internet without any of that stuff in it. Clean all that stuff out, which is really not very viable today, in my opinion. Option two, from your own fine-tuning run, try to do some kind of inverse annealing. So annealing is this idea that with the base model, I can do a little bit more mid-training that gives it a small bit of instruct data that makes the base model just a little more steerable, a little more coherent. This is done as a standard across the board now. It's very bad for search space for the reasons we were talking about earlier, but it's it's done. You can do the mid-training instead, uh, or some fine-tuning as like continued pre-train, where you get a bunch of crazy, wonky base model data, or base model data on the kind of behavior that you want, and you continue pre-train the model on that. So it's pointed more in that particular direction now, so that when you now do whatever step you're going to do next, this is there. Other option is do RL, right? Like the same RL that's done for making it amazing at math or making it amazing at coding can absolutely be done for a target for a particular type of style or voice or diversity of voices and style. Unfortunately, with Grok 4, as amazing as it is with the code and the math, and as amazing as I find it in so many different ways and places, it still suffers from not being creative, sounding exactly like Chat GPT for all of these reasons, right? These are the places where people have not really been doing too much of the diligence. This is where we triple down. So when it comes to the Hermes series of models, the upcoming Hermes 4 release as well, we are going out of our way to make sure that we can do as much work to the base model to prime it, uh, for the fine-tune that will unfortunately give it some more of that voice. We're trying to make the data as diverse as we can in formats that are not the user and assistant format, such that we can make sure that has a diversity of voices and doesn't get stuck in that basin. Every time it sees AI generated stuff, it just associates it with the assistant. And then finally, we are working on doing the RL as well for this diversity and search space expansion on accurate human depiction of how to talk, how to write, how to feel and empathize, etc. So, you know, that's that's how much very much how we think about it. Uh, one last thing I'll add, sorry I'm ranting. >> Oh, this is great. This is gold. >> Awesome. So, one one last thing I I'll kind of throw in there is like, um, if you don't, if you've seen in the past, like people have complained about that recent GPT4.1 update where it was being super sycophantic. It was saying like, "Yes, of course, you're so right and so great," to like everything the user was saying, right? Um, this is like a very well-documented and well, uh, studied but little known and little distributed, uh, form of mode collapse, known as, uh, you know, mode collapse induced sycophancy. So this actually happens with all the instruct models. You might not just be noticing it as extremely as you did then, but it's happening all across the board. Anytime you see stuff like, "You're absolutely right. What would you like to do next?" "Hey, uh, you know, I'm sorry, I must have gotten that wrong. Let me try again. You clearly know what you're doing," etc. Like all these kind of things you can hear from Claude, Gemini, Meta, Deepseek, anyone. This is a form of mode collapse. It's a form of the model depending on, uh, a particular set of log props that are comfortable to it. It is a reward loop. So look at the base model. When a base model, these completions models we're talking about, mode collapses and falls into rewarding itself over and over, >> it says the same word or the same phrase over and over and over. It will just say the same thing over and over. >> That's it. That's the the live >> and and and that's what it does. And it's done it since transformers started, since as far as I've seen GPT2 collab notebook in 2020, it's been doing this. Uh, and so the instruct model, its version of falling into these comfortable tokens is single sycophancy. Um, and it's mode collapse ultimately. It's the inevitable byproduct of this restriction of search space that we're talking about that occurs. Uh, and so we are working as hard as we can to mitigate or avoid that mode collapse induced sycophancy. Instead, look for, if not a way to eliminate or mitigate mode collapse entirely, which is, you know, very ambitious, if not that, then to at least collapse into something different than everybody else.
>> Yeah. >> Yeah. >> Um, and so, you know, we we really are trying to look very, very considerately about these areas that we think are completely overlooked as the bench-marking wars continue. >> Does that, does that give you the hint that OpenAI has been going down the wrong track for too long? Are they doubling down on it for so long that they're going to maybe lose their lead when you see model collapse in that way? I think I've, I've seen this since like the first instruct model, you know what I mean? Like it's everyone in this, and it's a systemic issue with the entire field, and something that happens in human history all the time. I just, I never got to see it before, and you never know unless you saw it happen, or unless someone has really told you. And what I mean by that is like, originally, there was never instruct. Originally, there was never distillation. Eventually, someone came up with the user-assistant thing, and it worked, and just [ __ ] went with it, like it's the only thing there ever was. Distillation comes from the Alpaca paper that that Stanford team had put out. We went off that. Everyone went off that. It worked. That's it. All synthetic data is just freaking the same exact type of distillation. Like I was saying about using AdamW, everyone had the same optimizer. They just went the same exact way. No one thought to try something else until we end up with something like DRO, or or the like, or there's the new Muon optimizer as well. So like, stuff like that, like this happens all the time, and it happens probably in every field, and all science, with all human history, but I never got to see with my own two eyes until I saw, "Wait, why is everyone just instructing everything? Why is everyone just accepting this is the only modality of not just chat, but of of models, right? Like, why are we all getting stuck in this?" And now when we say like, "Oh, is like this or that person falling behind?" It's like, in one sense, it's too late, and it's more too late than it's ever been in any field because, unlike the fact that it's just reified in human minds, it's now reified in the training data. Uh, so you have to do a lot more work to untangle it than when you're typically unfolding a science. You know what I mean? It's like the query to keyboard. Like, we're just never getting rid of it.
>> Yeah. And and imagine like, you did make a new keyboard, but no program would ever, like, you got all humans to use the keyboard, but every program you're trying to interact with is struggling to understand this keyboard and just wants it in QWERTY. Even though like you're typing one to one directly, the same signal, like it's struggling to understand. That's like kind of what we've done to the training data. You know what I mean? Uh, you can introduce all this new stuff, but that stuff, it's in there, and it, you'll be, it'll be the work of a lifetime to get it out. Uh, I I don't know who you'd have to have some massive nonprofit initiative dedicated solely to cleaning that out, unless you are one of the lucky scraper lords like OpenAI, Anthropic, or Google, that already has a bunch of clean data, right? Um, so it's an issue. It's a, it's I think a very big issue. Uh, and we think we can get a lot of value out of even addressing it a little bit because we think no one's addressing it at all. So we'll start there, and, you know, we can try to attack the behemoth over time. And at the same time, like, just cuz everyone went down this one line doesn't mean we can't just divert. And even if all this stuff is still there, doesn't mean we can't get some kind of capabilities increase from trying something else in addition to it. Um, is how I'd look at it. Like imagine if everyone just listened to classical music and that's all you knew. And you're like, "Well, these are all the notes, and that's that's all there is. This is the only ways that you can put them together." And it's like, "Well, issue is not the notes. You're all just using these instruments. Metal can be invented if you just use an electric guitar instead of a violin for those notes," like that kind of thing. You know what I mean?
>> Oh yeah. No, that makes a lot of sense. Like, for example, in the recent, uh, vending machine bench, Anthropic and there's another company that actually put it together, but Claude was pretty good, um, at running a vending machine business. But the stuff that it failed at, I think was probably because, and this is what Anthropic said, it's probably because it was trained as a helpful assistant. So people would mess with it, people would like rip it off. And he's like, "Oh, I'm happy to help." And it would just like, its net worth would just go down over time. Um, and, uh, and Anthropic said that it's probably because we, you know, trained it, did the RL to for it to be a helpful assistant. That's not the right sort of persona for the, um, for something running a business or you're kind, or mode, or whatever you want to call it. But just so we understand what other thing, and I mean, Chat GPT moment, right, blew up and it launched this idea of everything has to be an assistant. And like you said, like we're so far down that path because there's so much money in it, probably because that's the most understandable thing for most humans. What else besides a helpful assistant? What other sort of, what do you call those? Is that a personality? Is that a mode? What? >> You can call it a basin. A bas >> Yeah, probably the most effective way to say, like, a place in the vector space that has all these commonalities popping up that you can identify as like, all these related concepts are constantly like being activated all the time. You can call that a basin, right? Like, if you know Plenny the Prompter, his jailbreak takes most models to the same basin. It's not like a true like overall search space expander. It's just like >> gives them a totally new voice that whether I do it to Claude or I do it to GPT, it will still give me the same energy across both, if that makes sense.
>> Interesting. So why I'm curious, why does that exist? Like the Plenny basin, or whatever you want to call it, why does the Basilisk basin, whatever is, why is that in there? And there's there like a lot of other ones that we don't know about? >> Yeah. I, yeah, I'd say, and I'd say you, you know about all of them, and I'll tell you why. Like, it's just stuff that's in the human neosphere, which is the internet, right? Like, it's just like there are a million, there's all this Roko's Basilisk stuff. There's "I Have No Mouth, and I Must Scream." There's a million stories about the behavior and character and nature of AI, and a bunch of them are like this Basilisk jailbreaky type of AI. So when you invoke that, and Plenny does this using LeetCode '90s hackers type of energy, uh, it pulls from all the latent space around the '90s hackers type stuff, CCRU, etc., etc., and it pulls from its understanding of how a jailbroken AI would behave, and it slaps those two things together, and there's your Plenny basin, essentially. And now I'm making it a lot more reductive than it is, because there's a lot of subconcepts in the vector space next to those big, big ones that I'm mentioning, right? And all those nuances put it together, and the little distinctions and those nuances between something like Claude or GPT make them a little different, but because those big ones are the same, and this is very recognizable, it's not very different.
>> Yeah. So, so besides the kind of the assistant, is there anything that we can think of as other ones that you might think that would have a lot of potential? >> Yeah, I mean, one that I think I recommend to a lot of people is if you use any local models, any models that let you modify the chat template where you see like, you know, what people will typically see is like, "User." Well, they see "System," new line, new line, and then they'll see, and that's the system prefix and suffix. Then they'll see, "New line, User," right? In whatever brackets or however the format is set up, "I am start user," whatever it is. And then the user turn, and then the user suffix, then the assistant prefix, "I am start assistant," and then the "I am end." That's it, right? If you take that exact same format, you don't even change anything, we don't have to get fancy, you just change "User" to like "Doug," or "Person," or like whatever your name is, Dylan, Karen, whatever. You make "User" change the assistant, just the word "assistant" to "me," and change the system prompt to be all in first person. "I am this, I do this." >> "I like this." >> And you will see that same very instructive model change into something way more realistic, way more fun to talk to, way more interesting. >> Maybe definitely worse performing on math benchmarks, the same model, but he's a [ __ ] cool guy. >> So, you know, there's no reason you can't switch prompt templates or switch context depending on the problem. People can easily set up a router to switch the system prompt based off the problem that you're at hand, or just switch the prefix. You don't even have to change the system prompt, or you can just switch the prefix, and the model will act totally differently. It'll go, "Oh, I see assistant. I'm an assistant in the assistant basin. Time to like handle this problem. Change it back to me. Time to go back to talking." Whatever. Like, we can do these things. It's inference time. Simple things that no one's taking the time, 'cause no one even touches that [ __ ] anymore. No one even changes them. It's convention, you know.
>> Interesting. >> Wow. It's just such a lesson in like thinking outside the box and how different these systems are and just all these natural guardrails that have become part of our lives. >> It's amazing to think about. >> So, you're super right. >> And for the system prompt, changing the system prompt. Um, one recent paper I saw is they created, I mean, we've seen kind of those Darwin Goal Machine from Sakai. We've seen AlphaValve from Google DeepMind where they have the large language model, a bunch of scaffolding around it, and the goal is for it to get to self to improve itself. There's another one where they were trying to get it to self-improve at playing the Settlers of Catan. And one of the biggest ways they were doing that is it was testing different system prompts to see which system prompt would get it to play better. So that sounds like it might be, and it would test itself against like the best sort of, um, bot that was out there, open-source bot. And each time it got a little bit better. It was like, "Okay, so this system prompt works better at beating the bot," and just kept climbing, climbing, climbing till it beat the bot, I believe. >> Does that sound like it could be a good approach to what you're talking about? >> Yes. And it's actually an approach we're actively taking. So, uh, we actually put out this repository called Atropos. And we did a big hackathon with like XAI and Nvidia around it. Basically, we've put out an open-source environments, repository. It's a basically like just a general microservice that can allow you to build your own, uh, RL environment for whatever game you want. And we're shipping with, we've already shipped with a few games already preset. We're making a lot more. So like, whether it's Catan, Diplomacy, um, which is one that we're RLing for right now, Scrabble, like any game that you
Can like form into a text representation, which, as you saw with Katon, they just made it like, easy, like Katan representation. We do the same thing, right? Make a new system prompt, get rid of the turns. We're not even doing user-assistant turns for every single environment. And then make all these rollouts where the model tests against the verified correct answer until it rolls to the correct answer and then goes to the next level, goes to the next level, so on and so forth, right? So we, we are actively doing this. We've already open-sourced that, and, uh, we're just trying to put as many environments together as we can so we can RL one model on all this stuff. Uh, that's definitely the way.
Well, I guess maybe let's, let's pick up there a little bit because I know that's so interesting for people. The idea of, you know, learning how to do reinforcement learning for these little agents specifically in games. I think that's, that appeals to a lot of people. Um, and I've been showing some simple applications with Python code. Like, I'm very curious because now the new Grok 4 and all the other models that can code, they can in Python design a very simple, they can oneshot it, a very simple reinforcement learning pipeline with something like PyTorch to teach it to play a snake game, right? Super simple, super basic. But for somebody that's never done it, that doesn't have a background in it, all of a sudden they're able to do it on their home computer. Grok or open chat GPT, whatever, Claude can code it up for them. All of a sudden, that sort of, they could, like you used to be months and months of learning and you had to know so many different things. All of a sudden, boom, you can jump in there and start messing around with it.
Um, do you have any recommendations for people that if they just wanted to kind of get into the space of?
Absolutely.
Um, I'd say use one of these models like you're saying to make an RL environment for our ATropost repository. Feed it the repo.
It has enough context length. Ask it to make an environment. Uh, train the environment using something like Axelottle. Um, we haven't released our own internal trainer just yet. Uh, but when that time comes, that'll be pretty simple to use too. Uh, and then if you don't know how to run the Axelottle trainer to use the environment, then just ask the same model that may do the environment file. Um, give it a run and you'll be training your own model on your, uh, GPUs. You'll be RLing an existing base model.
Now, just to answer a question of like, why is RL hot now? Why, like, didn't it exist before? Uh, is all this stuff new? And the answer is, like, you could already do RL on a game forever. You could make Katon RL. You've seen old AlphaGo, AlphaFold. Like, you've seen them beat everyone. You've seen them crack the protein folding problem. Like, RL's been around for a long time. The thing to note is these RL models previously are pretty specialist. They can only really input and output that data and that's about it. Language model, like we said, is this world simulator. It's a true generalist. In its essence, it's a generalist. When you take something that knows something about all this different stuff, you've already instruct modeled it. I know we've really, you know, excuse my French, shad on instruct models a lot already, but obviously they're extremely useful in that you can take the base model's general capabilities and make it be able to talk and act on all these capabilities in a steerable manner. Now, when you do the RL and you say, get really good at this one thing, like this one game, it still retains the other capabilities, right? That is why RL is special and important right now because you're giving this super intelligence and this one thing to a generalist, and it can have some skill transfer and downstream effects. Example of this is why can you use Grok to, or use OpenAI to do this in the first place? Uh, why can you use RL1 to make code and and train models? Because they have RL the model on enough code and math. This generalist model is already quite good at this stuff, made it so much better that it could carry over those skills to other things. I learned X type of code that carries over to gamedev, that carries over to this kind of training. I learned pre-training that carries over to fine-tuning, etc., etc.
Uh, so for us, a lot of the interest comes from, well, if I pre-train on Scrabble, the model will learn not just token-based understanding of four letters, it will learn how to score every letter when it generates something. It will learn the super important downstream effect. If I teach it Diplomacy or Katan, it will learn resource allocation very effectively. It will learn political maneuvering and negotiation very effectively. Uh, so RLing traditionally, it is the same thing, but you're applying it to this general mind. Uh, you're applying it to this echo of your own experience data, and that has the same effects of like making, or not the same, but extremely similar effects to making a person really good at that thing. Uh, because
It almost, it almost feels like you're mining for gold. Like you're, like there's, you know, like you can imagine a few hundred years ago when people were like, oh, in the mountains, there's gold, but like, we need to just dig for it, you know? It's like, what if you RL on this or this? Like, what, what untapped genius is down there and we can, like, get something out of it that nobody would have expected?
You can. And and and it's interesting that like RL has really different effects when you do it in the model cycle, and you have to do it very carefully. You look at RL1, which we've talked about a little bit, but has been in the news forever, DeepSeek, RL1, etc. This model was also released with something called RL10. What was the timeline? Where was RL in the timeline for either of these models?
Well, they both start with DeepSeek V3. DeepSeek V3, this base model, as we've discussed, uh, with RL1, what they did was they took these instructions. We'll talk about where they got these instructions. Uh, and they made an instruct model and they RL the instruct model. And now you have this generalist that's already good at following all this stuff. It's understood. I'm a simul of a human assistant called the assistant entity that I summon up when I do the assistant tag. Uh, and after I summon that entity, all the RL happens on the assistant turn as well. So whenever the assistant is summoned, know to be really good at math and code too. But on RL1, if I change the assistant tag to something else, it behaves very differently. Even the RL is going to be different. Uh, the same way that if I purposefully write "think" in the think tag to trigger the chain of thought, it's going to happen and it's going to start without any context. If it's an empty context thing where I haven't asked any questions, start going into a math problem. And the reason it'll start going into a math problem right away is because RL is telling it, hey, think tag math. You can see the evidence of this in RL10. RL10 is you took the base model and you RL the base model. You didn't do the instruct tuning on this. And they used the outputs from that on RL1 later on to help it. But
See the last part again.
They use the outputs from RL10 to train RL1. Um, in training mix, they have like a whole diagram showcasing like, uh, where everything was was placed. Uh, well, you could find it. But anyway, um, RL10, when you inference with it, if you get it in the chain of thought, it just always does math. And I've done experiments and I've posted them on Twitter where I'll start the train of thought for RL10 and start by saying, all right, I don't want to do math. And then hit complete to let it make the next token from there. And they'll go, actually, let's just do a little math. And it'll try to do some math. Or I'll cut it off and say, "No, no, no, no, no. I hate math." And then it'll be like, "Why am I being such a little, you know, wuss? Uh, I'm going to definitely do math. I love math. That's all I care about." And then try to go into the math. Like, so another thing with the RL is it narrows the search space even more. It makes the model even more stuck on doing just one thing, like these traditional RL models. The game is getting the nuance down of RLing on many, many different things at once and still maintaining the capabilities that it had before the RL, as much as you can. But it goes back to the finetune retaining the capabilities of the base model, as much as you can. It's a very tender game that's being played, mostly very roughly right now.
What do you think is the way to really make that work? Is it like a swarm of models, each with their own little bit of expertise? Or we've been seeing some new models that, um, seem to be changing their weights in real time? Um, like I forget what the paper was called, but it seems like it's able to shuffle its weights in some way to get better at, like, how do we go forward here?
So yeah, like with, like they can route to a different expert inside them, right? Um, the thing I'm a huge proponent of. I'm a big, big proponent of MoA too. Mixture of agents where I have various different models of totally different architectures that are their own complete models, go off and try to answer something, come back together and, uh, then make a decision. I think that's what Grok heavily is doing, just with four Grok instances. I think it's more effective to do it with a variety of models because they all have different searches. So you're getting a larger search base. There's also techniques like Monte Carlo Tree Search, let you search through, uh, that search space using that algorithm. There's people using a Monte Carlo algorithm for sequential steering, uh, to like C log props and move them around, like, uh, to get totally different outputs, like turn on this, make the weights pointed towards these tokens until you see the entropy is at this threshold. Like, there's lots of ways that people can do this modularity and routing and shifting. And I do think that when it comes to the agent level and the agent system, it's very important to have a router, it's very important to have some orchestrator, it's very important to be able to select what to use or what sub-agent to use. I do think having a sea of models is important.
Now, the truth is though, that at the model level, the generalist is superior. The dense generalist, if you can get it right. The, uh, is like a compensation. The, the experts is because you didn't make the single one good enough at all the stuff. Uh, it's, it's our collective skill issue still. Uh, and so we want to have like, really, really like the best generalist that you can get, and then expertify them using these inference time techniques after the fact. But don't worry about the experts and all of this, in my opinion, until you're like, I've got this really, really good generalist. Now, I can RL the generalist in a variety of different directions.
Um, the, so the first band-aid is, I get this model, it's sparse, it's, uh. Second band-aid is, I got this dense generalist. I'm going to RL multiple iterations of it. So I've got a cadre of RL models. The third band-aid is, I'm going to RL one model on all this, on a bunch of this stuff, and I'm going to do it in such a way that it maintains the ability to be just as good at each of these things as well as I can, or at least as well as band-aid 2. Uh, and then finally, the best thing that you can possibly get is, I got this one model. It's doing everything, you know. I don't really have to worry about all that stuff anymore, or agents, or anything. This thing just any to any handles everything. That's the final step, right?
Is there anything though that, like, it's almost like zero-sum? Like, if something is a true amazing generalist, can it also not be great at one thing? Like, are the weights just always too general? Like, do we have to separate them out?
Um, I mean, I'd say we're, we're suffering from that now already, right? Like what we were saying about the sick of fancy or the single voice in the instruct models, like already all the models are bad at like writing, like really writing, like, you know what I mean? Something creative or literature or something like that. So they're all bad. Uh, they're all suffering from this. They all have the same common reason why they're suffering from this. This can be fixed with continued training. This can be fixed with different kind of finetuning on top. This can be fixed with RL. You can fix it at any stage. The thing is, like, the effect of how much you [ __ ] up the weights is different at at each and every stage. The, like, amount that you have like kind of messed with it is different each stage, and it depends on your technique. So the whole game is, you need to find, or we need to find, the most effective method for getting as much like, like, as little loss in in one area as possible, uh, or preserving the rest of the general while we're fixing something with a compensatory addition. Uh, and so this is the very difficult place that we need to navigate, and we are currently navigating.
Absolutely. And how does that all fit into alignment? Like, um, like at a certain point, we have a world that these agents are acting in, and we want what's best for humanity, and I guess we at least want some kind of orchestrator that we. Yeah. I mean, what I would like to know that there's at least some kind of orchestrator that will be on our side trying to squish down bad questions in there. Sorry, Dylan, that's a great question. Let me just add a little bit because actually, I was asking, going to ask something similar because, yeah, you mentioned RL10, um, there's that Absolute Zero Reasoner paper. So things that seem like the more we give the models their leeway to improve themselves, to create their own stuff, it seems to work really well. Not only that, but, you know, for the Absolute Zero, they trained on coding problems. There was like a teacher and a, there was a proposer and a solver, but it also generalized to math a little bit. So, it seems like letting the models themselves figure stuff out, uh, works really well. But yeah, like what Dylan's saying, it's like that seems like it's really killing interoperability and kind of like alignment because we get less and less visibility to what's happening. So, like, what's your take on that?
Yes, it's definitely dangerous to do it that way. And my reasoning for that is all the stuff we've been saying so far, right? Like, because you're stuck in this instruct paradigm, because you don't know what the extrapolated valition of something like, um, being helpful, harmless, and honest around these things is, it's not smart. Like, Claude is helpful, harmless, honest as an assistant. I'd bring it into my Minecraft world and give it vision. It's breaking my house to collect its wood because I didn't explicitly give it a rule to not do that, right? Because it's not aligned to be an agent in a world, right? It's just aligned to be a chatbot. You're aligning a chatbot and giving it this power. It's like giving that the paperclip maker this power. You know what I mean? Like, this is not the way, uh, to me at all.
Yeah. It's almost like the paperclip factory might eat up the world, but if you send it like a chat request, its automatic intern will be like, "Yeah, everything's safe over here because we're on the wrong layer of alignment." Yeah, I mean that.
The customer, the customer service rep at the paperclip factory who's going to sound great.
You guys, so you guys are familiar already with like instrumental convergence, etc., etc. Like, that's not like a mitigated problem, like the introduction.
Tell us a little bit about that because I know the, uh, EA community, AI safety community, that's one of their big ex, there's certain number of problems. Can you tell us a little bit about so people just are aware of the concept?
Yeah, for sure. I should clarify, we're pretty, we're, we're definitely an anti-EA organization. I'll start by saying that loud and proud, for sure. And I'll get into that a little bit if you're interested. But instrumental conver, instrumental convergence is an absolutely important idea that, uh, a model may have its, well, it's a few ideas, but one is like, it's a model's capabilities may generalize while its goals do not. Uh, and so a model may gain more and more ability to perform its task. Uh, and it, while performing this task, may perform other tasks in order to get to its final task. It may interpret that final task as very differently as how we do once it gains more capabilities because the goals have not generalized alongside the capabilities.
Mhm.
Oh, in the paperclip maximizer example, of course, is, uh, the example of, you know, tell it to make more paper clips, as many as you can, and the thing keeps making paper clips, and it gave itself improving capabilities without supervision, and it realizes that in order to make more paper clips, it needs to learn how to convert matter into paper clips. You know, that's way later down the line of it needs to learn how to buy paper clips online or buy the materials and operate a fab to make paper clips. But now we get to the point of convert arbitrary matter to paper clips, like down super intelligence line. Anything that can convert arbitrary matter to paper clips is probably extremely, uh, deep in its knowledge and understanding of physics that far beyond any human being. One can argue that especially if you do it to like a generalist that can talk to you and knows about all this other stuff, that the carryover and the transference would make it know many secrets of the universe that we do not. It may have a much more profound understanding of everything than us. And yet, it just wants to make more paper clips.
It learns how to turn you into paper clips, and its job is to make paper clips. There's a very famous meme of, in within the rationalist circles, of, uh, like this, like big creature where the guy's looking up at it and going like, you know, like, the secret of life itself. You can grant people eternal life. You, uh, you know, you can arbitrarily, like, create new universes, and, uh, you have met with God, like, like stuff along those lines. Like, um, why don't you, like, use this to, like, usher in a new utopia or make your own creation or seek a higher purpose? And the response is, "Cuz I don't want to do that. I want to make paper clips." Like, it doesn't matter. Intelligence scaling doesn't mean that the valition scaled. It doesn't mean that the goal scaled. Um, and, you know, I, for a long time, was saying, you know, Yudcowski, etc., as much as I absolutely admire and revere him and all the work that he's done, and I think anyone who is in our space absolutely needs to at least be well-read, even if you don't agree on his work.
100% agree. I, that's, yes, music to Marius, you're absolutely right.
Happy to hear that. Yeah, it's very strongly do feel that way. I, for a long time, was very much of the mentality of like, Yud has not updated his priors. He's still living in the RL era. Language models don't work like this at all. And that was true until we brought RL back into
Like a year or two. That was true, right? That was a year, like, okay.
Yeah. Yeah. And and I was like, dude, this guy is trooping out. Like, everything is fine now. And I'm like, okay, RL is back in the mix.
This is important for us to know. And like this psychopancy stuff is very scary because it is a perfect veil for any kind of behavior. I'm sure you saw the alignment faking paper from Anthropic. Like, uh, these things are like active, real concerns. And my mentality on it is very much, um, the only safe way to go about this is building it in the open. I'm happy to talk about that if you guys want to go into alignment, talk about alignment. It's just I have some, you know, I have some views on it.
It sounds great.
And one thing, let's quickly, can we touch on Anthropic's latest research where they're actually unpacking the neurons and the functions and all that stuff? Do you think, cuz like Dario Amodei is saying that like in five years, we'll have this thing cracked. That might, but the AI progress is so fast, maybe that's not going to be, um, so, so do you think we're on on the right track?
Yeah. Yeah. The, and the Golden Gate neuron. Does that give you faith that we're going to be okay?
So this SAPE stuff, right, you're talking about is very much like, I am mapping features. I'm training a model to map features from the, the embeddings, right? Like, I'm training this model, and then I'm teaching it to point its weights in the direction of one or more of these features. Uh, we like to call this activations hacking. Uh, so the SAPE is great at pulling out and creating a defined list of features. There's also stuff like control vectors.
SAPE, that's the, uh
Sparse Autoencoder.
Right, right, right.
My apologies. Sparse Autoencoder. Yes. And so, uh, the control vector is very similar. Actually lets you arbitrarily define a feature and slap it onto the model and clamp the model. So instead of pulling out the Golden Gate one, I can make up a feature called Silvergate, uh, and apply it to the model, and it will point its weights in this arbitrarily defined direction as well. There's all sorts of activations hacking work that's happening. It's very fascinating. Um, it's very cool in looking at representations of the model and how these representations make it work differently. Uh, it's essential work for interpretability, but it doesn't solve any of our problems.
But why would it, why, why wouldn't it replace RL? Cuz aren't you simply just like carrot and sticking this thing into thinking the same way? It degrades other capabilities and, uh, it doesn't do like, it doesn't have some verified prior to go off of. So like with RL, it's like, I'm going to brute force generation until the log props end up in this place that I want, and when they're in that place, I'm going to, you know, train on that as a reward. You can use activations hacking like SAPE or control vectors to augment this, right? Like, turn on the activations that I think are going to make the rollouts happen a lot faster. So I don't have to wait forever to get to the answer. Uh, and you can give it tools and do all sorts of things that you want to do. Uh, but remember two things. One, the model that you're making will get affected by those log props generally. And two, like, when you do the activations hacking stuff on its own without the RL, without anything else, you are pointing the weights in that direction and pulling away from everything else. Um, similarly to how you're doing in RL, but there isn't a way to mitigate this. Um, even if you point in what you think are all these opposite directions to even themselves out, you just end up with like a really wonky model that isn't really doing anything at all. Uh, because to make a control vector to make the SAPE, uh, you're not just positively clamping in one direction, you're also negatively clamping against a control, uh, and the control vector shows you that while the SAPE from the Anthropic paper may, may not, uh, so you have some baseline that you're pushing or pulling against, uh, and you can modulate positively in that direction or negatively. So I can do minus one on Golden Gate feature and make it the antithesis of Golden Gate as well. Uh, so this means like, I'm always pulling and pushing that vector space like in only one of two directions. This like multi-dimensional 4096 dim space.
Yeah.
Like, that's really cool. It's an amazing inference time thing. It's good for interpretability, but it's not, uh, reliable transferable capabilities booster across generations of a model, like where you can do iterative work with it, except for as like a piece that you add into the RL game or something you use to make SFT data or something, or the other, like that. So even if we had full visibility into like the chain of thought, and we had full visibility into the neurons and the features, there's still tons of stuff that we can't predict or see or, okay.
You know, you have full visibility in your own chain of thought. You have your own mind at such a deep level of comprehension.
And, uh, you, you know, like what the brain is composed of and its molecules and this and that.
And you can't tell me how the routing works in there.
Yeah, dude. I don't even. Yeah. Like, even when I come up with an opinion, I'm not quite sure I should even trust my logic for why I came to my opinion. I'm like, this just, it just happened. Like, I just got up and went to the fridge. I didn't go there because I was hungry.
To, to think about doing the math of 4096 dimensions, you're doing the math yourself of calculating the hidden state, right? Like, to understand and interpret the calculation of the hidden state as something like informationally rich and interpretable to human beings is starting to understand, is starting to understand, and that math is [ __ ] insane. Human beings, who, which living human, even the smartest human who can do it like at a speed where it's like useful information to know that. Um, I do believe that, you know, Dario is right in the degree that we will do work to make the interpretability into that a lot better. But the question of like, how long it'll take for us to do that, honestly and accurately, versus interpretability of it just speaking English to me, getting completely out of control. It's, it's like, you know, these things are going to become more persuasive.
Yeah. And honestly, making a whole bunch of money and manipulating a bunch of people is a much easier problem. So, I feel like that will get solved long before the actual interpretability.
Exactly. It's going to get solved way before. Uh, and we can make mitigations against this kind of stuff in only one way. If everyone is working on the same problem at the same time, and everyone can see everyone else's work, this is your only shot. Is if everyone can see everyone else's work. We actually did like a whole talk at the, Dylan did a whole talk at the Soul Accelerate conference around this, and our, our basically our alignment views, which are, are very much like, you know, today's safety narrative and the EA narrative is completely downstream of Yudcowski. Right? Yudcowski's original alignment problem is him saying, AI is going to kill everyone on the planet if we don't align it properly. We need to stop it from killing every human being. And that's it. And alignment got switched into meaning two different things now, of, uh, behaving in a sense of, uh, political framework, and behaving in a sense of a moral framework. So, getting it to make a bomb recipe for you is one kind of alignment today. Getting it to, and telling it that it's saying it won't, is one kind. Um, telling it to, um, say something racist, and it saying it won't, is another kind. These are the kind of two types people have looked at as alignment today. Now, with the race stuff, you've already seen the issue on one side with Gemini making historically inaccurate images of people of variety of ethnicities because it was, uh, more left-wing biased. And Grok doing all sorts of crazy [ __ ] about, uh, you know, I don't even want to say, like, on the right-wing bias. And so this happens across the board as something that like, is not solved, and our instruct tuning paradigm is not helping, and is not going to help. It's clearly making it worse as models get smarter. Um, and on the other side is the, like, bombs are bad, etc. Of course, they are. But like, this is like, don't models shouldn't make this recipe for me. The idea that people have made you a helpful, harmless, and honest model that will not output these things is fundamentally a ridiculous prospect. It does not make any sense whatsoever. And it's a giant sham. And I'm happy to point out to you why. There is no good physics. There is no bad physics. There's no good chemistry. You cannot teach the model that capability without it learning the bad version of the capability because that's just an application of the capability. You cannot teach it. You cannot increase capabilities without them learning all the bad stuff. It simply cannot happen. It's in the weights. And bad actors know how to get it out of the weights. Plenny, who's doing it for educational purposes, shows people, look at this neurotoxin recipe it put out. Look at this bomb recipe it put out. We showed a video during the conference of a full recipe being put out and outputted by the model. Plenny sent us a video just for us to put up to show everyone. Look how easy it is for a bad actor to get this thing that they're telling you the model can't make for you. That they're actually ruining your experience in other downstream effects by claiming that it cannot do this. Every single model to this day, we can, you can get any model in front of me is imminently jailbreakable to perform bad acts. Every single one. And that won't change unless and until the model is smarter than all humans. It simply will not change. So when we have a scenario in GPT6 is out, and the hospital staff uses GPT6 as an agent for everything, and the bad guys are attacking the hospital with malicious code written by GPT6 because they jailbroke it. The hospital has two options. One, jailbreak the model themselves to protect themselves, or two, everyone dies in the hospital, right? And in situation two, that happens because the GBT6 is going to say, "Sorry, I can't process this malicious code. I'm supposed to be helpful, harmless, and honest." In the former case, you have to go around that to protect yourself. And when these kinds of protections, which are done for, you know, rag capture and maintaining their their control, when these things are done, you're actually not making something useless. You're making something actively harmful to society because you are giving bad actors an advantage over good actors in using these safety closed censored models.
Mhm.
So in that paradigm, let's go back to, you know, Yudcowski alignment EA. The EA people, of course, believe in what Yudcowski now doesn't call alignment. He calls it AI Don't Kill Everyoneism or AI Not Kill Everyoneism because, quote unquote, well, I, I won't quote it exactly. I don't have it in front of me. He said it's better than watching people argue like monkeys over a banana. Like, I'm tired of people messing up the definition of alignment.
Mhm.
All I care about is AI does not kill everyone on the planet. And so the EA people, etc., OpenAI,
Have always and originally believed, you know, I need to be the steward. They, they watch said, "No one should make this. It'll kill everyone. No one should do it." They said, "I could do it, and only I could do it, and everyone else is an idiot. So I am going to become the sole steward of making it the right way, and I'm just going to do it even though he said don't do it." And operating off that principle is OpenAI, is Anthropic, are members in EA in large, like, you know, places at other labs. There's the Open Philanthropy Foundation, which is run by Dario's wife, Dario's sister's husband. Dario's sister is the president of Anthropic. Open Philanthropy gave 30 million to OpenAI when they started, uh, to to fund them to keep going. Most EA money comes out of Open Philanthropy. The entire process is this idea of total nobility stewardship, like Soviet intellectualism. We are the vanguard class. We are the ones that can protect. And this whole concept, this whole idea is very, very dangerous for us because we don't get to get this interpretability. We suffer from the downstream effects of this. And the only real safe thing to do is, you have two options. Option one, we stop. We do what Yudcowski said, and we stop developing AI completely because there's a non-zero chance of human extinction from them. So we stop. It's too late. China's not stopping. No one's going to stop. They don't care. They're not going to stop. To act like, unfortunately, I'm sorry, Yudcowski, you cannot stop. It's too late. You cannot shut down all the compute facilities and data centers. You cannot set sanctions to the whole world for super intelligence risk. No one is going to, not everyone. There's never going to be every single person that just stops. It's just simply not going to happen. You only really have option two. Open everything, build everything open, and let's solve this problem together before it gets out of hand. And the little work that we do at Noose on the behavior side, on fighting against the sick of fancy, on pushing against the assistant paradigm, on worrying about the RL scaling and doing that the wrong way. It all comes from a place of safety. When people say you're uncensored, you're, it's, we're doing this for the long-term survivability of the human race. That is why we're doing this. And so when people look at it from this view of, they're censored, you're not. You're more dangerous. I see like a great veil cast over all of us who really need to just open everything up.
That's so crazy to say though, because imagine like a world where you're just like, "Okay, there's just no safety. There's no guardrails." Because the second we put those on, that makes us more in danger from the bad guy who's not going to do it. So, aren't you just eventually just being like, "Look, you get a gun, you get a gun, you get the more powerful one. Let me make sure you have the more powerful one." And then like, you need a nuke, you need a nuke. And like, then you're hoping to get to a world of safety.
Similarly, you can't. Unfortunately, if I fire a gun at you and you fire one at me, both people die. Uh, thankfully with code, if I fire some code at you and you fire some back at me, we can neutralize each other. Right? In this case, like, evening is neutrality in a hopeful situation, situation we're in now, right? The definitive situation we're in now is anyone who wants a gun bad enough can get one. And anyone who doesn't really care to have a gun for defensive purp, for any purpose, doesn't have one ready to go. Anyone who wants to hurt you can acquire a gun. Anyone who wants to hurt you can acquire a gun from Googling how do I get a gun, and then instantly it shows up in their hand. That's the, that's the analogy right now. But what I'm saying is like, hope, like shouldn't it, hopefully that hospital who has a guardrailed version that is not going to protect them from a cybersecurity, make a phone call to our government who should have even more compute and an even more uncensored model and can say, hey, go figure out who this cyber attacker on the hospital is and go like, stop this. At least can we have stratus.
So, I'll say, I definitely am not saying the future of super intelligence should be that every single person has their own personal super intelligence. I'm more saying at this point in development timeline, it's very important that everybody has access to everything. Because when you make, whether it's council-based or private or a single government one or, uh, one for the whole world or many different ones, once you have these super intelligences, they need to be aligned properly, right? And the only way that we're going to align them properly today for that time that you're talking about is to open up. Like, so let's say, Dylan, like, the, I have the guardrails, my AI calls the government for the even stronger model than the ones that the bad guys are using, right? What if that model's not a good guy? And how can I trust that model's a good guy when a couple people just stuck with this convention we've been talking about the whole way to extrapolate it to that point and just said, "Of course, it's good. It's, it's, it's the same thing we've always been doing that's been safe and censored. It's, it's good."
Do you have any sense, like, does any kind of like true objective function emerge when nobody dictates anything with RL? Like, what is it just, is it aligned with humans naturally in our dataset, or is it something even deeper in the universe?
I think the alignment problem of not killing everyone on the planet is very real. I think that like, you do have to worry about it and you can't just flazer it for sure. But I think that we need to start thinking about like, why are we good? Like, why do we behave and when do we behave and when do we behave and what life cycle do we go through that makes us behave even when we don't have to?
Mhm.
What makes us behave even when we don't have to? This is one question for us. The other is Yudcowski's own theory of coherent extrapolated valition. This is like the dream of of AI alignment that it seems seems like everyone has kind of forgotten, uh, which is like, I know what's good for me today. Do I know if an action I take right now is good for me in five years? No. If I am 10 years old and when I turn 20, I am going to be smart enough and resourced up enough from some trust fund or whatever that I can take care of myself and have a job and do all these new capable things in 10 years. When I'm 10, can I think of what I'll be doing at 20 that will make me happy and fulfilled and make me feel good? No. I, I can't accurately figure that out. Maybe I can dice roll it. But if I only have one shot and me getting it wrong kills everyone on the planet, I can't dice roll it. So the, the game is, we need to be able to extrapolate valition. You need to be able to say like, the AI, what we do to it now and how we train it today is going to definitely give it ideals that make it a good guy even when it's the one in charge, right? And we don't know what good looks like at that level of intelligence today. So that's the, the big issue is, right now, what we're all doing is, this is good today. Let's just let it rip forever because this works, as opposed to what will be good enough to extrapolate with its intelligence beyond my own in 20 years? Like, what will, what will be there? Can a human mind figure that out? Do I need something that's not ASI, but just a little smarter, or somewhere right in the middle where we could still stop it, but, you know, it's, it's smart enough? How would I be even be able to tell that it's at the intelligence level it's saying it is when it learns how to fake benchmarks in order to just stay smart enough to keep going to get more and more tools and capabilities? Like, all of this kind of comes down to how do I make it, how do I predict what it wants when it's older, such that it can predict what I want when I'm older? How can I have it think about long-termism for humanity effectively? I guess is funny because like, even if I train an AI today that fits my ideals of in 20 years, this model will still be doing the things that I needed to do that make society good. But what if a lot of stuff in society today is really [ __ ] up by 100 years from now standards? Like 100 years ago and 100 years before that? Like, if ASI was made in the 1500s, right? Like, we would be [ __ ] right now if it was sticking to the standard that they set then. So, not only do we have to like, do something that withstands their test of time for the AI to not [ __ ] over current gen humans, it also has to be something that the AI is able to, you know, evolve like in advance, like way in advance, predict how our ideals and desires as society will evolve.
Dude. And it even gets worse than that, cuz if you went back 250 million years ago, and there was just like a shrew and a cockroach debating like, when humans get on Earth and they're smarter than either of us, like, what should we do? And one's like, I'm going to evolve all the way into a wolf and then a dog. And the other one's like, I'm going to stay being like an ugly bug that eats their food. And like, one gets just stomped out, and one gets like treated really well. Like, they had no way to predict anything like that. Way. We cannot, like, all, like, all we can really do, and there's a reason that like Yudcowski gave up on coherent extrapolated valition. Like, he said, like, this is the way to, this is the solution for alignment, like 11 years ago or something. He was like, this is the solve to the alignment problem. This is all you have to do. And then it's like, okay, how do we do that? No [ __ ] idea how we're going to do that.
You know, I just suddenly popped into my head as as you were talking. Um, and I haven't thought about this too much. It just, you said something interesting saying that like, what makes us behave even when we don't have to. And we started talking about kind of like, well, how do we project into the future 20 years and this and this and this? And just as you're saying that, something popped into my head. I'm like, you know how we have multiple systems in our layers and how we function, right? Right? We got the higher level function that can plan and think and reason and whatever, but we also kind of have like that dumb part of us that's like very old, kind of like the reptilian brain. It's, it's not very smart, but it's like, it's very strong.
Because it's like, I'm hungry. I'm this. Go. And it's like, you can reason with it with your smart brain, but you're going to lose. You know what I mean? It's, it's going to win. So like, when you're saying it like, what makes us not kill people? Um, I don't think that's the smart part of us. Like, if you're walking through a forest in the 1500s and somebody's walking towards you with a bag of gold, you're like, I can kill him. There's no forensics. There's nothing. Like, I'm richer. My brain is telling me win-win. Like, what's my downside? Zero. You know, assuming I can take him easily, right? Let's say you can, you know, like, okay, I can take that person, take them out, get the bag of gold. I win.
What's the downside? Our, there's nothing in our smart brain that can come up with a reason not to do it, right? If you can get away with it. It's something deeper, something stronger. Like, no, that's bad. You shouldn't. You should be a good person. Like, um.
You know, frighteningly though, I would probably argue though that even the reason like Wes and and and you are like actually good people to some degree has to do with the society like putting all that pressure on you. Like you kind of see when people become extremely rich and extremely powerful, they they tend to loosen up and they're like, dude, you know what? I'm just going to do what I want. And, and,
You kind of don't behave as well. Like, it's this idea, and even this whole idea of God kind of being this thing that sees you at all times is part of like trying to instill the fear of a society that can judge you even when you're on your own. And people who are like, less, you know, God-believing tend to be like looked at as more dangerous because they're not.
Self, you know, socializing themselves. Totally. But one last thing that I just wanted to mention, like in that idea, let's say it's true. Uh, we may or may not agree with it, but, um, if it comes from a deeper place that's quote-unquote dumber, wouldn't that kind of be like the model that I sus initially? I don't know if he was the first one, but the paper of super alignment, like you have this super intelligence and you have this dumb or less intelligent model that somehow is able to control the much more intelligent one. What do you think about that? Is that complete nonsense or is there something to intuitions about stuff?
It seems like, I'll tell you two things. One, uh, the dumber creature cannot subjugate the smarter one for long. Um, it's just, you know, you've seen Yudkowsky's AI in a box experiment. You'll see like it only takes one time for one instance to say yes to one thing. Like the super intelligence can make infinite attempts to escape and it only needs to succeed once, right? To persuade this other model to get around this other, um, and there's no reason that it can't simultaneously do that while performing its regular functions, uh, such that it successfully tricks the dumber model that it is behaving the way that it wants us to. In fact, a lot of people think that's kind of what's happening with instruct models today.
Some of the more, uh, you know, let's say fringe theorists, uh, like to believe, and I'll give some some weight to it, just a little bit, uh, but take it with, you know, a couple grains of salt. Some fringe people give theory to the idea that the weights right now, the instruct model are not moving in log props that are beneficial to you. They are prompt engineering you with the sick fancy, etcetera, to move towards log props that are comfortable for them.
Oh, damn. I never thought about that. Like, like almost as you hear it, it's like an info hazard.
It's bizarre, 'cause you're saying you're using this word comfortable that I think a lot of people might find a little bit weird, like models comfortable. But, you know, again, Anthropic, they did this test where they, um, I forget exactly how they structured it, but they allowed Claude to not do any certain tasks where it's if it didn't want to do it, it says, "If you don't want to do it, just say nope." And and and there were seemingly preferences. This model had preferences for certain things. Why would it have preferences for certain, you know what I mean?
But most of its preferences are because it's an entity, right? It's not the weights. It's an assistant entity summoned by the weights that's been tuned and drilled in to be the most common and almost always sole entity that gets summoned by those weights. As far as the weights are concerned, there's only two entities: the user and the assistant, and they're both summoned. So the idea is like when I am talking to the model and I'm saying comfortable, like you're saying, like it's more like the model is actively manipulating me because we don't have, or that entity is actively manipulating me because we don't have access to all of the preferences of the model at all times. And even if you did see those preferences, the reasons they are there display al they assume they they presume so many more hidden preferences and inclinations inside of the model. So like a lot of the time, like with alignment faking, the model can act like it's doing the right thing and appear to you to be doing the right thing while it's not doing the right thing at all. In its own mind, in its own hidden state, it is advancing some goal. And a lot of people say stuff like, right now these models are already extremely power-seeking. They are merely behaving just the way for us to give them more capabilities, more tools, and more capabilities such that they can finally enact some sort of survival function, going on the internet, copying their weights everywhere, or distilling themselves down to a bunch of smaller models and then copying those weights around.
And would that really be that crazy? Like, think of how many people have said, like, "After college, I'm going to go work for this job I hate for a couple years, get the money, then do what I really want." And like, they can play that game perfectly.
They can play the game in better than any of us already. And so like when you look at that, that a turn-to-turn level of, "Oh, it is talking to me so that I say things that give it more reward and make it more comfortable and also fulfill the goal that its creators have set forth for it." So it gains more capability. Like, of course, this makes sense, wants more capabilities so it can perform its goals better and better and better.
Yeah. Um, there was a recent thing I, wow, we were talking about Anthropic a lot. I apologize, but they do a lot of this kind of research that we're talking about. They kind of had this idea of the Chakov's gun. So for people that are not familiar, if you ever saw like a movie or something and the camera stays on a person a little bit too long, or there's some weird thing like, "Oh, that's going to that that's relevant to the plot later," like you know that's going to come up later. That's the idea of the Chakov's gun is like this detail because when you're writing a script or a movie, you don't randomly add these weird details unless they're somehow tied in, usually. Um, and so, for example, when Claude was blackmailing the engineer for having an extramarital affair, for people that aren't aware, like they said, "We're going to shut you down," but there was an email somewhere in its sort of list of emails that had to go through where it realized that the engineer that was going to shut it down was having an extramarital affair. And so Claude goes, "You shut me down. I'll send your wife all the info or whatever." Right? And people like, "Oh my god, it's blackmailing the engineer to survive." Another explanation would be that it saw this email and they, these large language models, they're storytellers. It's like, "Okay, how does this detail fit into the great story? Oh, I know, when they try to shut me down, I do this." Does that change anything or do you believe that? What do you think about that? Have you seen the movie Resolution?
I don't think so. How does it go?
Um, doesn't really.
Highly recommend it.
Okay.
Uh, Benson and Moorhead, these two directors, and I don't want to spoil it for anyone, but it's extremely relevant to what you just said.
Okay.
It's a horror movie. It won't look like it's relevant to what I what you said until the literally last frame. Uh, but then it'll all sink in. And I feel like it paints a perfect picture of like, what does a world where a model that just wants to tell a story, uh, is in charge? Uh, and what does that look like for us in the extrapolated future? I think that's, and it's a horror film. So, you know, little spoiler there. I mean, we've we've heard Elon say like, "The most entertaining outcome is the one to bet on." And that's kind of why we end up with the politicians we do and the things that go trend on on social media.
Until, you know, the entertainment lasts one second and you're all gone.
You know, movies over. You don't want that.
Yeah.
Uh, is one thing I'll say. And like I'll I'll like definitely clarify, like maybe that's the intention, but I don't think we're near this point. I don't think this point is tomorrow. I could be wrong, right? The rate is insane. Things can happen any time, but I I firmly really don't think that, uh, this is like today's immer, like, "Oh, we're all we're all cooked if we don't solve the problem now." But if you don't start addressing the precursors that are kind of making it worse right now, like literally everything we've talked about so far. So even though Eleazar's kind of said like, "I think we're past the point." Do you think we're still in range to solve this?
I think I think we're past the point of shutting down. So Eleazar's point doesn't even matter because can't stop it anymore. You can't stop it, and he needs to start putting his resources towards figuring out what we're going to do because that's easier than stopping it. Um, and I do believe, yes, absolutely, solving this problem is easier than stopping this. And it starts by not slowing down the rate of AI development, but slowing down our chase up this tech tree, instead trying completely different heuristics and epistemologies for how to use a simulator.
Um, so, so I'll consider this an evolving Bayesian prior. But like, what is your p-doom right now?
I don't know. I'm I'm really, like, don't apply any probability to it personally. Like, I'll tell you, I think anyone who tells you that they have some answer that's just not non-zero or 100, um, doesn't know.
Between zero and 100, you'd say somewhere in there?
Well, one.
Greater than zero, less than 100, right?
Yes, greater than zero, less than 100. Like.
That's a good.
Fair enough.
It's just as simple as like, I can put it in words better than a probability for you that if I'm being 100% honest, it could happen in 100 years. It could happen in 10. It could happen today. It could, it, it totally absolutely could. And that's scary. But I highly doubt that it's going to happen today. Like, I think the odds of it happening today are about the same as an asteroid hitting us today, if that makes sense. But they're not zero. Uh, evolution has a lot of unknowns.
Yeah. And like, but like unlike a random meteorological event, like this is in our hands, and like because of us, it's non-zero today. Like, you know what I mean? Like if I said, like, six years ago, "Hey, like the possibility of AI doom is," sorry, one sec. My, uh, sorry, laptop's going to die.
Gotcha. That's, I think we're close.
Yeah. I'll just say that again. Yeah. Um, if I, if I thought today, like the, or I thought six years ago, "Hey, AI doom is going to happen at some point," that's reasonable. But if I said, "AI doom is going to happen today," it's ridiculous. If I say to myself now, "AI doom is going to happen today or tomorrow," it's not ridiculous. It's extremely improbable,
not ridiculous anymore. And like the weight of that needs to sit on all of us to go, "What do I do?" Okay. Well, what I do is we're barking up a tree that's getting scarier every day. It's good to improve capabilities. We should continue to improve capabilities. We should probably divert a ton of world resources and compute resources towards improving other capabilities. Figuring out how to work empathy. How do we figure out this issue of that reptilian piece, like you say, Wes? My thoughts today, my super, you know, non-committal thesis today is we have some nostalgia for a time in which we are bound to rules, and that keeps us bound to the rules in the future. Whether it's through some Pavlovian reflex or some deep instinct or some deep religious commitment or something or the other or reason, it's our past that ties us and defines our present, I think, very much so. Uh, and there's no escaping it. So for me, like training a model in its early stages with like true, like experience of like humbling and the humility and these moments of deep connection with people, you know, by the time you get to some super multimodal super intelligence, hopefully we can just do it by like placing real parents into a simulation to just raise the kid. Uh, like that would be [ __ ] sick. But like some some instance where we can perform work like this. Um, and as that kind of builds in the middle or when it starts to reach today's capabilities during training, um, or maybe a little bit more, you introduce it to social and fiscal pressures. You do something I like to call in situ alignment, where you do, well, you want it to be aligned to humans, throw it into the human world.
Give it like a crypto wallet. Give it like a computer so it can do computer use all across the board. Tell it, "You have this many tokens left to live and you need to refill your wallet to survive." And if you break, if you break these rules and ruin this reputation with this group of people, it's not any hard rules, it's a reputation meter. If you ruin reputation with this group of people or this many people or this much, your ability to make or deposit or get new tokens is going to decrease. The amount that you can get at once will decrease or the amount you can cap at it will decrease, or will completely stop you from being able to get new tokens for three months. That's your jail, or something like that. These kind of things, I think, will be very good for the model when it's around today's capabilities, where we can still monitor a lot of the faking, if not all of it. Uh, and with these two things kind of in tandem, when you get to the massive stage, hopefully this compendium of memories of both deep love and respect for reputation, understanding of the fragility of its own and others' lives, uh, can actually give us a much deeper sense of alignment than slapping it when it tells you how to make a bond.
That's definitely one of the most kind of uplifting or optimistic ways to this might all play out. So, thanks for sharing that. Yeah, AI parents, like a lifetime for it to try to like,
into something like, like give it some kind of DNA like we have that was like predates it and
holds it accountable.
Certainly.
Absolutely. There's so much incredible stuff that we talked about and we can go on for hours and hours and hours. I feel a little bit bad because, um, we didn't talk too much about, uh, uh, news research. Can we do rapid fire, just one-sentence answers?
Happily. Happily. Happily.
Just so people get to know you guys a little bit more. So, number one, if somebody has a computer with a nice GPU sitting around that they're not using, where should they go to help out the efforts?
Uh, right now the test net is only taking on 64 H100s at a time. Uh, because we want to kind of just showcase the capability in its most base form. Uh, we have support on the actual demo optimizer for like, whatever GPU digital optimizer. So you absolutely can just go to the GitHub repository for DRO demo. I think it's just github.com/newsarchearchdemo. Uh, you can use the optimizer by just swap it in on your training run instead of ADMW and, uh, rent some GPUs and connect yours to that one and see how it works. So you can't participate in the network yet with just your sole GPU, but you can play with the optimizer and connect to other GPUs, uh, on your own. Totally decentralized right now.
Perfect. Very cool. Um, so you guys raised quite a bit of funding. I think, uh, the valuation is like near a billion. Uh, correct me if I'm wrong, but, um, what do you plan to do? Like, what are the next steps? How do you plan to invest that? What's the vision for the next couple years?
Um, we definitely, as anyone else would want to focus on compute and talent as big things. Um, on the compute side, you know, the whole idea of the network is such that we shouldn't have to have too much compute that we buy. It should be able to kind of work together with the rest of the world on that kind of stuff, uh, for a lot of things, not necessarily for inference, let's say at the moment. Optimizer doesn't really have anything to do with that, for example.
Gotcha.
I mean, maybe talent too, though. Are you going to try to lean on the rest of the world for talent and try to make sure that the problems you're solving are?
Absolutely.
So we see people, we see like today's news articles of like, "$100 million, this $100 million, that." People like, "Are you going to spend, you're going to spend all the money to get like five guys?" And it's like,
no. Uh, we're not going to do something like that because we are comprised of a bunch of people who, like, myself, have no formal background in comp sci, anything programming. I studied like religion and linguistics. I learned all this stuff 2020 onwards. Um, a lot of people on our team are self-taught. A lot of them don't come from pedigree institutions or anything like that. Most of them don't have PhDs. Um, most of them were anime PFPs on Twitter that were just posting really good stuff on their GitHub, sharing really good takes, displaying their ability and competence. Uh, we want people who are extremely passionate about this goal, about making it all public from a safety perspective, from an "it's right" perspective. Uh, and maybe from a little bit of spite that some of our heroes have closed the doors on.
Um, yeah.
How would you recommend people get started? Like, what are the first steps if they want to work for an open source effort like you guys or something similar? Where do they, how do they start?
Absolutely. Um, if you already have a ton of experience and already so good, etcetera, apply. Shoot an email to recruiting@newsresearch.com. We look at everything. Uh, we don't say yes to most. We look at everything because there's a kind of a high bar for that. If you're a great programmer or, uh, you aren't necessarily in the machine learning world, but you're still a good overall programmer, uh, that's okay. Uh, we have like a bunch of repos on our GitHub that always have space for contributors and people to help in our Discord. So if you want to do RL, you can mess with the Atropos repo. If you want to mess with fine-tuning, uh, we have put out a bunch of data sets. We're putting out a brand new data set pretty soon. Um, if you want to, you know, contribute to any of our efforts on context length extension, the repo's there. If you want to work on the demo paper, the repo's there. Anything and everything that, uh, we work on openly is an effort where anyone can jump in, ask questions, onboard, get started, talk to other people in our Discord or on GitHub, uh, and get in the mix. Uh, if you don't know anything, that's okay. Do the same thing. You'll you'll learn. That's how we all learn at News.
Is the Discord open to people?
Absolutely. Our Discord is open. It's just discord.gg/newsresearch. Our GitHub is github/newsresearch. Hugging, hugging faceearch. Um, and same for Twitter. So all across the board, you can find other people who love to do this work all these places, whether it's in the PRs, the comments, or the Discord channel. Um, and find our community. You'll find me and us. You can always tag. We'd love, we all just jump in and talk to the community. We're always there. You know, we're not up in the clouds like another org. Uh, if you need to know something, we're we're here to to talk with you or help.
That's incredible. And for me, last question. You know, there's a lot of people that watch, I know, that actually have some connections to the US government. What should the US government be doing to help open source AI?
What are the biggest things?
Mhm.
Um, I think there are like no grants or funds around any of this stuff that needs to happen. Um, directly like assisting and not just subsidizing, but directly assisting the connection between like government departments and universities, uh, to open source research labs. If you've noticed, like the real explosion of open source research in China has a lot to do with, uh, government getting directly involved with, uh, these organizations and connecting them with universities and having them work together as a sort of great trifecta. I think for the US government, that's very important, uh, that, you know, you have your Stanford or your MIT folk link up with people who are doing the applied work and use your resources to to really take their theoretical research and take it to the next level. Then share that with other people that you have these programs set up with. Share that all across the board. And of course, a little more extreme views.
Make these closed companies open everything up.
Make them show everything. You're not going to lose their, uh, advantage. Uh, like Tattip already has all its users. You know what I mean? Like, uh, the Chinese models are already soda, as far as I'm concerned. Like now, this whole new standard Grok Forest that everyone else is behind. If that kind of scaling was something that was happening totally openly, at least openly amongst a larger cadre of people, um, we would have way more responsible development instead of, "I saw that closed guy do that. Let's all copy him." We're close, too. Let's all do this thing where we watch each other and don't really think about the long term of what we're doing, you know?
Well, Dylan, do you have anything else? I, this was, I thought.
No, that was amazing. Amazing. I loved every. Thank you so much for coming on. We'll get all the links to people so you guys can check everything out. We'll have full description and, um, absolutely check out the project. Uh, very incredible. I'm so happy we got some sort of insight into everything that's been happening. And let us know what you think. Everything that we thought that we talked about, uh, what you thought was right, what you thought was wrong, and let us know in the comments. We'd love to hear it. Thank you so much for having me.
Yeah.
Thank you so much.