📱

Get Our Mobile App

Take your business learning on the go!

Download on the App StoreGet it on Google Play

Why bad memory is good for intelligence – Andrey Karpathy

Dwarkesh Clips2:34

Transcription

The best learners that we are aware of, which are children, are extremely bad at recollecting information. In fact, in the very earliest stages of childhood, you will forget everything. You're just an amnesiac about everything that happens before a certain, uh, year date. But you're like extremely good at picking up new languages and learning from the world. And maybe there's some element of like being able to see the forest for the trees.

Whereas if you compare it to the, uh, opposite end of the spectrum, you have LLM pre-training, which these models will literally be able to regurgitate word for word what is the next thing in a Wikipedia page, but their ability to learn abstract concepts really quickly, the way a child can, is much more limited. And then adults are somewhere in between, where they don't have the flexibility of childhood learning, but they can, you know, adults can memorize facts and information in a way that is harder for kids. And I don't know if there's something interesting about that. I think there's something very interesting about that. Yeah, 100%.

I do think that humans actually, um, they do kind of like have a lot more of an element compared to like seeing the forest for the trees, and we're not actually that good at memorization, which is actually a feature. Um, because we're not that good at memorization, we actually are kind of like forced to, uh, find the patterns, uh, um, like more in a more general sense. I think LLMs, for in comparison, are extremely good at memorization. They will recite passages from all these, uh, training sources. Uh, you can give them completely nonsensical data, like you can take, um, you can hash some amount of text or something like that. You get a completely random sequence. If you train on it, even just, I think, a single iteration or two, it can suddenly regurgitate the entire thing. It will memorize it. There's no way a person can read a single sequence of random numbers and recite it to you. Um, and that's a feature, not a bug, almost. Uh, because it forces you to like only learn the generalizable components, whereas LLMs are distracted by all the memory that they have of the pre-trained documents, and it's probably very distracting to them, uh, in a certain sense.

So that's why when I talk about the cognitive core, I actually want to remove the memory, which is what we talked about. I'd love to have them have less memory so that they have to look things up, uh, and that they only maintain the algorithms for like thought, uh, and the idea of an experiment, and all this cognitive glue of, um, of acting. And this is also relevant to preventing model collapse. >> Um, let me think. Um, I'm not sure. I think it's almost like a separate axis. It's almost like the models are way too good at, uh, memorization, and somehow we should, we should remove that. And I think people are much worse, but it's a good thing.

If you enjoyed this clip, you can watch the full episode here and subscribe for more clips. Thanks.