Transcription
You pulled the 9B because it fits your laptop. But a comment under my last video nailed the catch. The model fits, the context for your session likely won't. So, let's check it.
The model file is 5.6 GB. A short chat adds a quarter GB of cache. It fits. That's the rent. Because agents don't chat, they accumulate. Watch what one session loads. A file, 12,000 tokens, a diff, three more. Test output, seven more. Real sessions blow past 50,000 and long ones push 128,000.
Here's the cache math. The model's config prices every token at about 32 KB of KV cache. At 128,000 tokens, that's 4 GB. If your sessions crawl, this is why. Weights plus cache plus overhead, be kind. About 10 GB max the context and the cache alone hits eight. And that's the efficient case. Just eight of 32 layers grows the cache. A dense model needs 16 GB. The benchmark ran at 256,000 tokens.
But the cache isn't the sneakiest part. Stop blaming the model for forgetting. It supports 262,000 tokens, but Ollama's default is 4,096, buried in the FAQ. It was never remembering your files.
So, here's what actually works. Turn on flash attention and one environment variable will store the cache at Q8. Half the memory, very small precision loss per Ollama's own FAQ. 4 GB of cache becomes two. A 16 GB machine is back in play. Size the context to your RAM. Quantize the cache or cap the context. Which one are you running? One word below. And subscribe for more local model deep dives.