Transcription
Nvidia CEO [music] did not sleep last night. Because one developer in Italy just proved you do not need his GPUs. He took GLM 5.2, the 744 billion parameter model that beats the ones you pay for, and ran it on [music] a laptop with 25 gigs of RAM. Zero graphics card. And he gave the code away.
So, welcome back guys. This is day 159 building you 100 X. [music] The project is called Colibri. Zero dependencies. Only about 10 gigs of that model is actually thinking at any moment. The rest is 21,000 experts just sitting there waiting to be called. So, he keeps the thinker in your RAM and streams the experts off your SSD on demand. Like Netflix streams a movie instead of downloading it.
Now, here is how you run it. Step one, clone the repo, go into the C folder, and run setup.sh. It builds itself. No Python, no CUDA, no PyTorch. Step two, download the model, GLM 5.2, the INT4 version. 370 gigs on your drive. Step three, point Colai at that folder and type Colai chat. And boom, you are talking to a 744 billion parameter model on your own laptop, [music] offline.
Step four, type Colai serve instead and it turns into an OpenAI endpoint. Change one line in your code, the base URL. Same app, zero token bill. Nothing ever leaves your machine. Nvidia sells you the GPU. One developer just made your hard drive do the job. And he gave it away for free. Comment AI and I will send you the GitHub link with the full install guide. Follow for more such videos.