Transcription
Gemma 4 launches keep coming. Yesterday a 12B model. Today Google dropped QAT checkpoints for the whole lineup on Hugging Face. And one number stands out. The smallest model now fits in under 1 GB of memory.
Here is how they did it. QAT means quantization aware training. Normally you train a model in full precision, then squeeze it down afterwards. That squeeze cost quality. Google baked the quantization process directly into training instead. The model learns to perform at low precision from So quality holds in the popular Q40 format and in a custom schema built for mobile chips.
That mobile schema is aggressive. Token generation layers get compressed all the way down to 2-bit. Activations use pre-calculated scaling, so the phone's processor does less work. The KV cache is optimized, so long conversations stay light. Put together Gemma 4 E2B now runs in under 1 GB.
And this covers every Gemma 4 size. The 2-bit and 4-bit. The new 12B. And the 26B mixture of experts. Plus their drafters for faster decoding. It runs where you already work. Ollama, llama.cpp, and LM Studio on desktop. MLX on Apple Silicon. vLLM, Transformers, and Unsloth for developers. Even in the browser through transformers.js. Full model quality, a fraction of the memory.
So here is the question. Do you run your models locally or is the cloud just easier? Tell me in the comments.