125 Billion Parameters in 32GB RAM: Benchmarking Sushi, Qwen, and Thinking Tokens on Apple M4
Let’s pause for a moment to appreciate the sheer, glorious absurdity of local AI in late 2026.
Just three years ago, if you wanted to run a model with over 100 billion parameters, you needed a server chassis the size of a mini-fridge, an electrical circuit that could power a laundromat, and a bank loan to pay for four NVIDIA A100 GPUs. If someone told you that you would soon run a 125-billion parameter reasoning model on a standard 32GB consumer Mac, you would have politely recommended they seek medical attention.
