125 Billion Parameters in 32GB RAM: Benchmarking Sushi, Qwen, and Thinking Tokens on Apple M4

Let’s pause for a moment to appreciate the sheer, glorious absurdity of local AI in late 2026.

Just three years ago, if you wanted to run a model with over 100 billion parameters, you needed a server chassis the size of a mini-fridge, an electrical circuit that could power a laundromat, and a bank loan to pay for four NVIDIA A100 GPUs. If someone told you that you would soon run a 125-billion parameter reasoning model on a standard 32GB consumer Mac, you would have politely recommended they seek medical attention.

Transducers and Reducers, Finally Explained (For People Like Me)

Let’s be completely honest with each other.

If you have spent any time around Clojure, functional programming, or modern Lisp communities over the last decade, you have almost certainly encountered people speaking about Transducers and Reducers in hushed, reverent tones—as if they were forbidden alien technology recovered from a crashed UFO in Roswell.

You open the documentation, and you are immediately greeted by sentences like:

“Transducers are composable algorithmic transformations decoupled from their input or output sources. A transducer is a function that accepts a reducing function and returns a new reducing function.”

10ms Decision Intelligence: Running Julia-1 on Apple Metal with Coni's Native CSP Engine

A week ago, I wrote about how Coni’s native CSP meets Ollama 0.35’s System One API. The premise was simple: traditional generative LLMs are excruciatingly slow when all you need is a routing decision. By replacing token-by-token autoregressive generation with single-token classification models (like nimble), we achieved instant decision-making and orchestrated massive agent swarms using Coni’s Go-backed Goroutines (spawn) and channels (chan).

And then, SupersonicLabs dropped Julia-1.

Julia-1 is not just another decision model. It is an architectural leap forward: a 144-million parameter bidirectional decision model based on ModernBERT (mmBERT-small). Instead of forcing causal left-to-right attention onto a classification task, Julia-1 evaluates entire prompts bidirectionally with specialized marker tokens, pre-norm transformer heads, and learned type embeddings.

The Ultimate AI Superpower: Coni's Native CSP Meets Ollama 0.35's System One API

If you’ve been working with LLMs recently, you already know the pain of conversational generation: it’s inherently slow. Waiting for a model to auto-regressively spit out tokens one by one is fine for writing poetry or generating code, but when you’re trying to build autonomous agents that need to make thousands of logical routing decisions a second, traditional chat models are a massive bottleneck.

Enter Ollama 0.35 and the System One API.

Building a Zero-Dependency AST Profiler in Pure Coni

Performance optimization is critical, but profiling dynamic languages often comes with a massive caveat: runtime overhead. When dealing with interpreters, developers usually resort to mocking or wrapping functions dynamically at runtime using constructs like with-redefs.

However, when you need your profiler to work flawlessly in both a Native interpreter and Ahead-Of-Time (AOT) compiled WebAssembly environments, dynamic runtime rebinding simply isn’t an option. AOT compilers resolve functions statically.

So, how do you profile an application without modifying the compiler itself or suffering runtime penalties?

True AOT: Compiling Lisp Shaders directly to Wasm-GC

In our quest to build a pure, high-performance Lisp ecosystem for the browser, we’ve hit a monumental milestone. Today, we’re showing off our new Infinity 999Hz WebGL application, and more importantly, how it’s completely Ahead-of-Time (AOT) compiled into a native Wasm-GC binary.

/infinity-999hz.png

As you can see in the screenshot above, the app features a mesmerizing, audio-reactive particle swarm with dynamic magenta and cyan hues, complete with a reactive UI slider to control the color frequency. But the real magic isn’t just on the screen—it’s how the screen is being drawn.