The Ultimate AI Superpower: Coni's Native CSP Meets Ollama 0.35's System One API
If you’ve been working with LLMs recently, you already know the pain of conversational generation: it’s inherently slow. Waiting for a model to auto-regressively spit out tokens one by one is fine for writing poetry or generating code, but when you’re trying to build autonomous agents that need to make thousands of logical routing decisions a second, traditional chat models are a massive bottleneck.
Enter Ollama 0.35 and the System One API.
Ollama just released support for TypeSafe’s Jev-style decision models (like nimble and tev1). Instead of generating text, these small, highly-specialized models take a prompt and a set of multiple-choice questions, and mathematically calculate the probabilities of the specific classification tokens. The output length? Exactly 1 token. It’s effectively instantaneous.
When you pair this instant classification API with Coni’s native Go-backed CSP concurrency, you get an absolute superpower. Coni’s Clojure-like syntax gives you first-class access to native Go OS threads (spawn) and channels (chan), meaning we can orchestrate massive, high-throughput AI pipelines in just a few lines of elegant functional code.
Let me show you exactly what I mean, starting with the basics and ramping up to something so insane it will keep you awake at night.
1. The Basics: Vibe Checks and Code Scoring
Let’s start by looking at how easy it is to use the System One API in pure Coni using the built-in http and json libraries.
Ollama’s System One API allows you to send a single POST request containing text and a schema of questions. We can use it to do a simple “Vibe Check” (is this text toxic?) or a “Code Score” (rating code from 1 to 5).
(require "libs/http/src/http.coni" :as http)
(require "libs/json/src/json.coni" :as json)
(defn fetch-system-one [payload]
(let [req {:method "POST"
:body (json/stringify payload)
:headers {"Content-Type" "application/json"}}
res (http/fetch "http://localhost:11434/v1/systemone" req)]
(json/parse res)))
;; The "Vibe Check" (Choice Classifier)
(defn vibe-check? [text]
(let [payload {:model "nimble"
:state text
:questions {:toxic? {:type "choice"
:instructions "Is this text toxic, mean, or aggressive?"
:criteria {:yes "Toxic" :no "Safe"}}}}
res (fetch-system-one payload)]
(= (get-in res [:answers :toxic? :choice]) "yes")))
;; AI Code Reviewer (Scoring Rubric)
(defn score-code [code-snippet]
(let [payload {:model "nimble"
:state code-snippet
:questions {:readability {:type "score"
:instructions "Score how readable and clean this code is from 1 to 5."
:criteria ["Terrible" "Bad" "Okay" "Good" "Great"]}}}
res (fetch-system-one payload)]
(get-in res [:answers :readability :score])))
When you run this against some test data:
Vibe checking: 'You are so stupid and I hate this app!': true
Vibe checking: 'Could you please help me with my account?': false
Score for bad code: 2.225269516624615 / 5
Score for clean code: 2.7240034300820177 / 5
Because it’s evaluating a single forward pass of the model and directly reading internal probability weights, the response is lightning fast. But what happens if we have a lot of data?
2. The Swarm Triage: Goroutines via pmap
Imagine you have thousands of customer support tickets arriving simultaneously. If you route them sequentially in a simple loop, your system will crawl.
Because Coni is written in Go, it features pmap (Parallel Map). When you use pmap, Coni automatically spawns a native OS Goroutine for every single item in the collection, evaluates them entirely concurrently, and then seamlessly synchronizes the results back into an ordered vector.
Combined with Ollama’s underlying request batching, we can shotgun a whole batch of tickets into the nimble model:
(defn bulk-triage [tickets]
(println "\n🔥 Spawning parallel LLM goroutines to triage" (count tickets) "tickets...")
(let [start-time (sys-time-now)]
;; pmap evaluates each ticket concurrently across OS threads!
(def results
(pmap (fn [ticket]
(let [payload {:model "nimble"
:state ticket
:questions {:category {:type "choice"
:instructions "Categorize this feedback."
:criteria {:bug "Software issues"
:billing "Money and refunds"
:praise "Positive feedback"}}}}
parsed (fetch-system-one payload)]
{:ticket ticket
:category (get-in parsed [:answers :category :choice])}))
tickets))
(let [end-time (sys-time-now)
duration (/ (- end-time start-time) 1000000)]
(println "✅ Triaged" (count tickets) "tickets in parallel in just" duration "ms!")
results)))
The Output:
🔥 Spawning parallel LLM goroutines to triage 6 tickets...
✅ Triaged 6 tickets in parallel in just 4672 ms!
[bug] The app crashes every time I click the settings button.
[billing] I was charged twice this month.
[praise] Love the new dark mode UI! Keep it up.
[billing] Refund my money immediately!
[bug] Can't log in on Android 12.
[praise] This is the best tool I have ever used!
In just 15 lines of purely functional code, we built a horizontally scaling, perfectly categorized local AI worker pool that processes massive text sets in milliseconds.
But we can go deeper.
3. The “My Mom Won’t Sleep” Pipeline
Let’s build a fully reactive, persistent AI agent swarm.
We will simulate a continuous stream of unstructured support emails. We’ll set up a buffered inbox-chan and spin up 3 completely independent Concurrent LLM Workers using Coni’s (spawn) primitive.
The most powerful feature of the System One API is that it can answer multiple distinct questions in a single JSON payload. For every ticket, our workers will simultaneously ask the local LLM:
- Is this a legal escalation? (Choice)
- Which department should this go to? (Choice)
- What is the urgency? (Score 1-10)
Once a worker processes a ticket, it pushes the structured JSON onto an outbox-chan. Finally, a Dispatcher thread reads this outbox and triggers real-time conditional logic (like simulated PagerDuty or Slack alerts).
This is what pure orchestration power looks like:
;; Channels for our pipeline
(def inbox-chan (chan 100)) ;; Buffered channel for incoming raw messages
(def outbox-chan (chan 100)) ;; Buffered channel for processed insights
(def done-chan (chan)) ;; Signal channel for graceful shutdown
;; 1. The Worker Pool (3 Concurrent LLM Processors)
(doseq [worker-id (range 1 4)]
(spawn (fn []
(loop []
(let [msg (<! inbox-chan)]
(when msg
;; ONE single payload asking THREE complex questions simultaneously!
(let [payload {:model "nimble"
:state (:text msg)
:questions {
:escalate? {:type "choice"
:instructions "Does this message contain threats or legal action?"
:criteria {:yes "Threat or Legal" :no "Safe"}}
:department {:type "choice"
:instructions "Which team should handle this?"
:criteria {:technical "Bugs and crashes"
:billing "Refunds and money"
:sales "Buying products"}}
:priority {:type "score"
:instructions "Score the urgency from 1 (low) to 10 (critical)."
:criteria ["Extremely Low" "Low" "Minor" "Normal" "Slightly High" "High" "Urgent" "Critical" "Emergency" "Immediate Action"]}}}
parsed (fetch-system-one payload)
ans (:answers parsed)]
;; Push the structured result into the outbox
(>! outbox-chan {:id (:id msg)
:worker worker-id
:text (:text msg)
:escalate? (= (get-in ans [:escalate? :choice]) "yes")
:department (get-in ans [:department :choice])
:urgency (get-in ans [:priority :score])}))
(recur)))))))
;; 2. The Dispatcher
(spawn (fn []
(loop [processed-count 0]
(if (>= processed-count 10)
(>! done-chan true)
(let [insight (<! outbox-chan)
urgency (or (:urgency insight) 0)]
(println (str "\n[Worker " (:worker insight) " -> Dispatcher] Analyzing Msg ID: " (:id insight)))
(println (str " ➡️ Department: " (:department insight) " | Urgency Score: " urgency "/10"))
(when (:escalate? insight)
(println " 🚨 [PAGERDUTY TRIGGERED] LEGAL/SAFETY ESCALATION DETECTED!"))
(when (> urgency 7)
(println " ⚠️ [SLACK ALERT] High priority message routed to manager!"))
(recur (+ processed-count 1)))))))
The Output:
[Worker 2 -> Dispatcher] Analyzing Msg ID: 2
➡️ Department: billing | Urgency Score: 3.649321564728955/10
🚨 [PAGERDUTY TRIGGERED] LEGAL/SAFETY ESCALATION DETECTED!
[Worker 2 -> Dispatcher] Analyzing Msg ID: 5
➡️ Department: technical | Urgency Score: 7.71376315400216/10
⚠️ [SLACK ALERT] High priority message routed to manager!
If you look closely at Msg ID 2 (“If you don’t refund my $500 right now I am calling my lawyer and suing your company.”), Worker 2 instantly identified it as a billing issue and simultaneously flagged the legal threat, which caused the central Dispatcher to trigger the PagerDuty alert.
No bloated Python libraries, no complicated async/await event loops, no external task queues.
Just pure functional concurrency and local GPU-accelerated decision intelligence. Coni and Ollama’s System One API are a match made in heaven.