Ollama has introduced local support for decision models based on TypeSafe's Jev API. The new /v1/systemone endpoint, available from Ollama 0.35, lets you send text together with a set of named questions and get answers from a model running on your own computer – at no extra cost and without sending data over the network.
New /v1/systemone interface and real-time control
Ollama announced support for decision models in a post on X and in a post on its blog dated September 29, 2026. These models are based on TypeSafe's Jev API, designed for fast, typed decisions. The /v1/systemone endpoint is available from Ollama 0.35: you send text as the state (state) together with a set of named questions, and a model running on your computer answers all of them in a single request.
To try it, you can download the Nimble model with ollama pull nimble. In the video from the X post, the model made decisions in real time in the game Ollama racer, and in the Pac-Man example on Ollama's blog the average decision time of Nimble 9B was 91 ms (MacBook Pro with an M5 Max chip).
Main infrastructure use cases
Instead of generating text, decision models return a result for questions of three types: choosing one of several options (choice), yes/no questions (noul) and scoring on a scale (score). According to Ollama, they work well for tasks that need a fast decision:
- Ticket triage – choosing the right category or team for an incoming ticket.
- Model routing – deciding which model a given request should be sent to.
- Content and safety moderation – assessing whether a piece of content needs intervention.
Requests do not travel over the network, so latency is low, although the result depends on hardware and configuration. The documentation also lists limitations: the API returns a single JSON response, does not support streaming, images or tools, and a single request can be at most 64 KiB.
Three models are available at launch: nimble – an open-source 9B model from Bespoke Labs, tev1 – an experimental 4B model from Together AI, and tev1:0.8b – an experimental 0.8B model from Together AI. Ollama says more are coming, including models served by its cloud.
What this change means for automation projects
This is an additional tool, not a replacement for large models: for simple classification, yes/no decisions or scoring on a scale, a specialized model like Nimble is enough, instead of running a large generative model that would return a single word. The model runs locally, at no extra cost and without sending data over the network.
Administrators and developers should look at Ollama's documentation and examples before deploying – this is a new feature, and the tev1 models are marked as experimental.
Sources
- https://x.com/ollama/status/2105152056382345544
- https://ollama.com/blog/ollama-now-supports-jev-style-decision-models
- https://docs.ollama.com/capabilities/decision
- https://docs.ollama.com/api/systemone
- https://ollama.com/library/nimble
- https://ollama.com/library/tev1:4b
- https://ollama.com/library/tev1:0.8b
Comments