tag:github.com,2008:https://github.com/ollama/ollama/releases

Release notes from ollama

2026-09-28T20:31:38Z tag:github.com,2008:Repository/658928958/v0.35.0 2026-09-29T01:54:15Z

v0.35.0

<h2>Decision models</h2> <p>Ollama now supports decision models through <code>/v1/systemone</code>, based on <a href="https://typesafe.ai" rel="nofollow">TypeSafe’s Jev API</a>.</p> <p>Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.</p> <p>Available models:</p> <ul> <li><a href="https://ollama.com/library/nimble" rel="nofollow"><strong>Nimble</strong></a> from Bespoke Labs</li> <li><a href="https://ollama.com/library/tev1" rel="nofollow"><strong>Tev1</strong></a> from Together AI</li> </ul> <div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama pull nimble"><pre>ollama pull nimble</pre></div> <p>Send context and one or more questions:</p> <div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="curl http://localhost:11434/v1/systemone \ -H 'Content-Type: application/json' \ -d '{ "model": "nimble", "state": "Our checkout has returned 500 errors since 9am.", "questions": { "label": { "type": "choice", "instructions": "Which label fits this ticket?", "criteria": { "billing": "Payments and refunds", "bug": "Software errors", "account": "Login and account access" } } } }'"><pre>curl http://localhost:11434/v1/systemone \ -H <span class="pl-s"><span class="pl-pds">'</span>Content-Type: application/json<span class="pl-pds">'</span></span> \ -d <span class="pl-s"><span class="pl-pds">'</span>{</span> <span class="pl-s"> "model": "nimble",</span> <span class="pl-s"> "state": "Our checkout has returned 500 errors since 9am.",</span> <span class="pl-s"> "questions": {</span> <span class="pl-s"> "label": {</span> <span class="pl-s"> "type": "choice",</span> <span class="pl-s"> "instructions": "Which label fits this ticket?",</span> <span class="pl-s"> "criteria": {</span> <span class="pl-s"> "billing": "Payments and refunds",</span> <span class="pl-s"> "bug": "Software errors",</span> <span class="pl-s"> "account": "Login and account access"</span> <span class="pl-s"> }</span> <span class="pl-s"> }</span> <span class="pl-s"> }</span> <span class="pl-s"> }<span class="pl-pds">'</span></span></pre></div> <p>Example response:</p> <div class="highlight highlight-source-json notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="{ "model": "nimble", "answers": { "label": { "type": "choice", "choice": "bug", "probabilities": { "billing": 0.0125, "bug": 0.9781, "account": 0.0093 }, "confidence": 0.8906 } }, "usage": { "input_tokens": 174, "output_tokens": 1 } }"><pre>{ <span class="pl-ent">"model"</span>: <span class="pl-s"><span class="pl-pds">"</span>nimble<span class="pl-pds">"</span></span>, <span class="pl-ent">"answers"</span>: { <span class="pl-ent">"label"</span>: { <span class="pl-ent">"type"</span>: <span class="pl-s"><span class="pl-pds">"</span>choice<span class="pl-pds">"</span></span>, <span class="pl-ent">"choice"</span>: <span class="pl-s"><span class="pl-pds">"</span>bug<span class="pl-pds">"</span></span>, <span class="pl-ent">"probabilities"</span>: { <span class="pl-ent">"billing"</span>: <span class="pl-c1">0.0125</span>, <span class="pl-ent">"bug"</span>: <span class="pl-c1">0.9781</span>, <span class="pl-ent">"account"</span>: <span class="pl-c1">0.0093</span> }, <span class="pl-ent">"confidence"</span>: <span class="pl-c1">0.8906</span> } }, <span class="pl-ent">"usage"</span>: { <span class="pl-ent">"input_tokens"</span>: <span class="pl-c1">174</span>, <span class="pl-ent">"output_tokens"</span>: <span class="pl-c1">1</span> } }</pre></div> <p>The API supports three question types:</p> <ul> <li><code>choice</code>: Select an option and return probabilities for each.</li> <li><code>noul</code>: Return the probability that a condition is true.</li> <li><code>score</code>: Return a score across an ordered set of criteria</li> </ul> <h2>What's Changed</h2> <ul> <li>Settings now opens without waiting for model discovery.</li> <li>Fixed the macOS update menu and icon not reflecting an available update at startup.</li> <li>Fixed stalled MLX model downloads hanging indefinitely.</li> <li>Requests containing the deprecated <code>typical_p</code> parameter now log a warning instead of failing.</li> </ul> <p><strong>Full Changelog:</strong> <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.4...v0.35.0"><tt>v0.34.4...v0.35.0</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.35.0-rc1 2026-09-28T20:31:38Z

v0.35.0-rc1

<p>mlx: bound pull stall retries and let the watchdog interrupt them (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="1777625644" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/1" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/1/hovercard" href="https://github.com/ollama/ollama/pull/1">#1</a>…</p> dhiltgen tag:github.com,2008:Repository/658928958/v0.35.0-rc0 2026-09-28T20:21:01Z

v0.35.0-rc0

<p>feat: add System One scoring API (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5550448056" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18606" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18606/hovercard" href="https://github.com/ollama/ollama/pull/18606">#18606</a>)</p> ParthSareen tag:github.com,2008:Repository/658928958/v0.40.0-rc0 2026-09-25T15:22:00Z

v0.40.0

<h2>What's Changed</h2> <p><strong>Models run on MLX on Apple Silicon by default</strong></p> <p>In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.</p> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama pull qwen3.8 ollama run qwen3.8"><pre class="notranslate"><code>ollama pull qwen3.8 ollama run qwen3.8 </code></pre></div> <p>During the pre-release we will be testing and enabling additional models.</p> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.4...v0.40.0-rc0"><tt>v0.34.4...v0.40.0-rc0</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.4 2026-09-24T04:45:43Z

v0.34.4

<h2>What's Changed</h2> <ul> <li>Structured outputs on thinking models now apply in a single pass, making them faster and more reliable.</li> <li>Fixed intermittent "model not found" errors with a large local library</li> <li>Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running.</li> <li>Qwen 3.8 prompt processing is faster on Apple Silicon.</li> <li>Gemma 4 on Apple Silicon now picks the best image resolution per image, keeping more detail in high-resolution images.</li> <li>Updated llama.cpp, MLX, and XGrammar.</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.3...v0.34.4"><tt>v0.34.3...v0.34.4</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.4-rc1 2026-09-23T23:36:46Z

v0.34.4-rc1: mlxrunner: Update XGrammar to 0.2.7 for structured outputs

<p>We pick up schema fixes for typed dictionary values and short arrays.</p> jessegross tag:github.com,2008:Repository/658928958/v0.34.4-rc0 2026-09-23T00:53:23Z

v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550)

<ul> <li>mlx: speed up Qwen 3.8 prompt processing</li> </ul> <p>Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU.</p> <ul> <li>address comments</li> </ul> dhiltgen tag:github.com,2008:Repository/658928958/v0.34.3 2026-09-22T20:42:48Z

v0.34.3

<h2>What's Changed</h2> <p><code>GET /api/show</code> now advertises each model's thinking controls and default:</p> <p>Available in the CLI with:</p> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama show gemma4"><pre class="notranslate"><code>ollama show gemma4 </code></pre></div> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content=" thinking levels false, true default true"><pre class="notranslate"><code> thinking levels false, true default true </code></pre></div> <p>Available in the API with:</p> <div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}'"><pre>curl http://localhost:11434/api/show -d <span class="pl-s"><span class="pl-pds">'</span>{"model": "glm-5.3-flash:cloud"}<span class="pl-pds">'</span></span></pre></div> <div class="highlight highlight-source-json notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="{ "thinking": { "values": ["low", "high", "max"], "default": "max" } }"><pre>{ <span class="pl-ent">"thinking"</span>: { <span class="pl-ent">"values"</span>: [<span class="pl-s"><span class="pl-pds">"</span>low<span class="pl-pds">"</span></span>, <span class="pl-s"><span class="pl-pds">"</span>high<span class="pl-pds">"</span></span>, <span class="pl-s"><span class="pl-pds">"</span>max<span class="pl-pds">"</span></span>], <span class="pl-ent">"default"</span>: <span class="pl-s"><span class="pl-pds">"</span>max<span class="pl-pds">"</span></span> } }</pre></div> <p>Also available on ollama.com directly for cloud models.</p> <ul> <li><strong>Nemotron H</strong> vision models are now supported on Apple Silicon with MLX</li> <li>Ollama's macOS app will now no longer reopen windows you've closed when activating it</li> <li>Fix for model pulls from HuggingFace</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.2...v0.34.3"><tt>v0.34.2...v0.34.3</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.3-rc1 2026-09-19T00:15:35Z

v0.34.3-rc1

<p>server: allow registry cross-host redirects among allowlisted hosts (…</p> pdevine tag:github.com,2008:Repository/658928958/v0.34.3-rc0 2026-09-18T21:10:53Z

v0.34.3-rc0

<p>api: expose model thinking levels and defaults (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5465487978" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18473" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18473/hovercard" href="https://github.com/ollama/ollama/pull/18473">#18473</a>)</p> ParthSareen