tag:github.com,2008:https://github.com/ollama/ollama/releases

Release notes from ollama

2026-09-25T02:25:52Z tag:github.com,2008:Repository/658928958/v0.40.0-rc0 2026-09-25T15:22:00Z

v0.40.0

<h2>What's Changed</h2> <p><strong>Models run on MLX on Apple Silicon by default</strong></p> <p>In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.</p> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama pull qwen3.8 ollama run qwen3.8"><pre class="notranslate"><code>ollama pull qwen3.8 ollama run qwen3.8 </code></pre></div> <p>During the pre-release we will be testing and enabling additional models.</p> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.4...v0.40.0-rc0"><tt>v0.34.4...v0.40.0-rc0</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.4 2026-09-24T04:45:43Z

v0.34.4

<h2>What's Changed</h2> <ul> <li>Structured outputs on thinking models now apply in a single pass, making them faster and more reliable.</li> <li>Fixed intermittent "model not found" errors with a large local library</li> <li>Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running.</li> <li>Qwen 3.8 prompt processing is faster on Apple Silicon.</li> <li>Gemma 4 on Apple Silicon now picks the best image resolution per image, keeping more detail in high-resolution images.</li> <li>Updated llama.cpp, MLX, and XGrammar.</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.3...v0.34.4"><tt>v0.34.3...v0.34.4</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.4-rc1 2026-09-23T23:36:46Z

v0.34.4-rc1: mlxrunner: Update XGrammar to 0.2.7 for structured outputs

<p>We pick up schema fixes for typed dictionary values and short arrays.</p> jessegross tag:github.com,2008:Repository/658928958/v0.34.4-rc0 2026-09-23T00:53:23Z

v0.34.4-rc0: mlx: speed up Qwen 3.8 prompt processing (#18550)

<ul> <li>mlx: speed up Qwen 3.8 prompt processing</li> </ul> <p>Use MLX's gated-delta kernel for long scans and fold dense MLP global scales into SwiGLU.</p> <ul> <li>address comments</li> </ul> dhiltgen tag:github.com,2008:Repository/658928958/v0.34.3 2026-09-22T20:42:48Z

v0.34.3

<h2>What's Changed</h2> <p><code>GET /api/show</code> now advertises each model's thinking controls and default:</p> <p>Available in the CLI with:</p> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama show gemma4"><pre class="notranslate"><code>ollama show gemma4 </code></pre></div> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content=" thinking levels false, true default true"><pre class="notranslate"><code> thinking levels false, true default true </code></pre></div> <p>Available in the API with:</p> <div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}'"><pre>curl http://localhost:11434/api/show -d <span class="pl-s"><span class="pl-pds">'</span>{"model": "glm-5.3-flash:cloud"}<span class="pl-pds">'</span></span></pre></div> <div class="highlight highlight-source-json notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="{ "thinking": { "values": ["low", "high", "max"], "default": "max" } }"><pre>{ <span class="pl-ent">"thinking"</span>: { <span class="pl-ent">"values"</span>: [<span class="pl-s"><span class="pl-pds">"</span>low<span class="pl-pds">"</span></span>, <span class="pl-s"><span class="pl-pds">"</span>high<span class="pl-pds">"</span></span>, <span class="pl-s"><span class="pl-pds">"</span>max<span class="pl-pds">"</span></span>], <span class="pl-ent">"default"</span>: <span class="pl-s"><span class="pl-pds">"</span>max<span class="pl-pds">"</span></span> } }</pre></div> <p>Also available on ollama.com directly for cloud models.</p> <ul> <li><strong>Nemotron H</strong> vision models are now supported on Apple Silicon with MLX</li> <li>Ollama's macOS app will now no longer reopen windows you've closed when activating it</li> <li>Fix for model pulls from HuggingFace</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.2...v0.34.3"><tt>v0.34.2...v0.34.3</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.3-rc1 2026-09-19T00:15:35Z

v0.34.3-rc1

<p>server: allow registry cross-host redirects among allowlisted hosts (…</p> pdevine tag:github.com,2008:Repository/658928958/v0.34.3-rc0 2026-09-18T21:10:53Z

v0.34.3-rc0

<p>api: expose model thinking levels and defaults (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5465487978" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18473" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18473/hovercard" href="https://github.com/ollama/ollama/pull/18473">#18473</a>)</p> ParthSareen tag:github.com,2008:Repository/658928958/v0.34.2 2026-09-17T23:04:24Z

v0.34.2

<h2>What's Changed</h2> <ul> <li>Added first-run setup when running <code>ollama</code>, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.</li> <li>Added <code>ollama://apps</code> to open the desktop app’s Apps page directly on macOS and Windows.</li> <li>Fixed excessive memory growth during long generations with MLX speculative decoding.</li> <li>Updated llama.cpp.</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.1...v0.34.2"><tt>v0.34.1...v0.34.2</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.2-rc3 2026-09-17T18:49:24Z

v0.34.2-rc3

<p>cli: add first-run onboarding shared with the desktop app (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5480870631" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18495" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18495/hovercard" href="https://github.com/ollama/ollama/pull/18495">#18495</a>)</p> hoyyeva tag:github.com,2008:Repository/658928958/v0.34.2-rc2 2026-09-17T16:45:01Z

v0.34.2-rc2: mlxrunner: Release freed KV buffers during speculative decode

<p>The decode loop releases MLX's pool of freed buffers every 256 generated<br> tokens, which is also how often the KV cache grows and drops its previous,<br> smaller buffers. The check fires only when the token count lands exactly on<br> a multiple of 256. Speculative decoding emits several tokens per round, so<br> most rounds step over the boundary and the pool is never released. Each<br> growth at a long context leaves several GB of buffers that no later<br> allocation can reuse, so the runner's footprint keeps climbing over a long<br> generation until the system runs out of memory.</p> <p>We now release the pool whenever a round crosses a multiple of 256 tokens,<br> which is what a single-token round already did. With qwen3.8:27b-mlx at a<br> 98k-token context on a 128 GB machine, a long speculative generation<br> previously grew the runner past 90 GB and panicked the kernel; it now stays<br> flat at 30 GB.</p> jessegross