tag:github.com,2008:https://github.com/ollama/ollama/releases

Release notes from ollama

2026-09-23T00:53:23Z tag:github.com,2008:Repository/658928958/v0.34.4-rc0 2026-09-23T02:24:43Z

v0.34.4

<h2>What's Changed</h2> <ul> <li>server: fix intermittent "model not found" errors. by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/rick-github/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/rick-github">@rick-github</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5453632971" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18438" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18438/hovercard" href="https://github.com/ollama/ollama/pull/18438">#18438</a></li> <li>server: apply structured outputs in a single pass on thinking models by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/jessegross/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/jessegross">@jessegross</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5468031118" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18479" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18479/hovercard" href="https://github.com/ollama/ollama/pull/18479">#18479</a></li> <li>app: avoid System Events for ChatGPT/Codex detection by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/hoyyeva/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/hoyyeva">@hoyyeva</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5544210665" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18601" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18601/hovercard" href="https://github.com/ollama/ollama/pull/18601">#18601</a></li> <li>llama.cpp: version update by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/dhiltgen/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dhiltgen">@dhiltgen</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5532309403" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18577" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18577/hovercard" href="https://github.com/ollama/ollama/pull/18577">#18577</a></li> <li>MLX: version bump by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/dhiltgen/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dhiltgen">@dhiltgen</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5531817294" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18576" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18576/hovercard" href="https://github.com/ollama/ollama/pull/18576">#18576</a></li> <li>mlx: select Gemma 4 image resolution dynamically by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/dhiltgen/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dhiltgen">@dhiltgen</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5545364467" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18603" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18603/hovercard" href="https://github.com/ollama/ollama/pull/18603">#18603</a></li> <li>mlx: speed up Qwen 3.8 prompt processing by <a class="user-mention notranslate" data-hovercard-type="user" data-hovercard-url="/users/dhiltgen/hovercard" data-octo-click="hovercard-link-click" data-octo-dimensions="link_type:self" href="https://github.com/dhiltgen">@dhiltgen</a> in <a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5514757532" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18550" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18550/hovercard" href="https://github.com/ollama/ollama/pull/18550">#18550</a></li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.3...v0.34.4-rc0"><tt>v0.34.3...v0.34.4-rc0</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.3 2026-09-22T20:42:48Z

v0.34.3

<h2>What's Changed</h2> <p><code>GET /api/show</code> now advertises each model's thinking controls and default:</p> <p>Available in the CLI with:</p> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama show gemma4"><pre class="notranslate"><code>ollama show gemma4 </code></pre></div> <div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content=" thinking levels false, true default true"><pre class="notranslate"><code> thinking levels false, true default true </code></pre></div> <p>Available in the API with:</p> <div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="curl http://localhost:11434/api/show -d '{"model": "glm-5.3-flash:cloud"}'"><pre>curl http://localhost:11434/api/show -d <span class="pl-s"><span class="pl-pds">'</span>{"model": "glm-5.3-flash:cloud"}<span class="pl-pds">'</span></span></pre></div> <div class="highlight highlight-source-json notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="{ "thinking": { "values": ["low", "high", "max"], "default": "max" } }"><pre>{ <span class="pl-ent">"thinking"</span>: { <span class="pl-ent">"values"</span>: [<span class="pl-s"><span class="pl-pds">"</span>low<span class="pl-pds">"</span></span>, <span class="pl-s"><span class="pl-pds">"</span>high<span class="pl-pds">"</span></span>, <span class="pl-s"><span class="pl-pds">"</span>max<span class="pl-pds">"</span></span>], <span class="pl-ent">"default"</span>: <span class="pl-s"><span class="pl-pds">"</span>max<span class="pl-pds">"</span></span> } }</pre></div> <p>Also available on ollama.com directly for cloud models.</p> <ul> <li><strong>Nemotron H</strong> vision models are now supported on Apple Silicon with MLX</li> <li>Ollama's macOS app will now no longer reopen windows you've closed when activating it</li> <li>Fix for model pulls from HuggingFace</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.2...v0.34.3"><tt>v0.34.2...v0.34.3</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.3-rc1 2026-09-19T00:15:35Z

v0.34.3-rc1

<p>server: allow registry cross-host redirects among allowlisted hosts (…</p> pdevine tag:github.com,2008:Repository/658928958/v0.34.3-rc0 2026-09-18T21:10:53Z

v0.34.3-rc0

<p>api: expose model thinking levels and defaults (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5465487978" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18473" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18473/hovercard" href="https://github.com/ollama/ollama/pull/18473">#18473</a>)</p> ParthSareen tag:github.com,2008:Repository/658928958/v0.34.2 2026-09-17T23:04:24Z

v0.34.2

<h2>What's Changed</h2> <ul> <li>Added first-run setup when running <code>ollama</code>, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.</li> <li>Added <code>ollama://apps</code> to open the desktop app’s Apps page directly on macOS and Windows.</li> <li>Fixed excessive memory growth during long generations with MLX speculative decoding.</li> <li>Updated llama.cpp.</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.1...v0.34.2"><tt>v0.34.1...v0.34.2</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.2-rc3 2026-09-17T18:49:24Z

v0.34.2-rc3

<p>cli: add first-run onboarding shared with the desktop app (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5480870631" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18495" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18495/hovercard" href="https://github.com/ollama/ollama/pull/18495">#18495</a>)</p> hoyyeva tag:github.com,2008:Repository/658928958/v0.34.2-rc2 2026-09-17T16:45:01Z

v0.34.2-rc2: mlxrunner: Release freed KV buffers during speculative decode

<p>The decode loop releases MLX's pool of freed buffers every 256 generated<br> tokens, which is also how often the KV cache grows and drops its previous,<br> smaller buffers. The check fires only when the token count lands exactly on<br> a multiple of 256. Speculative decoding emits several tokens per round, so<br> most rounds step over the boundary and the pool is never released. Each<br> growth at a long context leaves several GB of buffers that no later<br> allocation can reuse, so the runner's footprint keeps climbing over a long<br> generation until the system runs out of memory.</p> <p>We now release the pool whenever a round crosses a multiple of 256 tokens,<br> which is what a single-token round already did. With qwen3.8:27b-mlx at a<br> 98k-token context on a 128 GB machine, a long speculative generation<br> previously grew the runner past 90 GB and panicked the kernel; it now stays<br> flat at 30 GB.</p> jessegross tag:github.com,2008:Repository/658928958/v0.34.2-rc1 2026-09-16T21:06:08Z

v0.34.2-rc1: mlxrunner: lay out model by contract, checkpoint and construction

<p>model is one package with three jobs: the contract between the runner<br> and the architectures, the opened checkpoint, and building nn layers<br> from checkpoint tensors. Its files did not say which was which. base.go<br> carried the folded package's name over the interfaces and the registry,<br> root.go held the safetensors header scan next to Root, and quant.go<br> mixed the nvfp4 global-scale helpers with quant parameter resolution.</p> <p>base.go becomes model.go, named for what it holds. root.go keeps Root<br> and Open; TensorQuantInfo and the header scan join quant.go, so<br> everything the checkpoint says about quantization is read and resolved<br> in one file. The global-scale helpers move to globalscale.go with their<br> tests. Root.Close, a no-op with one caller, goes. No code changes<br> otherwise.</p> jessegross tag:github.com,2008:Repository/658928958/v0.34.2-rc0 2026-09-15T20:13:31Z

v0.34.2-rc0: llama.cpp: version bump b10969 (#18446)

<p>llama.cpp build changes resulted in duplicate symbols between libllama and libmtmd. This moves the compat patch into libllama with exported symbols.</p> dhiltgen tag:github.com,2008:Repository/658928958/v0.34.1 2026-09-15T20:10:42Z

v0.34.1

<h2>What's Changed</h2> <ul> <li>MLX safetensors <code>ollama create</code> no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.</li> <li>Improved MLX memory handling on Apple Silicon</li> <li>Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR)</li> <li><code>/api/tags</code> is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently.</li> <li>Deprecated <code>typical_p</code>: it can no longer be set when creating new models, existing GGUF models retain support.</li> <li>MLX and llama.cpp updates</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.0...v0.34.1-rc1"><tt>v0.34.0...v0.34.1-rc1</tt></a></p> github-actions[bot]