tag:github.com,2008:https://github.com/ollama/ollama/releases

Release notes from ollama

2026-09-17T21:21:28Z tag:github.com,2008:Repository/658928958/v0.34.2 2026-09-17T22:27:21Z

v0.34.2

<h2>What's Changed</h2> <ul> <li>Added first-run setup when running <code>ollama</code>, with options to sign in or continue locally. Setup completion is shared with the desktop app on macOS and Windows.</li> <li>Added <code>ollama://apps</code> to open the desktop app’s Apps page directly on macOS and Windows.</li> <li>Fixed excessive memory growth during long generations with MLX speculative decoding.</li> <li>Updated llama.cpp.</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.1...v0.34.2-rc0"><tt>v0.34.1...v0.34.2-rc0</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.2-rc3 2026-09-17T18:49:24Z

v0.34.2-rc3

<p>cli: add first-run onboarding shared with the desktop app (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5480870631" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18495" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18495/hovercard" href="https://github.com/ollama/ollama/pull/18495">#18495</a>)</p> hoyyeva tag:github.com,2008:Repository/658928958/v0.34.2-rc2 2026-09-17T16:45:01Z

v0.34.2-rc2: mlxrunner: Release freed KV buffers during speculative decode

<p>The decode loop releases MLX's pool of freed buffers every 256 generated<br> tokens, which is also how often the KV cache grows and drops its previous,<br> smaller buffers. The check fires only when the token count lands exactly on<br> a multiple of 256. Speculative decoding emits several tokens per round, so<br> most rounds step over the boundary and the pool is never released. Each<br> growth at a long context leaves several GB of buffers that no later<br> allocation can reuse, so the runner's footprint keeps climbing over a long<br> generation until the system runs out of memory.</p> <p>We now release the pool whenever a round crosses a multiple of 256 tokens,<br> which is what a single-token round already did. With qwen3.8:27b-mlx at a<br> 98k-token context on a 128 GB machine, a long speculative generation<br> previously grew the runner past 90 GB and panicked the kernel; it now stays<br> flat at 30 GB.</p> jessegross tag:github.com,2008:Repository/658928958/v0.34.2-rc1 2026-09-16T21:06:08Z

v0.34.2-rc1: mlxrunner: lay out model by contract, checkpoint and construction

<p>model is one package with three jobs: the contract between the runner<br> and the architectures, the opened checkpoint, and building nn layers<br> from checkpoint tensors. Its files did not say which was which. base.go<br> carried the folded package's name over the interfaces and the registry,<br> root.go held the safetensors header scan next to Root, and quant.go<br> mixed the nvfp4 global-scale helpers with quant parameter resolution.</p> <p>base.go becomes model.go, named for what it holds. root.go keeps Root<br> and Open; TensorQuantInfo and the header scan join quant.go, so<br> everything the checkpoint says about quantization is read and resolved<br> in one file. The global-scale helpers move to globalscale.go with their<br> tests. Root.Close, a no-op with one caller, goes. No code changes<br> otherwise.</p> jessegross tag:github.com,2008:Repository/658928958/v0.34.2-rc0 2026-09-15T20:13:31Z

v0.34.2-rc0: llama.cpp: version bump b10969 (#18446)

<p>llama.cpp build changes resulted in duplicate symbols between libllama and libmtmd. This moves the compat patch into libllama with exported symbols.</p> dhiltgen tag:github.com,2008:Repository/658928958/v0.34.1 2026-09-15T20:10:42Z

v0.34.1

<h2>What's Changed</h2> <ul> <li>MLX safetensors <code>ollama create</code> no longer experimental. GGUF model creation now requires using llama.cpp tooling for safetensor conversion and quantization.</li> <li>Improved MLX memory handling on Apple Silicon</li> <li>Runaway repeat token detection now requires 100 repeat tokens for reduced false positives (e.g. OCR)</li> <li><code>/api/tags</code> is much faster on large model libraries (3.1 s → 294 ms cold in testing), and model capabilities are now reported consistently.</li> <li>Deprecated <code>typical_p</code>: it can no longer be set when creating new models, existing GGUF models retain support.</li> <li>MLX and llama.cpp updates</li> </ul> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.0...v0.34.1-rc1"><tt>v0.34.0...v0.34.1-rc1</tt></a></p> github-actions[bot] tag:github.com,2008:Repository/658928958/v0.34.1-rc2 2026-09-15T04:24:26Z

v0.34.1-rc2: API: Deprecate typical_p (#18448)

<p>Drop support for creating new models with typical_p parameters, while<br> retaining support for existing GGUF models with the setting.</p> dhiltgen tag:github.com,2008:Repository/658928958/v0.34.1-rc1 2026-09-14T20:34:03Z

v0.34.1-rc1

<p>mlx: add mlx patch to docker build context (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5453819882" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18440" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18440/hovercard" href="https://github.com/ollama/ollama/pull/18440">#18440</a>)</p> dhiltgen tag:github.com,2008:Repository/658928958/v0.34.1-rc0 2026-09-14T16:49:46Z

v0.34.1-rc0: MLX: version bump (#18235)

<ul> <li> <p>MLX: version bump</p> </li> <li> <p>mlx: support ModelOpt global scales in MoE models</p> </li> <li> <p>address comments</p> </li> <li> <p>address comments</p> </li> </ul> dhiltgen tag:github.com,2008:Repository/658928958/v0.34.0 2026-09-10T06:18:35Z

v0.34.0

<h2>Use Ollama models in ChatGPT Desktop</h2> <p>Ollama models can now be used directly in ChatGPT Desktop, so you can keep your existing workflow while running open models. Setup is available from the Ollama app on MacOS.</p> <a target="_blank" rel="noopener noreferrer" href="https://private-user-images.githubusercontent.com/29360864/649245808-e10e299d-c11f-447d-9234-afa855824efe.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODk2ODQ0MDgsIm5iZiI6MTc4OTY4NDEwOCwicGF0aCI6Ii8yOTM2MDg2NC82NDkyNDU4MDgtZTEwZTI5OWQtYzExZi00NDdkLTkyMzQtYWZhODU1ODI0ZWZlLnBuZz9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNjA5MTclMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjYwOTE3VDIyMjgyOFomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPTFiNWMyZDkwMDEwMDkyODQ5Y2Q3MGQ2OGJlOWNkMmY5YmE1ODgxYTkxZjEwMWQyNjdhMjA4NDY1MjYwZWQwNzEmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0JnJlc3BvbnNlLWNvbnRlbnQtdHlwZT1pbWFnZSUyRnBuZyJ9.sqJcm0Mzm8iiDjbhON-6jC685dEYhMPMVNF8qX3evw4"><img width="1374" height="1300" alt="CleanShot 2026-09-08 at 11 04 07 AM@2x" src="https://private-user-images.githubusercontent.com/29360864/649245808-e10e299d-c11f-447d-9234-afa855824efe.png?jwt=eyJ0eXAiOiJKV1QiLCJhbGciOiJIUzI1NiJ9.eyJpc3MiOiJnaXRodWIuY29tIiwiYXVkIjoicmF3LmdpdGh1YnVzZXJjb250ZW50LmNvbSIsImtleSI6ImtleTUiLCJleHAiOjE3ODk2ODQ0MDgsIm5iZiI6MTc4OTY4NDEwOCwicGF0aCI6Ii8yOTM2MDg2NC82NDkyNDU4MDgtZTEwZTI5OWQtYzExZi00NDdkLTkyMzQtYWZhODU1ODI0ZWZlLnBuZz9YLUFtei1BbGdvcml0aG09QVdTNC1ITUFDLVNIQTI1NiZYLUFtei1DcmVkZW50aWFsPUFLSUFWQ09EWUxTQTUzUFFLNFpBJTJGMjAyNjA5MTclMkZ1cy1lYXN0LTElMkZzMyUyRmF3czRfcmVxdWVzdCZYLUFtei1EYXRlPTIwMjYwOTE3VDIyMjgyOFomWC1BbXotRXhwaXJlcz0zMDAmWC1BbXotU2lnbmF0dXJlPTFiNWMyZDkwMDEwMDkyODQ5Y2Q3MGQ2OGJlOWNkMmY5YmE1ODgxYTkxZjEwMWQyNjdhMjA4NDY1MjYwZWQwNzEmWC1BbXotU2lnbmVkSGVhZGVycz1ob3N0JnJlc3BvbnNlLWNvbnRlbnQtdHlwZT1pbWFnZSUyRnBuZyJ9.sqJcm0Mzm8iiDjbhON-6jC685dEYhMPMVNF8qX3evw4" content-type-secured-asset="image/png" style="max-width: 100%; height: auto; max-height: 1300px;"></a> <p>This release also improves structured output performance on Apple Silicon, adds support for OpenAI-compatible client tool search and response compaction.</p> <p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.33.3...v0.34.0"><tt>v0.33.3...v0.34.0</tt></a></p> github-actions[bot]