tag:github.com,2008:https://github.com/ollama/ollama/releasesRelease notes from ollama2026-10-02T01:05:01Ztag:github.com,2008:Repository/658928958/v0.35.12026-10-02T18:28:37Zv0.35.1<h2>Clef decision models</h2>
<p>Ollama now supports <a href="https://ollama.com/library/clef" rel="nofollow">Clef</a> and <a href="https://ollama.com/library/clef" rel="nofollow">Clef Flash</a>, Cloudflare's new open-source decision models, through <code>/v1/systemone</code>.</p>
<p>Clef (27B) and Clef Flash (9B) are multimodal: requests can now include images alongside the text state, shared by all questions and scored jointly with it.</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="curl http://localhost:11434/v1/systemone -d '{
"model": "clef-flash",
"state": "The user took this screenshot.",
"images": ["<base64-encoded image>"],
"questions": {
"has_ollama": {"type": "noul", "instructions": "Does this image contain Ollama?"}
}
}'"><pre>curl http://localhost:11434/v1/systemone -d <span class="pl-s"><span class="pl-pds">'</span>{</span>
<span class="pl-s"> "model": "clef-flash",</span>
<span class="pl-s"> "state": "The user took this screenshot.",</span>
<span class="pl-s"> "images": ["<base64-encoded image>"],</span>
<span class="pl-s"> "questions": {</span>
<span class="pl-s"> "has_ollama": {"type": "noul", "instructions": "Does this image contain Ollama?"}</span>
<span class="pl-s"> }</span>
<span class="pl-s">}<span class="pl-pds">'</span></span></pre></div>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="{
"model": "clef-flash",
"answers": {
"has_ollama": {
"type": "noul",
"noul": 0.958
}
},
"usage": {
"input_tokens": 548,
"output_tokens": 0
}
}"><pre class="notranslate"><code>{
"model": "clef-flash",
"answers": {
"has_ollama": {
"type": "noul",
"noul": 0.958
}
},
"usage": {
"input_tokens": 548,
"output_tokens": 0
}
}
</code></pre></div>
<h2>What's Changed</h2>
<ul>
<li>Models using web search can now perform up to ten searches per response, up from three</li>
<li>Modelfiles now support <code>CAPABILITY</code> declarations, so model creators can explicitly declare what a model can do. Declarations are preserved when creating from GGUF or safetensors, through model inheritance, and on Modelfile export</li>
<li><code>ollama show</code> and the model list now report only <code>decision</code> as the capability for decision models, so clients no longer offer them for general chat, tools, or thinking</li>
<li>Updated llama.cpp and the MLX engine</li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.35.0...v0.35.1"><tt>v0.35.0...v0.35.1</tt></a></p>github-actions[bot]tag:github.com,2008:Repository/658928958/v0.35.1-rc22026-10-02T01:05:01Zv0.35.1-rc2<p>ci: fix missing build context (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5671220102" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18742" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18742/hovercard" href="https://github.com/ollama/ollama/pull/18742">#18742</a>)</p>dhiltgentag:github.com,2008:Repository/658928958/v0.35.1-rc12026-10-01T22:52:04Zv0.35.1-rc1<p>models: add clef support (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5669656657" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18741" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18741/hovercard" href="https://github.com/ollama/ollama/pull/18741">#18741</a>)</p>jmorgancatag:github.com,2008:Repository/658928958/v0.35.1-rc02026-09-29T18:10:25Zv0.35.1-rc0: create: support explicit model capabilities (#18708)<p>Add CAPABILITY declarations to Modelfiles and an additive capabilities<br>
field to create requests. Preserve declarations across GGUF and safetensors<br>
creation, inheritance, and Modelfile export.</p>
<p>Require decision capability before scheduling System One requests instead<br>
of matching Qwen architecture/renderer metadata. Retain main's GGUF-only<br>
scoring restriction until the separate MLX runtime work lands.</p>
<p>Extracted from the capability foundation in <a class="commit-link" data-hovercard-type="commit" data-hovercard-url="https://github.com/ollama/ollama/commit/36d46a0c33b6e4bbb5daf319f46f71433b70b96b/hovercard" href="https://github.com/ollama/ollama/commit/36d46a0c33b6e4bbb5daf319f46f71433b70b96b"><tt>36d46a0</tt></a> on system_one_mlx;<br>
MLX scoring and manifest-list changes are intentionally separate.</p>dhiltgentag:github.com,2008:Repository/658928958/v0.35.02026-09-29T23:56:54Zv0.35.0<h2>Decision models</h2>
<p>Ollama now supports decision models through <code>/v1/systemone</code>, based on <a href="https://typesafe.ai" rel="nofollow">TypeSafe’s Jev API</a>.</p>
<p>Decision models return choices, probabilities, and scores instead of text. Use them for tasks such as ticket triage, model routing, and content classification.</p>
<p>Available models:</p>
<ul>
<li><a href="https://ollama.com/library/nimble" rel="nofollow"><strong>Nimble</strong></a> from Bespoke Labs</li>
<li><a href="https://ollama.com/library/tev1" rel="nofollow"><strong>Tev1</strong></a> from Together AI</li>
</ul>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama pull nimble"><pre>ollama pull nimble</pre></div>
<p>Send context and one or more questions:</p>
<div class="highlight highlight-source-shell notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="curl http://localhost:11434/v1/systemone \
-H 'Content-Type: application/json' \
-d '{
"model": "nimble",
"state": "Our checkout has returned 500 errors since 9am.",
"questions": {
"label": {
"type": "choice",
"instructions": "Which label fits this ticket?",
"criteria": {
"billing": "Payments and refunds",
"bug": "Software errors",
"account": "Login and account access"
}
}
}
}'"><pre>curl http://localhost:11434/v1/systemone \
-H <span class="pl-s"><span class="pl-pds">'</span>Content-Type: application/json<span class="pl-pds">'</span></span> \
-d <span class="pl-s"><span class="pl-pds">'</span>{</span>
<span class="pl-s"> "model": "nimble",</span>
<span class="pl-s"> "state": "Our checkout has returned 500 errors since 9am.",</span>
<span class="pl-s"> "questions": {</span>
<span class="pl-s"> "label": {</span>
<span class="pl-s"> "type": "choice",</span>
<span class="pl-s"> "instructions": "Which label fits this ticket?",</span>
<span class="pl-s"> "criteria": {</span>
<span class="pl-s"> "billing": "Payments and refunds",</span>
<span class="pl-s"> "bug": "Software errors",</span>
<span class="pl-s"> "account": "Login and account access"</span>
<span class="pl-s"> }</span>
<span class="pl-s"> }</span>
<span class="pl-s"> }</span>
<span class="pl-s"> }<span class="pl-pds">'</span></span></pre></div>
<p>Example response:</p>
<div class="highlight highlight-source-json notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="{
"model": "nimble",
"answers": {
"label": {
"type": "choice",
"choice": "bug",
"probabilities": {
"billing": 0.0125,
"bug": 0.9781,
"account": 0.0093
},
"confidence": 0.8906
}
},
"usage": {
"input_tokens": 174,
"output_tokens": 1
}
}"><pre>{
<span class="pl-ent">"model"</span>: <span class="pl-s"><span class="pl-pds">"</span>nimble<span class="pl-pds">"</span></span>,
<span class="pl-ent">"answers"</span>: {
<span class="pl-ent">"label"</span>: {
<span class="pl-ent">"type"</span>: <span class="pl-s"><span class="pl-pds">"</span>choice<span class="pl-pds">"</span></span>,
<span class="pl-ent">"choice"</span>: <span class="pl-s"><span class="pl-pds">"</span>bug<span class="pl-pds">"</span></span>,
<span class="pl-ent">"probabilities"</span>: {
<span class="pl-ent">"billing"</span>: <span class="pl-c1">0.0125</span>,
<span class="pl-ent">"bug"</span>: <span class="pl-c1">0.9781</span>,
<span class="pl-ent">"account"</span>: <span class="pl-c1">0.0093</span>
},
<span class="pl-ent">"confidence"</span>: <span class="pl-c1">0.8906</span>
}
},
<span class="pl-ent">"usage"</span>: {
<span class="pl-ent">"input_tokens"</span>: <span class="pl-c1">174</span>,
<span class="pl-ent">"output_tokens"</span>: <span class="pl-c1">1</span>
}
}</pre></div>
<p>The API supports three question types:</p>
<ul>
<li><code>choice</code>: Select an option and return probabilities for each.</li>
<li><code>noul</code>: Return the probability that a condition is true.</li>
<li><code>score</code>: Return a score across an ordered set of criteria</li>
</ul>
<h2>What's Changed</h2>
<ul>
<li>Settings now opens without waiting for model discovery.</li>
<li>Fixed the macOS update menu and icon not reflecting an available update at startup.</li>
<li>Fixed stalled MLX model downloads hanging indefinitely.</li>
<li>Requests containing the deprecated <code>typical_p</code> parameter now log a warning instead of failing.</li>
</ul>
<p><strong>Full Changelog:</strong> <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.4...v0.35.0"><tt>v0.34.4...v0.35.0</tt></a></p>github-actions[bot]tag:github.com,2008:Repository/658928958/v0.35.0-rc12026-09-28T20:31:38Zv0.35.0-rc1<p>mlx: bound pull stall retries and let the watchdog interrupt them (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="1777625644" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/1" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/1/hovercard" href="https://github.com/ollama/ollama/pull/1">#1</a>…</p>dhiltgentag:github.com,2008:Repository/658928958/v0.35.0-rc02026-09-28T20:21:01Zv0.35.0-rc0<p>feat: add System One scoring API (<a class="issue-link js-issue-link" data-error-text="Failed to load title" data-id="5550448056" data-permission-text="Title is private" data-url="https://github.com/ollama/ollama/issues/18606" data-hovercard-type="pull_request" data-hovercard-url="/ollama/ollama/pull/18606/hovercard" href="https://github.com/ollama/ollama/pull/18606">#18606</a>)</p>ParthSareentag:github.com,2008:Repository/658928958/v0.40.0-rc02026-09-25T15:22:00Zv0.40.0<h2>What's Changed</h2>
<p><strong>Models run on MLX on Apple Silicon by default</strong></p>
<p>In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.</p>
<div class="snippet-clipboard-content notranslate position-relative overflow-auto" data-snippet-clipboard-copy-content="ollama pull qwen3.8
ollama run qwen3.8"><pre class="notranslate"><code>ollama pull qwen3.8
ollama run qwen3.8
</code></pre></div>
<p>During the pre-release we will be testing and enabling additional models.</p>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.4...v0.40.0-rc0"><tt>v0.34.4...v0.40.0-rc0</tt></a></p>github-actions[bot]tag:github.com,2008:Repository/658928958/v0.34.42026-09-24T04:45:43Zv0.34.4<h2>What's Changed</h2>
<ul>
<li>Structured outputs on thinking models now apply in a single pass, making them faster and more reliable.</li>
<li>Fixed intermittent "model not found" errors with a large local library</li>
<li>Fixed the macOS app becoming unresponsive when checking if ChatGPT or Codex is running.</li>
<li>Qwen 3.8 prompt processing is faster on Apple Silicon.</li>
<li>Gemma 4 on Apple Silicon now picks the best image resolution per image, keeping more detail in high-resolution images.</li>
<li>Updated llama.cpp, MLX, and XGrammar.</li>
</ul>
<p><strong>Full Changelog</strong>: <a class="commit-link" href="https://github.com/ollama/ollama/compare/v0.34.3...v0.34.4"><tt>v0.34.3...v0.34.4</tt></a></p>github-actions[bot]tag:github.com,2008:Repository/658928958/v0.34.4-rc12026-09-23T23:36:46Zv0.34.4-rc1: mlxrunner: Update XGrammar to 0.2.7 for structured outputs<p>We pick up schema fixes for typed dictionary values and short arrays.</p>jessegross