--- snapshot-1781342047+++ snapshot-1789251377@@ -2,44 +2,44 @@ Release notes from whisper.cpp -2026-06-01T11:56:20Z tag:github.com,2008:Repository/541269386/v1.8.6 2026-06-02T06:22:27Z +2026-09-11T05:31:53Z tag:github.com,2008:Repository/541269386/v1.9.4 2026-09-11T05:31:55Z + +v1.9.4 + +

Overview

New version has been released.

Nightly build: b5130
More info: dist : releases and versioning of ggml-org projects

Changelog since v1.9.3

927cfce metal : remove leftover ggml-metal.metal kernels file (#4051)
dd80eb2 scripts : use sed instead of grep for version parsing [no ci] (#4052)
1fa6dfa ci : update WoA CUDA 13.4 to use 13.4.1 GA (#4053)
a2b36eb whisper : bump version to 1.9.4 (#4050)
6fb4cd6 ci : add Windows On ARM support to release job (#4048)
c44b60b whisper : call encoder_begin_callback before language auto-detect (#3936)
6d0ed91 whisper : re-seed decoder 0 between calls (#4025)
cec4dbe ci : update close-issue job to not close issues (#4045)
79f2d92 tests : load backends before init when built with GGML_BACKEND_DL (#4031)
61e6cca server : return language in detect response (#4035)
52a939a sync : ggml
a937f4e ggml : bump version to 0.23.0 (ggml/1618)
11d4eec metal : add remaining fa-vec tunings for M3 Max (llama/28373)
140e57a ggml : replace compile definitions with version.h.in (llama/28364)
e2389eb ggml : rename and make private ggml_op_alloc_size_may_expand() (ggml/0)
1b37bea ggml : don't crash when backend search path can't be read (llama/28271)
f32e6fa ggml : remove GGML_CUDA_PEER_MAX_BATCH_SIZE (llama/28177)
e1bbe40 ggml-cpu(s390x) : fix q5_1 uninitialized v_acc (llama/28332)
d1e0e64 sycl: fuse rms_norm+mul+add and add+add residual chains (llama/27610)
36f170e SYCL: Refactor GGML_SYCL_ENABLE_MKL_FA to global var (llama/26863)
d784add opencl: quant lm_head / decode GEMV and medium-batch GEMM optimizations (speculative decoding/MTP) (llama/26477)
0a4a95c tune MMVQ to MMQ crossover for SM87 (llama/28285)
4dd48dd metal : add sparse FA (llama/28098)
d55d345 metal : fix glu dispatch with ne00 = 1 (llama/28306)
25350b5 CUDA: Allow concurrent streams per split for multi-GPU (llama/28198)
47d348a vulkan: fix FA dequant path engagement (llama/28190)
f24a386 sycl : enhance the api to support peer-to-peer copy (llama/27550)
a704770 sycl: reduce redundant work in Q4_K multi-column MMVQ (llama/27062)
e560569 finetune: fix no KV cache (llama/27199)
37f0f44 ggml-hexagon: add F16 support for unary ops (llama/28228)
1bdda1e metal : add fa-vec tunings for M3 (llama/28236)
3a1c7d6 metal : fix memory query under low-memory conditions (llama/27701)
4d343d7 ggml-cuda : remove unused vars (llama/28235)
519df61 CUDA + ggml: add sparse-fa for DSV4/GLM (llama/27970)
c2b4007 ggml: avoid KleidiAI buffer type init on dispatch (llama/27891)
dc70853 hexagon: MUL_MAT and MUL_MAT_ID fusion and fixes (llama/28202)
1c7d35e vulkan: handle larger batch sizes (>4) efficiently for IQ3_S mat-vec (llama/27449)
a9e5861 vulkan : only request VK_KHR_shader_bfloat16 extension if supported (llama/28155)
d57ae98 ggml-cpu : conditionally add SpacemiT IME kernel sources (llama/27961)
35133c9 opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632)
c94921f hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (llama/28217)
fcc2fee metal : add metallib build support for xcframework (llama/28163)
2c48678 cuda: fuse MoE weighted expert reduction (llama/25952)
f162a19 Revert "sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…" (#28184)
408faaa sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 1280 (llama/28016)
5f07f85 metal : add fa-vec tuning for M2 Pro (llama/28122)
8cca1a3 metal : add fa-vec tunings for A18 Pro (MacBook Neo) (llama/28152)
a245a8f metal : fix more leaks due to missing autoreleasepools (llama/27883)
870db2a metal : add fa-vec tuning for M2 Max (llama/28015)
4f3a2a4 sycl : support limit max alloc memory within 2GB for host-pinned memory (llama/27559)
5032008 metal: enable Metal 4.0 tensor API on M5+/A19+ (llama/27461)
8e54c65 metal : add fa-vec tunings for M1 Ultra (llama/28088)
f22bb2e CUDA: XOR swizzle flash attn K,V smem fp16 tiles (llama/25635)
dbc40ef metal : add concat support for quantized types (llama/28116)
2f608ab AVX2: Speed up large batch size prompt processing of IQ models (llama/27402)
c6934d0 metal : add top-k radix implementation (llama/28073)
088c603 opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG and PP performance (llama/26438)
c648b9a webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation (llama/28045)
c1be45b ROCm: add radix TOP_K for long rows (llama/27466)
7614a4c metal : add fa-vec tunings for M1 (llama/28078)
b0f4bc0 CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-router fusion were restricted to 1 token (llama/27621)
76a51e8 sycl : Enhance to get the free memory of Intel GPU (llama/27968)
6ce7b89 vulkan: tune mat-vec rows for batched inference on Strix Halo (llama/27909)
96dddd8 ggml : add MUL_MAT to the list of ops that may need additional memory (for WebGPU) (llama/28071)
db00b01 vulkan: top_k radix select for k >= 1024 for Qwen 3.8 Flash Next (llama/28032)
01ebd22 hexagon: fix CPY fence bug (llama/28033)
e5c96ca metal : add remaining Q4_1/Q5_0/Q5_1 fa-vec tunings for M2 (llama/28017)
4089fa6 rpc: avoid serializing buffers from other servers (llama/26500)
749683d ggml : fix ggml_backend_buft_get_alloc_size() guard (llama/28038)
e9583f0 ggml: add SWIGLU_CLAMP (llama/27930)
e900a73 CUDA: use the fast mm_ids_helper path for any n_expert_used (llama/27978)
35d9e22 hip: tune rdna 3 mmq config (llama/26284)
e5c9e3e hip : optimize Q2_0 dot-product path for gfx1201 (llama/26753)
4b2243a ggml : add ggml_backend_op_alloc_size_may_expand, use it in RPC (llama/27960)
43acf3d rpc: fix apple rdma error spew on teardown (llama/27908)
1e0f382 metal: add fa-vec tunings for M3 Ultra (llama/27999)
b66593e metal : Add fa-vec tuning for M3 Pro (llama/27963)
5e49459 rpc : fix pre-rdma macOS versions (llama/27815)
3ad8b9b hexagon: support for device discovery and create sessions on demand (llama/27785)
3d4e0e9 sycl: split long rows in TOP_K instead of one work-group per row (llama/27847)
c68f205 metal : fix null-pipeline crash for F16 src1 mul_mat/mul_mat_id (llama/25648)
c969c68 ggml: allow passing alloc dependencies in graph_optimize (llama/27301)
b33bbc5 metal : add fa-vec tunings for M2 (llama/27940)
2a11026 opencl: use a better matmul path on two Adreno GPU generations (llama/27640)
285f1ff metal : assert shared memory padding (llama/27951)
308fa4f metal : add remaining fa-vec tunings for M4 Pro (llama/27915)
325c8d1 sycl: make --fit respect --fit-target better (llama/27629)
e644752 vulkan: combine duplicated fastdiv functions, rename the one optimizing small divs (llama/27526)
590fe18 metal : add fa-vec tunings for M1 Max (llama/27932)
d501a0a vulkan: Change mul_mat_id to pad K rather than N (llama/27925)
4c38040 vulkan: fix missing view-alias dependencies in ggml_vk_graph_optimize (llama/27812)
caea96f ggml : fix conv_transpose_2d for multiple batches (llama/26132)
ba99c09 Vulkan: add hoisting support for row IDs and expert count in shaders (llama/26686)
0a15087 metal : add fa-vec tunings for M4 (llama/27875)
7f78e1b OpenVINO: Update OV to 2026.3.1, whisper.cpp support, Qwen3.5 on NPU, and new ops (llama/27843)
fa4d244 sycl: use TILE for quantized KV decode on BMG (llama/26689)
97d0da2 sycl: bind the f16 KV cache in place for the oneDNN SDPA path (llama/27468)
530e3f4 metal : add fa-vec tunings for M3 Max, M5 and M5 Pro (llama/27863)
ff38b98 metal : add fa-vec tunings for M4 Pro (llama/27824)
b6571e4 ggml-hexagon: add HTP unary ops for ABS and LOG (llama/27786)
a0614d9 hex-unary: fix RMS_NORM_MUL weight-offset bugs for grouped/broadcast norms (llama/27798)
8529971 opencl: add bin kernels kernel_gemm_moe_q4_0_q8_1_dp4a_bin, kernel_gemm_moe_mxfp4_q8_1_dp4a_bin (llama/27768)
55ab1e5 Feature: Added LIGHTNING_INDEXER support for Deepseek V4 ops on Vulkan Backend (llama/27453)
a5db1d6 metal : fix memory leaks due to missing autoreleasepools (llama/27758)
3fea10d hexagon: support for multi-NPU devices (IQ9, IQ10) and fully asynchronous backend (llama/26501)
5271734 vulkan: warptiles currently assume warp sizes <= 64, clamp to work around larger warps (llama/27726)
0a02697 Implemented vulkan cross_entropy_loss and cross_entropy_loss_back (llama/27216)
82f5f85 rpc : implement event and async backend APIs (llama/18626)
9d8e6b9 cuda: unblock mmq for MoE on sm_60 (llama/26264)
8c0adb0 ggml-metal: add chunked SSD MMA for Mamba-2 prefill optimization (llama/26647)
8df657a ggml-meta: propagate buffer usage and call init on the new tensors (llama/27586)
482956e kleidiai: Rework KleidiAI Build System/Integration (llama/26077)
e820c28 rpc: support apple RDMA as an RPC transport (llama/26421)
be12d39 metal : null-check buffer alloc to fix OOM crash (llama/25371)
642b5d3 ruby : Add #free method, check MemoryView strictly (#4032)
eacbd82 whisper : default-initialize whisper_mel to avoid uninitialized read (#3981)
c4ac001 parakeet : fix TDT decode by outputting raw logits from the joint graph (#4017)
9781133 talk-llama : sync llama.cpp
3680f66 pi : init
0414519 sync : ggml
d470c9d ggml : bump version to 0.22.0 (ggml/1607)
322a77c sycl : mark tq2_0 as not supported (llama/27660)
17a522a webgpu : fix handling of infinity values during ARGSORT and TOP_K (llama/27538)
fa3b87c metal : per-device tuned (Q, NE) for flash-attn vec (llama/26570)
fe52277 sync : ggml
b15d31d metal: per-op source split + parallel compile (llama/26561)
4257445 scripts : update ggml-am
1d8e052 sync : ggml
aa25d33 ggml : shorten virtual device naming in CUDA and Metal (llama/27608)
103305e webgpu : reorder includes since V that appears in common_decls.tmpl may be defined as K in flash_attn_decls.tmpl if KV_OVERLAP (llama/27545)
32d0f49 ggml : fix ggml_clamp (llama/27644)
20209c2 Deepseek 4: -sm tensor (llama/26490)
3b89b37 Fix meta tensor split state propagation (llama/27574)
c8a4009 cuda : add POOL_1D support (llama/27573)
13a7856 vulkan : added the PAD_REFLECT_1D operation (llama/26586)
1efb31e ggml: optimize concat op by replacing per-element memcpy with row-level memcpy (llama/24575)
21a67dd sycl : add Q2_K reordered MMVQ and ESIMD kernels (again) (llama/27490)
5f7bd9d opencl: fold the gpt-oss MoE per-expert bias adds into the epilogue (op/kernel fusion) (llama/26431)
a722846 whisper : guard null source in buffer loader read callback (#3982)
c122757 docs : center badges in README.md [no ci] (#4012)
2569409 devops : add main-rocm Dockerfile (#3975)
52dec9d vitisai : add VitisAI Plugin for AMD Ryzen AI NPU encoder offload (#3608)
233fe1f whisper : bypass cross-attention scaling for OpenVINO backend (#3997)
51de5e8 openvino : update model conversion and README.md (#4003)
3391d6b scripts : add release.sh script (#4010)
a4610c7 docs : add release badge and remove stable/roadmap [no ci] (#4009)
ab57887 make : add --parallel to cmake build command (#4007)
45f1593 sync : ggml
ce77728 Revert "sycl : add Q2_K reordered MMVQ and ESIMD kernels (llama/26336)" (llama/27486)
0d9ba28 ggml : bump version to 0.21.0 (ggml/1597)
d6c416e kleidiai : add SME2 F32 GEMV kernel support (llama/26891)
d60ef65 sycl : add Q2_K reordered MMVQ and ESIMD kernels (llama/26336)
b1cb805 sycl : Add Q5_K ESIMD kernel (llama/26376)
2cb52dd opencl: keep the vocab-scale K-quant lm_head on the CPU for Adreno A7X (compiler issue workaround) (llama/26440)
af74f97 sycl: Update gate logic for Alchemist GPUs regarding OneDNN features. (llama/26635)
f19250e sycl: fix multiple warnings in compiling sycl backend (llama/26713)
12137c3 sycl : fix load model with mlock issue (llama/27250)
5656e44 ggml: support ggml_rope_set_offset on opencl, sycl, wgpu, hexagon (llama/27345)
d68216a metal : clamp K extent in tensor API mat-mat kernel for K not a multiple of 32 (llama/27450)
c16cb42 opencl: fix q6_K flat mul_mat for Adreno A6x/A7x GPUs with older E031 compilers (llama/26476)
73c2b7e opencl: fix local size for norm (llama/27339)
8539d72 vulkan: FA MMQ should use fp32 for Q quantization calculations (llama/27413)
60f14a5 metal : dequant kv cache only for large batches (llama/27438)
2922418 CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows (llama/26678)
c2dc946 CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover (llama/26079)
283775e metal : dequantize quantized KV to F16 before flash attention (llama/27390)
02be8f5 Revert "tensor-split meta backend fixes (ggml/26502)" (llama/27433)
13fa860 ggml: fix backend split scheduler race condition (llama/26040)
2648a70 ggml-cuda: provide static workspace for cuBLAS handles (llama/26574)
acfad32 vulkan : add source groups for shaders (llama/26666)
a3386d4 opencl: make the MoE expert scatter deterministic (llama/26464)
135f348 tensor-split meta backend fixes (llama/26502)
9f4b18a hexagon: fix FA HMX queue ordering and pack the rescale D matrices (llama/27042)
cd340ea opencl: port fused ssm_scan kernel (Mamba-2, d_state in {128, 256}) to GPU (llama/26439)
3d27742 ggml-cpu: gate __fp16 on __ARM_FP16_FORMAT_IEEE (llama/26860)
11e58f3 vulkan : dequant q8_0 KV once in coopmat1 (llama/25494)
4ef3e87 vulkan: add null checks in ggml_vk_queue_command_pools_cleanup (llama/27353)
689ad69 sycl: report zero devices instead of aborting when the host has none (llama/27291)
8442c74 ggml: add ggml_rope_set_offset (+ metal support) (llama/27120)
d830bd2 metal : dequantize q8_0 using packed types (llama/27370)
1c882a8 vulkan: tiled transpose for 0<->2 permuted CONT (llama/26585)
7df5fa8 ggml-webgpu: add mulmat with overlapping src0/src1 (e.g., for minimax-01) (llama/27321)
fa1e2bc opencl: fix WAR race in the generic FA tile kernels when the WG spans subgroups (llama/26434)
4be3101 RPC: populate use_count to enable fusion inside backends (llama/27142)
dbab353 sycl: honor GGML_HINT_SRC0_IS_HADAMARD (llama/27298)
c6a2bd0 devops : stop shadowing proper cuda libraries in runtime docker build (#3989)
ec73735 tests: add check for nullptr for wctx in test-vad-full (#3998)
a4ad15f ci : release clean-up (#4005)
81a3fad server : only enable token timestamps when the response needs them (#3990)
d61695d docs : fix typos in readme files (#4004)
b78df3d ci : move close-issue.yml to .github/workflows [no ci] (#4002)
339f2b4 bindings-javascript : remove package.json from git (#4001)

github-actions[bot] tag:github.com,2008:Repository/541269386/b5130 2026-09-11T05:30:01Z + +b5130 + +

metal : remove leftover ggml-metal.metal kernels file (#4051)

This commit removes the ggml-metal.metal file that used to contain all
the metal kernels. These have now been separated into separate kernels
in ggml/src/ggml-metal/kernels and it seems that this file was not
removed when synced with ggml.

This is currently preventing whisper.cpp to be released as the release
check fails.

github-actions[bot] tag:github.com,2008:Repository/541269386/b5127 2026-09-10T11:56:28Z + +b5127 + +

Note: the Windows arm64 (CUDA) build (whisper-bin-win-cuda-*-arm64.zip) uses a preview edition of the CUDA Toolkit for Windows on Arm and should be considered experimental.

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.9.3 2026-08-20T11:42:23Z + +v1.9.3 + +

Note

Semantic versioning is still work in progress.
More info can be found in https://github.com/ggml-org/ggml/discussions/1579

Nightly build: b4938

Change log since v1.9.2

371b5a7 release : v1.9.3 (#4000)
81c1905 cmake : update semver and release process [no ci] (#3996)
4834a23 talk-llama : sync llama.cpp
6b014cf sync : ggml
8189458 ggml : bump version to 0.20.2 (ggml/1589)
51319a2 CUDA: MMVQ nwarps=8 for bs=1 for dense models on DGX Spark (llama/26843)
71759b7 cuda : skip UMA override for HIP builds (llama/27083)
964bb1b ggml : bump version to 0.20.1 (ggml/1587)
4257f47 sycl: fix thread/block count in quantized cpy kernel launches (llama/27160)
0f38613 support OP OPT_STEP_ADAMW, OPT_STEP_SGD (llama/25268)
9a0d190 vulkan: add SHMEM_STRIDE_PAD/APPLY_SLM_A_RESHAPE for coopmat1 on Intel Xe (llama/25380)
ab71410 fixed indent
9b98e57 Fixed gating logic for problematic Intel driver version
1fe009c talk-llama : fix build (#0)
733f281 sync : ggml
667da04 ggml : bump version to 0.20.0 (ggml/1584)
b7ea8b1 ggml : recurrent state rollback for ggml_ssm_scan (llama/26623)
2aef2a0 sycl: fuse mul_mat(gate) + mul_mat(up) + GLU for q4_K dense FFN (llama/26779)
5bc0c85 ggml: force single thread on wasi (llama/25686)
2d05b6e sycl: fuse the gated-delta-net state writeback cpy (llama/26643)
43cbe41 OpenVINO: Qwen3.5, memory optimization, and test-recurrent-state-rollback (llama/26952)
94d2d35 Support host pinned mem to improve SYCL Host-to-Device Memory Access (llama/26789)
34009e8 metal : add TQ2_0 support (llama/26980)
62031fe ggml-cpu/ops: vectorize flash-attention V-cache F16 to F32 conversion (llama/26947)
ac9a74f sycl: remove separate fp32 type promotion in gemm non-oneDNN path (llama/26372)
3425d13 sycl: fuse UNARY(silu|sigmoid|softplus) + MUL (llama/26411)
90d4ed1 sycl : Add DMMV ESIMD Q3_K kernel (llama/26251)
f701160 sycl : enhance concat to support Q4_0, Q4_1, Q5_0, Q5_1, Q8_0 (llama/26800)
c7be8b3 ggml-hip : remove -funsafe-math-optimizations (llama/26696)
406b116 ggml : fix arm builds, unused var (llama/26991)
b0e3297 gguf : harden loader against malformed tensor dims and metadata types (llama/25596)
92423af kleidiai: Add runtime feature detection mechanism for aarch64/kleidiai (llama/26076)
1b93067 opencl: default FA c8 cluster width to 16 on X1E (llama/26433)
6329040 vulkan: add TQ2_0 (ternary) support (llama/25850)
db3687b opencl: use flat mv q5_k when weight exceeds image1d_buffer_t limit (llama/26880)
cdbe455 CUDA: only disable CUDA graphs when mul_mat_id actually needs a stream sync (llama/26802)
030656f cuda : add warp-per-row wkv7 kernel for single-token decode (llama/26111)
d958968 llama: add default load-mode auto, which avoids mmap on iGPUs (llama/26081)
0f60f8e ggml-webgpu: fix CI errors from #25025 and #25262 (llama/26566)
a6e2630 opencl: transpose the K tile in local memory for FA prefill kernels (llama/26428)
830ec22 ggml-cpu : fix CPU affinity mask being ignored on Android (llama/26838)
b3bc904 ggml : require contiguous src for ROLL on CUDA and Metal (llama/25928)
f09a97c ggml-webgpu : refactor several wgsl files and simplify flash_attn wgsl. (llama/26134)
877761c ggml-cpu : fix missing Q5_0 dispatch in SpaceMiT backend (llama/26792)
10791af CUDA: fuse rms_norm + mul + rope (+ view + set_rows) (llama/26767)
068d3b0 CUDA: fix thread/block count in quantized cpy kernel launches (llama/26731)
eb3296f sycl: coalesce the ssm_conv window loads (llama/26612)
8a5ba01 metal : fix NORM/RMS_NORM for row lengths that leave a partial simdgroup (llama/26708)
06248ca cmake : add config version support (ggml/1582)
592feef talk-llama : sync llama.cpp
8770492 sync : ggml
84cdcad ggml : bump version to 0.19.0 (ggml/1581)
8587ad3 ggml : add aarch64 HWCAP fallbacks and fix fp16 variant detection (llama/25554)
56cb154 sycl: fix UE4M3 parsing (llama/25608)
077c5d4 sycl: *glu flat path (llama/26354)
9faa9ee sycl : Support DSv4 OPs: LIGHTNING_INDEXER,DSV4_HC_COMB,DSV4_HC_POST,DSV4_HC_PRE (llama/26568)
fb9e8ca sycl : fix error Error OP FLASH_ATTN_EXT on arc770 (llama/26441)
89d45af sycl : enhance OP set_rows to support all missed data types (llama/26515)
79ab70c cuda: fix warnings for unused variable/function (llama/26688)
5693378 metal : avoid threadgroup matrix array instantiation in kernel_lightning_indexer (llama/26646)
69bd0a9 ci : onboard AMD ROCm CI with gfx1151 fixes (llama/26544)
5a80d0a vulkan: fix submission batching size, add debug tools for diagnosing causes of DeviceLost drivers errors (llama/26371)
60894f1 mtmd/ggml: add ggml_build_forward_order (llama/26649)
8731021 vulkan backend ops: implemented GATED_LINEAR_ATTN (llama/25601)
8631825 whisper : heap out-of-bounds read in log_mel_spectrogram on very short audio (#3956)
df1547b whisper,parakeet : reject invalid n_dims in tensor header to prevent stack-buffer-overflow on malformed model files (#3957)

github-actions[bot] tag:github.com,2008:Repository/541269386/b4938 2026-08-20T11:33:53Z + +b4938 + +No content. github-actions[bot] tag:github.com,2008:Repository/541269386/v1.9.2 2026-08-04T15:31:37Z + +v1.9.2 + +

What's Changed

New Contributors

Full Changelog: v1.9.1...v1.9.2

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.9.1 2026-06-19T05:53:19Z + +v1.9.1 + +

What's Changed

Full Changelog: v1.9.0...v1.9.1

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.9.0 2026-06-17T10:19:21Z + +v1.9.0 + +

What's Changed

Full Changelog: v1.8.7...v1.9.0

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.8.7 2026-06-16T10:59:32Z + +v1.8.7 + +

What's Changed

New Contributors

Full Changelog: v1.8.6...v1.8.7

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.8.6 2026-06-02T06:22:27Z v1.8.6 -

What's Changed

Full Changelog: v1.8.5...v1.8.6

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.8.5 2026-05-29T12:07:04Z - -v1.8.5 - -

Overview

Maintenance release + performance improvements all around:

https://github.com/ggml-org/whisper.cpp/blob/master/scripts/bench-all-gg.txt

What's Changed

New Contributors

Full Changelog: v1.8.4...v1.8.5

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.8.4 2026-03-19T15:08:38Z - -v1.8.4 - -

Overview

Maintenance release, latest ggml, some performance gains across the board.

What's Changed

New Contributors

Full Changelog: v1.8.3...v1.8.4

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.8.3 2026-01-15T18:47:18Z - -v1.8.3 - -

Overview

Maintenance release, latest ggml, minor improvements in the tools/server/bindings.

What's Changed

New Contributors

Full Changelog: v1.8.2...v1.8.3

ggerganov tag:github.com,2008:Repository/541269386/v1.8.2 2025-10-15T08:32:18Z - -v1.8.2 - -

Overview

What's Changed

Full Changelog: v1.8.1...v1.8.2

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.8.1 2025-10-12T10:18:45Z - -v1.8.1 - -

Overview

What's Changed

New Contributors

Full Changelog: v1.8.0...v1.8.1

github-actions[bot] tag:github.com,2008:Repository/541269386/v1.8.0 2025-10-01T06:59:18Z - -v1.8.0 - -

Overview

M1 Pro

CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M1 Pro METAL tiny 1 0 32.44 1.71 0.43 0.04 8a67c55
M1 Pro METAL base 1 0 63.54 2.62 0.71 0.06 8a67c55
M1 Pro METAL small 1 0 200.30 5.34 1.72 0.17 8a67c55
M1 Pro METAL medium 1 0 580.06 11.71 4.18 0.45 8a67c55
CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M1 Pro METAL tiny 1 1 22.09 1.84 0.43 0.03 8a67c55
M1 Pro METAL base 1 1 40.57 2.22 0.44 0.04 8a67c55
M1 Pro METAL small 1 1 135.15 4.23 0.95 0.12 8a67c55
M1 Pro METAL medium 1 1 395.18 9.14 2.21 0.30 8a67c55

M2 Ultra

CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M2 ULTRA METAL tiny 1 0 8.63 1.09 0.27 0.01 b57b9d3
M2 ULTRA METAL tiny-q5_0 1 0 9.04 1.06 0.28 0.01 b57b9d3
M2 ULTRA METAL tiny-q5_1 1 0 8.98 1.06 0.28 0.01 b57b9d3
M2 ULTRA METAL tiny-q8_0 1 0 8.69 1.06 0.27 0.01 b57b9d3
M2 ULTRA METAL base 1 0 15.39 1.54 0.43 0.02 b57b9d3
M2 ULTRA METAL base-q5_0 1 0 16.50 1.50 0.42 0.02 b57b9d3
M2 ULTRA METAL base-q5_1 1 0 16.45 1.49 0.43 0.02 b57b9d3
M2 ULTRA METAL base-q8_0 1 0 15.62 1.51 0.42 0.02 b57b9d3
M2 ULTRA METAL small 1 0 45.99 2.99 0.90 0.05 b57b9d3
M2 ULTRA METAL small-q5_0 1 0 50.65 2.98 0.92 0.06 b57b9d3
M2 ULTRA METAL small-q5_1 1 0 50.74 2.96 0.92 0.06 b57b9d3
M2 ULTRA METAL small-q8_0 1 0 47.16 2.83 0.89 0.06 b57b9d3
M2 ULTRA METAL medium 1 0 132.78 6.46 2.02 0.13 b57b9d3
M2 ULTRA METAL medium-q5_0 1 0 149.35 6.11 2.09 0.14 b57b9d3
M2 ULTRA METAL medium-q5_1 1 0 149.11 6.09 2.11 0.14 b57b9d3
M2 ULTRA METAL medium-q8_0 1 0 137.37 6.05 2.03 0.13 b57b9d3
M2 ULTRA METAL medium-dis 1 0 121.60 0.90 0.25 0.02 b57b9d3
M2 ULTRA METAL large-v2 1 0 231.19 9.40 3.10 0.22 b57b9d3
M2 ULTRA METAL large-v2-q5_0 1 0 265.90 8.98 3.11 0.25 b57b9d3
M2 ULTRA METAL large-v2-q5_1 1 0 265.18 8.92 3.13 0.25 b57b9d3
M2 ULTRA METAL large-v2-q8_0 1 0 240.23 9.06 2.98 0.23 b57b9d3
M2 ULTRA METAL large-v2-dis 1 0 210.25 0.99 0.28 0.02 b57b9d3
M2 ULTRA METAL large-v3-turbo 1 0 211.72 1.52 0.46 0.03 b57b9d3
M2 ULTRA METAL large-v3-turbo-q5_0 1 0 242.17 1.40 0.47 0.04 b57b9d3
M2 ULTRA METAL large-v3-turbo-q8_0 1 0 219.75 1.40 0.45 0.04 b57b9d3
CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M2 ULTRA METAL tiny 1 1 6.28 0.96 0.22 0.01 a77d11d
M2 ULTRA METAL tiny-q5_0 1 1 6.69 0.92 0.22 0.01 a77d11d
M2 ULTRA METAL tiny-q5_1 1 1 6.67 0.91 0.22 0.01 a77d11d
M2 ULTRA METAL tiny-q8_0 1 1 6.34 0.92 0.21 0.01 a77d11d
M2 ULTRA METAL base 1 1 10.77 1.30 0.32 0.02 a77d11d
M2 ULTRA METAL base-q5_0 1 1 11.84 1.23 0.33 0.02 a77d11d
M2 ULTRA METAL base-q5_1 1 1 11.95 1.24 0.33 0.02 a77d11d
M2 ULTRA METAL base-q8_0 1 1 11.14 1.23 0.32 0.02 a77d11d
M2 ULTRA METAL small 1 1 32.12 2.43 0.65 0.04 a77d11d
M2 ULTRA METAL small-q5_0 1 1 36.95 2.42 0.68 0.04 a77d11d
M2 ULTRA METAL small-q5_1 1 1 37.40 2.42 0.68 0.04 a77d11d
M2 ULTRA METAL small-q8_0 1 1 33.48 2.30 0.65 0.04 a77d11d
M2 ULTRA METAL medium 1 1 89.28 5.05 1.46 0.09 a77d11d
M2 ULTRA METAL medium-q5_0 1 1 105.24 4.89 1.48 0.11 a77d11d
M2 ULTRA METAL medium-q5_1 1 1 105.28 4.98 1.49 0.11 a77d11d
M2 ULTRA METAL medium-q8_0 1 1 93.61 4.89 1.43 0.10 a77d11d
M2 ULTRA METAL medium-dis 1 1 78.44 0.81 0.20 0.01 a77d11d
M2 ULTRA METAL large-v2 1 1 165.69 7.50 2.16 0.17 a77d11d
M2 ULTRA METAL large-v2-q5_0 1 1 199.40 7.37 2.18 0.20 a77d11d
M2 ULTRA METAL large-v2-q5_1 1 1 199.29 7.37 2.21 0.20 a77d11d
M2 ULTRA METAL large-v2-q8_0 1 1 174.60 6.87 2.16 0.18 a77d11d
M2 ULTRA METAL large-v2-dis 1 1 145.80 0.90 0.22 0.02 a77d11d
M2 ULTRA METAL large-v3-turbo 1 1 146.98 1.31 0.34 0.03 a77d11d
M2 ULTRA METAL large-v3-turbo-q5_0 1 1 176.77 1.19 0.35 0.03 a77d11d
M2 ULTRA METAL large-v3-turbo-q8_0 1 1 154.73 1.20 0.33 0.03 a77d11d

M4 Max

CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M4 Max METAL tiny 1 0 10.51 0.86 0.23 0.01 47fcd7d
M4 Max METAL tiny-q8_0 1 0 10.73 0.84 0.24 0.01 47fcd7d
M4 Max METAL base 1 0 19.50 1.34 0.36 0.02 47fcd7d
M4 Max METAL base-q8_0 1 0 20.17 1.25 0.36 0.02 47fcd7d
M4 Max METAL small 1 0 61.91 2.77 0.78 0.06 47fcd7d
M4 Max METAL small-q8_0 1 0 64.17 2.43 0.78 0.06 47fcd7d
M4 Max METAL medium 1 0 181.50 6.44 1.85 0.15 47fcd7d
M4 Max METAL medium-q8_0 1 0 187.71 5.80 1.84 0.15 47fcd7d
M4 Max METAL large-v2 1 0 335.49 10.49 3.01 0.26 47fcd7d
M4 Max METAL large-v2-q8_0 1 0 349.89 8.65 2.97 0.27 47fcd7d
M4 Max METAL large-v3-turbo 1 0 301.34 1.83 0.49 0.04 47fcd7d
CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M4 Max METAL tiny 1 1 8.23 0.71 0.16 0.01 47fcd7d
M4 Max METAL tiny-q8_0 1 1 8.47 0.67 0.16 0.01 47fcd7d
M4 Max METAL base 1 1 15.47 1.12 0.26 0.02 47fcd7d
M4 Max METAL base-q8_0 1 1 15.70 1.05 0.27 0.02 47fcd7d
M4 Max METAL small 1 1 49.82 2.37 0.53 0.05 47fcd7d
M4 Max METAL small-q8_0 1 1 51.76 1.99 0.53 0.05 47fcd7d
M4 Max METAL medium 1 1 147.76 5.52 1.27 0.12 47fcd7d
M4 Max METAL medium-q8_0 1 1 153.98 4.59 1.24 0.13 47fcd7d
M4 Max METAL large-v2 1 1 282.89 9.06 2.11 0.22 47fcd7d
M4 Max METAL large-v2-q8_0 1 1 296.43 7.44 2.09 0.23 47fcd7d
M4 Max METAL large-v3-turbo 1 1 249.91 1.65 0.38 0.04 47fcd7d

RTX 5090

GPU Config Model Th FA Enc. Dec. Bch5 PP Commit
RTX 5090 CUDA tiny 1 0 2.06 0.55 0.13 0.00 e4bf87b
RTX 5090 CUDA tiny-q8_0 1 0 2.50 0.55 0.14 0.01 e4bf87b
RTX 5090 CUDA base 1 0 3.72 0.81 0.19 0.01 e4bf87b
RTX 5090 CUDA base-q8_0 1 0 4.35 0.79 0.20 0.01 e4bf87b
RTX 5090 CUDA small 1 0 11.24 1.55 0.38 0.02 e4bf87b
RTX 5090 CUDA small-q8_0 1 0 12.69 1.69 0.40 0.02 e4bf87b
RTX 5090 CUDA medium 1 0 31.16 3.19 0.79 0.04 e4bf87b
RTX 5090 CUDA medium-q8_0 1 0 32.74 3.43 0.80 0.05 e4bf87b
RTX 5090 CUDA large-v2 1 0 50.09 4.55 1.14 0.05 e4bf87b
RTX 5090 CUDA large-v2-q8_0 1 0 52.44 4.76 1.11 0.07 e4bf87b
RTX 5090 CUDA large-v3-turbo 1 0 46.78 0.70 0.17 0.01 e4bf87b
RTX 5090 CUDA large-v3-turbo-q8_0 1 0 48.57 0.70 0.16 0.01 e4bf87b
GPU Config Model Th FA Enc. Dec. Bch5 PP Commit
RTX 5090 CUDA tiny 1 1 1.39 0.47 0.11 0.00 e4bf87b
RTX 5090 CUDA tiny-q8_0 1 1 1.83 0.48 0.12 0.01 e4bf87b
RTX 5090 CUDA base 1 1 2.17 0.70 0.16 0.01 e4bf87b
RTX 5090 CUDA base-q8_0 1 1 2.78 0.68 0.17 0.01 e4bf87b
RTX 5090 CUDA small 1 1 5.02 1.33 0.32 0.01 e4bf87b
RTX 5090 CUDA small-q8_0 1 1 6.39 1.46 0.34 0.02 e4bf87b
RTX 5090 CUDA medium 1 1 13.89 2.68 0.64 0.03 e4bf87b
RTX 5090 CUDA medium-q8_0 1 1 15.40 2.92 0.67 0.04 e4bf87b
RTX 5090 CUDA large-v2 1 1 21.24 3.88 0.96 0.04 e4bf87b
RTX 5090 CUDA large-v2-q8_0 1 1 23.54 4.01 0.93 0.05 e4bf87b
RTX 5090 CUDA large-v3-turbo 1 1 18.18 0.62 0.15 0.01 e4bf87b
RTX 5090 CUDA large-v3-turbo-q8_0 1 1 19.89 0.61 0.14 0.01 e4bf87b

What's Changed

New Contributors

Full Changelog: v1.7.6...v1.8.0

ggerganov tag:github.com,2008:Repository/541269386/v1.7.6 2025-06-26T04:30:08Z - -v1.7.6 - -

Overview

M2 Ultra

Flash Attention ON:

CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M2 ULTRA METAL tiny 1 1 7.72 1.05 0.32 0.01 dc8dda6
M2 ULTRA METAL tiny-q5_0 1 1 8.20 0.98 0.31 0.01 dc8dda6
M2 ULTRA METAL tiny-q5_1 1 1 8.13 0.99 0.31 0.01 dc8dda6
M2 ULTRA METAL tiny-q8_0 1 1 7.96 0.93 0.30 0.01 dc8dda6
M2 ULTRA METAL base 1 1 13.52 1.39 0.35 0.02 dc8dda6
M2 ULTRA METAL base-q5_0 1 1 14.88 1.31 0.34 0.02 dc8dda6
M2 ULTRA METAL base-q5_1 1 1 14.76 1.33 0.34 0.02 dc8dda6
M2 ULTRA METAL base-q8_0 1 1 14.04 1.28 0.34 0.02 dc8dda6
M2 ULTRA METAL small 1 1 38.78 2.72 0.67 0.04 dc8dda6
M2 ULTRA METAL small-q5_0 1 1 44.01 2.64 0.69 0.05 dc8dda6
M2 ULTRA METAL small-q5_1 1 1 44.02 2.66 0.69 0.05 dc8dda6
M2 ULTRA METAL small-q8_0 1 1 40.79 2.49 0.67 0.05 dc8dda6
M2 ULTRA METAL medium 1 1 104.48 5.57 1.61 0.10 dc8dda6
M2 ULTRA METAL medium-q5_0 1 1 122.24 5.00 1.58 0.12 dc8dda6
M2 ULTRA METAL medium-q5_1 1 1 121.99 5.02 1.59 0.12 dc8dda6
M2 ULTRA METAL medium-q8_0 1 1 111.68 4.99 1.52 0.11 dc8dda6
M2 ULTRA METAL medium-dis 1 1 93.23 0.87 0.21 0.01 dc8dda6
M2 ULTRA METAL large-v2 1 1 189.82 8.36 2.35 0.19 dc8dda6
M2 ULTRA METAL large-v2-q5_0 1 1 225.73 7.34 2.40 0.22 dc8dda6
M2 ULTRA METAL large-v2-q5_1 1 1 225.88 7.60 2.40 0.22 dc8dda6
M2 ULTRA METAL large-v2-q8_0 1 1 203.55 7.32 2.26 0.20 dc8dda6
M2 ULTRA METAL large-v2-dis 1 1 168.20 0.98 0.24 0.02 dc8dda6
M2 ULTRA METAL large-v3-turbo 1 1 170.22 1.46 0.37 0.03 dc8dda6
M2 ULTRA METAL large-v3-turbo-q5_0 1 1 201.88 1.27 0.38 0.04 dc8dda6
M2 ULTRA METAL large-v3-turbo-q8_0 1 1 182.37 1.24 0.36 0.03 dc8dda6

Flash Attention OFF:

CPU Config Model Th FA Enc. Dec. Bch5 PP Commit
M2 ULTRA METAL tiny 1 0 10.15 1.20 0.36 0.01 dc8dda6
M2 ULTRA METAL tiny-q5_0 1 0 10.21 1.15 0.39 0.01 dc8dda6
M2 ULTRA METAL tiny-q5_1 1 0 9.26 1.15 0.38 0.01 dc8dda6
M2 ULTRA METAL tiny-q8_0 1 0 9.00 1.12 0.37 0.01 dc8dda6
M2 ULTRA METAL base 1 0 15.77 1.73 0.45 0.02 dc8dda6
M2 ULTRA METAL base-q5_0 1 0 16.90 1.63 0.44 0.02 dc8dda6
M2 ULTRA METAL base-q5_1 1 0 16.93 1.64 0.44 0.02 dc8dda6
M2 ULTRA METAL base-q8_0 1 0 16.13 1.63 0.43 0.02 dc8dda6
M2 ULTRA METAL small 1 0 45.15 3.45 0.92 0.05 dc8dda6
M2 ULTRA METAL small-q5_0 1 0 50.63 3.36 0.94 0.06 dc8dda6
M2 ULTRA METAL small-q5_1 1 0 50.56 3.36 0.94 0.06 dc8dda6
M2 ULTRA METAL small-q8_0 1 0 47.52 3.20 0.92 0.05 dc8dda6
M2 ULTRA METAL medium 1 0 122.55 7.38 1.95 0.12 dc8dda6
M2 ULTRA METAL medium-q5_0 1 0 140.61 6.73 2.02 0.14 dc8dda6
M2 ULTRA METAL medium-q5_1 1 0 140.48 6.76 2.04 0.14 dc8dda6
M2 ULTRA METAL medium-q8_0 1 0 131.00 6.57 1.96 0.13 dc8dda6
M2 ULTRA METAL medium-dis 1 0 110.85 1.00 0.24 0.02 dc8dda6
M2 ULTRA METAL large-v2 1 0 222.28 10.96 3.03 0.21 dc8dda6
M2 ULTRA METAL large-v2-q5_0 1 0 258.64 9.79 3.04 0.25 dc8dda6
M2 ULTRA METAL large-v2-q5_1 1 0 258.32 9.87 3.05 0.24 dc8dda6
M2 ULTRA METAL large-v2-q8_0 1 0 236.55 9.61 2.87 0.23 dc8dda6
M2 ULTRA METAL large-v2-dis 1 0 199.84 1.14 0.27 0.02 dc8dda6
M2 ULTRA METAL large-v3-turbo 1 0 201.52 1.77 0.45 0.03 dc8dda6
M2 ULTRA METAL large-v3-turbo-q5_0 1 0 233.14 1.56 0.47 0.04 dc8dda6
M2 ULTRA METAL large-v3-turbo-q8_0 1 0 214.23 1.53 0.44 0.04 dc8dda6

What's Changed

New Contributors

Full Changelog: v1.7.5...v1.7.6

ggerganov tag:github.com,2008:Repository/541269386/danbev-testing-xcframework-release 2025-05-07T05:50:01Z - -danbev-testing-xcframework-release: ci : add zip extension to xcframework artifact name - -

This commit add the .zip extension to the xcframework artifact name in
the GitHub Actions workflow.

The motivation for this that the release job will look for .zip files
and will not find the xcframework artifact without the extension, and
hence will not upload it to the release.

danbev tag:github.com,2008:Repository/541269386/danbev-java-jar-artifact-test 2025-05-07T11:31:41Z - -danbev-java-jar-artifact-test: ci : add bindings-java jar artifact to release - -

This commit adds the jar artifact from bindings java to the release
process.

danbev+

What's Changed

Full Changelog: v1.8.5...v1.8.6

github-actions[bot]