tag:github.com,2008:https://github.com/OHF-Voice/piper1-gpl/releases Release notes from piper1-gpl 2026-09-04T16:40:50Z tag:github.com,2008:Repository/956792596/v1.8.0 2026-09-04T16:47:48Z v1.8.0 <ul> <li>Add Thai phonemizer using TLTK in the new <code>th</code> extra <ul> <li><code>--data.phoneme_type thai</code> for training; <code>"phoneme_type": "thai"</code> in a voice config for synthesis</li> <li>espeak-ng's Thai voice is a placeholder: its <code>th_dict</code> holds no lexicon, so unspaced Thai is never segmented; the leading vowels เ แ โ ใ ไ are not reordered; and a tone mark deletes the syllable's vowel, collapsing ป่า/ป้า/ป๊า/ป๋า to the same phonemes</li> <li>TLTK does dictionary-based word segmentation and emits a tone digit (1-5) per syllable, in the same style <code>phonemize_chinese</code> uses for Mandarin tones</li> <li>The whole inventory already has ids in the default IPA map, so Thai voices stay compatible with the IPA-based (espeak) warmstart</li> </ul> </li> <li>Add <code>script/setup --th</code>, and install the <code>th</code> extra in CI so the Thai tests run</li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.7.0 2026-08-15T18:51:35Z v1.7.0 <ul> <li>Add Japanese phonemizer using OpenJTalk (<code>pyopenjtalk-plus</code>) in the new <code>ja</code> extra <ul> <li><code>--data.phoneme_type japanese</code> for training; <code>"phoneme_type": "japanese"</code> in a voice config for synthesis</li> <li>espeak-ng has no kanji coverage (it reads out Unicode character names) and no pitch accent</li> <li>Full-context labels are parsed for pitch accent and mapped to IPA, so Japanese voices stay compatible with the IPA-based (espeak) warmstart</li> </ul> </li> <li>Add <code>script/setup --ja</code>, and install the <code>ja</code> extra in CI so the Japanese tests run</li> <li><code>libpiper</code>: add <code>piper_create_options</code> and <code>piper_create_with_options()</code>, with <code>piper_create()</code> kept as a wrapper for ABI compatibility</li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.6.1 2026-08-13T15:24:25Z v1.6.1 <ul> <li>Run the g2pW model through <code>piper.g2pw_onnx</code> instead of <code>g2pw.api</code>, dropping <code>torch</code> (~750 MB installed) and <code>requests</code> from the <code>zh</code> extra <ul> <li><code>g2pw.api</code> imports torch only to build padded tensors and iterate batches; the model itself already ran under onnxruntime</li> <li>Also 1.5-2x faster, since it no longer forks DataLoader worker processes on every call</li> <li><code>g2pW</code> is still required, for its pinyin/bopomofo lookup tables</li> </ul> </li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.6.0 2026-07-23T16:12:10Z v1.6.0 <ul> <li>Add Hebrew phonemizer using Nakdimon</li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.5.0 2026-07-17T20:06:41Z v1.5.0 <ul> <li>Add <code>libpiper</code> C++ CLI executable ported from the legacy Piper repository, plus a C++ test suite</li> <li>Fix <code>libpiper</code> builds on Windows (MSVC, MSYS2-GCC) and Windows CI</li> <li>Bump embedded espeak-ng version</li> <li>Add default speaker id for multi-speaker voices</li> <li>Add vowel clustering support (<code>--data.vowel_clusters</code>)</li> <li>Add in-memory patching for alignments</li> <li>Training: add MRD (Multi-Resolution STFT) discriminator, loss/MOS tracking with UTMOS, silence-trim fixes, and dataloader performance improvements</li> <li>Pass custom phoneme id map when training</li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.4.2 2026-04-02T21:17:31Z v1.4.2 <ul> <li>Fix <code>pathvalidate</code> dependency</li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.4.1 2026-02-05T09:58:41Z v1.4.1 <ul> <li>Add missing wheels</li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.4.0 2026-01-30T17:00:19Z v1.4.0 <ul> <li>Add Chinese phonemizer based on <a href="https://github.com/GitYCC/g2pW/">g2pW</a> <ul> <li>Using a quantized version of the original model with <code>quantize_dynamic</code></li> </ul> </li> <li>Add <code>--data.phoneme_type pinyin</code> for Chinese phonemization using g2pW</li> <li>Add <code>--data.phoneme_type text</code> for using IPA phonemes directly (no espeak-ng)</li> <li>Add <code>--model.vocoder_warmstart_ckpt &lt;CHECKPOINT&gt;</code> to restore vocoder params only</li> <li>Add <code>--data.dataset_type 'phoneme_ids'</code> to train with pre-generated phoneme ids <ul> <li>Use <code>--data.num_symbols &lt;N&gt;</code> to set number of phonemes</li> <li>Use <code>--data.phonemes_path "/path/to/phonemes.json"</code> for phoneme/id map</li> </ul> </li> <li>Add <code>--output-dir-naming</code> option with <code>timestamp</code> (default) and <code>text</code></li> </ul> github-actions[bot] tag:github.com,2008:Repository/956792596/v1.3.0 2025-07-10T21:14:16Z v1.3.0 <ul> <li>Moved development to OHF-Voice org</li> <li>Removed C++ code for now to focus on Python development <ul> <li>A C API <code>libpiper</code> written in C++ is planned</li> </ul> </li> <li>Embed espeak-ng directly instead of using separate <code>piper-phonemize</code> library</li> <li>Change license to GPLv3</li> <li>Use Python stable ABI (3.9+) so only a single wheel per platform is needed</li> <li>Change Python API: <ul> <li><code>PiperVoice.synthesize</code> takes a <code>SynthesisConfig</code> and generates <code>AudioChunk</code> objects</li> <li><code>PiperVoice.synthesize_raw</code> is removed</li> </ul> </li> <li>Add seperate <code>piper.download_voices</code> utility for downloading voices from HuggingFace</li> <li>Allow text as CLI argument: <code>piper ... -- "Text to speak"</code></li> <li>Allow text from one or more files with <code>--input-file &lt;FILE&gt;</code></li> <li>Excluding any file output arguments will play audio directly with <code>ffplay</code></li> <li>Support for raw phonemes in text with <code>[[ &lt;phonemes&gt; ]]</code></li> <li>Adjust output volume with <code>--volume &lt;MULTIPLIER&gt;</code> (default is 1.0)</li> </ul> synesthesiam