tag:github.com,2008:https://github.com/OHF-Voice/piper1-gpl/releasesRelease notes from piper1-gpl2026-09-04T16:40:50Ztag:github.com,2008:Repository/956792596/v1.8.02026-09-04T16:47:48Zv1.8.0<ul>
<li>Add Thai phonemizer using TLTK in the new <code>th</code> extra
<ul>
<li><code>--data.phoneme_type thai</code> for training; <code>"phoneme_type": "thai"</code> in a voice config for synthesis</li>
<li>espeak-ng's Thai voice is a placeholder: its <code>th_dict</code> holds no lexicon, so unspaced Thai is never segmented; the leading vowels เ แ โ ใ ไ are not reordered; and a tone mark deletes the syllable's vowel, collapsing ป่า/ป้า/ป๊า/ป๋า to the same phonemes</li>
<li>TLTK does dictionary-based word segmentation and emits a tone digit (1-5) per syllable, in the same style <code>phonemize_chinese</code> uses for Mandarin tones</li>
<li>The whole inventory already has ids in the default IPA map, so Thai voices stay compatible with the IPA-based (espeak) warmstart</li>
</ul>
</li>
<li>Add <code>script/setup --th</code>, and install the <code>th</code> extra in CI so the Thai tests run</li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.7.02026-08-15T18:51:35Zv1.7.0<ul>
<li>Add Japanese phonemizer using OpenJTalk (<code>pyopenjtalk-plus</code>) in the new <code>ja</code> extra
<ul>
<li><code>--data.phoneme_type japanese</code> for training; <code>"phoneme_type": "japanese"</code> in a voice config for synthesis</li>
<li>espeak-ng has no kanji coverage (it reads out Unicode character names) and no pitch accent</li>
<li>Full-context labels are parsed for pitch accent and mapped to IPA, so Japanese voices stay compatible with the IPA-based (espeak) warmstart</li>
</ul>
</li>
<li>Add <code>script/setup --ja</code>, and install the <code>ja</code> extra in CI so the Japanese tests run</li>
<li><code>libpiper</code>: add <code>piper_create_options</code> and <code>piper_create_with_options()</code>, with <code>piper_create()</code> kept as a wrapper for ABI compatibility</li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.6.12026-08-13T15:24:25Zv1.6.1<ul>
<li>Run the g2pW model through <code>piper.g2pw_onnx</code> instead of <code>g2pw.api</code>, dropping <code>torch</code> (~750 MB installed) and <code>requests</code> from the <code>zh</code> extra
<ul>
<li><code>g2pw.api</code> imports torch only to build padded tensors and iterate batches; the model itself already ran under onnxruntime</li>
<li>Also 1.5-2x faster, since it no longer forks DataLoader worker processes on every call</li>
<li><code>g2pW</code> is still required, for its pinyin/bopomofo lookup tables</li>
</ul>
</li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.6.02026-07-23T16:12:10Zv1.6.0<ul>
<li>Add Hebrew phonemizer using Nakdimon</li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.5.02026-07-17T20:06:41Zv1.5.0<ul>
<li>Add <code>libpiper</code> C++ CLI executable ported from the legacy Piper repository, plus a C++ test suite</li>
<li>Fix <code>libpiper</code> builds on Windows (MSVC, MSYS2-GCC) and Windows CI</li>
<li>Bump embedded espeak-ng version</li>
<li>Add default speaker id for multi-speaker voices</li>
<li>Add vowel clustering support (<code>--data.vowel_clusters</code>)</li>
<li>Add in-memory patching for alignments</li>
<li>Training: add MRD (Multi-Resolution STFT) discriminator, loss/MOS tracking with UTMOS, silence-trim fixes, and dataloader performance improvements</li>
<li>Pass custom phoneme id map when training</li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.4.22026-04-02T21:17:31Zv1.4.2<ul>
<li>Fix <code>pathvalidate</code> dependency</li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.4.12026-02-05T09:58:41Zv1.4.1<ul>
<li>Add missing wheels</li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.4.02026-01-30T17:00:19Zv1.4.0<ul>
<li>Add Chinese phonemizer based on <a href="https://github.com/GitYCC/g2pW/">g2pW</a>
<ul>
<li>Using a quantized version of the original model with <code>quantize_dynamic</code></li>
</ul>
</li>
<li>Add <code>--data.phoneme_type pinyin</code> for Chinese phonemization using g2pW</li>
<li>Add <code>--data.phoneme_type text</code> for using IPA phonemes directly (no espeak-ng)</li>
<li>Add <code>--model.vocoder_warmstart_ckpt <CHECKPOINT></code> to restore vocoder params only</li>
<li>Add <code>--data.dataset_type 'phoneme_ids'</code> to train with pre-generated phoneme ids
<ul>
<li>Use <code>--data.num_symbols <N></code> to set number of phonemes</li>
<li>Use <code>--data.phonemes_path "/path/to/phonemes.json"</code> for phoneme/id map</li>
</ul>
</li>
<li>Add <code>--output-dir-naming</code> option with <code>timestamp</code> (default) and <code>text</code></li>
</ul>github-actions[bot]tag:github.com,2008:Repository/956792596/v1.3.02025-07-10T21:14:16Zv1.3.0<ul>
<li>Moved development to OHF-Voice org</li>
<li>Removed C++ code for now to focus on Python development
<ul>
<li>A C API <code>libpiper</code> written in C++ is planned</li>
</ul>
</li>
<li>Embed espeak-ng directly instead of using separate <code>piper-phonemize</code> library</li>
<li>Change license to GPLv3</li>
<li>Use Python stable ABI (3.9+) so only a single wheel per platform is needed</li>
<li>Change Python API:
<ul>
<li><code>PiperVoice.synthesize</code> takes a <code>SynthesisConfig</code> and generates <code>AudioChunk</code> objects</li>
<li><code>PiperVoice.synthesize_raw</code> is removed</li>
</ul>
</li>
<li>Add seperate <code>piper.download_voices</code> utility for downloading voices from HuggingFace</li>
<li>Allow text as CLI argument: <code>piper ... -- "Text to speak"</code></li>
<li>Allow text from one or more files with <code>--input-file <FILE></code></li>
<li>Excluding any file output arguments will play audio directly with <code>ffplay</code></li>
<li>Support for raw phonemes in text with <code>[[ <phonemes> ]]</code></li>
<li>Adjust output volume with <code>--volume <MULTIPLIER></code> (default is 1.0)</li>
</ul>synesthesiam