

Confucius4-TTS
#2 v Syntéza řeči (TTS)Unknown · v4 · tts · 2× · naposledy 24. 8. 2026
36
Momentum
Confucius4-TTS is an Open Source speech synthesis model from NetEase Youdao based on a combination of speech encoder and large language model (LLM). The system enables Zero-Shot voice cloning without reference text and transfers a cloned voice accent-free across 14 languages while preserving emotional expression. It was released on June 23, 2026 as part of the "Zi Yue 4.0" model system and made freely available under the Apache 2.0 license with complete model weights.
Vývoj momenta
26.05.24.08.
Vlastnosti
| Real-Time Streaming | Supports both streaming and non-streaming inference (PCM audio stream via API) |
| Latency | First-packet latency of 85 ms (output streaming) / 54 ms (dual-streaming mode) |
| License | Apache 2.0, no restrictions on commercial use |
| Platform | GitHub, Hugging Face, online demo (Gradio) at confucius4-tts.youdao.com |
| Price | Free (open source, model weights freely downloadable) |
| Release Date | June 23, 2026 |
| Languages | 14 languages: Chinese, English, Japanese, Korean, German, French, Spanish, Indonesian, Italian, Thai, Portuguese, Russian, Malay, Vietnamese |
| Voice Cloning | Zero-shot voice cloning without reference transcript, from about 3 seconds of reference audio |