

MOSS-TTS
#28 v Syntéza řeči (TTS)Moss · od 10. Februar 2026 (MOSS-TTS Family); MOSS-TTS-Nano: 10. April 2026 · 3× · naposledy 29. 6. 2026
MOSS-TTS is an open-source family of speech and sound generation models from MOSI.AI and the OpenMOSS team (Shanghai Innovation Institution, Fudan NLP Lab). The family includes several specialized models for high-fidelity long-form narration, multi-speaker dialogue, voice design, environmental sound effects, and real-time streaming TTS. It also includes MOSS-TTS-Nano, a very small model (0.1B parameters) for CPU-based real-time speech generation with 48kHz stereo output in up to 20 languages. All models are released under the Apache 2.0 license and are approved for commercial use.
Vlastnosti
| Real-Time Streaming | Yes, streaming output with low first-token latency, including automatic chunking of long text |
| Latency | MOSS-TTS-Realtime: ~180ms first-byte latency; Nano: real-time factor < 1.0 on 4 CPU cores |
| License | Apache License 2.0 |
| Platform | Open-source on GitHub/Hugging Face; runs locally (GPU for flagship 8B, CPU for Nano); supported by vLLM-Omni, SGLang, ComfyUI, mlx-audio, ONNX, llama.cpp |
| Price | Free, open-source model weights (Apache 2.0) for self-hosting |
| Release Date | MOSS-TTS Family: February 10, 2026; MOSS-TTS-Nano: April 10, 2026 |
| Languages | 20 languages (incl. Chinese, English, German, Spanish, French, Japanese, Korean, Arabic, Persian) |
| Voice Cloning | Zero-shot voice cloning from short reference audio (3-15 seconds), no fine-tuning required |