

VoiceStudio
#7 v Syntéza řeči (TTS)Voicestudio · 3× · naposledy 07. 9. 2026
VoiceStudio (formerly OmniVoice-Studio) is an Open Source, fully locally-running desktop application for voice cloning, voice design, video dubbing, dictation, transcription, and audiobook creation, developed as a free alternative to cloud services like ElevenLabs. The software supports 16 TTS and 11 ASR engines as well as a voice catalog of 646 languages and runs on macOS, Windows, and Linux (including Docker image), without account, API key, or usage limits in local workflows. It is licensed under AGPL-3.0, with the speech models used having separate license terms. Voice cloning works via Zero-Shot procedure starting from a 3-second reference recording, and the app offers an OpenAI-compatible local API as well as WebSocket streaming with measured latency.
Vlastnosti
| Echtzeit-Streaming | Ja, über WebSocket-Endpunkt /ws/tts mit Zeitangaben zu erstem Audio-Frame und Generierungsdauer |
| Latenz | Gemessene Time-to-First-Audio (ttfa_ms) und Real-Time-Factor (rtf) werden pro Streaming-Anfrage im /ws/tts-Endpunkt zurückgegeben; ca. 28s Ladezeit-Untergrenze pro Aufruf bei Subprocess-Engines |
| Lizenz | AGPL-3.0 (Anwendungslizenz); genutzte Sprachmodelle unterliegen eigenen, separaten Lizenzen |
| Plattform | Desktop-App für macOS (Apple Silicon, ab 13.3), Windows 10/11 x64, Linux x86_64, sowie Docker-Image |
| Preis | Kostenlos (Open Source); kommerzielle/Closed-Source-Nutzung nur auf Anfrage (Pro, enquiry-only, kein Online-Checkout) |
| Release-Datum | Umbenennung von OmniVoice-Studio zu VoiceStudio mit Release v0.5.0 |
| Sprachen | 646 Sprachen (TTS-Katalog), abhängig von der gewählten Engine; 16 TTS- und 11 ASR-Engines |
| Voice-Cloning | Zero-Shot-Klonen bereits ab 3 Sekunden Referenzaudio möglich, 5-15 Sekunden empfohlen für bessere Qualität |