

VoiceStudio
#5 in Spraaksynthese (TTS)Unknown · 2× · tolest 29. Sept. 2026
VoiceStudio (formerly OmniVoice-Studio) is an Open Source, fully locally-running desktop application for voice cloning, voice design, video dubbing, dictation, transcription, and audiobook creation, developed as a free alternative to cloud services like ElevenLabs. The software supports 16 TTS and 11 ASR engines as well as a voice catalog of 646 languages and runs on macOS, Windows, and Linux (including Docker image), without account, API key, or usage limits in local workflows. It is licensed under AGPL-3.0, with the speech models used having separate license terms. Voice cloning works via Zero-Shot procedure starting from a 3-second reference recording, and the app offers an OpenAI-compatible local API as well as WebSocket streaming with measured latency.
Features
| Real-Time Streaming | Yes, via WebSocket endpoint /ws/tts reporting time-to-first-audio and generation duration |
| Latency | Measured time-to-first-audio (ttfa_ms) and real-time factor (rtf) returned per streaming request via /ws/tts endpoint; ~28s floor per call for subprocess engines |
| License | AGPL-3.0 (application license); bundled speech models carry their own separate licenses |
| Platform | Desktop app for macOS (Apple Silicon, 13.3+), Windows 10/11 x64, Linux x86_64, plus Docker image |
| Price | Free (open source); commercial/closed-source use enquiry-only (Pro tier, no online checkout) |
| Release Date | Renamed from OmniVoice-Studio to VoiceStudio with release v0.5.0 |
| Languages | 646 languages (TTS catalogue), coverage depends on selected engine; 16 TTS and 11 ASR engines |
| Voice Cloning | Zero-shot cloning possible from as little as 3 seconds of reference audio, 5-15 seconds recommended for better quality |