

Fish Audio
#6 v Syntéza řeči (TTS)Fish Audio · 3× · naposledy 04. 8. 2026
Fish Audio is a platform for AI-powered text-to-speech and voice cloning, developed by the team behind the Open Source project "Fish Speech". The current flagship model S2 (Pro) uses a dual-autoregressive architecture (4B-parameter Slow-AR plus 400M-parameter Fast-AR) and was trained on over 10 million hours of audio data in more than 80 languages. It supports fine-grained emotion and prosody control via free-text tags, multi-speaker dialogue, and voice cloning from short reference recordings (10–30 seconds). The model weights are freely available under a proprietary "Fish Audio Research License" for research and non-commercial use; commercial use requires a separate license. API access is available on a pay-per-use basis.
Vlastnosti
| Real-Time Streaming | Yes, audio is streamed as it generates (SGLang-based streaming inference engine) |
| Latency | Time-to-first-audio ~100ms; Real-Time Factor of 0.195 on a single NVIDIA H200 GPU (S2 Pro) |
| License | Fish Audio Research License: free for research/non-commercial use; commercial use requires a separate license from Fish Audio |
| Platform | Web app (fish.audio), REST API, Python SDK, self-hostable (GitHub/Hugging Face), iOS and Android companion app |
| Price | API: $15 per 1M UTF-8 bytes (s2-pro model); Subscriptions: Free $0, Plus $5.50–15/month, Pro $37.50–100/month, Max up to $749/month |
| Release Date | OpenAudio S1 launched June 2025 (rebrand to Fish Audio); S2 Pro open-sourced March 2026 |
| Languages | 80+ languages (S2 Pro); some marketing front-ends cite 30+ or 40+ languages |
| Voice Cloning | Voice cloning from 10–30 seconds of reference audio, capturing timbre, speaking style, and emotional tendencies without additional fine-tuning |