

Fish Audio S2.1 Pro
#3 en Synthèse vocale (TTS)Fish Audio · pro · 4× · vu le 31 juil. 2026
53
Momentum
Fish Audio S2.1 Pro is a production-ready Text-to-Speech model from Fish Audio, released in June/July 2026 as the successor to S2-Pro. The model supports 83 languages, offers Voice-Cloning from short reference recordings, and provides word-level control of emotion and prosody via natural language tags. It is designed for low latency in real-time dialogue applications. Access is provided through the Fish Audio Cloud-API, with a free developer tier offering the same model quality without SLA guarantees and a paid production tier with latency and availability guarantees.
Historique du momentum
08.05.06.08.
Fonctionnalités
| Real-Time Streaming | Yes, WebSocket streaming designed for real-time conversational voice output |
| Latency | ~70–100 ms time-to-first-audio (TTFA) for single requests |
| License | Proprietary hosted model via API; earlier open models (S2) released under Fish Audio Research License with separate commercial license for self-hosting |
| Platform | Fish Audio Cloud API (REST/WebSocket, msgpack), Python & TypeScript SDKs, web playground at fish.audio |
| Price | Free API through July 24, 2026 (per another source through Aug 31, 2026); afterwards pay-as-you-go from $15/1M characters; subscription plans $0–$749/month |
| Release Date | Public launch July 28, 2026; free API availability announced starting June 2026 |
| Languages | 83 languages |
| Voice Cloning | Yes, voice cloning from 5 seconds of reference audio via API |