

Fish Audio S2.1 Pro
#2 in Text-to-Speech (TTS)Fish Audio · pro · 3× · last seen Jul 30, 2026
75
Momentum
Fish Audio S2.1 Pro is a production-ready Text-to-Speech model from Fish Audio, released in June/July 2026 as the successor to S2-Pro. The model supports 83 languages, offers Voice-Cloning from short reference recordings, and provides word-level control of emotion and prosody via natural language tags. It is designed for low latency in real-time dialogue applications. Access is provided through the Fish Audio Cloud-API, with a free developer tier offering the same model quality without SLA guarantees and a paid production tier with latency and availability guarantees.
Momentum trend
01.05.30.07.
Features
| Real-Time Streaming | Yes, WebSocket streaming designed for real-time conversational voice output |
| Latency | ~70–100 ms time-to-first-audio (TTFA) for single requests |
| License | Proprietary hosted model via API; earlier open models (S2) released under Fish Audio Research License with separate commercial license for self-hosting |
| Platform | Fish Audio Cloud API (REST/WebSocket, msgpack), Python & TypeScript SDKs, web playground at fish.audio |
| Price | Free API through July 24, 2026 (per another source through Aug 31, 2026); afterwards pay-as-you-go from $15/1M characters; subscription plans $0–$749/month |
| Release Date | Public launch July 28, 2026; free API availability announced starting June 2026 |
| Languages | 83 languages |
| Voice Cloning | Yes, voice cloning from 5 seconds of reference audio via API |