

Sonic-3.5
#13 v Syntéza řeči (TTS)Cartesia · v3.5 · od 2026-06-16 · 9× · naposledy 30. 6. 2026
14
Momentum
Sonic-3.5 is Cartesia's Text-to-Speech model for real-time speech synthesis in Voice-Agents. It was released on June 16, 2026 together with the transcription model Ink-2 as a joint Voice-Stack. According to the manufacturer, the model achieves speech output latency of under 90ms, natively supports 42 languages, and is listed as a leading TTS model on Benchmark platforms such as Artificial Analysis. It offers Instant Voice-Cloning, Voice-Changer, and localization features and is available through the Cartesia API as well as partners like LiveKit.
Vývoj momenta
16.05.14.08.
Vlastnosti
| Real-Time Streaming | Yes, designed for real-time voice agents with bidirectional streaming |
| Latency | Sub-90ms speech latency (time-to-first-audio); around 82ms per Artificial Analysis |
| License | Commercial use license included from the Pro plan (paid); Free tier not intended for commercial use |
| Platform | Cartesia's cloud API (model ID sonic-3.5), also available via partners like LiveKit; deployable in cloud, on-premise, and on-device |
| Price | From $5/month (Pro plan, ~133 min TTS/month); Free tier $0/month (20,000 credits); Startup $49/month; Scale $299/month; API usage e.g. $0.03/min or $50 per 1M characters (via LiveKit) |
| Release Date | June 16, 2026 |
| Languages | 42 languages natively, including English, Hindi, Spanish, French, German, Japanese, Hebrew |
| Voice Cloning | Instant voice cloning with just 10 seconds of audio; Professional Voice Cloning also available for higher quality |