

Sonic 3.6
#2 in Spraaksynthese (TTS)Cartesia · v3.6 · siet 17. August 2026 · 2× · tolest 20. Aug. 2026
36
Momentum
Sonic-3.6 is Cartesia's current real-time text-to-speech model, released on August 17, 2026 as the successor to Sonic-3.5. The model runs on state space models rather than transformers and achieves a claimed sub-90ms time-to-first-audio latency. It supports 44 languages, offers instant and professional voice cloning, and currently leads both Artificial Analysis Speech Arena leaderboards (Provider Voice and Controlled Voice). It is available as a hosted API in beta, not as self-hostable weights.
Momentum-Verloop
22.05.20.08.
Features
| Real-Time Streaming | Yes, streaming TTS model based on state space models rather than transformers |
| Latency | Sub-90ms time-to-first-audio (vendor claim) |
| License | Commercial use included from Pro plan ($5/month); Free tier has no commercial license |
| Platform | Hosted API (beta), not self-hosted weights |
| Price | From $5/month (Pro plan, commercial license); Free tier $0 (20,000 credits/month, non-commercial); Startup $49/month; Scale $299/month |
| Release Date | August 17, 2026 |
| Languages | 44 languages |
| Voice Cloning | Instant voice cloning (from Pro plan) and professional voice cloning (from Startup plan) available |