

Cartesia · v2 · od 16. Juni 2026 · 6× · naposledy 29. 6. 2026
11
Momentum
Ink-2 is Cartesia's Speech-to-Text (STT) model, developed specifically for real-time voice agents. It is based on a State-Space Model architecture and offers, according to the manufacturer, the lowest word error rate and the most accurate built-in Turn-Detection among streaming STT models. The model reliably recognizes structured data such as phone numbers, dates, and email addresses, and signals via Turn-Events when a speaker begins and ends without requiring separate Voice-Activity-Detection. Currently, Ink-2 supports English only, with additional languages announced for the future.
Vývoj momenta
16.05.14.08.
Vlastnosti
| Real-Time Streaming | Yes – streaming STT processing partial audio in real time, including turn events (turn.start, turn.update, turn.eager_end, turn.resume, turn.end) |
| Latency | Time-to-Final-Transcript (TTFT) of 0.1 seconds; ~100ms transcript latency in the full stack |
| Platform | Available via API, Cartesia Playground (play.cartesia.ai), and integrations with LiveKit, Vapi, and Pipecat |
| Price | STT usage billed at 1 credit per second of audio; ~$0.06/min call duration on Startup plan or $0.014/min on Scale plan |
| Release Date | June 16, 2026 (launched together with Sonic-3.5) |
| Languages | Currently English only; multilingual support announced |