Synthszr Charts — die großen AI-Marken im Wettkampf ums Podium
synthszr charts
cartesia

Cartesia · v2 · siet 16. Juni 2026 · 6× · tolest 29. Juni 2026

11
Momentum

Ink-2 is Cartesia's Speech-to-Text (STT) model, developed specifically for real-time voice agents. It is based on a State-Space Model architecture and offers, according to the manufacturer, the lowest word error rate and the most accurate built-in Turn-Detection among streaming STT models. The model reliably recognizes structured data such as phone numbers, dates, and email addresses, and signals via Turn-Events when a speaker begins and ends without requiring separate Voice-Activity-Detection. Currently, Ink-2 supports English only, with additional languages announced for the future.

Momentum-Verloop
16.05.14.08.

Features

Real-Time StreamingYes – streaming STT processing partial audio in real time, including turn events (turn.start, turn.update, turn.eager_end, turn.resume, turn.end)
LatencyTime-to-Final-Transcript (TTFT) of 0.1 seconds; ~100ms transcript latency in the full stack
PlatformAvailable via API, Cartesia Playground (play.cartesia.ai), and integrations with LiveKit, Vapi, and Pipecat
PriceSTT usage billed at 1 credit per second of audio; ~$0.06/min call duration on Startup plan or $0.014/min on Scale plan
Release DateJune 16, 2026 (launched together with Sonic-3.5)
LanguagesCurrently English only; multilingual support announced

Mehr Produkten in disse Kategorie: Spraaksynthese (TTS)

Belege (6)

Subscribe free. Unsubscribe the second it sucks.

High-signal news across AI, business, UX, and tech. Every morning.