

StepAudio 3 ASR
#10 v Přepis řeči (STT)Stepaudio · v3 · asr · 2× · naposledy 24. 9. 2026
StepAudio 3 ASR is a cloud-based Speech-to-Text model developed by StepFun, built on LLM technology, released on September 15, 2026 as part of the five-model StepAudio 3 family. It achieved 1st place in Artificial Analysis' non-streaming AA-WER Index with 1.7% WER, significantly improving accuracy compared to its predecessor StepAudio 2.5 ASR (4.7% WER). The model supports Chinese and English, with preview support for Japanese, Korean, French, and Spanish, recognizes domain-specific vocabulary across 20+ industries, and handles challenging audio content including whispers, singing, and background noise. It is exclusively accessible via the hosted StepFun API; there are no Open Source weights or on-premise option.