

Stable Audio Open
#22 in AI Music GenerationStability Ai · since 2024-06-05 · 2× · last seen Jun 30, 2026
1
Momentum
Stable Audio Open 1.0 is an open-weights text-to-audio model by Stability AI with approximately 1.21 billion parameters, built on a latent diffusion architecture with DiT components and T5-based text conditioning. It generates variable-length stereo audio of up to 47 seconds at 44.1 kHz. The model was trained exclusively on Creative Commons-licensed audio data (Freesound and Free Music Archive) and is primarily intended for research, sound design, and non-commercial use. Vocal or speech generation is explicitly not supported by the model.
Momentum trend
19.05.17.08.
Features
| Real-Time Streaming | No real-time streaming; generation is offline/batch-based (max 47s output) |
| Latency | Autoencoder latent rate of 21.5Hz for music and audio generation |
| License | Stability AI Community License (research, non-commercial, limited commercial) |
| Max Music Duration (Seconds) | 47 seconds (variable stereo audio at 44.1 kHz) |
| Platform | Model weights on Hugging Face; code/training via stable-audio-tools on GitHub |
| Price | Model weights free to download; commercial use free up to $1M annual revenue, Enterprise license required above |
| Release Date | June 5, 2024 (announcement); research paper July 19/22, 2024 |
| Languages | English only (prompt understanding); other languages perform worse |
| Supported Input Formats | Text input (text prompts in English) with optional time conditioning (seconds_start, seconds_total); audio variations and style transfer from audio samples also possible |
| Voice Cloning | Not supported – model is not able to generate realistic vocals |
| Vocal/Singing Quality | Not supported – the model is unable to generate realistic vocals or intelligible speech/singing |