

FastH3 v1
#36 in Open-Source LLMsHao Ai Lab · v3 · v1 · since 27. August 2026 (Blog-Announcement "FastH3 Preview v1"), vorherige Preview-Version v0.2 am 23. August 2026 · 3× · last seen Aug 31, 2026
FastH3 v1 ("FastH3 Preview v1") is an open-weight text-to-video-and-audio model released by Hao AI Lab (UC San Diego) as part of the FastVideo framework, in collaboration with Nuva Lab and the NVIDIA FastGen team. It is a 4-step DMD2 distillation of the MiniMax H3 base model (a 33B-parameter dual-modality video+audio diffusion transformer) using 90% sparse attention (VSA), reducing the original 49-50 transformer forward passes to 4. On a single NVIDIA Blackwell GPU (B200) it achieves up to 14x speedup versus the dense base model, generating a 15s 768p video with synchronized audio in under 13-47 seconds depending on GPU count. The model inherits the MiniMax H3 Community License, which restricts local use in the US, EU, UK and South Korea.
Features
| Key Benchmark (%) | 14.38x faster than dense base model (15s clip on 1x B200, 1344x768, 24 FPS); 45% preferred/neutral vs. MiniMax H3 in user study |
| License | MiniMax H3 Community License (inherited from base model; restricts local use in US, EU, UK, South Korea) |
| Multimodality | Text-to-video-and-audio (T2VA): generates synchronized video and audio from text in a single pipeline call |
| Platform | Runs via FastVideo framework on NVIDIA GPUs (incl. B200/Blackwell); checkpoints on Hugging Face, code on GitHub |
| Release Date | August 27, 2026 (Preview v1); v0.2 on August 23, 2026 |