

DreamX-Creator
#22 in Text-to-VideoUnknown · 2× · last seen Sep 04, 2026
35
Momentum
DreamX-Creator 1.0 is an Open Source research project published by Alibaba's DreamX-Team (AMAP-ML) for native joint audio-video generation. Based on a 7B-parameter generator, the system produces synchronized video and audio streams from a start image and text prompt using Gated Cross-Modal Attention. A downstream "Autoregressive 1-Step 2K Refiner" (SR-DiT, 5B parameters) subsequently upscales the generated video to 2K resolution without altering content, motion, or audio synchronization. Model weights, code, and technical report are publicly available on GitHub, Hugging Face, and arXiv.
Momentum trend
08.06.06.09.
Features
| Fine-Tuning | Progressive Joint Training with two audio-video pre-training stages plus High-Quality Finetuning; additionally Audio-Video Reinforcement Learning with Modality-Aware Multimodal Feedback |
| Generation Time | 2K refiner requires only one denoising evaluation per temporal chunk (one-step refinement) |
| License | Apache License 2.0 |
| Max Resolution | 2K (via Autoregressive 1-Step 2K Refinement) |
| Platform | Model weights distributed via Hugging Face and ModelScope; code on GitHub |
| Price | Free / open source (no commercial pricing page found) |
| Release Date | arXiv report published August 31, 2026; GitHub repo initialized September 1, 2026 |