

Unknown · depuis 31. Juli 2026 (Launch), 3. August 2026 (Open-Weights-Release) · 3× · vu le 05 août 2026
MiniMax H3 is a general-purpose omni-modal generation model developed by MiniMax that processes text, images, video and audio within a single unified context. It generates video with natively synchronized stereo audio at up to 2K resolution and durations of up to 15 seconds from any combination of input modalities. Launched July 31, 2026 via API and the Hailuo app, the base weights (H3-Base, ~33B parameters) were open-sourced on August 3, 2026 under the MiniMax Community License, while the H3-Context-IR and H3-Regenerate-2K components required for the full 2K pipeline remain hosted/proprietary. The model reached top rankings in independent blind comparisons by Artificial Analysis for video editing, text-to-video and image-to-video.
Fonctionnalités
| Key Benchmark (%) | Artificial Analysis Video Arena: #1 Video Editing, #2 Text-to-Video (Elo 1238), #3 Image-to-Video (Elo 1351 without audio) |
| Context Window (Tokens) | No classic token context window; multimodal reference: up to 9 images, 3 video clips, 3 audio clips per generation |
| License | MiniMax H3 Community License – free for non-commercial use and companies with annual revenue below $20M (with attribution); excludes US, EU, UK, South Korea |
| Multimodality | Unified context for text, image, video and audio as input and output (video with native stereo audio) |
| Platform | API (platform.minimax.io), Hailuo AI App, Hugging Face (MiniMaxAI/MiniMax-H3), fal.ai, ComfyUI |
| Price per 1M Tokens | No per-token price; API: $0.13 per second at 2K ($7.80/minute); 768p tier $0.09/second |
| Release Date | July 31, 2026 (launch); open weights (H3-Base) released August 3, 2026 |