

MiMo-V2.5
#24 in Multimodale ModelleUnknown · v2.5 · siet 22. April 2026 · 2× · tolest 31. Aug. 2026
22
Momentum
MiMo-V2.5 is an open-weight multimodal Mixture-of-Experts model developed by Xiaomi, with 310 billion total parameters (15 billion active), natively processing text, image, video, and audio within a unified architecture. It is built on the MiMo-V2-Flash backbone with hybrid sliding-window attention and extended with dedicated vision and audio encoders. The model supports a context window of up to 1 million tokens and was trained on roughly 48 trillion tokens. It was announced on April 22, 2026, and later fully open-sourced under the MIT license.
Momentum-Verloop
02.06.31.08.
Features
| Key Benchmark (%) | 62.3 on Claw-Eval (agentic benchmark, general subset) |
| Context Window (Tokens) | 1,048,576 tokens (1M native context window) |
| License | MIT (fully open-source, commercial use permitted) |
| Multimodality | Native omnimodal: text, image, video, and audio in unified architecture (understanding) |
| Platform | Xiaomi MiMo API Platform, Hugging Face, OpenRouter, vLLM/SGLang-compatible |
| Price per 1M Tokens | $0.14 input / $0.28 output per 1M tokens (Xiaomi first-party API, per Artificial Analysis) |
| Release Date | April 22, 2026 |