

Ling-3.0-flash
#57 in Open-Source-SpraakmodelleUnknown · v3.0 · flash · 2× · tolest 08. Aug. 2026
Ling-3.0-flash is an Open-Weight language model developed by Ant Group's InclusionAI (Ling series) featuring a hybrid linear Mixture-of-Experts architecture: 124 billion total parameters with approximately 5.1 billion activated per Token (1/64 Sparse-MoE-Routing combined with Kimi Delta Attention and Gated MLA). It supports a native context window of 262,144 Token (expansion to 1M Token announced), is text-based (no multimodal input), and optimized for agentic applications such as coding, tool calls, and long-context tasks. The model was released around July 23, 2026, with weights available via Hugging Face among other sources; independent Benchmark evaluations (e.g., Artificial Analysis) became available shortly after release.
Features
| Key Benchmark (%) | GPQA Diamond: 59.3%; Artificial Analysis Intelligence Index: 26 |
| Context Window (Tokens) | 262,144 tokens native, extendable up to 1M tokens |
| License | MIT license (per derivative model on Hugging Face inheriting license from base model) |
| Multimodality | Text-only (input and output), no image/audio support |
| Platform | Available via Hugging Face (weights), API on InclusionAI/Ant Group, OpenRouter, ZenMux, Kilo Code, Puter |
| Price per 1M Tokens | $0.07 input / $0.22 output per 1M tokens (InclusionAI API, per Artificial Analysis) |
| Release Date | July 23, 2026 (per OpenRouter model creation date); Artificial Analysis lists August 4, 2026 |