

DeepSeek V4 Flash 0731
#3 in Local LLM RuntimesDeepSeek · v4 · flash 0731 · 6× · last seen Aug 01, 2026
DeepSeek-V4-Flash-0731 is the official, production-ready version of DeepSeek-V4-Flash, replacing the previously used Preview version. It is a Mixture-of-Experts model with 284 billion total parameters and 13 billion active parameters per Token, a 1-million-Token context window, and significantly improved agentic and coding capabilities compared to the Preview, achieved through intensive Post-Training rather than architectural changes. The weights are openly available under MIT license for download (including from Hugging Face) and can be deployed locally via vLLM or llama.cpp/Unsloth-GGUF quantizations; in parallel, the model is also accessible through the official, OpenAI-compatible DeepSeek API.