

DeepSeek multimodal model
#20 in Multimodal ModelsDeepSeek · 2× · last seen Sep 15, 2026
DeepSeek-V4-Flash-Vision-Exp was an experimental multimodal model from DeepSeek, released on the DeepSeek-API platform on August 21, 2026. It is based on the DeepSeek-V4-Flash architecture (Mixture-of-Experts, 284 billion total parameters, 13 billion active parameters) and extends the text-based V4-Flash with image and screenshot understanding while maintaining text, Reasoning, and agent capabilities. The model has since been superseded by DeepSeek-V4.1-Flash, which integrates native multimodal capabilities into the official main line; the old model ID is currently automatically redirected to V4.1-Flash. An "uncensored" FP8 variant with removed safety filters exists, but does not originate from DeepSeek itself, rather from an independent third party (d
Features
| Key Benchmark (%) | ApexBench: 36.5% (vs. 26.2% for V4-Flash); Agents' Last Exam: 27.3% (vs. 25.2%) |
| Context Window (Tokens) | 1,048,576 tokens (1M) |
| License | No separate open license known for Vision-Exp; base V4 models (Pro/Flash) available under MIT License on Hugging Face |
| Multimodality | Text + image input (screenshots, charts, documents), text output; no image output |
| Platform | DeepSeek API Platform (model name: deepseek-v4-flash-vision-exp, now migrated to deepseek-flash) |
| Price per 1M Tokens | $0.2156 input / $0.6468 output (same rate as V4-Flash) |
| Release Date | August 21, 2026 (V4-Flash-Vision-Exp); successor V4.1-Flash on September 10, 2026 |