

DeepSeek V4-Flash-Vision
#8 en Modèles multimodauxDeepSeek · v4 · flash vision · 8× · vu le 03 sept. 2026
DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal AI model from DeepSeek built on the DeepSeek-V4-Flash architecture and extended with visual modules. It was initially released as an API model on August 21, 2026, and subsequently made available as Open Source weights under MIT license on Hugging Face on August 31, 2026. The model achieves performance comparable to DeepSeek-V4-Flash on text tasks (agents, Reasoning, world knowledge) and demonstrates significant progress on multimodal agent Benchmarks, which according to the manufacturer approaches Anthropic's Claude Opus-4.8. It is based on a Mixture-of-Experts architecture with 284 billion total parameters and approximately 13 billion active parameters, along with a 1-Million-Token context window.
Fonctionnalités
| Key Benchmark (%) | ApexBench (Pass@1): 36.5% (vs. 26.2% for text-only V4-Flash) |
| Context Window (Tokens) | 1,048,576 tokens (1M), max 384,000 tokens output |
| License | MIT License (model weights on Hugging Face) |
| Multimodality | Text and image (JPEG, PNG, GIF, WebP) input, text output; native vision integrated into V4-Flash architecture |
| Platform | DeepSeek API (OpenAI- and Anthropic-compatible), Hugging Face (open weights), inference via vLLM/SGLang |
| Price per 1M Tokens | Same as V4-Flash: $0.22 input (cache miss) / $0.66 output, off-peak; images capped at 384 tokens each |
| Release Date | API launch: August 21, 2026; open weights: August 31, 2026 |