Skip to content

DeepSeek V4 Flash Vision Exp

deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Open Source · chat · open-weights
Open Alert me on changes
Context
1M
Max output
384K
Weights
Open
API $/1M
$0.22 / $0.66
Modalities
text · image
Released
31 Aug 2026
Download image Share on X Share on LinkedIn
AI summary
● machine-written

DeepSeek V4 Flash Vision Exp adds image understanding to efficient MoE model

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash, adding image understanding capabilities while maintaining text performance in agents, reasoning, and world knowledge. The model is a sparse mixture-of-experts architecture with 13B active parameters out of 284B total and supports a 1M-token context window. It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images.

What's new
  • Adds image understanding to V4 Flash base model
  • Supports multimodal input (text and images)
  • Sparse MoE architecture with 13B active/284B total parameters
  • 1M-token context window with up to 384K output tokens
Best for
Document and chart understandingVisual question answeringMultimodal agent workflowsImage-text reasoning tasks
Sources

Source: https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp