DeepSeek has released “V4-Flash-Vision-Exp,” its first multimodal model that understands images. Available via the API, it is said to come close to top models on visual agent tasks – with a context window of over one million tokens.
Seeing, not just reading
On pure-text capabilities (agents, reasoning, world knowledge), the model is on par with the existing V4-Flash. The leap comes in visual understanding: on agent benchmarks that require image recognition, DeepSeek says it approaches the multimodal capabilities of Opus 4.8.
Technical details
The context window spans 1,048,576 tokens, with a maximum output of 384,000 tokens. “Thinking” mode is on by default. The model is addressed via the identifier “deepseek-v4-flash-vision-exp” – through either OpenAI-compatible or Anthropic-format endpoints.
What it means
With a powerful, affordable vision model, DeepSeek again raises the pressure on the big Western providers. The release went live on 21 August via the DeepSeek API; the “Exp” in the name signals its experimental status.
Source: DeepSeek API Docs.



















