DeepSeek-V4-Flash-Vision-Exp
DeepSeek's experimental vision variant of V4 Flash: a 305B-param FP8 MoE with image input, 1M context, reasoning, and tool calling for multimodal agent tasks.
Model details
View repositoryDeepSeek-V4-Flash-Vision-Exp is the first experimental multimodal member of the DeepSeek V4 family, extending the DeepSeek-V4-Flash architecture with visual modules and continued training to unlock visual understanding. It retains the same Mixture-of-Experts backbone along with the 1M-token context window and hybrid attention architecture that keeps long-context processing efficient, plus configurable reasoning-effort levels for trading off latency against deeper deliberation.
Compared to DeepSeek-V4-Flash-0731, DeepSeek-V4-Flash-Vision-Exp achieves substantial improvements on its multimodal agent capabilities, while maintaining comparable performance on text-only agent tasks. This makes it well suited to applications that combine visual and textual reasoning while preserving the efficiency and cost profile that made the 0731 release attractive for latency-sensitive use cases.