DeepSeek-V4-Flash-Vision-Exp
DeepSeek · China · 2026
DeepSeek's first model that reads images — and the first the company sells only as a service, with no weights to download.
DeepSeek-V4-Flash-Vision-Exp went live on the DeepSeek API on 21 August 2026 as the company's first publicly available model that accepts pictures. It takes text and images together — JPEG, PNG, GIF and WebP, up to 600 images in a single request — and answers in text: describing photographs, reading screenshots, working through charts and driving agents that have to look at a screen before they act. On text alone DeepSeek says it is level with DeepSeek-V4-Flash, the open-weights model it is built beside; on agent benchmarks that require seeing something, the company reports a clear jump over that sibling, and puts the result close to Anthropic's Opus 4.8. The interesting part is what the release does not include. DeepSeek built its reputation on publishing weights under permissive licences — V4-Flash sits on Hugging Face under MIT — but a week after launch there is no repository for the vision model at all, and the documentation offers only an API name. The catalogue therefore lists it as a hosted service, not an open model, and keeps the pilot status the vendor's own 'exp' label implies: there is no model card, no safety report and no commitment that the endpoint will survive in this form. Pricing is the same as the text model, and images are cheap in a way that matters for anyone processing archives: whatever their original size, pictures are resized and billed at a maximum of 384 input tokens each. A request stuffed with the maximum 600 images therefore costs about ten US cents at peak rates and half that off-peak, before the text is counted. Put differently, roughly six thousand images fit into one dollar of peak-rate input. The context window is the same 1,000,000 tokens as the rest of the V4 line, output reaches 384,000 tokens, and a free Files API lets one uploaded picture be reused across requests instead of re-sent. Images are accepted only in user messages; system and assistant messages return an error, and code-completion (FIM) is not supported on this model at all.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!