DeepSeek-V4-Flash
DeepSeek · China · 2026
DeepSeek's efficient V4 model — a 284-billion-parameter MoE with only 13 billion active per token, released under MIT with open weights.
DeepSeek-V4-Flash is the compact member of DeepSeek's fourth-generation line, first published as a preview in April 2026 alongside the much larger V4-Pro. It is a sparse mixture-of-experts model with 284 billion total parameters, of which roughly 13 billion are activated per token, and it handles a context window of one million tokens with outputs of up to 384,000 tokens. The technical report describes a hybrid attention scheme combining compressed and hyper-connected attention, manifold-constrained hyper-connections and the Muon optimiser, over a training run of more than 32 trillion tokens. The model is text-only: it has no vision or audio input. On 31 July 2026 DeepSeek shipped the 0731 revision, a re-post-trained build of the same architecture aimed at agentic work and software engineering. The gains there were unusually large: Terminal-Bench 2.1 rose from 61.8 to 82.7 and DeepSWE from 7.3 to 54.4, while independent testing by Artificial Analysis put the model at 50 points on its Intelligence Index, ten points above the previous Flash build. Weights are published on Hugging Face under the MIT licence, and the hosted API is priced at 0.14 US dollars per million input tokens and 0.28 per million output tokens — among the lowest rates for a model at this capability level. The earlier DeepSeek-R1 remains available and is documented in its own profile.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!