DeepSeek-V4-Pro
DeepSeek · China · 2026
DeepSeek's flagship: a 1.65-trillion-parameter MoE with open weights under MIT, whose final 0813 build lifts agentic coding to frontier level at a fraction of closed-model prices.
DeepSeek-V4-Pro is the flagship of DeepSeek's fourth-generation line and the larger of the two V4 variants the company shipped. It is a sparse mixture-of-experts model with 1.6 trillion total parameters, of which roughly 49 billion are activated per token, and it shares the one-million-token context window of its smaller sibling V4-Flash. Like the rest of the line it is text-only — no vision, audio or video input. The model was first published as a preview on 24 April 2026 and reached general availability on 19 July 2026. The weights did not change between the two milestones; what changed was the commercial packaging. DeepSeek introduced peak-time pricing for the first time, doubling rates during Beijing business hours, and retired the legacy deepseek-chat and deepseek-reasoner API endpoints on 24 July 2026, forcing callers onto explicit model IDs. On DeepSeek's published GA figures V4-Pro scores 80.6% on SWE-bench Verified — the strongest open-weights result on that benchmark — alongside a Codeforces rating of 3,206, 57.9% on SimpleQA Verified and 37.7% on Humanity's Last Exam. Independent measurement is less flattering to the trend line: Artificial Analysis placed V4-Pro at 44 points on its Intelligence Index at GA, down from 52 at preview, not because the model regressed but because the frontier moved during the three-month gap. The economics are the real argument. Off-peak API pricing is 0.435 US dollars per million input tokens and 0.87 per million output, with cached input at roughly a hundredth of that — an order of magnitude below Western frontier models of comparable coding ability. The weights are on Hugging Face under the MIT licence, so the model can also be self-hosted outright. DeepSeek-V4-Flash and the earlier DeepSeek-R1 remain available and have their own profiles in this catalogue. On 13 August 2026 DeepSeek published DeepSeek-V4-Pro-0813 and described it as the official release of V4-Pro, superseding the preview weights that had carried the name since April. The architecture is unchanged, but a DSpark speculative-decoding module is attached and the reasoning_effort control now offers low, high and max. The gains the company reports fall almost entirely in agentic and coding work: DeepSWE rises from 12.8% to 62.7%, NL2Repo from 38.5% to 61.5%, Cybergym from 52.7% to 83.3% and Terminal Bench 2.1 from 72.1 to 87.9. The release also drops the Jinja chat template in favour of a documented Python encoder, which means work for anyone running the weights locally. These figures are the vendor’s own.
▸News
DeepSeek's most downloaded model is 19 months old — its newest flagship build is pulled 32 times less often a day
8/25/2026 · Hugging Face (model repositories)
DeepSeek's clock is a Beijing clock: the same working day costs 75% more there than in San Francisco
8/19/2026 · DeepSeek — dokumentacja API (Models & Pricing)
DeepSeek ships the final V4-Pro: open weights, 1M context, agentic scores up sharply
8/13/2026 · DeepSeek (Hugging Face model card)
DeepSeek's 1.6-trillion-parameter flagship goes GA — and starts charging by the clock
7/19/2026 · AIToolsReview
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!