DeepSeek-V3.2
DeepSeek · China · 2025
The last model of DeepSeek's third generation: 685 billion parameters, a 163,840-token window and sparse attention, released with open weights under the MIT licence.
DeepSeek-V3.2 is the closing model of DeepSeek's third generation, published on Hugging Face on 1 December 2025 — eleven months after DeepSeek-R1 and five months before the V4 line began. It is a sparse mixture-of-experts transformer: 685 billion parameters in total across 61 layers, 256 routed experts plus one shared, eight of them active per token, a vocabulary of 129,280 and a context window of 163,840 tokens. The weights ship quantised to FP8 with 128×128 blocks, and the model card does not state how many parameters are activated per token. The lab describes three changes over the previous release. DeepSeek Sparse Attention (DSA) cuts the cost of attention, which is what makes the long window affordable; a scaled reinforcement-learning stage lifts reasoning and tool use; and an agentic task synthesis pipeline generates the training data for that. DeepSeek claims the model performs comparably to GPT-5, and that a high-compute variant, DeepSeek-V3.2-Speciale, goes past it — with gold-medal results at the 2025 International Mathematical Olympiad and the International Olympiad in Informatics. The competition submissions themselves were published in the repository, which is unusual and lets anyone check the claim rather than take it on trust. Speciale is the deep-reasoning build and, by DeepSeek's own note, does not support tool calling. The model has a switchable thinking mode in its chat template and a separate developer role reserved for search agents. Weights and repository are MIT-licensed, as with the rest of DeepSeek's catalogue up to V4-Flash. Read on 25 August 2026, nine months after publication, the repository still records 1.18 million downloads in thirty days — more than the current V4-Pro flagship build.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!