DeepSeek-V3.1
DeepSeek · China · 2025
One model with a thinking switch: DeepSeek folded its reasoning line back into the main model and trained it in a data format aimed at domestic chips.
DeepSeek-V3.1, published on 21 August 2025, is the release in which DeepSeek stopped shipping its reasoning models separately. It is a hybrid: the same weights answer immediately or think first, and the choice is made by the chat template rather than by picking a different model. The company's own comparison is blunt — the thinking build matches the answer quality of DeepSeek-R1-0528 while replying faster. Post-training also targeted tool calling and agent tasks, and tool use is available in the non-thinking mode. Underneath it is still the V3 body: 671 billion parameters with 37 billion activated per token, 61 layers, 256 routed experts plus one shared, eight active per token, and a 128K window (163,840 tokens). What changed is how that window was earned. V3.1 is post-trained on a base checkpoint built from the original V3 base through a two-phase long-context extension: the 32K phase was increased tenfold to 630 billion tokens and the 128K phase 3.3-fold to 209 billion tokens. That is a lot of compute spent on long documents alone. The quieter detail is the arithmetic. DeepSeek trained V3.1 using the UE8M0 FP8 scale format for both weights and activations, stating the goal as compatibility with microscaling data formats. In plain terms, the model was shaped to run on the accelerators its home market can actually buy. Unlike the original V3, the weights here are MIT-licensed, with a single LICENSE file covering both code and model.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!