DeepSeek-R1-0528
DeepSeek · China · 2025
The May 2025 rebuild of DeepSeek-R1: accuracy on AIME 2025 went from 70 to 87.5 per cent, and the model now spends 23,000 tokens thinking about each question instead of 12,000.
DeepSeek called this a minor version upgrade. The numbers say otherwise. Published on 28 May 2025 on the same 671-billion-parameter body as DeepSeek-R1, R1-0528 raised accuracy on the AIME 2025 mathematics set from 70 to 87.5 per cent, its Codeforces rating from 1530 to 1930, SWE-bench Verified from 49.2 to 57.6 per cent and Aider-Polyglot from 53.3 to 71.6 per cent. Humanity's Last Exam, the hardest of the general benchmarks, more than doubled: 8.5 to 17.7 per cent. The company is unusually direct about where the gain comes from, and it is not a new architecture. R1-0528 simply thinks longer. On the AIME set the earlier model used about 12,000 tokens per question; this one averages 23,000. The improvement was bought with post-training compute and algorithmic changes to the reasoning stage, which makes this release a clean, publicly documented data point on how far the same weights can be pushed by spending more tokens at answer time. One figure moved the other way and DeepSeek published it anyway: SimpleQA, which measures factual recall, fell from 30.1 to 27.8 per cent. A model that reasons harder is not automatically a model that remembers better. The release also added function calling, JSON output and support for a system prompt — the practical gaps that kept the first R1 out of production stacks — and dropped the requirement to prefix answers with a thinking token. Weights and repository are MIT-licensed. Maximum generation length in DeepSeek's own evaluation was 64,000 tokens.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!