News

What's happening in robotics and AI — curated by the wujec.ai editors.

Releases9/12/2026 · DeepSeek model card, DeepSeek-V4.1-Flash

DeepSeek's new model beats Opus and GPT on the agent test everyone has already saturated. On the two harder versions of that same test, it loses by twenty points

DeepSeek published V4.1-Flash on 10 September 2026 with a comparison table that reads like a clean sweep — until you follow one test through three of its versions. On Terminal-Bench 2.1 the new model scores 90.6, ahead of Opus-5.0 at 89.1 and GPT-5.6 Sol at 88.8. On Terminal-Bench 3.0 it scores 30.0 against Opus-5.0's 43.3. On version 4.0, 31.2 against 51.8. The three numbers describe the same model on the same family of tasks, and the spread is a lesson in how to read a benchmark. Version 2.1 has been in circulation long enough for every serious laboratory to score in the high eighties; a 1.5-point lead there separates models that are, in practice, equally capable. Versions 3.0 and 4.0 are the harder revisions written precisely because the old one stopped telling models apart — and there the distance between DeepSeek's model and the most expensive Western one is roughly twenty points. The same pattern repeats elsewhere in the table. On Humanity's Last Exam, a test still far from saturated, V4.1-Flash reaches 36.8 against Opus-5.0's 56.3. On ProgramBench, 20.3 against 37.0. What the model does win is worth stating plainly, because it is not a small thing. Its Codeforces rating of 3,471 is the highest in the table. It resolves 74.2 percent of DeepSWE v1.1 tasks, leads on CyberGym with 88.1 and on AutomationBench with 54.8. And it does this while activating 8 billion parameters per token when reading a prompt and 16 billion when writing — out of 552 billion in the backbone. The efficiency is the actual headline. DeepSeek rebuilt the part of the model that stores the conversation, cutting the memory held per token to 890 bytes, about a quarter of what the July model needed. Through the company's API the new model costs $0.30 per million input tokens at peak and $1.20 per million output, against $0.44 and $1.32 for its predecessor, with off-peak hours at half price. Weights are MIT-licensed and were downloaded 140,636 times in the following weeks. Read together, the table says something more useful than 'best model'. For agent work at scale, where a million-token session runs on a budget, this is a remarkable instrument at a fraction of frontier pricing. For the hardest reasoning a buyer can currently pose, the expensive models remain ahead, and DeepSeek's own numbers say so.

DeepSeek-V4.1-Flash