News

What's happening in robotics and AI — curated by the wujec.ai editors.

Releases9/12/2026 · DeepSeek model card, DeepSeek-V4.1-Flash

DeepSeek's new model beats Opus and GPT on the agent test everyone has already saturated. On the two harder versions of that same test, it loses by twenty points

DeepSeek published V4.1-Flash on 10 September 2026 with a comparison table that reads like a clean sweep — until you follow one test through three of its versions. On Terminal-Bench 2.1 the new model scores 90.6, ahead of Opus-5.0 at 89.1 and GPT-5.6 Sol at 88.8. On Terminal-Bench 3.0 it scores 30.0 against Opus-5.0's 43.3. On version 4.0, 31.2 against 51.8. The three numbers describe the same model on the same family of tasks, and the spread is a lesson in how to read a benchmark. Version 2.1 has been in circulation long enough for every serious laboratory to score in the high eighties; a 1.5-point lead there separates models that are, in practice, equally capable. Versions 3.0 and 4.0 are the harder revisions written precisely because the old one stopped telling models apart — and there the distance between DeepSeek's model and the most expensive Western one is roughly twenty points. The same pattern repeats elsewhere in the table. On Humanity's Last Exam, a test still far from saturated, V4.1-Flash reaches 36.8 against Opus-5.0's 56.3. On ProgramBench, 20.3 against 37.0. What the model does win is worth stating plainly, because it is not a small thing. Its Codeforces rating of 3,471 is the highest in the table. It resolves 74.2 percent of DeepSWE v1.1 tasks, leads on CyberGym with 88.1 and on AutomationBench with 54.8. And it does this while activating 8 billion parameters per token when reading a prompt and 16 billion when writing — out of 552 billion in the backbone. The efficiency is the actual headline. DeepSeek rebuilt the part of the model that stores the conversation, cutting the memory held per token to 890 bytes, about a quarter of what the July model needed. Through the company's API the new model costs $0.30 per million input tokens at peak and $1.20 per million output, against $0.44 and $1.32 for its predecessor, with off-peak hours at half price. Weights are MIT-licensed and were downloaded 140,636 times in the following weeks. Read together, the table says something more useful than 'best model'. For agent work at scale, where a million-token session runs on a budget, this is a remarkable instrument at a fraction of frontier pricing. For the hardest reasoning a buyer can currently pose, the expensive models remain ahead, and DeepSeek's own numbers say so.

DeepSeek-V4.1-Flash
Regulation8/25/2026 · Hugging Face (model repositories and licence files)

Open weights stopped asking permission — of the forty most downloaded models, only Meta's still need an account

We read the licence of every repository in the forty most downloaded text-generation models on Hugging Face today, 25 August 2026, and checked whether the weights can actually be fetched. Two of the forty are gated: Meta's Llama-3.2-1B-Instruct and Llama-3.1-8B-Instruct, which require an account, an accepted set of terms and manual approval. A third, Google's Gemma 3 1B, is gated for the same reason. Everything else in the list downloads without an account. The change that made this true is recent and belongs to Google. Gemma 3, published under Google's own Gemma Terms of Use, could not be downloaded — or even have its licence text read — without logging in; a request for the file returns HTTP 401. Gemma 4, published on 11 March 2026, is under plain Apache 2.0 and is not gated at all. It is not a marginal model: the 31B instruction-tuned repository alone records 8.79 million downloads in thirty days, more than any other frontier-class open-weight model we track. What remains of bespoke licensing is narrower than its reputation. Moonshot's Kimi K2 carries a "Modified MIT" licence whose only modification is a display duty: a product with more than 100 million monthly active users or more than 20 million dollars in monthly revenue must show the words "Kimi K2" in its interface. DeepSeek-V3, the model on which that lab's whole line stands, is still routinely described as MIT-licensed and is not — its repository ships MIT for the code and a separate DeepSeek License Agreement for the weights. Only from March 2025 did the company put the weights themselves under MIT. One habit is worth flagging for anyone who takes a label at face value. Several heavily downloaded repositories declare a licence in the card metadata and ship no licence file at all, and quantised repacks inherit the label from a model card rather than from the original terms. The label is metadata; the file is the contract. Where the two disagree, only one of them is enforceable.

Gemma 4 31B
Business8/19/2026 · DeepSeek — dokumentacja API (Models & Pricing)

DeepSeek's clock is a Beijing clock: the same working day costs 75% more there than in San Francisco

DeepSeek's price list, in force since 16 August 2026, is the first from a major model provider in which the hour of the day is a full pricing variable rather than a promotion: peak hours are 01:00–04:00 and 06:00–10:00 UTC, and everything outside them costs exactly half. The published windows say nothing about where those hours fall. Converted, they say a great deal. Beijing keeps UTC+8 all year, with no daylight saving. The two peak windows land at 09:00–12:00 and 14:00–18:00 local time — Chinese office hours. The gap between them, 04:00–06:00 UTC, is 12:00–14:00 in Beijing: the lunch break, priced at half rate. The tariff is not described in geographic terms anywhere in the documentation, but it is drawn around one country's working day. That has an arithmetic consequence nobody has published. Take a team using DeepSeek-V4-Pro evenly through a 09:00–17:00 local working day, and price a million output tokens ($3.96 at peak, $1.98 off-peak). In San Francisco the local working day is 16:00–24:00 UTC and misses both peak windows entirely, so every hour bills at $1.98. In Warsaw it is 07:00–15:00 UTC, of which three hours are peak: $2.72 per million, 37.5% above San Francisco. In Beijing it is 01:00–09:00 UTC, of which six hours are peak: $3.47 per million — 75% above San Francisco, for identical work on identical weights. Bangalore lands in between at $3.09, or 56% above. European bills also move with the clock change. In winter Warsaw shifts to UTC+1, only two working hours stay inside the peak, and the same million falls to $2.48 — a 9% discount granted by nothing but the end of daylight saving. What this does not show: DeepSeek is not charging anyone by location. The tariff is identical worldwide and time-based, and the company presents it as load management — the peak windows are simply when its servers are busiest, which is when its home market is at work. Batch and overnight jobs can be moved into the cheap hours by anyone, anywhere, and for most production workloads cached input, billed at a fiftieth of a cache miss, matters far more than the clock. The figures above assume usage spread evenly across office hours, which no real team does exactly. But the direction is not an artefact: the further a user's working day sits from Beijing's, the less DeepSeek's price rise costs them.

DeepSeek-V4-Pro
Business8/15/2026 · DeepSeek — dokumentacja API (cennik)

The cheapest lab stops being cheap: DeepSeek raises API prices from 16 August and starts charging by the clock

DeepSeek has published a new price list for its V4 models, effective 16 August 2026. The company that built its reputation on undercutting everyone else is raising rates across the board — by between 50 percent and more than 1,100 percent, depending on the model, the type of token and the hour of the day. The numbers for DeepSeek-V4-Flash: output tokens go from $0.28 to $1.32 per million at peak, cache-miss input from $0.14 to $0.44, and cached input from $0.0028 to $0.014 — a fivefold rise on the cheapest line in the catalogue. DeepSeek-V4-Pro goes from $0.87 to $3.96 per million output tokens at peak, while its cached input rises from roughly $0.0036 to $0.044 per million: the 1,100 percent figure comes from that one line, not from the headline rate. The structural change matters as much as the numbers. From Sunday the price depends on when the request is made: peak hours are 01:00–04:00 and 06:00–10:00 UTC, everything else is off-peak and costs exactly half. DeepSeek says the tiered structure is meant to "allocate resources more reasonably" and push developer workloads towards less congested windows. In practice it is the first time a major model provider has made the hour of the day a first-class pricing variable rather than a promotional discount. The timing is what makes this striking. In the same week Anthropic cancelled a planned 50 percent rise for Claude Sonnet 5 and Google put an expiry date on its Gemini 3.7 Flash discount, DeepSeek moved in the opposite direction — and moved further than either. Off-peak V4-Flash output at $0.66 per million is still cheap by Western standards, but the gap that made DeepSeek an obvious default has narrowed by a factor of four or five. Both V4 models keep their 1M-token context window and 384K maximum output; the concurrency limits (2,500 simultaneous requests for Flash, 500 for Pro) are unchanged. We have updated the pricing in both profiles in the catalogue.

DeepSeek-V4-Flash
Releases8/13/2026 · DeepSeek (Hugging Face model card)

DeepSeek ships the final V4-Pro: open weights, 1M context, agentic scores up sharply

DeepSeek published DeepSeek-V4-Pro-0813 on Hugging Face on 13 August 2026, describing it as the official release of V4-Pro and the successor to the preview weights that had carried the name since April. The repository went up under the MIT licence, so the full model — 1.65 trillion parameters in a sparse mixture-of-experts layout, 61 layers, 384 routed experts with six active per token, shipped in FP8 — can be downloaded and self-hosted. The context window stays at 1,048,576 tokens. The architecture is unchanged from the preview; what is new is a DSpark speculative-decoding module attached to the model and a reworked reasoning control. The `reasoning_effort` parameter now takes three levels — low, high and max — letting callers trade deliberation time against cost. DeepSeek also dropped the Jinja chat template in favour of a documented Python encoder, a change that will require work from anyone running the weights locally. The gains DeepSeek reports are concentrated in agentic and coding work, and they are large. On the company's own table the new build reaches 62.7% on DeepSWE against 12.8% for the preview, 61.5% on NL2Repo against 38.5%, 83.3% on Cybergym against 52.7% and 87.9 on Terminal Bench 2.1 against 72.1. On Humanity's Last Exam it scores 42.7% without tools and 60.0% with them, up from 37.7%. Against rivals in the same table the picture is closer. Kimi K3 edges it on Terminal Bench (88.3) and DeepSWE (67.5), and Anthropic's Fable 5 leads on Humanity's Last Exam (53.3 without tools). What DeepSeek keeps is the position it has held all year: comparable results at frontier level, with the weights published rather than rented. Benchmark figures here are the vendor's own and have not yet been reproduced independently.

DeepSeek-V4-Pro
Releases7/19/2026 · AIToolsReview

DeepSeek's 1.6-trillion-parameter flagship goes GA — and starts charging by the clock

DeepSeek-V4-Pro left preview on 19 July 2026, three months after its first public release. The weights are the same ones published on 24 April: a sparse mixture-of-experts model with 1.6 trillion total parameters, roughly 49 billion of them active per token, a one-million-token context window and an MIT licence that allows anyone to download and self-host it. What general availability changed was the commercial side. DeepSeek introduced peak-time pricing for the first time, doubling rates during Beijing business hours; off-peak the API costs 0.435 US dollars per million input tokens and 0.87 per million output. Cached input is charged at roughly a hundredth of the off-peak input rate. On 24 July the company retired the legacy deepseek-chat and deepseek-reasoner endpoints, so integrations now have to name a model explicitly. On DeepSeek's own GA figures the model scores 80.6% on SWE-bench Verified — the strongest published result for an open-weights model — with a Codeforces rating of 3,206, 57.9% on SimpleQA Verified and 37.7% on Humanity's Last Exam. Independent measurement tells a subtler story: Artificial Analysis rated V4-Pro at 44 on its Intelligence Index at GA, down from 52 at preview. The model did not get worse; the field moved during the three months it spent in preview. The smaller DeepSeek-V4-Flash and the earlier DeepSeek-R1 remain available, and all three now have profiles in the wujec.ai catalogue.

DeepSeek-V4-Pro