Nex-N2.5-Pro
Nex AGI · China · 2026
Agent model for operating computers and browsers: 397B parameters with 17B active, sight used as feedback rather than input, and a Qwen3.5 base that the maker names on its website but not in the model card.
Nex-N2.5-Pro is the middle and most widely used member of the Nex-N2.5 family, published with open weights on 8 September 2026 by Nex AGI, a Chinese laboratory that trains no foundation models of its own. This is the single most important fact for reading every number below: the company's own website lists Qwen3.5-397B-A17B as the base of this model, Qwen3.5-35B-A3B-Base for the mini, and DeepSeek-V4-Pro-Base for the Max. The Hugging Face model card omits those names and speaks only of building on the multimodal foundations of the earlier Nex-N2. Nex AGI is therefore a post-training laboratory: it buys none of the pre-training glory and all of the agent behaviour. What the post-training targets is visible in the benchmark selection. The model is built for long-horizon work in real environments — operating a computer, driving a browser, running and testing code — and vision is treated as the channel through which the agent checks whether its own action worked, not merely as a way of reading pictures. On the maker's published results it scores 82.7 on Terminal-Bench 2.1, 61.2 on SWE-Bench Pro, 56.4 on OSWorld-2 and 87.4 on OSWorld-G, the last of which beats every model in the comparison, including Claude Opus 5 and GPT-5.6 Sol. On broader knowledge work it falls clearly behind those two: 41.4 on Job Bench against 65.7 for Opus 5. The architecture is the Qwen3.5 sparse mixture of experts: 60 layers, 512 routed experts with 10 activated per token, hidden size 4096, vocabulary 248,320, hybrid attention with a linear-attention path, and a declared window of 262,144 tokens. Weights are published quantised to FP8 in blocks, which is why a single node with eight H100 cards is enough to serve the model — an unusually modest requirement for a model of this size and a deliberate part of the offer. Thinking is adaptive by default: the reasoning_effort parameter has three settings, and in the middle one the model itself decides whether to think before answering. Function calling uses the Qwen tool-call format, which is another visible trace of the base. At publication the model was hosted on OpenRouter free of charge, with no paid listing, so no market price exists yet. All benchmark figures here come from the maker, obtained partly with its own evaluation harnesses (NexAU for coding, NexCUA for computer use); the competitors' scores were copied from their makers' reports rather than re-measured.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!