All newsReleases

Nex AGI's own benchmark page changes the owner of its third column halfway down — and the model that disappears is the flagship

Published: 9/13/2026 · Source: Nex AGI — Nex-N2.5 model card (Hugging Face)

Nex AGI published three models with open weights on 8 September 2026: the small mini, the mid-sized Pro and the 1.6-trillion-parameter Max. The results for all three sit on one page, in two tables placed a few dozen lines apart — and the tables do not have the same columns. In the upper table, covering coding and agent work, the third column is Nex-N2.5-Max. In the lower one, covering multimodal benchmarks and computer use, the third column belongs to MiniMax-M3, a competitor's model. Nothing but the header marks the change. A reader who takes in both tables with one sweep of the eye will credit the flagship with 22.3 points on OSWorld-2, the benchmark for actually operating a desktop — a figure that belongs to somebody else's, much smaller model. The flagship has no result there at all, and the reason is stated plainly by the maker a few paragraphs higher: Max is built on a text-only foundation, while the mini and the Pro continue the multimodal line. A model with no sight is not put through tests that consist of looking at a screen, so it drops out of the second table entirely instead of scoring badly in it. What follows is practical rather than academic. Anyone picking a Nex model to drive a computer is choosing between the Pro (OSWorld-2 56.4, OSWorld-G 87.4 — the best figure in the maker's own comparison) and the mini (30.5 and 82.9, the latter ahead of Claude Opus 5 at 76.8 and GPT-5.6 Sol at 77.7). Size is not the axis here; sight is. And the largest model in the family, by far the most expensive to run at two nodes and sixteen H200 cards, is the one that cannot do this job at all. Editorial note: our own profile of Nex-N2.5-Max carried that 22.3 for a day, taken from the column in good faith. We corrected it on 17 September, which is also how this story came about.