NVIDIA's robot brain wins its own comparison table by thirty points in one category. That category is a single test, and NVIDIA publishes it
Published: 9/13/2026 · Source: NVIDIA — Cosmos-Reason2 model card, Hugging Face ↗
NVIDIA compares all three sizes of Cosmos Reason 2 with the Alibaba model each of them was post-trained from. In three of the four domains the gain is measured in single digits. In the fourth, Smart Spaces, it is enormous: 77.79 against 47.55 for the 32B size, 69.96 against 42.66 for the 8B, 64.14 against 36.63 for the 2B. Read one row lower and the reason appears. Smart Spaces is not a domain with several tests in it. It is one test, Warehouse AI, and the dataset behind it sits on NVIDIA's own Hugging Face account.
The pattern repeats, more quietly, in Robotics. The overall gain there is real but modest, and it comes almost entirely from the two sets that carry the model family's own name: CR Common and CR Embodied, where every size beats its base model. On the two academic benchmarks in the same block the picture changes. On ERQA the 8B and the 2B score exactly what their base models score, to the decimal, and the 32B scores below its base, 45.25 against 46.50. On Where2Place the 8B scores 50.00 against 53.00 for the model it was trained from.
None of this is hidden. Every figure above comes from NVIDIA's own published table, with the source of each dataset linked from the same page, and the company deserves credit for printing rows where its model loses. It is also a reminder of what a vendor comparison table can and cannot tell a reader. A category built from one dataset is a single measurement, not a domain, and when the measurement and the model come from the same house, the margin says as much about the choice of test as about the model.
Our three Cosmos Reason 2 profiles carry these figures and we re-checked each number against the column it belongs to today, as part of an audit of every profile in the catalogue whose benchmark field mentions a rival model.