MDL-6995EST.2025 · IDX.838
Robotics AIIn production

NVIDIA Cosmos Reason 2 2B

NVIDIA · USA · 2025

A 2B vision-language model that watches video and reasons about physics, space and time — the planning half of NVIDIA's Cosmos stack, post-trained from Alibaba's Qwen3-VL-2B-Instruct.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

NVIDIA Cosmos Reason 2 2B is the reasoning layer of NVIDIA's Cosmos stack for physical AI — and it is worth being precise about what that means, because the Cosmos name now covers two different kinds of model. Cosmos 3 Super, Nano and Edge are world models: they take a movement and predict what the scene will look like afterwards. Cosmos Reason does not generate anything visual. It watches video and writes text about it: what is happening, in what order, whether it is physically plausible, and what an embodied agent should do next. The checkpoint holds 2,438,696,960 parameters. NVIDIA published this checkpoint on 19 December 2025, together with the 8B tier, and refreshed it on 10 March 2026 with what it describes as improved benchmark scores and fewer hallucinations. A 32B tier followed on 29 April 2026. The architecture is not NVIDIA's own. Cosmos Reason 2 2B is post-trained from Qwen3-VL-2B-Instruct, Alibaba's open vision-language model, and keeps its architecture unchanged — a vision transformer feeding a dense transformer language model. NVIDIA's contribution is the post-training, and the benchmark table in the model card is honest about what that buys. Against the very base model it was built from, general vision scores move barely at all: 62.21 against 59.60 overall. On robotics questions the gap is wider, 45.52 against 42.07. But on the domains NVIDIA cares about commercially the difference is a different order entirely: 57.37 against 42.73 on self-driving benchmarks, and 64.14 against 36.63 on warehouse spatial intelligence. Post-training here does not make a smarter model — it makes a model that knows a specific world. Practically, the model reads text plus video (MP4) or images (JPG) and returns text only. The context window is 256,000 tokens, which is what makes long-video analysis possible; NVIDIA recommends feeding video at 4 frames per second to match the training setup, and reserving at least 4,096 output tokens, because the model is trained to think in a visible chain of thought before answering. Beyond describing a scene, it can point: 2D and 3D point localisation and bounding boxes, each with a text explanation of why that object was picked. NVIDIA names three uses. Video analytics agents that search and summarise recorded or live footage. Data curation, where the model filters and labels the sensor data used to train other physical-AI models. And robot planning, where it sits above a vision-language-action model as the deliberate, slow half of the brain — reading a complex instruction, breaking it into steps, and handing them down. Licensing is permissive with one sharp edge. The NVIDIA Open Model License allows commercial use, derivative models and unrestricted use of outputs — but the agreement terminates automatically if you disable or weaken the model's safety guardrails without putting a comparable one in place. The weights sit behind an automatic licence gate on Hugging Face — you accept the terms and the download starts, no human review. Judged by use rather than by size, this is the most popular model NVIDIA has in the entire Cosmos line: 969,959 downloads from Hugging Face in the 30 days to 26 August 2026, against roughly 121,000 for the 64-billion-parameter Cosmos 3 Super flagship. The small reasoner is the one people actually run.

#reasoning VLM#physical AI#open weights#video understanding#robot planning
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review