MDL-7984EST.2025 · IDX.153
Robotics AIIn production

NVIDIA Cosmos Predict 2.5 2B

NVIDIA · USA · 2025

A 2-billion-parameter diffusion world model that turns a text prompt, a photo or a clip into five seconds of physically plausible 720p video — the small, widely used tier of NVIDIA's Cosmos prediction layer.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

NVIDIA Cosmos Predict 2.5 2B is the prediction layer of the Cosmos stack for physical AI, and the smaller of its two tiers. Where Cosmos Reason watches video and writes text about it, Predict does the opposite: it is handed a text description, a first frame, or a short clip, and it generates what happens next — a five-second 720p video at 16 frames per second. The point is not entertainment. The generated video is training data and a rehearsal space for robots and self-driving cars, worlds that a machine can practise in without anyone building them. The checkpoint holds 2,059,174,912 parameters. It is a diffusion transformer that denoises video in latent space, with cross-attention layers letting the text prompt steer the whole denoising process; when an image or a video is supplied, its latent frames are concatenated with the frames being generated. NVIDIA states plainly that the model was developed from Cosmos-Predict2-2B-Video2World, the previous generation's video-to-world checkpoint. The most practical change in this generation is consolidation. In Cosmos Predict 2, text-to-image, text-to-world and video-to-world lived in separate repositories and separate downloads. Predict 2.5 puts one base checkpoint and its post-trained variants under a single roof: a pre-trained and a post-trained general model; an automotive multiview variant that predicts the same scene across seven vehicle cameras; two robot multiview variants, one of them post-trained specifically on AgiBot data; an action-conditioned variant that takes a robot's action sequence as the condition and answers at 256p and 4 fps; and a policy variant post-trained on the LIBERO and RoboCasa benchmarks that returns robot actions and future proprioception rather than just pictures. NVIDIA is candid about the limits. The models struggle with long, high-resolution video without artefacts; temporal inconsistency, unstable camera and object motion and imprecise interactions are named in the card, along with objects that disappear or morph and motions that are not physically plausible. Applications that need law-abiding physics or complex multi-agent dynamics remain hard. The weights are a free download under the NVIDIA Open Model License: commercial use allowed, derivative models allowed, no NVIDIA claim on outputs — with the standard Cosmos sting that the licence terminates automatically if you disable the model's safety guardrail without an equivalent replacement. The download sits behind an automatic licence gate on Hugging Face. One number puts the generation in perspective. In the 30 days to 26 August 2026, this checkpoint was downloaded 10,247 times, while the model it replaced — Cosmos Predict 2 2B Video2World — was downloaded 229,143 times. Nearly a year after release, the newer world model is still the less used one.

#world model#physical AI#open weights#video generation#robot data
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review