MDL-3959EST.2025 · IDX.268
Robotics AIIn production

NVIDIA Cosmos Predict 2.5 14B

NVIDIA · USA · 2025

The 14-billion-parameter tier of NVIDIA's Cosmos prediction layer: the same five-second 720p world generation as the 2B model, with seven times the parameters and none of the specialised robot variants.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

NVIDIA Cosmos Predict 2.5 14B is the large tier of the prediction layer in NVIDIA's Cosmos stack — the part that imagines what happens next. Given a text description, a first frame or a short clip, it generates a five-second 720p video at 16 frames per second, meant not for viewing but as synthetic training material and as a world in which robots and self-driving software can rehearse. The checkpoint holds 14,368,048,004 parameters — seven times the 2B tier — and shares its design: a diffusion transformer denoising video in latent space, with cross-attention carrying the text prompt through the whole process and conditioning frames concatenated along the time axis. NVIDIA states that the model was developed from Cosmos-Predict2-14B-Video2World, the corresponding checkpoint of the previous generation. It arrived on 4 December 2025, two months after the 2B tier. The difference between the tiers is not only size. The 2B checkpoint carries a whole shelf of post-trained variants — seven-camera automotive multiview, robot multiview including one trained on AgiBot data, action-conditioned prediction, and a policy variant that outputs robot actions. The 14B checkpoint ships as two things only: pre-trained and post-trained. The large model is where output quality lives; the specialised, deployable work happens one tier down. The download figures say the same thing more bluntly. In the 30 days to 26 August 2026, the 14B model was downloaded 1,971 times, the 2B model 10,247 times, and the previous generation's Cosmos Predict 2 2B Video2World 229,143 times. World models are expensive to run, and the field votes with its GPUs. NVIDIA's stated limits apply here as well: artefacts in long, high-resolution video, temporal inconsistency, unstable camera and object motion, objects that disappear or morph, and physically implausible movement. The company tested inference on H100, A100 and B200 hardware, supports the Ampere, Hopper and Blackwell microarchitectures, and has tested only BF16 precision and only Linux. Licensing is the standard Cosmos arrangement: the NVIDIA Open Model License permits commercial use and derivative models and claims nothing over generated outputs, but terminates automatically if the model's safety guardrail is bypassed or weakened without an equivalent replacement. Unlike the 2B checkpoint, the 14B weights are not behind a licence gate — the model card and files are open to read and download directly.

#world model#physical AI#open weights#video generation#simulation
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review