NVIDIA Cosmos Transfer 2.5 2B
NVIDIA · USA · 2025
A 2-billion-parameter control model that repaints a simulation: give it a depth map, a segmentation mask, an edge map or a blurred clip and it returns photorealistic 720p video with the same geometry — NVIDIA's bridge from simulator to training data.
NVIDIA Cosmos Transfer 2.5 2B is the third layer of the Cosmos stack, and the one with the clearest job. Cosmos Predict imagines what happens next; Cosmos Reason watches video and explains it. Transfer does neither — it takes a video whose geometry is already decided and repaints it into photorealism. Feed it a depth map, a segmentation mask, a Canny edge map or a blurred clip, plus a text prompt describing the look you want, and it returns a 720p video at 16 frames per second with the same shapes in the same places, but rendered as if a camera had filmed it. That is the sim-to-real problem stated as a product. A robotics team can build a scene in a simulator, where every object position is known and every label is free, then hand the geometry to Transfer and ask for the same scene at dusk, in rain, in a different warehouse, with different flooring. The result keeps the annotations that make the data useful for training and loses the plastic look that makes simulator footage a poor substitute for reality. The checkpoint holds 2,358,047,744 parameters. Architecturally it is the Cosmos Predict 2.5 base model with control branches bolted on: a few transformer blocks are replicated to process each control video, extract its signal and inject it back into the matching blocks of the base network. Up to four control inputs can run at once — they must be derived from the same source video and share identical spatio-temporal dimensions — and are merged through spatio-temporal weight maps, which is what lets one part of the frame follow the depth map while another follows the segmentation mask. NVIDIA states that the model was developed from Cosmos-Predict2.5. Edge and blur controls can be extracted automatically when only an RGB video is supplied. A second variant serves autonomous driving: given seven control videos from a vehicle's camera rig — front centre, front left, front right, rear left, rear right, rear tele and front tele — it generates 29 view-consistent frames for each of the seven cameras at 1280x720. That variant was trained on 720p video at 10 fps. The cost is not hidden. NVIDIA states the model needs 65.4 GB of GPU memory, and publishes generation times for a single segmentation-controlled clip: 285.83 seconds on a B200, 719.4 seconds on an H100 NVL, 870.3 seconds on an H100 PCIe and 2,326.6 seconds — nearly 39 minutes — on an H20, the cut-down Hopper card sold into the Chinese market. Synthetic data at this quality is bought in GPU-hours. The licence is the usual Cosmos arrangement: commercial use and derivative models allowed under the NVIDIA Open Model License, no claim on outputs, automatic termination if the safety guardrail is disabled without an equivalent replacement. In use, this is the second most downloaded model in the Cosmos catalogue's control and prediction layers: 84,673 downloads in the 30 days to 26 August 2026, eight times the Predict 2.5 model it is built on.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!