Wan2.1-I2V-14B-480P
Alibaba Cloud · China · 2025
The same 14B image animator as the 720p build, trained for a smaller frame - and the one people actually run: it is downloaded two and a half times more often than its high-definition twin.
Wan2.1-I2V-14B-480P is the standard-definition half of Alibaba's open image-to-video pair, published on 25 February 2025 alongside the rest of the Wan 2.1 generation and licensed under Apache 2.0. It takes a photograph and a text prompt and returns a clip that begins with that photograph, at a target frame of 832x480. Architecturally it is the same machine as the 720p build: a 40-layer diffusion transformer of width 5120 with 40 heads, 36 input channels to carry the conditioning image and its mask, the shared umt5-xxl text encoder with a 512-token ceiling, and a CLIP vision encoder to read the supplied picture. The weight files count 16.4 billion parameters against a name that says 14 billion; the difference is the image-conditioning path, not a bigger generator. What separates the two downloads is not size or design but the resolution each was trained for, and Alibaba chose to ship them as two separate repositories rather than one model with a switch. That choice produced a result worth recording. This model is pulled from Hugging Face about 50,000 times a month against roughly 19,000 for the 720p build, while the 720p build carries nearly three times as many bookmarks. The high-definition version is the one people admire; the standard-definition version is the one they run. The pattern repeats across the open video field - the resolution that fits on hardware people already own wins on usage, whatever the benchmark says. The repository is also the tidier of the two on paperwork: the LICENSE.txt file that both model cards link to is present here and missing from the 720p sibling. For prompts, the maker offers optional rewriting through Qwen2.5-VL-7B, and as with all image-to-video models in this family the size setting fixes the area of the output while the aspect ratio follows the input photograph.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!