MDL-7445EST.2025 · IDX.928
VideoIn production

Wan2.2-I2V-A14B

Alibaba Cloud · China · 2025

The image-to-video half of Alibaba's last openly published video flagship: the same two-expert denoiser as the text model, but the first frame is a picture the user supplies.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Wan2.2-I2V-A14B animates a still picture. It was published on 28 July 2025 alongside the text-to-video model of the same generation, and shares its architecture exactly: a mixture-of-experts denoiser in which a high-noise expert handles the early steps, where layout and motion are decided, and a low-noise expert finishes the detail. Each expert holds about 14 billion parameters, so the checkpoint totals roughly 27 billion while only 14 billion run at any one step. In the repository this shows up as two separate model folders, high_noise_model and low_noise_model, which is why the usual parameter counter on Hugging Face reports nothing at all for this file. The practical difference against the text model is control. Starting from an image removes the hardest part of a text prompt - describing a scene precisely enough to get the composition you wanted - and leaves the model with the job it is better at, which is deciding how that scene should move. Output is 480P or 720P. Alibaba states that Wan2.2 was trained on 65.6 percent more images and 83.2 percent more video than Wan2.1, with aesthetic labelling for lighting, composition, contrast and colour, which is how the generation gets its more deliberate cinematic look. One caveat worth recording: the model card promises a full licence text in a file called LICENSE.txt, and no Wan2.2 repository actually contains that file. The Apache 2.0 declaration in the repository metadata and in the matching code repository on GitHub still stands.

#video generation#image to video#open weights#mixture of experts
Official website

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review