Wan2.1-T2V-14B
Alibaba Cloud · China · 2025
The model that opened Alibaba's video line in February 2025: 14 billion parameters, 480P and 720P clips from text, and the first video generator able to render legible Chinese and English lettering inside the frame.
Wan2.1-T2V-14B is the founding model of Alibaba's video family, published on 22 February 2025 together with its inference code under a plain Apache 2.0 licence. Eighteen months later it is still the most bookmarked file on the company's Hugging Face account, which says something about how few openly published video models of this class exist. Technically it is a diffusion transformer trained with flow matching: 40 layers, 40 attention heads, a hidden width of 5120, an umT5 encoder that reads prompts in Chinese and English, and a purpose-built 3D causal autoencoder the team calls Wan-VAE. That autoencoder is the part worth understanding, because it is what makes the rest affordable - it compresses video in space and time while preserving temporal causality, and it can encode and decode 1080P footage of unlimited length without losing what came earlier in the clip. The model itself outputs five-second clips at 480P or 720P. One capability was genuinely new at the time and is still uncommon: the model can generate readable text inside the video, in both Chinese and Latin script. Most video generators of that period produced letter-shaped smudges. Alibaba's own comparison put the model ahead of the open and closed competition of early 2025 on an internal set of 1,035 prompts scored across 14 dimensions - a vendor scoreboard, and worth reading as one. The licence deserves a note. Wan2.1 repositories ship the full Apache 2.0 text as a file, and the model card adds that Alibaba claims no rights over what users generate. From Wan2.2 onwards that file disappears from the weight repositories, even though the cards still link to it.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!