Wan2.2-S2V-14B
Alibaba Cloud · China · 2025
Give it one photograph and a recording of a voice, and it returns a filmed character speaking - Alibaba's open model for audio-driven performance, including long takes and lip-sync editing of existing footage.
Wan2.2-S2V-14B, published on 26 August 2025, takes a portrait and an audio track and produces video of that character performing the audio. The category is not new - Hunyuan-Avatar and OmniHuman were doing it first - but Alibaba's stated target is different. Existing systems handle a talking head or a singer well and fall apart in anything resembling a film scene, where characters interact, bodies move plausibly and the camera does not stand still. The published research is aimed squarely at that gap, and the accompanying paper reports wins over both of those competitors. Two secondary uses are arguably more practical than the headline one. The model does long-form generation, so the output is not limited to a single short take, and it does precise lip-sync editing of existing video - replacing the mouth movement in footage that already exists to match a different audio track, which is the dubbing problem. The repository ships an audio encoder alongside the video weights, a wav2vec2 model, which is why the parameter counter reports 16.3 billion for a checkpoint named 14B; the video model itself is the 14-billion part. On 5 September 2025 Alibaba added the CosyVoice speech synthesiser to the pipeline, so the driving audio can be generated from text rather than recorded. Output is 480P or 720P. The licence is Apache 2.0 by repository metadata, though as with the rest of Wan2.2 the licence file the card points at is missing.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!