Fun-CineForge
Alibaba Cloud · China · 2026
Open dubbing model from Alibaba's FunAudioLLM lab: it speaks a translated or rewritten line in a chosen voice, timed to the actor's lips, across monologue, narration, dialogue and crowded multi-speaker scenes. Ships as a complete pipeline, face recognition included.
Fun-CineForge is an open model for dubbing film and television. Given a clip, a script with timings and a reference voice, it speaks the line so that it fits the actor's mouth, and it is built specifically for the cases where dubbing usually falls apart: narration over a scene, two people talking over each other, a crowd where the speaker changes every few seconds. The unusual part is what comes in the package. This is not a single model but a working pipeline, 12.68 GiB across 42 files. The voice itself comes from a language model of about 5.8 GB on a Qwen2-0.5B CosyVoice backbone, a 3.8 GB flow-matching module and a vocoder. Around them sit the tools needed to understand the footage: vocal separation to strip the music off the original track, voice activity detection, speaker verification, active speaker detection, face detection, facial landmarks, face quality assessment and a 249 MB face recognition network. A dubbing model has to know which face on screen is talking, so it ships with the machinery to recognise faces — a capability worth noticing in a package many people would install without reading the file list. Alongside the model the team published its dataset pipeline, which is arguably the larger contribution: a documented route from raw television episodes to labelled dubbing data, and with it CineDub-CN, described as the first large-scale Chinese television dubbing dataset, later extended to English as CineDub-EN. One step in that route is easy to miss and says something about the state of the field. The cleaning stage, which corrects the output of the specialist models, calls a general multimodal model through an API — and the documented default is Google's Gemini 3 Pro. An open Alibaba pipeline reaches for a competitor's commercial model to check its own work. The producer reports the payoff: character error rate falls from 4.53 to 0.94 percent, and speaker diarisation error from 8.38 to 1.20 percent, which they describe as matching or beating manual transcription. The timeline runs from the pipeline toolkit in December 2025, through English support in February 2026, to the inference code and checkpoints on 16 March 2026. Chinese and English only. The repository carries an Apache 2.0 label, but the card states plainly that this is a research artifact rather than a commercial Tongyi Lab product, and that the dataset samples come with their own terms. The producer says inference runs on a consumer-grade GPU; the interface for dubbing a raw video straight from a subtitle file is still described as under development.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!