Releases8/11/2026 · BytePlus ModelArk documentationByteDance's new video model holds one shot for thirty seconds - and its API card quietly says 720p, not 4K
ByteDance Seed's Seedance 2.5 is the first of the company's video models built for scenes rather than clips. The ceiling on a single generation rose from fifteen seconds to thirty, and the documentation is explicit that those thirty seconds are one coherent take - not segments stitched together, which is how most tools reach that length. Audio is produced in the same pass as the picture, and a reference track of speech, music or sound effects can be the sole input, with pacing and lip movement fitted to it.
The capability the company puts first is what it calls omni reference. One request may carry up to fifty assets - thirty images, ten video clips and ten audio files - and the model reads them together. Practically, that turns the endpoint into four tools at once: text-to-video, generation from a first frame or a first-and-last frame pair, timestamp-level editing of an existing video (replacing a subject, removing an object, repainting part of a frame while keeping the original aspect ratio and duration), and extending a clip forwards or backwards. Output was widened to mov with H.264 video and PCM audio so that colour and sound survive an edit.
One figure in circulation does not hold up. Much of the coverage credits Seedance 2.5 with 4K output; the model card on BytePlus ModelArk lists 480p and 720p for dreamina-seedance-2-5-260628, and it is the older Seedance 2.0 entry that is documented up to 1080p and 4K. The pricing table is equally specific - USD 10.70 per million tokens without video input and USD 6.40 with it - and applies to 480p and 720p outputs only. As with the rest of ByteDance's video line, no weights, parameter count or technical report were published.
Seedance 2.5 →