MDL-5562EST.2026 · IDX.827
VideoIn production

Grok Imagine Video 1.5

xAI · USA · 2026

xAI's video model: up to 15 seconds at native 1080p, with reference images and a preset voice for the subject.

wujec.ai score

/10

Community score

no votes yet
Sign in to rate

Grok Imagine Video 1.5 is xAI's second-generation video model, reachable through the same Imagine API as the company's image models. It generates clips of one to fifteen seconds from a text prompt, from a still image, or from a set of reference images, and it is the first xAI video model to render text-to-video and image-to-video at native 1080p rather than upscaling from a lower resolution. Text-to-video on this model is not a single pass: xAI states in its documentation that the model first generates a first frame from the prompt and then animates it, with the intermediate image never returned to the caller. Reference-to-video works differently from image-to-video, guiding the clip with up to several images without locking the opening frame - the company points to virtual try-on, product placement and character-consistent storytelling as the intended uses. The model can also give its subject a voice: up to three preset voices from the same roster as the company's text-to-speech models can be attached to a request and tagged in the prompt. Voice cloning from a customer's own audio is not generally available and is offered only to selected partners on request. Two further endpoints extend an existing clip from its last frame and edit an existing clip, the latter capped at 720p and about 8.7 seconds. At $0.08 per second of output the model costs sixty percent more than the first-generation Grok video model, and unlike that older model it cannot be used through the discounted batch API - xAI rejects both Imagine 1.5 models with an explicit "not supported for batch processing" error.

#video generation#image-to-video#voice
Official website

News

Videos

No videos yet.

Reviews

No reviews yet. Be the first!

Sign in to write a review