Grok Imagine Video 1.5
xAI · USA · 2026
xAI's video model: up to 15 seconds at native 1080p, with reference images and a preset voice for the subject.
Grok Imagine Video 1.5 is xAI's second-generation video model, reachable through the same Imagine API as the company's image models. It generates clips of one to fifteen seconds from a text prompt, from a still image, or from a set of reference images, and it is the first xAI video model to render text-to-video and image-to-video at native 1080p rather than upscaling from a lower resolution. Text-to-video on this model is not a single pass: xAI states in its documentation that the model first generates a first frame from the prompt and then animates it, with the intermediate image never returned to the caller. Reference-to-video works differently from image-to-video, guiding the clip with up to several images without locking the opening frame - the company points to virtual try-on, product placement and character-consistent storytelling as the intended uses. The model can also give its subject a voice: up to three preset voices from the same roster as the company's text-to-speech models can be attached to a request and tagged in the prompt. Voice cloning from a customer's own audio is not generally available and is offered only to selected partners on request. Two further endpoints extend an existing clip from its last frame and edit an existing clip, the latter capped at 720p and about 8.7 seconds. At $0.08 per second of output the model costs sixty percent more than the first-generation Grok video model, and unlike that older model it cannot be used through the discounted batch API - xAI rejects both Imagine 1.5 models with an explicit "not supported for batch processing" error.
▸News
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!