xAI on May 31 launched Grok Imagine Video 1.5 as an API preview. Within days, this model topped the Artificial Analysis Image-to-Video Arena 720p leaderboard with 1473 Elo — 52 points higher than the previous generation Grok Imagine Video (1421), crossing ByteDance's Seedance 2.0 (1467) and Google Veo 3.1 (1397). For xAI, this is a release that didn't make the front page of the media, but pushed the engineering frontier of generative video forward a notch.
The architecture is the real variable this time. The model's internal code name is Aurora, taking an autoregressive Mixture of Experts (MoE) path — diverging from the current mainstream diffusion path, using "frame-by-frame generation" for temporal extension, rather than one-shot denoising of the whole video. xAI trained on Colossus supercomputer with 110,000 GB200 cards, and the Hotshot video team acquired in March 2025 contributed key capabilities. On the technical specs, the output is 480p/720p, fixed 24 FPS, single segment 6-15 seconds, priced at $0.08-0.14 per second, an order of magnitude lower than Runway Gen-4, Kling 2.x, and Veo 3.1 at the same tier.
What really changes the workflow is "native same-frame audio." 1.5 puts voice, sound effects, ambient sound, and music into the same forward inference, with ambient sound following the spatial position of objects in the picture. Coupled with $0.01 per image input and "multi-shot storyboard" splicing capability, the advertising team gets an industrial-grade pipeline from still-image assets directly to short film with sound — a single 6-second 720p video is about $0.85, image and audio included.
It should be noted that the current preview only supports image-to-video, not T2V, editing, or multi-image editing; X Premium full rollout is still in progress.
Commentary: Grok Imagine 1.5's significance is not "another AI video competitor," but that it slams the invisible cost curve of generative video downward — $0.85 per segment, zero post-production audio, and the natural advantage of autoregressive MoE on temporal consistency. Together, these mean that "dynamic advertising assets" have the possibility of industrial-scale deployment for the first time. The next thing to see is whether Aurora can maintain consistency at minute-level long temporal scales, and whether xAI will merge T2V into the same inference endpoint.