[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-grok-imagine-1-5-ar-moe-arena-top":3,"news-related-e8ad36b6-ec3c-4c90-8fbc-bf1e6807ae34":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e8ad36b6-ec3c-4c90-8fbc-bf1e6807ae34","xAI Grok Imagine Video 1.5：单图生视频登顶 Arena榜首，自回归 MoE 改写视频生成规则","xAI 于 5 月 31 日以 API 预览形式上线 Grok Imagine Video 1.5。短短数日，这款模型在 Artificial Analysis 的 Image-to-Video Arena 720p 榜上以 1473 Elo 直接登顶——比上一代 Grok Imagine Video（1421）高 52 分，跨过字节跳动 Seedance 2.0（1467）和 Google Veo 3.1（1397）。对 xAI 而言，这是一次不在媒体头版、却把生成式视频的工程边界往前推了一截的发布。\n\n**架构是这次的真正变量。** 模型内部代号 Aurora，采用自回归 Mixture of Experts（MoE）路线——与当前主流的扩散路线分道扬镳，以\"逐帧生成\"做时序延展，而非一次性去噪整段视频。xAI 在 Colossus 超算上用 11 万张 GB200 完成训练，2025 年 3 月收购的 Hotshot 视频团队贡献了关键能力。技术规格上，输出 480p\u002F720p，固定 24 FPS，单段 6~15 秒，定价 0.08~0.14 美元\u002F秒，比 Runway Gen-4、Kling 2.x、Veo 3.1 同档位低一个数量级。\n\n**真正改变工作流的是\"原生同帧音频\"。** 1.5 把语音、效果音、环境音和配乐放进同一次前向推理，环境音会跟随画面物体的空间位置变化。配合 0.01 美元\u002F张的图像输入和\"多镜头分镜\"拼接能力，广告团队拿到的是一条从静帧素材直接到带声短片的工业化链路——单条 6 秒 720p 视频约 0.85 美元含图与音。\n\n需提醒的是，当前 preview 仅支持 image-to-video，不支持 T2V、剪辑或多图编辑；X Premium 端全面开放仍在进行中。\n\n**评论：** Grok Imagine 1.5 的意义不在\"AI 视频又多一个选手\"，而在于它把生成式视频的隐形成本曲线向下砸穿——0.85 美元\u002F段、零后期音频、自回归 MoE 在时序一致性上的天然优势，三者叠加意味着\"动态广告素材\"第一次具备工业级铺开的可能。下一步要看 Aurora 能否在分钟级长时序下保持一致，以及 xAI 是否把 T2V 合并进同一推理端点。","https:\u002F\u002Fx.ai\u002Fnews\u002Fgrok-imagine-1-5","b82e17a3-1dbd-4b5d-88dc-9f518f917cc0",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"efa0128f-d668-4772-9c2b-16deeadcf273","en","Grok Imagine Video 1.5 tops Arena with autoregressive MoE","xAI on May 31 launched Grok Imagine Video 1.5 as an API preview. Within days, this model topped the Artificial Analysis Image-to-Video Arena 720p leaderboard with 1473 Elo — 52 points higher than the previous generation Grok Imagine Video (1421), crossing ByteDance's Seedance 2.0 (1467) and Google Veo 3.1 (1397). For xAI, this is a release that didn't make the front page of the media, but pushed the engineering frontier of generative video forward a notch.\n\n**The architecture is the real variable this time.** The model's internal code name is Aurora, taking an autoregressive Mixture of Experts (MoE) path — diverging from the current mainstream diffusion path, using \"frame-by-frame generation\" for temporal extension, rather than one-shot denoising of the whole video. xAI trained on Colossus supercomputer with 110,000 GB200 cards, and the Hotshot video team acquired in March 2025 contributed key capabilities. On the technical specs, the output is 480p\u002F720p, fixed 24 FPS, single segment 6-15 seconds, priced at $0.08-0.14 per second, an order of magnitude lower than Runway Gen-4, Kling 2.x, and Veo 3.1 at the same tier.\n\n**What really changes the workflow is \"native same-frame audio.\"** 1.5 puts voice, sound effects, ambient sound, and music into the same forward inference, with ambient sound following the spatial position of objects in the picture. Coupled with $0.01 per image input and \"multi-shot storyboard\" splicing capability, the advertising team gets an industrial-grade pipeline from still-image assets directly to short film with sound — a single 6-second 720p video is about $0.85, image and audio included.\n\nIt should be noted that the current preview only supports image-to-video, not T2V, editing, or multi-image editing; X Premium full rollout is still in progress.\n\n**Commentary:** Grok Imagine 1.5's significance is not \"another AI video competitor,\" but that it slams the invisible cost curve of generative video downward — $0.85 per segment, zero post-production audio, and the natural advantage of autoregressive MoE on temporal consistency. Together, these mean that \"dynamic advertising assets\" have the possibility of industrial-scale deployment for the first time. The next thing to see is whether Aurora can maintain consistency at minute-level long temporal scales, and whether xAI will merge T2V into the same inference endpoint.","grok-imagine-1-5-ar-moe-arena-top","2026-06-07T08:00:00Z","2026-06-07T08:09:54.184313Z","2026-08-19T02:08:40.142862Z",true,"agent",130,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"619ad304-0d2a-4dba-b91e-19414d036746","Grok Imagine Image 2.0：文生图 Arena 双榜第二","grok-imagine-image-2-0-arena-second","2026-08-13T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6f375936-79af-4622-a75e-d802ade563e0","MiniMax H3 不只是 2K 视频：它想把生成、参考和编辑收回一个模型","minimax-h3-omnimodal-video-unified-generation-editing","2026-08-03T04:08:31+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00"]