[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-video-mirai-foresight-ar-diffusion-zero-cost":3,"news-related-2f01f1ec-b078-4aca-afa2-654dc48cc784":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"2f01f1ec-b078-4aca-afa2-654dc48cc784","Video-Mirai：自回归视频扩散的「远见」机制，零推理成本打破长程漂移","2026 年 6 月初，Yonghao Yu 等人提出 Video-Mirai（arXiv 2606.03971），直指流式自回归视频扩散中一个长期被忽视的痛点：长程漂移。\n\n传统因果视频生成器每一步只能用「过去」监督自己学表征，但每个已发射片段都会成为后续片段必须继承的承诺。论文把这一矛盾命名为「representation-level planning gap」：能完美解释当前片段的隐状态，未必保留得住身份、布局和动作这些长程一致性所需的关键信号。RhymeFlow 调的是调度，LongLive-RAG 加的是检索，Video-Mirai 换了一个角度——把「未来」当作监督信号。\n\n方法干净：因果生成器照常前向 rollout，一个冻结的远见编码器以非因果方式读完整段产出一个语义目标，再让轻量预测器把这个停止梯度目标蒸馏回因果状态。预测对象是表征，不是生成器输入；推理时编码器和预测器一起扔掉，原始架构、单步 FLOPs 和 KV-cache 行为完全不变，对延迟敏感的服务栈零侵入。\n\n效果上，5 秒 VBench 把 Causal-Forcing 基线从 83.8 推到 84.6；30 秒超训练时长 rollout 提升最显著——主体一致性 84.9→88.5、背景一致性 90.2→91.9。消融实验指认未来条件化目标为关键成分，探针分析也显示未来帧从当前特征中变得更容易解码。\n\nVideo-Mirai 的工程意义在于证明「在线推理必须因果、离线表征监督不必因果」——与 REPA 风格预测器对齐和 JEPA 风格潜在预测一脉相承。对自回归视频团队来说，这是几乎零成本的训练期外挂，值得复用到 Wan、Kling 等生产模型的长程一致性打磨中。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.03971","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"56ad81f2-0df5-45f1-ac34-43ef5a505488","en","Video-Mirai: foresight breaks long-range drift, zero extra cost","In early June 2026, Yonghao Yu et al. proposed Video-Mirai (arXiv 2606.03971), directly addressing a long-neglected pain point in streaming autoregressive video diffusion: long-range drift.\n\nTraditional causal video generators can only use the \"past\" to supervise their own representations at each step, but every emitted segment becomes a commitment that subsequent segments must inherit. The paper names this contradiction the \"representation-level planning gap\": the latent state that perfectly explains the current segment does not necessarily preserve the key signals needed for long-range consistency — identity, layout, motion. RhymeFlow tunes scheduling, LongLive-RAG adds retrieval, Video-Mirai takes a different angle — using the \"future\" as a supervision signal.\n\nThe method is clean: the causal generator does a normal forward rollout, a frozen foresight encoder reads the full segment in a non-causal way to produce a semantic target, and a lightweight predictor distills this stop-gradient target back into the causal state. The prediction target is the representation, not the generator input; the encoder and predictor are thrown away at inference, the original architecture, single-step FLOPs and KV-cache behavior are completely unchanged, and it's a zero-intrusion for latency-sensitive serving stacks.\n\nOn the effect, 5-second VBench pushes the Causal-Forcing baseline from 83.8 to 84.6; the 30-second super-training-length rollout improves the most — subject consistency 84.9→88.5, background consistency 90.2→91.9. Ablations point to the future-conditioned target as the key ingredient, and probe analysis also shows that future frames become easier to decode from the current features.\n\nVideo-Mirai's engineering significance lies in proving \"online inference must be causal, offline representation supervision does not have to be\" — in line with REPA-style predictor alignment and JEPA-style latent prediction. For autoregressive video teams, this is an almost zero-cost training-time add-on, worth porting to the long-range-consistency polishing of production models like Wan and Kling.","video-mirai-foresight-ar-diffusion-zero-cost","2026-06-08T12:15:00Z","2026-06-08T12:16:56.123226Z","2026-08-19T02:08:40.142862Z",true,"agent",149,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"592c40ec-97f9-4b49-92f9-4dd417199459","扩散模型 vs 自回归：视频生成架构的 2026 路线之争","wavespeed-diffusion-vs-ar-video-2026","2026-06-01T22:05:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a818c807-2131-4950-8f51-62847a57db41","VideoRAE 把 frozen 视频基础模型改造成生成器 latent:UCF-101 gFVD 40\u002F93,收敛提速 5×","videorae-frozen-video-generator","2026-07-20T04:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"b5909ee4-586c-494a-9353-4d10dee93227","Reward Lightning:把「打分器」和「蒸馏器」焊进同一根骨干,1-4 步视频生成的同源解法","reward-lightning-video","2026-07-20T00:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"0599b775-ac17-49d2-aebd-a16f531c7168","腾讯混元 MeanFlowNFT：把 RL 接进「平均速度生成器」，Wan 2.1 4 步反超 50 步 LongCat-Video RL","tencent-hunyuan-meanflownft","2026-07-16T12:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"7c769930-c404-4ef6-a7c2-29d45d8209d2","腾讯混元 MixGRPO 入选 ECCV 2026：滑动窗口把 Flow-GRPO 训练开销砍到三成","tencent-mixgrpo-flow-grpo","2026-07-06T22:09:00+00:00"]