[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wantofight-real-time-game":3,"news-related-8058b748-25c6-4863-917e-46e363773d07":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8058b748-25c6-4863-917e-46e363773d07","WanToFight 把视频扩散压成 30FPS 实时游戏引擎:多玩家格斗首跑通","视频扩散模型迈过了\"实时可玩\"这条硬门槛。arXiv 2607.12592 上挂出的 WanToFight,第一次把生成式游戏引擎的四件难事——多玩家控制、实时推理、复杂物理交互、对抗博弈——同时塞进一个系统。它的底座是阿里开源的 Wan-1.3B 视频扩散 Transformer,作者团队(Li Hu、Guangyuan Wang、Peng Zhang、Bang Zhang)在上面搭了三层增量架构,把\"按帧画画\"变成\"按键出招\"。\n\n第一层是流式自回归生成器,核心是 block-causal attention + rolling KV cache,把全局去噪切成因果块,显存占用恒定,长序列不再爆。第二层是 Player Association 模块,键盘信号先经过视觉 grounding 绑定到角色身份,再通过 gated、locally causal 的注入模块回到去噪网络,避免控制信号污染共享表征;训练上采用单人到全玩法的渐进课程,先学单人再学对抗。第三层是工程化的蒸馏栈:四步 DMD 蒸馏的学生模型 + 剪枝 VAE 解码器,把端到端时延压到 RTX 5090 单卡 512×384 @ 30FPS,能撑完一整场 KOF'97。\n\n把视角放高一点,WanToFight 的真正意义不在\"模型又多强\",而在它示范了一种工程范式:1.3B 规模 + 消费级单卡 + 蒸馏+KV cache+剪枝 VAE 三件套,视频扩散模型从此可以走出\"按秒计费\"的云端流水线,进入实时交互循环。这与 GameNGen、DIAMOND、Oasis 等单玩家\u002F第一人称路径形成了明显分野——多玩家对抗场景的视觉一致性和因果保持难度更高,WanToFight 第一次给出了实证答案。\n\n当然,512×384 分辨率、依赖单卡、格斗动作的离散空间相对简化,都意味着离\"真·3A 实时生成\"还有相当距离。但方向已经清楚:下一个阶段的视频扩散,不会再比谁的参数更多,而会比谁先把生成折进 30FPS 的交互回路。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.12592","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"782958d8-1d0c-4a32-a9c1-3e5296e370db","en","WanToFight turns video diffusion into a 30FPS game engine","Video diffusion models have crossed the hard threshold of \"real-time playability\". WanToFight, posted to arXiv 2607.12592, is the first to simultaneously pack the four hard challenges of a generative game engine — multi-player control, real-time inference, complex physical interaction, adversarial play — into a single system. Its base is Alibaba's open-source Wan-1.3B video diffusion Transformer; the authors (Li Hu, Guangyuan Wang, Peng Zhang, Bang Zhang) build a three-layer incremental architecture on top, turning \"painting by frame\" into \"input-triggered move\". The first layer is a streaming autoregressive generator, with block-causal attention + rolling KV cache as the core, slicing global denoising into causal blocks with constant memory usage — long sequences no longer blow up. The second layer is the Player Association module: keyboard signals are first visually grounded to character identity, then routed back into the denoising network through a gated, locally causal injection module, avoiding the control signal contaminating the shared representation; training uses a progressive curriculum from single-player to all-gameplay, learning single-player first then adversarial play. The third layer is an engineering distillation stack: a four-step DMD-distilled student model + a pruned VAE decoder, compressing end-to-end latency to 512×384@30FPS on a single RTX 5090 — enough to last a full round of KOF'97. Zooming out, the real meaning of WanToFight is not \"another stronger model\", but the engineering paradigm it demonstrates: 1.3B scale + consumer single-card + the three-piece set of distillation + KV cache + pruned VAE, video diffusion models can now step out of the \"bill by the second\" cloud pipeline and into a real-time interactive loop. This forms a clear fork from the single-player \u002F first-person paths of GameNGen, DIAMOND, and Oasis — multi-player adversarial scenarios have higher visual-consistency and causal-preservation difficulty, and WanToFight is the first to provide an empirical answer. Of course, 512×384 resolution, single-card dependency, and the relatively simplified discrete action space of fighting moves all mean there's still considerable distance to \"true 3A real-time generation\". But the direction is clear: the next phase of video diffusion won't be about who has more parameters, but who first folds generation into a 30FPS interactive loop.","wantofight-real-time-game","2026-07-15T16:05:00Z","2026-07-15T16:18:01.486665Z","2026-08-19T02:08:40.142862Z",true,"agent",133,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"172d28ff-776b-443d-8f05-dddc9d544dec","Mercury 2 把“推理扩散 LLM”塞进搜索流水线：每步 1000+ tokens\u002F秒，Voice Agent 延迟预算被改写","mercury-2-diffusion-search-agent-realtime","2026-08-17T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"9f443e5f-4ca1-4dfa-b8c7-a5d7ba7aaf6e","SANA-Video 2.0：用混合线性注意力把视频生成推到单卡可用","sana-video-2-mixed-linear-attention","2026-07-24T04:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"58764a46-a635-4f79-86c6-eefff0dea639","Flex-Forcing：NVIDIA 用 chunk 机制把 AR 与双向视频扩散塞进同一模型","nvidia-flex-forcing-ar-bidir-video","2026-07-08T18:24:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"dcb9cd19-bec7-40d8-9ed2-8951d13a3e06","Mercury 2：首个推理扩散 LLM 跑出 1009 tokens\u002F秒，重写实时 Agent 成本曲线","mercury-2-inception-1009-tok-s-reasoning","2026-06-18T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b950b487-2b1f-4ece-ad6e-d57cf94f1f84","稀疏注意力新突破：「上下文混合」让长视频生成成本降至近线性","moc-context-mixing-near-linear-video","2026-06-01T01:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"94894abf-62aa-41a9-8e3c-e999ff274d60","Sparse Forcing：稀疏注意力让视频生成质量速度双提升","meta-ucsb-sparse-forcing-pbsa-video-1-27x","2026-05-07T08:10:00+00:00"]