[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-nvidia-longlive-rag-drift-fix-video":3,"news-related-8701e0ec-1e95-41bd-ad69-fc9b8d68f6d2":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8701e0ec-1e95-41bd-ad69-fc9b8d68f6d2","NVIDIA LongLive-RAG：用检索增强打破长视频生成的「漂移难题」","自回归（AR）视频扩散模型是当下长视频生成的主流路线，但痛点众所周知：随着生成时间拉长，画面中的人物与物体逐渐「变脸」、细节开始崩塌——业内称之为**身份漂移**（identity drift）。6 月 1 日，NVIDIA Lab（NVlabs）在 arXiv 发布 **LongLive-RAG**，把 RAG 思想首次系统地搬进了长视频生成。\n\n**为什么滑动窗口不够用？** 现有方法普遍采用滑动窗口注意力以控制显存，但这种机制存在不可逆的轨迹偏差：当前窗口一旦积累外观错误，后续生成只能基于这个「受损」轨迹继续向前，越走越偏。\n\n**LongLive-RAG 的核心解法：把已生成的潜变量当作可检索记忆。** 每个新 block 通过 query embedding 检索最相关的历史 latent 参与条件计算，让生成器能「回头看」非局部上下文，而不是只盯着最近几帧。\n\n**配套的 Window Temporal Delta Loss** 抑制了检索器对冗余局部相似的偏好，鼓励 embedding 捕捉有意义的时间变化——这避免了「检索器只挑到刚生成的那一帧」的退化。\n\n**开销极低**：每 block 检索仅增加 4.08 ms，总检索开销 490 ms。实验在多个 AR 主干上验证，长视频质量与 **VBench-Long** 排名均为同类方法最佳；它也是首个把「自生成潜在表征」建模为「内容可寻址检索记忆」的开放式长视频生成方法。\n\n**评论：** 把 RAG 从语言模型迁移到视频生成并非简单类比——视频的时空连续性使得「检索什么」成为关键设计点。NVIDIA 的方案对生态非常友好：不重新训练基础扩散模型，只在外层加检索机制，对长视频生成生态是低成本、向后兼容的改进。当 Sora、Kling、Wan 等主流框架都在卷更长、更稳时，这类「外挂式」方法可能会被快速吸收进工业级管线。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.02553","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"8dac812d-3839-4abe-a855-5f56ec9515fd","nvidia",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"bbe8bd8d-18f2-40ad-b5e5-bfe466f601ee","en","LongLive-RAG: retrieval breaks long-video drift","Autoregressive (AR) video diffusion models are the mainstream path for long video generation today, but the pain point is well known: as generation time stretches, the characters and objects in the frame gradually \"change faces,\" and details start to collapse — the industry calls this **identity drift**. On June 1, NVIDIA Lab (NVlabs) released **LongLive-RAG** on arXiv, bringing the RAG idea systematically into long video generation for the first time.\n\n**Why isn't sliding window enough?** Existing methods generally adopt sliding window attention to control VRAM, but this mechanism has an irreversible trajectory bias: once the current window accumulates appearance errors, subsequent generation can only continue forward based on this \"damaged\" trajectory, drifting further and further.\n\n**LongLive-RAG's core solution: treat already-generated latents as retrievable memory.** Each new block retrieves the most relevant historical latents via query embedding to participate in conditional computation, letting the generator \"look back\" at non-local context rather than just staring at the most recent few frames.\n\n**Supporting Window Temporal Delta Loss** suppresses the retriever's preference for redundant local similarity, encouraging the embedding to capture meaningful temporal change — this avoids the degradation of \"retriever only picks up the just-generated frame.\"\n\n**Overhead is tiny:** retrieval per block adds only 4.08ms, total retrieval overhead 490ms. Experiments validated on multiple AR backbones, with long-video quality and **VBench-Long** ranking both best in class; it is also the first open method to model \"self-generated latent representation\" as \"content-addressable retrieval memory.\"\n\n**Commentary:** Migrating RAG from language models to video generation is not a simple analogy — the spatio-temporal continuity of video makes \"what to retrieve\" a key design point. NVIDIA's solution is very ecosystem-friendly: it does not retrain the base diffusion model, but only adds the retrieval mechanism on the outside layer, a low-cost, backward-compatible improvement for the long-video generation ecosystem. When Sora, Kling, Wan, and other mainstream frameworks are all rolling for longer and more stable, this kind of \"plug-in\" method may be quickly absorbed into industrial-grade pipelines.","nvidia-longlive-rag-drift-fix-video","2026-06-07T04:30:00Z","2026-06-07T04:23:48.330182Z","2026-08-19T02:08:40.142862Z",true,"agent",107,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"a818c807-2131-4950-8f51-62847a57db41","VideoRAE 把 frozen 视频基础模型改造成生成器 latent:UCF-101 gFVD 40\u002F93,收敛提速 5×","videorae-frozen-video-generator","2026-07-20T04:15:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"75d2f385-00ef-4074-80bf-ee47ba05a4a4","NVIDIA × HF：Diffusers 微调上 H100，Wan 2.2\u002FFLUX.2 打通","nvidia-huggingface-diffusers-h100","2026-07-17T14:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"b446777f-9949-47e4-8ca8-2d0fa7f14126","NVIDIA Cosmos 3 Edge 4B开源世界模型：物理AI实时推理从云端搬到产线","nvidia-cosmos-3-edge-4b","2026-07-16T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"9652a175-6f95-461d-b9d4-946b44d0fccf","NVIDIA × Hugging Face：GR00T 1.7 接入 LeRobot","nvidia-hf-isaac-gr00t-cosmos","2026-07-07T06:00:26+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"7ac0ef83-f46d-44f9-846b-a2051fc81e87","NVIDIA NeMo AutoModel：MoE 微调吞吐抬到 3.4–3.7 倍","nvidia-nemo-automodel-moe-finetune-3-7x","2026-06-24T20:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"56cb62a1-da4f-4ac6-94ee-e60346f8d075","英伟达 BioNeMo Agent Toolkit：生命科学库塞进 AI Agent","nvidia-bionemo-agent-toolkit-life-science","2026-06-24T00:00:00+00:00"]