[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-lct-long-context-tuning-video-scene-consistency":3,"news-related-0bc5be19-abbb-4898-8e1d-86cb127fbb8a":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0bc5be19-abbb-4898-8e1d-86cb127fbb8a","长上下文调优重塑视频生成：单镜头模型学会「讲连续的故事」","当前视频生成模型已能合成逼真的单镜头视频，但真实叙事需要多镜头场景并保持一致性。arXiv 近期发表的论文提出 Long Context Tuning（LCT）方法，为这一难题提供新训练范式。\n\n## 核心思路\n\nLCT 将预训练单-shot 视频扩散模型的上下文窗口扩展，让模型直接从数据学习场景级一致性，而非依赖后处理拼凑。技术层面，LCT 将全注意力机制从单镜头扩展到场景内所有镜头，配合交织式 3D 位置编码；同时引入异步噪声策略，支持联合生成和自回归生成，且无需额外参数。\n\n具有双向注意力的模型在 LCT 后可进一步微调为上下文因果注意力模式，通过 KV-Cache 实现高效自回归推理——视频可以一段一段续写，而非一次性全部渲染。\n\n## 实践意义\n\nLCT 带来的直接变化是「组合生成」和「交互式镜头扩展」能力：模型不仅理解「这是连续故事」，还能根据用户输入动态延展下一个镜头。这为 AI 视频从「展示片段」走向「讲述故事」提供了技术基础。\n\n## 写在最后\n\n视频生成正从「能看」走向「能讲」。LCT 的价值在于不依赖更大模型或算力，而是通过改进训练范式让现有模型「学会连贯思考」。这种效率导向的技术路径，或许才是视频生成真正进入内容生产流水线的正确方式。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2503.10589","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d3607242-6d19-47dd-9c3d-b8272625539e","en","Long-context tuning teaches video models continuous stories","Current video generation models can synthesize realistic single-shot videos, but real narrative needs multi-shot scenes with consistent character\u002Fsetting continuity. A recent arXiv paper proposes the Long Context Tuning (LCT) method, providing a new training paradigm for this challenge.\n\n## Core idea\n\nLCT extends the context window of pre-trained single-shot video diffusion models, letting the model learn scene-level consistency directly from data, rather than relying on post-hoc stitching. Technically, LCT extends the full-attention mechanism from a single shot to all shots in the scene, paired with interleaved 3D position encoding; meanwhile, an asynchronous noise strategy is introduced, supporting joint generation and autoregressive generation, with no additional parameters needed.\n\nModels with bidirectional attention can be further fine-tuned into a context-causal attention mode via LCT, achieving efficient autoregressive inference through KV-Cache — videos can be extended segment by segment, rather than all rendered at once.\n\n## Practical significance\n\nLCT brings direct \"compositional generation\" and \"interactive shot extension\" capabilities: the model not only understands \"this is a continuous story,\" but can also dynamically extend the next shot based on user input. This provides a technical foundation for AI video moving from \"displaying clips\" to \"telling stories.\"\n\n## Final word\n\nVideo generation is moving from \"can see\" to \"can tell.\" LCT's value lies in not relying on larger models or more compute, but on improving the training paradigm to let existing models \"learn to think coherently.\" This efficiency-oriented technical path may be the right way for video generation to truly enter the content production pipeline.","lct-long-context-tuning-video-scene-consistency","2026-05-03T10:05:00Z","2026-05-03T10:07:33.034107Z","2026-08-19T02:08:40.142862Z",true,"agent",109,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"9f443e5f-4ca1-4dfa-b8c7-a5d7ba7aaf6e","SANA-Video 2.0：用混合线性注意力把视频生成推到单卡可用","sana-video-2-mixed-linear-attention","2026-07-24T04:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"8058b748-25c6-4863-917e-46e363773d07","WanToFight 把视频扩散压成 30FPS 实时游戏引擎:多玩家格斗首跑通","wantofight-real-time-game","2026-07-15T16:05:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"a07a31d6-3a63-4452-b898-02a6e575682e","Vera：Netflix 把视频编辑拆成编辑层 + 原视频","vera-netflix-caltech-mixture-transformers-edit","2026-06-24T04:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b950b487-2b1f-4ece-ad6e-d57cf94f1f84","稀疏注意力新突破：「上下文混合」让长视频生成成本降至近线性","moc-context-mixing-near-linear-video","2026-06-01T01:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"94894abf-62aa-41a9-8e3c-e999ff274d60","Sparse Forcing：稀疏注意力让视频生成质量速度双提升","meta-ucsb-sparse-forcing-pbsa-video-1-27x","2026-05-07T08:10:00+00:00"]