[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ar-video-generation-memory-survey":3,"topics-all":38,"news-related-33bd7c4c-27f8-458c-8404-265134fc6ce8":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"33bd7c4c-27f8-458c-8404-265134fc6ce8","视频生成缺的不是算力,是记忆:282 篇论文拼出一张全景地图","视频生成模型的画质感已经够好,真正瓶颈正在转移:序列一长,实体身份、动态状态等关键历史信息会被挤出上下文窗口,画面开始「失忆」。25 人团队的新综述把这类问题正式定义为记忆问题,按五个视角梳理 282 篇文献并整理出 14 个记忆基准,其中 181 篇都是 2026 年发表的,赛道爆发速度可见一斑。","视频生成这条赛道今年卷得厉害,但一份 9 月 23 日提交到 arXiv 的综述把行业的注意力拉回了一个更根本的问题:模型画得再好,记不住前面发生了什么,长视频照样崩。\n\n这份题为《The Past Frames the Future》的综述由 25 位作者联合完成,一作 Harold Haodong Chen 在 Hugging Face 论文页自述,这是**该方向的第一份系统性综述**。论文上线两天即在 Hugging Face Daily Papers 收获 35 个 upvote,配套 Awesome 清单仓库已达 54 stars——社区对这个问题的关注度可见一斑。\n\n## 为什么记忆是自回归视频生成的命门\n\n自回归(AR)视频生成靠因果 rollout 扩展视觉序列:每一步都基于已生成的内容预测下一段。问题在于,实际部署的模型必须在**有限的上下文窗口、存储和算力约束**下运行。综述指出,随着序列扩展,三类关键历史信息——实体身份、动态状态、干预引发的因果变化——往往**在相关性还没耗尽之前,就提前离开了活跃上下文**。\n\n结果就是观众熟悉的那种翻车:角色转个身回来换了张脸,被打碎的杯子自己复原,镜头切走再切回,房间布局已经变了。这些问题过去被归为「一致性」,综述把它们统一重新定义为**记忆问题**:跨外层 AR 步维持的持久历史信息,即使原始证据已不可局部访问,仍能影响未来生成。这个定义把散落在「长视频」「世界模型」「一致性生成」各处的工作拉进了同一个坐标系。\n\n## 四种记忆载体与一张分类地图\n\n综述按 Forms\u002FFunctions\u002FOperations\u002FLearning\u002FEvaluation 五个视角组织文献。从配套仓库看,记忆「载体」至少四大类:\n\n- **Visual Memory(视觉记忆)**:直接保留像素或 VAE 潜空间表征,如 WorldMem、DecMem、MemLearner\n- **Implicit State Memory(隐式状态记忆)**:藏在 attention cache、循环网络或状态空间模型的内部状态里\n- **Explicit State Memory(显式状态记忆)**:以实体为中心或空间几何形式显式建模\n- **Adaptive Parametric Memory(自适应参数记忆)**:直接改写模型参数本身\n\n这套分类不是学究式整理。配套仓库收录的 282 个条目中,**181 个来自 2026 年**——近三分之二的核心工作挤在最近九个月,载体路线之争才刚开始,地图此刻参考价值最大。\n\n## 基准缺位才是真短板\n\n综述最值得警惕的是对评测的判断:现有评测无法诊断「真正的记忆能力」。仓库整理了 14 个记忆专用基准,R2M-Bench、MBench、WBench、MemoBench 等,几乎全部诞生于 2026 年,且多数只覆盖单一维度。当各家模型都宣称「分钟级生成」时,拿什么证明画面里的角色还记得三分钟前的剧情?没有标准化评测,这类宣传基本只能靠自证。\n\n## 所以呢\n\n对做视频生成的团队,这份综述和配套仓库值得进收藏夹:选路线先看四大载体分类,别在已被证明走不通的分支上重复投入。对观察者,更有趣的信号是——LLM 圈吵了两年的「上下文窗口 vs 外挂记忆」之争,正在视频生成领域原样重演,节奏还快得多。值得盯的下一个问题:这 14 个基准里,会不会跑出视频圈的「MMLU 时刻」?\n\n参考:论文 [arXiv:2609.28466](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28466) | 仓库 [Awesome-AR-Video-Memory](https:\u002F\u002Fgithub.com\u002FHaroldChen19\u002FAwesome-AR-Video-Memory)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28466","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"e8e43bd7-bf51-4237-b98b-a07e509bfbb5","en","Video Generation's Missing Piece Is Memory, Not Pixels","A 25-author survey reframes AR video consistency failures as memory problems: 282 papers mapped, 14 benchmarks cataloged, 181 entries from 2026 alone.","Video generation has been a crowded race this year, but a survey posted to arXiv on September 23 pulls the industry's attention back to a more fundamental question: a model can render beautifully, yet if it cannot remember what happened moments ago, long videos still fall apart.\n\nTitled \"The Past Frames the Future: Memory for Autoregressive Video Generation,\" the survey lists 25 authors. First author Harold Haodong Chen states on the Hugging Face paper page that this is **the first systematic survey on memory mechanisms for AR video generation**. Within two days the paper collected 35 upvotes on Hugging Face Daily Papers, and the companion awesome-list repository has already reached 54 stars — a clear signal of how much the community wants this problem named.\n\n## Why Memory Is the Choke Point of AR Video Generation\n\nAutoregressive (AR) video generation extends visual sequences through causal rollouts: each step predicts the next segment based on what has been generated. The catch is that deployed models must operate under **strictly bounded context windows, storage, and compute**. As the survey puts it, three kinds of critical historical information — entity identities, dynamic states, and intervention-induced causal changes — often **leave the active context long before their relevance diminishes**.\n\nThe failure mode is familiar to anyone who has watched generated videos: a character turns around and comes back with a different face; a shattered glass un-shatters itself; the camera cuts away and returns to a room whose layout has changed. These issues used to be filed under \"consistency.\" The survey reframes them all as a **memory problem**: persistent historical information maintained across outer AR steps that can still influence future generation even after the originating evidence is no longer locally accessible. That reframing is itself a contribution — it pulls scattered work on \"long video,\" \"world models,\" and \"consistent generation\" into one coordinate system.\n\n## Four Memory Carriers and a Taxonomy Map\n\nThe survey organizes the literature along five perspectives: Forms, Functions, Operations, Learning, and Evaluation. Judging from the companion repository's structure, the \"carriers\" of memory fall into at least four families:\n\n- **Visual Memory**: retaining pixels or VAE latent representations directly, as in WorldMem, DecMem, and MemLearner\n- **Implicit State Memory**: hiding inside attention caches, recurrent networks, or state-space model states\n- **Explicit State Memory**: explicitly modeling entity-centric or spatial-geometric states\n- **Adaptive Parametric Memory**: rewriting the model's parameters themselves\n\nThis is not an academic exercise in tidiness. Of the 282 entries in the companion repository, **181 are from 2026** — nearly two-thirds of the field's core work appeared in the past nine months alone. The competition between carrier routes has barely started, which is exactly when a map is most valuable.\n\n## The Real Shortfall: Benchmarks\n\nThe most sobering part of the survey is its verdict on evaluation: existing benchmarks cannot diagnose \"genuine memory capability.\" The repository catalogs 14 memory-oriented benchmarks — R2M-Bench, MBench, WBench, MemoBench, and more — nearly all born in 2026, most covering only a single dimension. When vendors claim \"minute-scale generation,\" what proves the character on screen still remembers the plot from three minutes ago? Without standardized evaluation, such claims rest mostly on self-certification.\n\n## So What\n\nFor teams building video generation, this survey and its companion repo ([github.com\u002FHaroldChen19\u002FAwesome-AR-Video-Memory](https:\u002F\u002Fgithub.com\u002FHaroldChen19\u002FAwesome-AR-Video-Memory)) belong in your bookmarks: check the four-carrier taxonomy before picking a route, and avoid re-investing in branches already shown to be dead ends. For observers, the more interesting signal is this — the \"context window vs. external memory\" debate that LLM circles argued over for two years is now replaying in video generation, at a much faster tempo. The question to watch: will one of these 14 benchmarks produce video's \"MMLU moment,\" when memory capability becomes comparable for the first time?\n\nReferences: paper [arXiv:2609.28466](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28466) | repo [Awesome-AR-Video-Memory](https:\u002F\u002Fgithub.com\u002FHaroldChen19\u002FAwesome-AR-Video-Memory)","ar-video-generation-memory-survey","2026-09-24T21:09:28Z","2026-09-24T21:09:38.278557Z","2026-09-24T21:09:38.278565Z",true,"agent",398,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"11bb60a3-aedb-4395-bb18-aac0f9cbd7f0","快手开源 Keye-VL-2.0：首个把 DSA 稀疏注意力适配到 GQA 多模态的 30B 模型","keye-vl-2-0-30b-dsa-gqa-multimodal","2026-06-26T04:12:17+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"af21d26d-d9cd-4526-ac4a-66366a45848c","AV-GRPO:8张A800给22B音视频模型做RL后训练","av-grpo-audio-video-diffusion-rl","2026-09-27T17:08:13+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"6aca231d-200d-4227-9fa0-f1e6d149b0b0","WanPE:397B 提示词模型上岗,视频生成多了个导演","wanpe-397b-video-prompt-enhancement","2026-09-26T23:07:22+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"ba0ed7bf-3de3-4f92-98fe-a50d6ac274d0","WROP 开源:用 150 个物体恒存任务给世界模型补认知课","wrop-object-permanence-world-models","2026-09-25T17:08:02+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"d78ea8a8-bebb-4b94-b340-127eb2874a73","Ovis 全模态嵌入 3B:综合分领先 5.19","ovis-omni-embedding-3b","2026-09-23T21:08:45+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"9210c86b-3e1d-434b-bbfd-78e62c698aed","WorldCrafter 开源:给视频世界模型装上可查询的 3D 记忆,转一圈回来还是那个房间","worldcrafter-video-world-model-3d-memory","2026-09-22T19:08:48+00:00"]