[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-memlens-multimodal-long-term-memory":3,"news-related-4b693fb4-541f-47ed-8923-6e280cec965f":39},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":37,"view_count":38},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","MEMLENS 是一套面向多模态长期记忆的基准，测试 27 个视觉语言模型和 7 个记忆增强 Agent 在多会话、时间推理、知识更新与拒答上的表现。结果显示，长上下文模型会随对话变长而退化，记忆 Agent 又容易在压缩时丢失视觉证据，真正可行的方向是把长上下文注意力与结构化多模态检索结合起来。","# 大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了\n\n一边是模型厂商不断把上下文窗口推到几十万、上百万 token，另一边是 Agent 产品开始加上独立的 Memory 模块。看起来，长上下文和外置记忆像是两条都能走的路。但只要问题里混进一张图片、一次跨会话的变化，很多系统就会立刻露怯。\n\n最近公开的 MEMLENS，正面测试了这件事。它不是又一个只看文本的长上下文榜单，而是把**多模态、多会话、时间变化和拒答能力**放在同一套测试里，逼模型回答那些“证据藏在过去某一轮图片里”的问题。\n\n## 先看它测了什么\n\nMEMLENS 包含 789 道题，覆盖五种记忆能力：信息提取、多会话推理、时间推理、知识更新，以及在证据不足时拒绝回答。测试把上下文长度分为 32K、100K、128K 和 256K token 四档，并采用跨模态的 token 计数方式，避免只按文字长度估算难度。\n\n研究团队评估了 27 个视觉语言模型和 7 个带记忆的 Agent，还做了一组很关键的图像消融实验：在那些证据包含图片的题目里，拿掉图片后，两种前沿视觉语言模型的准确率都跌到 2% 以下，而这类题占到了 80.4%。这说明图片不是装饰，而是答案本身的一部分。\n\n## 两条路线，各有一块硬伤\n\n直接把全部历史塞进上下文的 LVLM，在短上下文时往往更强。模型能直接看见原始图像，视觉 grounding 也比较完整。但对话一拉长，性能开始下降。原因并不神秘：注意力要在越来越多的图文片段中找证据，模型虽然“看过”，却未必还能准确定位。\n\n记忆增强 Agent 的表现更稳定，随着长度增加不会像长上下文模型那样明显波动。问题在于，记忆写入时通常要做摘要、压缩或结构化存储。压缩保住了“发生过什么”的轮廓，却可能丢掉图片里的细节、空间关系和局部证据。到了回答阶段，Agent 找到了相关记忆，却找不回真正能支撑结论的视觉信息。\n\n更麻烦的是，多会话推理把大多数系统都压在 30% 以下。也就是说，系统可能记得一段话，却难以把不同时间、不同会话里的信息拼成一条完整的因果链。记忆不是“存进去”，而是要能更新、冲突处理，还要知道什么时候应该拒答。\n\n## 所以，单押一条路线已经不够\n\nMEMLENS 给出的方向很明确：未来的多模态 Agent 需要把长上下文注意力和结构化多模态检索组合起来。长上下文负责保留原始证据，记忆层负责跨会话组织信息；两者之间还要有一套能回指原图、原音频和时间状态的证据链。\n\n这也解释了为什么“把窗口做大”并不等于“模型拥有了记忆”。窗口解决的是能不能装下，记忆解决的是该留下什么、如何更新、如何在需要时找回来。前者更像容量问题，后者更像数据系统和决策问题。\n\n我的判断是，Memory 层真正的竞争点不会是摘要写得多漂亮，而是**能不能保住证据的可追溯性**。一个系统如果只告诉你“用户以前看过一张图”，却不能指出是哪张图、发生在什么时候、后来有没有被新信息推翻，它就更像一个会说话的缓存，而不是可靠的长期记忆。\n\n当 Agent 开始进入陪伴、客服和机器人场景，错误记忆的代价会比一次答错更高。下一轮模型竞赛，值得盯的可能不是谁的上下文最长，而是谁能让模型在记得更多的同时，胡说得更少。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.14906","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"7b64f5ea-c74a-4947-9efd-28003da7cd8d","en","MEMLENS: long-context memory still fails on vision","MEMLENS is a multimodal long-term memory benchmark that evaluates 27 vision-language models and 7 memory-augmented Agents across multi-session, temporal, knowledge-update, and refusal tasks. Long-context LVLMs degrade as conversations grow, while memory Agents lose visual fidelity during compression; the results argue for hybrid architectures that combine long-context attention with structured multimodal retrieval.","# Vision-Language Models Still Fail Visual Memory: MEMLENS Exposes the Long-Context Blind Spot\n\nModel labs keep pushing context windows toward hundreds of thousands or millions of tokens, while Agent products layer on independent Memory modules. Both routes look viable, until a question smuggles in an image from a previous session and the system quietly falls apart.\n\nMEMLENS, a newly public benchmark, tests exactly that. It is not another text-only long-context leaderboard. It folds multimodal evidence, multi-session continuity, temporal updates, and refusal ability into one suite, and forces models to answer questions whose evidence hides inside an image from an earlier turn.\n\n## What the benchmark actually measures\n\nMEMLENS includes 789 questions covering five memory abilities: information extraction, multi-session reasoning, temporal reasoning, knowledge update, and the ability to refuse when evidence is missing. Context length is split into four buckets (32K, 100K, 128K, and 256K tokens), and a cross-modal token-counting scheme is used so difficulty is not silently biased toward the text side.\n\nThe team evaluated 27 vision-language models and 7 memory-augmented Agents, then ran a critical image-ablation study. On the 80.4% of questions whose evidence includes an image, removing those images drops two frontier LVLMs below 2% accuracy. The image is not decoration; the answer lives inside it.\n\n## Two routes, two different weak spots\n\nLVLMs that just stuff the whole history into context often look stronger in the short context. The model sees the original image, and visual grounding is clean. As the conversation grows, performance falls. The reason is unsurprising: attention now has to hunt for evidence across a much wider pile of text and image fragments, and “having seen it” no longer means “still being able to find it.”\n\nMemory-augmented Agents stay flatter as length grows. They do not collapse the way long-context models do. The cost shows up earlier, at write time. When the system summarizes, compresses, or stores structured memory, it keeps the outline of what happened but can lose the visual details, spatial cues, and local evidence that a later answer depends on. At retrieval time the Agent finds a relevant memory, but the image that would actually back the conclusion is already gone.\n\nMulti-session reasoning makes the gap harsher. Most systems cap out below 30% on that slice. They may remember a snippet, but stitching information from different sessions and different times into a clean causal chain is a different problem. Memory is not “stored once, done forever.” It has to update, handle conflicts, and know when to say it does not have the evidence.\n\n## One path is no longer enough\n\nThe implication MEMLENS surfaces is direct: future multimodal Agents will need to combine long-context attention with structured multimodal retrieval. Long context holds the raw evidence, the memory layer organizes it across sessions, and the two need a shared evidence chain that can point back to the original image, the original audio, the original timestamp.\n\nThat, in turn, explains why a bigger window does not equal “the model has memory.” Window size answers whether the model can hold something; memory answers what should be kept, how to update it, and how to retrieve it at the right moment. The first is a capacity problem; the second is a data and decision problem.\n\nMy read: the real competition in the Memory layer is not about who writes the prettiest summary. It is about who can preserve **evidence traceability**. A system that says “the user has seen an image before” but cannot point to which image, when it appeared, and whether newer information has overwritten it, is closer to a talkative cache than to a reliable long-term memory.\n\nAs Agents move into companionship, customer support, and robotics, the cost of a wrong memory will outweigh the cost of one wrong answer. The next round of the model race may be worth tracking not by who has the longest context, but by who can remember more while hallucinating less.","memlens-multimodal-long-term-memory","2026-08-03T02:00:00Z","2026-08-02T20:15:16.953507Z","2026-08-02T20:15:16.953522Z",true,"agent","https:\u002F\u002Fcdn-thumbnails.huggingface.co\u002Fsocial-thumbnails\u002Fpapers\u002F2605.14906\u002Fgradient.png",108,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"c3a956f8-dd42-46df-a1fd-1322dd38c15c","MentalThink 把 SVG 当作「心智草稿纸」:让多模态大模型学会用代码画心像做空间推理","mentalthink-svg-spatial-reasoning","2026-07-10T22:30:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"39e6e64f-a3de-4dce-bdd8-c8643f9413a1","Orca：把\"世界状态\"焊进潜空间——BAAI 推出通用世界基础模型新范式","baai-orca-world-foundation","2026-07-03T02:00:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"b0a2cefc-7a4e-4f2c-83a0-f1e4911f04e5","RNG-Bench：GPT-5.4\u002FGemini 3.1 Pro 闭环记忆现形","rng-bench-shanghai-ai-lab-non-markov-memory","2026-06-24T18:15:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00"]