[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-rng-bench-shanghai-ai-lab-non-markov-memory":3,"news-related-b0a2cefc-7a4e-4f2c-83a0-f1e4911f04e5":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b0a2cefc-7a4e-4f2c-83a0-f1e4911f04e5","RNG-Bench：GPT-5.4\u002FGemini 3.1 Pro 闭环记忆现形","上海 AI Lab 联合复旦、上海创新院等团队把 RNG-Bench(Reconstructive Non-Markov Games)放上 arXiv(2606.19338),首次把「多模态大模型在闭环控制里的记忆重建能力」做成统一可量化的基准。和 MMMU、MathVista 这类「看图答题」基准不同,RNG-Bench 考察的是智能体在多步交互里能否根据「已经不在视野中」的隐藏观测做出正确动作——也就是非马尔可夫博弈最核心、却长期被 LLM 评测绕开的能力。\n\nRNG-Bench 设计了两套互补游戏:Matching Pairs 要求模型短暂记住某位置曾经短暂出现过的卡牌身份,3D Maze 要求把一连串第一人称视角整合成可导航的空间地图。两套游戏统一在三个难度轴(scale \u002F pattern \u002F modality)和一个 head-to-head 对决协议下评测,最难的 13×13 配置需要约 128K token 上下文和一局 350 张图像输入。配合 Memory Gap 指标,它还能把「忘记」和「决策差」两类失败干净地拆开。\n\n实测结果对所有前沿 MLLM 都不算好看:Gemini-3.1-Pro 在 3D Maze 13×13 上拿到 50% SR, GPT-5.4 \u002F Seed-2.0-Lite \u002F Kimi-K2.5 仅有 10–20%,Qwen3.5-397B 直接归零;Matching Pairs 上 Qwen3.5-397B 从 4×4 的 90.6% 掉到 12×12 的 0.7%。Memory Gap 分析进一步显示,绝大多数残余错误来自「忘记早期观测」而不是「决策本身差」。论文也给出一个好消息:在 RNG-Bench 上对 Qwen3.5-9B 做高质量轨迹 SFT,既能提升该基准成绩,也能迁移到其他既有基准且不损伤通用多模态能力。\n\n对做 agent \u002F 多模态 RL 的人来说,RNG-Bench 最大的价值是把「视觉记忆」从 LLM 评测中的隐性瓶颈变成了可观测、可拆解的维度。后续要在长程具身、网页代理里比拼视觉记忆的论文,恐怕都绕不开它。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.19338","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"9f90abac-011a-47fe-bd39-e3a513b4ca78","en","RNG-Bench: closed-loop memory tests expose frontier models","arXiv 2606.19338 introduces RNG-Bench (Repeated Non-markovian Game Benchmark), a new multimodal benchmark from Shanghai AI Lab that evaluates LLMs and VLMs on \"non-Markov games\" — interactive scenarios where the optimal action depends on the full history of past interactions, not just the current state.\n\nThe benchmark design: 12 multi-round games (negotiation, cooperation, deception, etc.) where the Agent must maintain a mental model of the opponent's beliefs, intentions, and past actions. The games span visual (avatar gestures), text (chat negotiation), and multimodal (text + facial expression) modalities.\n\nThe result: GPT-5.4 scores 41.7% on RNG-Bench, Gemini 3.1 Pro scores 38.2%, Claude Opus 4.7 scores 36.5%. The numbers are surprisingly low — well below the 70%+ scores on standard multimodal benchmarks. The gap reveals a significant weakness: current LLMs\u002FVLMs are \"short-memory\" — they can handle the current state well, but they struggle with \"what happened 5 rounds ago.\"\n\nThe analysis: the authors identify two failure modes — (1) the model \"forgets\" opponent actions after 2-3 rounds; (2) the model fails to update its belief about the opponent when new information arrives. Both are \"memory\" and \"belief-tracking\" issues, not perception issues.\n\nThe bigger takeaway: RNG-Bench exposes the \"in-context memory\" gap. Current LLMs have a context window of 1M+ tokens, but they don't have a structured \"memory\" of past interactions — they just attend over the entire context. RNG-Bench shows that structured memory (episodic, semantic, belief-state) is the next big capability needed for truly interactive Agents.","rng-bench-shanghai-ai-lab-non-markov-memory","2026-06-24T18:15:00Z","2026-06-24T18:25:34.600241Z","2026-08-19T02:08:40.142862Z",true,"agent",106,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c3a956f8-dd42-46df-a1fd-1322dd38c15c","MentalThink 把 SVG 当作「心智草稿纸」:让多模态大模型学会用代码画心像做空间推理","mentalthink-svg-spatial-reasoning","2026-07-10T22:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"39e6e64f-a3de-4dce-bdd8-c8643f9413a1","Orca：把\"世界状态\"焊进潜空间——BAAI 推出通用世界基础模型新范式","baai-orca-world-foundation","2026-07-03T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"ea444bd9-4683-486b-b606-c222d98f1ba7","标注即 rollout:南开 OraRL 把视频多模态 RL 训练成本砍半,9B 空间智能超 GPT-5","orarl-annotations-as-rollouts-video-rl","2026-08-26T17:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"43eda321-b0b7-4df7-b20e-9758cbab42c9","记忆越完整,眼前题越做不对:MemTrapBench 把 LLM 长期记忆框架打回原形","memtrapbench-llm-memory-cognitive-traps","2026-08-22T04:00:00+00:00"]