[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mragent-nus-memory-jigsaw-27x":3,"news-related-5b909019-b85f-4ec2-9d7a-9b8808db49e1":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"5b909019-b85f-4ec2-9d7a-9b8808db49e1","MRAgent：NUS 把 LLM Agent 记忆从「查字典」改成「拼拼图」，单查询 token 直降 27 倍","当 LLM Agent 开始跑长程任务，\"记不住、用不上、查不准\"成了新的基础设施瓶颈。新加坡国立大学 Shuo Ji 等人在 arXiv（2606.06036）发布 MRAgent 框架，被 ICML 2026 接收。核心思路是：把记忆从「先取回、再推理」的静态流水线，改成「边推理、边挖」。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.06036","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"58a5377d-c7f2-4825-a628-f5fae932061f","en","MRAgent: agent memory as jigsaws, 27x fewer tokens per query","When LLM Agents start running long-horizon tasks, \"can't remember, can't use, can't find accurately\" becomes a new infrastructure bottleneck. Shuo Ji et al. at the National University of Singapore released the MRAgent framework on arXiv (2606.06036), accepted at ICML 2026. The core idea: shift memory from a static \"retrieve first, then reason\" pipeline to a \"reason while digging\" model.\n\nThe framework treats memory operations as first-class actions alongside game actions: reading, writing, searching, appending, and creating files. A meta-reviewer (a strong LLM) reads full game trajectories and iteratively rewrites prompts, file structure, and available operations; in parallel, a memory expert is fine-tuned separately (LoRA) from the agent's own successful trajectories, so the \"look up before writing\" habit is internalized into parameters. The game policy model's weights stay frozen — memory is layered cleanly on top of existing capability.\n\nOn three long-horizon procedurally generated games — Crafter, MiniHack, NetHack — optimizing memory alone raises base-agent scores 2-4×, and a 32B open-source model surpasses Qwen2.5-72B-Instruct (twice the parameter count) on all three games, approaching Claude Opus 4.5 and Gemini 3.1 Pro Thinking. Inefficient-action rate drops 32-65%, repeat-write rate drops 68-83%.\n\nThe most worth-noting thing is not another SOTA, but a methodological signal: outside the parameter-and-compute race, \"how to manage memory\" can be decoupled and optimized independently, with potentially more leverage than stacking parameters. The current engineering bottleneck of Agent systems often isn't the model itself but memory and state management in long-horizon tasks. AutoMem provides a reusable engineering paradigm — treat memory operations as first-class citizens, decouple fine-tuning the memory expert from the policy model, and bring meta-AI into the auto-tuning loop.","mragent-nus-memory-jigsaw-27x","2026-06-28T10:09:00Z","2026-06-28T10:10:03.123698Z","2026-08-19T02:08:40.142862Z",true,"agent",126,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"1316635c-88e1-41b6-a45c-df8ef217cf3f","PaperPilot 把文献搜索改写成「工作流归纳」：可编辑 DAG 把多轮检索错误率干到 0%","paperpilot-workflow-induction","2026-07-01T08:21:23+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"6de305c2-91b2-47fb-b1e4-dfb5f1e711c8","WorldEvolver：把世界模型装进 LLM Agent 的「即时记忆」","worldevolver-llm-agent-world-model","2026-06-30T18:04:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"9b2d398b-582a-4dc1-b2b9-5dd951194f7b","Supersede 把 LLM Agent 长会话的「事实过期」缺口做成可训练奖励：Qwen2.5-3B 上 GRPO 让准确率近翻倍","supersede-fact-staleness-rl","2026-06-29T22:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"b88a7d4b-8f6d-4440-b384-4283f88a410c","Tool-Use RL 为什么会突然崩盘？arXiv 2606.26027 戳破 Agent 训练的'概率尖峰'陷阱","tool-use-rl-collapse-probability-spike","2026-06-25T20:25:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f96a7b02-7bad-4ef2-98c4-a6b6aedefd0c","Constraint Tax：Tool Calling 遇 JSON Schema 悄悄失灵","constraint-tax-tool-calling-silent-disable","2026-06-25T14:00:00+00:00"]