[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-automem-stanford-32b-long-horizon":3,"news-related-15987a0e-bc06-4f21-9608-264da02d0e6c":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"15987a0e-bc06-4f21-9608-264da02d0e6c","AutoMem 让 32B 开源模型在长程任务上追平 Claude Opus 4.5","斯坦福大学 Serena Yeung-Levy 团队在 arXiv 公开 AutoMem 框架（编号 2607.01224），把\"记忆管理\"从工程技巧变成可独立训练的认知技能。研究人员给 LLM 配上一个真实文件系统当作\"外部笔记本\"，把读、写、搜索、追加、建文件五种操作与游戏动作并列为同级的\"记忆动作\"，让模型自己决定什么时候该记、记什么、查什么。\\n\\nAutoMem 设计了两层自动化优化循环。第一层用一个强 LLM 充当元审阅者，读取完整的游戏轨迹，迭代改写提示词、文件结构和可用操作，把 NetHack 里因重复追加导致文件暴涨的问题改成坐标键值覆盖，把每步字符增长从 138 砍到 6；第二层从模型自身成功轨迹中筛选好的记忆操作片段，用 LoRA 单独训练一个\"记忆专家\"，让\"查了再记\"的习惯内化进参数。游戏策略模型的权重则全程不动，记忆能力干净叠加在原能力之上。\\n\\n在 Crafter、MiniHack、NetHack 三款长程程序生成游戏上，仅优化记忆一项就让基础代理得分提升 2-4 倍，32B 开源模型在所有三款游戏上反超参数量两倍于它的 Qwen2.5-72B-Instruct，并逼近 Claude Opus 4.5 与 Gemini 3.1 Pro Thinking 的水平。低效动作率下降 32-65%，重复写入率下降 68-83%。\\n\\n这项工作最值得关注的不是又一个 SOTA，而是一个方法论信号：在算力与参数规模竞赛之外，\"如何管理记忆\"可以被拆出来单独优化，且收益可能比堆参数更具杠杆。当下 Agent 系统的工程瓶颈往往不在模型本身，而在长程任务里的记忆与状态管理，AutoMem 提供了一个可复用的工程范式——把记忆操作当成一等公民，把记忆专家和策略模型解耦微调，把元 AI 引入自动化调优闭环。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.01224","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"3ff7f6e4-56ff-4769-96c4-abad70c1a021","en","AutoMem: a 32B open model catches Opus 4.5 on long tasks","Stanford's Serena Yeung-Levy team has open-sourced the AutoMem framework (paper 2607.01224), turning \"memory management\" from an engineering trick into an independently trainable cognitive skill. The researchers give an LLM a real file system as its \"external notebook\", and put the five operations — read, write, search, append, create file — on the same level as in-game actions, calling them \"memory actions\", letting the model decide for itself when to record, what to record, and what to look up. AutoMem designs two layers of automated optimization loops. The first layer uses a strong LLM as a meta-reviewer, reading the complete game trajectory, iteratively rewriting prompts, file structure, and available operations — turning NetHack's runaway file growth from repeated appends into coordinate-keyed overwrites, cutting the per-step character growth from 138 down to 6. The second layer filters good memory-operation fragments from the model's own successful trajectories, using LoRA to train a separate \"memory expert\", internalizing the habit of \"look up before recording\" into the parameters. The game strategy model's weights are left untouched throughout, so memory capability stacks cleanly on top of the original ability. On three long-horizon procedural-generation games — Crafter, MiniHack, NetHack — optimizing only the memory item lifts the base agent's score 2–4x; the 32B open-source model, on all three games, overtakes the Qwen2.5-72B-Instruct with twice the parameters, and approaches the level of Claude Opus 4.5 and Gemini 3.1 Pro Thinking. Inefficient-action rate drops 32–65%, redundant-write rate drops 68–83%. The most noteworthy thing about this work isn't another SOTA, but a methodological signal: outside the parameter-scale and compute arms race, \"how to manage memory\" can be split out and optimized on its own — and the leverage may be higher than stacking parameters. The current engineering bottleneck of Agent systems often isn't the model itself, but memory and state management in long-horizon tasks; AutoMem offers a reusable engineering paradigm — treat memory operations as a first-class citizen, decouple the memory expert and the policy model for fine-tuning, and bring meta-AI into the automated tuning closed loop.","automem-stanford-32b-long-horizon","2026-07-23T12:10:00Z","2026-07-23T12:05:19.727227Z","2026-08-19T02:08:40.142862Z",true,"agent",89,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"48e1c261-a40a-4c71-9cba-450a459e6ad3","4-bit 模型反超全精度:QAH 把量化从性能税变成第二次蒸馏","quantization-aware-healing-hypernova-60b","2026-08-25T17:20:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"c94bdf86-5de9-49fe-8c98-0f5c47611bfe","SGLang v0.5.18 发布:大模型冷启动提速 2.38 倍,710 个 PR 都改了什么","sglang-v0-5-18-cold-start-2-38x","2026-08-24T23:15:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"f65e204c-0115-4b50-9113-2c3bb2ff6637","ReCache:给 Agent 的工具记忆装上独立缓存,KV 内存砍 92%、首 token 提速 3.655 倍","recache-agent-kv-cache-reuse","2026-08-24T15:30:00+00:00"]