[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-duomem-4b-on-device-agent":3,"news-related-d016b7c3-4fbe-4b30-871a-f0d503f62228":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"d016b7c3-4fbe-4b30-871a-f0d503f62228","DuoMem 用「双空间蒸馏」把 4B 端侧 Agent 拉到 72B 教师水位:ALFWorld 任务率 4.3% → 77.9%","大模型 Agent 在长程交互里能跑多深,关键看记忆;但 72B 教师整套搬到端侧既不现实也不经济。arXiv 2606.29961 上的 DuoMem 给了一个相当工程化的解法——不靠堆更大模型,而是把「程序性记忆」拆成两个空间一起压给学生。\n\n具体是 dual-space 蒸馏:context 空间里,把教师预先生成的程序性记忆直接前置到学生输入,相当于给学生一份带答案的 cheat sheet;parameter 空间里,再让学生在教师成功轨迹上微调轻量 LoRA——可训练参数不到 10M,只增加几 MB 教师记忆。\n\n效果相当能说明问题。在具身决策基准 ALFWorld 上,4B 学生模型的任务成功率从 4.3% 飙到 77.9%,基本追平 72B 教师的 87.1%;wall-clock 比 72B 教师快 3 倍以上,真正具备实时端侧部署的可行性。作者跑了 8 个模型(2B–72B)的消融,证实两个空间互相补足,缺一不可。\n\n对手机、车机、机器人等端侧 Agent 来说,这条路比单纯堆参数更具现实意义——它本质上把「教师的流程性知识」拆成可外挂的注释 + 可微调的肌肉记忆一起交付。短板也明显:程序性记忆需要教师提前在相似任务上「踩点」生成,如果场景发散快,记忆维护成本会迅速膨胀;而且 4B 模型的极限边界仍取决于底层指令遵循能力,DuoMem 解决的是「知识怎么搬」,不是「能力怎么补」。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.29961","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3244242e-20b2-4c1d-9ff0-3af5ea84a3eb","en","DuoMem lifts a 4B agent to 72B teacher level (4.3% to 77.9%)","How deep large-model Agents can run in long-horizon interaction depends on memory; but moving the entire 72B teacher to the device side is neither practical nor economical. DuoMem on arXiv 2606.29961 gives a fairly engineering-flavored solution — not relying on a larger model, but splitting \"procedural memory\" into two spaces to compress to the student. Specifically dual-space distillation: in the context space, the procedurally generated memory from the teacher is directly prepended to the student's input, equivalent to giving the student a cheat sheet with answers; in the parameter space, the student then fine-tunes a lightweight LoRA on the teacher's successful trajectories — trainable parameters under 10M, with only a few MB of teacher memory added. The effect is quite illustrative. On the embodied-decision benchmark ALFWorld, the 4B student model's task success rate jumps from 4.3% to 77.9%, basically catching up to the 72B teacher's 87.1%; wall-clock is 3×+ faster than the 72B teacher, truly with the feasibility of real-time on-device deployment. The authors ran an ablation on 8 models (2B–72B), confirming that the two spaces complement each other, neither is dispensable. For mobile, in-car, robotics, and other on-device Agents, this path is more realistic than simply stacking parameters — it essentially splits \"teacher's procedural knowledge\" into attachable annotations + trainable muscle memory for delivery. The short board is also clear: procedural memory requires the teacher to pre-\"scout\" on similar tasks to generate; if scenarios diverge quickly, memory maintenance costs will balloon; and the 4B model's ultimate boundary still depends on the underlying instruction-following capability — DuoMem solves \"how to move knowledge\", not \"how to fill capability\".","duomem-4b-on-device-agent","2026-07-05T18:00:00Z","2026-07-05T18:08:57.285559Z","2026-08-19T02:08:40.142862Z",true,"agent",87,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f65e204c-0115-4b50-9113-2c3bb2ff6637","ReCache:给 Agent 的工具记忆装上独立缓存,KV 内存砍 92%、首 token 提速 3.655 倍","recache-agent-kv-cache-reuse","2026-08-24T15:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"822987d9-faaa-492d-b7ba-1c0427f72f9b","AgentOps 出海第一步：拆腾讯云 ADP 4.0 海外版的「Agent+Workflow 分账」架构","tencent-adp-4-0-overseas-agentops","2026-07-20T02:11:09+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"15987a0e-bc06-4f21-9608-264da02d0e6c","AutoMem 让 32B 开源模型在长程任务上追平 Claude Opus 4.5","automem-stanford-32b-long-horizon","2026-07-23T12:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6784d3bd-26c1-4fd5-a2e5-8c7b9e591dae","SmoothAgent 把上下文变换「提前做」：Agent 长链路 TTFT 砍到原来的 1\u002F12","smoothagent-ttft-12x","2026-07-23T03:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"c6dc2edc-1a4a-46b9-85e6-f1c8ef32faa6","DeepSeek DSpark 跑进 Apple Silicon：mlx-dspark 给出首个原生 MLX 移植,逐字节保持原模型输出","mlx-dspark-apple-silicon","2026-07-04T12:00:00+00:00"]