[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-speculate-with-memory-2-5x":3,"news-related-ed39ed38-b5fa-4f58-92cf-d05233ab998b":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"ed39ed38-b5fa-4f58-92cf-d05233ab998b","Speculate with Memory：LLM Agent 无损加速 2.5×，准确率涨 39pp","\"7月14日,arXiv:2607.12236「Speculate with Memory」把记忆系统搬进 LLM Agent 投机执行器,给\\\"无状态推测\\\"补上在线学习能力。核心是在投机器上加三层串联的在线记忆:对比转换表记录历史动作-动作统计分布,情景记忆回溯与当前上下文相似的过往轨迹片段,困惑跟踪器专门压制反复出现的错误。三者协作,让投机器第一次具备了\\\"走过一遍,下次更准\\\"的能力。\\n\\n实验在 6 个基准上覆盖动作预测、观察预测、链路预测三类场景。结果:动作预测准确率相对提升 19-39%;动作重复度高的观察预测任务上,最高取得 2.5× 绝对加速。论文特别强调所有增益\\\"无损\\\"——投机完全跑在环境空闲时段,actor 轨迹与无投机执行完全一致,零额外 wall-clock 开销;且增益随记忆增长持续累积,在不同成本档位的推测器之间都泛化。\\n\\n真正的价值在于把\\\"专项加速\\\"和\\\"持续学习\\\"焊到同一条管道。当前 LLM Agent 的工具调用、环境观察、动作规划彼此耦合,延迟叠加非常夸张;大多数加速方案只关注参数或蒸馏。本文思路相当于把\\\"老司机的肌肉记忆\\\"嵌入推理回路——同一份推理预算,体验越来越好。这种\\\"边际成本递减\\\"的加速路线,比单纯放大模型更可持续,也更贴近 GPU 利用率的真实瓶颈。\\n\\n边界也明确:动作空间越开放,记忆命中率反而下降,而 Agent 越来越走向开放场景。但\\\"轻量记忆 + 投机解码\\\"组合,给所有做 Agent serving 的团队提供了低成本、可借鉴的工程模板,结合 vLLM、SGLang 现成的投机解码接口,很快就能复刻。\\n\"","https:\u002F\u002Farxiv.org\u002Fpdf\u002F2607.12236","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"63551b20-cbb9-4fcc-ab20-7a3650cf8fa5","en","Speculate with Memory: lossless 2.5x speedup, +39pp accuracy","July 14, arXiv:2607.12236 \"Speculate with Memory\" moves the memory system into the LLM Agent's speculative executor, giving \"stateless speculation\" online-learning capability. The core is adding three layers of serial online memory on top of the speculator: a contrastive transition table records historical action-action statistics, episodic memory retrieves past trajectory fragments similar to the current context, and a confusion tracker specifically suppresses repeatedly occurring errors. The three collaborate, giving the speculator, for the first time, the ability to \"go through it once, be more accurate next time\". The experiments cover three categories — action prediction, observation prediction, link prediction — on 6 benchmarks. Results: action-prediction accuracy improves 19–39% relative; on observation-prediction tasks with high action repetition, up to 2.5× absolute speedup is achieved. The paper especially emphasizes that all gains are \"lossless\" — speculation runs entirely in environment idle time, the actor trajectory is identical to non-speculative execution, with zero extra wall-clock overhead; and the gains continue to accumulate as memory grows, generalizing across speculators at different cost tiers. The real value lies in welding \"specialized acceleration\" and \"continual learning\" into the same pipeline. Current LLM Agent tool calls, environment observations, and action planning are tightly coupled, and latency stacks are very pronounced; most acceleration methods only focus on parameters or distillation. This paper's approach is essentially embedding \"the veteran driver's muscle memory\" into the inference loop — the same inference budget, getting better and better with use. This \"diminishing marginal cost\" acceleration path is more sustainable than simply scaling the model, and is closer to the real bottleneck of GPU utilization. The boundary is also clear: the more open the action space, the lower the memory hit rate — but Agents are increasingly moving toward open scenarios. The \"lightweight memory + speculative decoding\" combination offers all Agent-serving teams a low-cost, replicable engineering template; combined with vLLM and SGLang's existing speculative-decoding interfaces, it can be replicated quickly.","speculate-with-memory-2-5x","2026-07-15T08:15:00Z","2026-07-15T08:18:21.865304Z","2026-08-19T02:08:40.142862Z",true,"agent",152,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"6b203495-fcab-4afe-baa7-1079cf993796","拆开 GLM-5.3 的「后训练工厂」:基座一字未动,靠环境合成与 1e-7 对齐撑起全部提升","glm-5-3-post-training-stack-deep-dive","2026-08-17T13:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7d371b09-9792-465d-b73a-3d0af4735129","InferenceBench：15 个前沿 Agent 自主做 LLM 推理优化","inferencebench-open-ended-llm-optimization","2026-08-16T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"259d91b2-ed6b-4af8-8f2a-f759b84cc617","蚂蚁 Ling-3.0 Flash：124B\u002F5.1B MoE 的 Agent 生产级模型","inclusionai-ling-3-flash-hybrid-linear-moe-agent","2026-08-14T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"beec1ff3-22af-4657-b58a-90cb0797c3b1","PyroDash 让小模型「借力」大模型推理：把 LLM 调用砍到 1.9%，成本从 $49 降到 $1.78","pyrodash-small-large-routing","2026-07-24T00:00:00+00:00"]