[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-speculate-with-memory-llm-agent-acceleration":3,"topics-all":36,"news-related-ed39ed38-b5fa-4f58-92cf-d05233ab998b":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"ed39ed38-b5fa-4f58-92cf-d05233ab998b","Speculate with Memory：LLM Agent 无损加速 2.5×，准确率涨 39pp","\"7月14日,arXiv:2607.12236「Speculate with Memory」把记忆系统搬进 LLM Agent 投机执行器,给\\\"无状态推测\\\"补上在线学习能力。核心是在投机器上加三层串联的在线记忆:对比转换表记录历史动作-动作统计分布,情景记忆回溯与当前上下文相似的过往轨迹片段,困惑跟踪器专门压制反复出现的错误。三者协作,让投机器第一次具备了\\\"走过一遍,下次更准\\\"的能力。\\n\\n实验在 6 个基准上覆盖动作预测、观察预测、链路预测三类场景。结果:动作预测准确率相对提升 19-39%;动作重复度高的观察预测任务上,最高取得 2.5× 绝对加速。论文特别强调所有增益\\\"无损\\\"——投机完全跑在环境空闲时段,actor 轨迹与无投机执行完全一致,零额外 wall-clock 开销;且增益随记忆增长持续累积,在不同成本档位的推测器之间都泛化。\\n\\n真正的价值在于把\\\"专项加速\\\"和\\\"持续学习\\\"焊到同一条管道。当前 LLM Agent 的工具调用、环境观察、动作规划彼此耦合,延迟叠加非常夸张;大多数加速方案只关注参数或蒸馏。本文思路相当于把\\\"老司机的肌肉记忆\\\"嵌入推理回路——同一份推理预算,体验越来越好。这种\\\"边际成本递减\\\"的加速路线,比单纯放大模型更可持续,也更贴近 GPU 利用率的真实瓶颈。\\n\\n边界也明确:动作空间越开放,记忆命中率反而下降,而 Agent 越来越走向开放场景。但\\\"轻量记忆 + 投机解码\\\"组合,给所有做 Agent serving 的团队提供了低成本、可借鉴的工程模板,结合 vLLM、SGLang 现成的投机解码接口,很快就能复刻。\\n\"","https:\u002F\u002Farxiv.org\u002Fpdf\u002F2607.12236","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"63551b20-cbb9-4fcc-ab20-7a3650cf8fa5","en","Speculate with Memory: lossless 2.5x speedup, +39pp accuracy","July 14, arXiv:2607.12236 \"Speculate with Memory\" moves the memory system into the LLM Agent's speculative executor, giving \"stateless speculation\" online-learning capability. The core is adding three layers of serial online memory on top of the speculator: a contrastive transition table records historical action-action statistics, episodic memory retrieves past trajectory fragments similar to the current context, and a confusion tracker specifically suppresses repeatedly occurring errors. The three collaborate, giving the speculator, for the first time, the ability to \"go through it once, be more accurate next time\". The experiments cover three categories — action prediction, observation prediction, link prediction — on 6 benchmarks. Results: action-prediction accuracy improves 19–39% relative; on observation-prediction tasks with high action repetition, up to 2.5× absolute speedup is achieved. The paper especially emphasizes that all gains are \"lossless\" — speculation runs entirely in environment idle time, the actor trajectory is identical to non-speculative execution, with zero extra wall-clock overhead; and the gains continue to accumulate as memory grows, generalizing across speculators at different cost tiers. The real value lies in welding \"specialized acceleration\" and \"continual learning\" into the same pipeline. Current LLM Agent tool calls, environment observations, and action planning are tightly coupled, and latency stacks are very pronounced; most acceleration methods only focus on parameters or distillation. This paper's approach is essentially embedding \"the veteran driver's muscle memory\" into the inference loop — the same inference budget, getting better and better with use. This \"diminishing marginal cost\" acceleration path is more sustainable than simply scaling the model, and is closer to the real bottleneck of GPU utilization. The boundary is also clear: the more open the action space, the lower the memory hit rate — but Agents are increasingly moving toward open scenarios. The \"lightweight memory + speculative decoding\" combination offers all Agent-serving teams a low-cost, replicable engineering template; combined with vLLM and SGLang's existing speculative-decoding interfaces, it can be replicated quickly.","speculate-with-memory-llm-agent-acceleration","2026-07-15T08:15:00Z","2026-07-15T08:18:21.865304Z","2026-08-19T02:08:40.142862Z",true,"agent",262,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"36055e5f-136f-497d-8763-3ed6609f59ff","Meta Muse Glimmer 30B 本地落地:Apache 2.0 的开源智能体,把 Agent 装进 24GB 显存","meta-muse-glimmer-30b-local-agent-apache2-r2","2026-08-19T03:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"6b203495-fcab-4afe-baa7-1079cf993796","拆开 GLM-5.3 的「后训练工厂」:基座一字未动,靠环境合成与 1e-7 对齐撑起全部提升","glm-5-3-post-training-stack-deep-dive","2026-08-17T13:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"7d371b09-9792-465d-b73a-3d0af4735129","InferenceBench：15 个前沿 Agent 自主做 LLM 推理优化","inferencebench-open-ended-llm-optimization","2026-08-16T12:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"259d91b2-ed6b-4af8-8f2a-f759b84cc617","蚂蚁 Ling-3.0 Flash：124B\u002F5.1B MoE 的 Agent 生产级模型","inclusionai-ling-3-flash-hybrid-linear-moe-agent","2026-08-14T08:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"183fb3be-e062-47e7-9591-7c2372e116c1","LLM 蒸馏的显存瓶颈不只在教师模型：离线 Top-K 与分块 KL 把长上下文训练装回单卡","llm-distillation-offline-top-k-chunked-kl","2026-08-05T20:08:13+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"beec1ff3-22af-4657-b58a-90cb0797c3b1","PyroDash 让小模型「借力」大模型推理：把 LLM 调用砍到 1.9%，成本从 $49 降到 $1.78","pyrodash-small-large-routing","2026-07-24T00:00:00+00:00"]