[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-speakermem-r1-multi-party-memory":3,"topics-all":38,"news-related-e98cf9a2-8348-4142-9a56-c11166774798":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","浙江大学 SpeakerMem-R1 用双轨记忆攻多方对话长期记忆:逐字轨道保留说话人标签,结构轨道维护个人与群体状态。Qwen2.5-3B 写入器经 RL 把 305 题准确率从 57.38% 提到 68.20%,EverMemBench 榜 62.33% 领先 EverOS,登 HF 日榜第一。","多方对话里的长期记忆,正在成为 LLM 记忆系统的照妖镜。浙江大学团队的最新论文点出一个尴尬现实:通用 LLM 记忆系统在多方对话基准上经常丢失人物与群体关系,某些场景下甚至跑不过最朴素的 BM25 检索——问题不在\"找不到相关内容\",而在分不清\"谁说了什么\"、这句话说的是谁、群体共享了哪些信息,以及状态如何随时间变化。这两个瓶颈,论文分别称为消息归因与状态重建。\n\n## 双轨记忆:逐字原文与结构化状态互补\n\nSpeakerMem-R1 的核心设计是把记忆拆成两条轨道。System 1 保留全部原始消息的逐字内容,每条都带说话人和时间标签,依赖原始措辞的问题随时能回到原话取证;System 2 由一个写入器把对话提炼成四层结构化状态——个人核心记忆、个人画像、群体交互、群体洞察,每条记录维护归属、来源、层级、时间戳和溯源链接,写入动作只有 ADD、UPDATE、NOOP 三种,更新是非破坏式的。查询时,两条轨道的证据再按实体、事件、时间组合起来交给回答模型。\n\n## 3B 小写入器,RL 也能练出来\n\n结构化记忆的难点在写入容易出错:归因错人、状态改错。团队没有依赖大模型写入,而是用 SpeakerLevenshtein 奖励加说话人条件化的 GRPO,训练出 Qwen2.5-3B 的 Writer-R1,让记忆写入组件可以整体本地部署。在 305 道留出题的受控评测里,RL 把 SFT 写入器的平均准确率从 57.38% 拉到 68.20%;同一设置下外部 LLM 写入器是 71.48%,小模型加专用奖励已经吃掉大部分差距。训练数据也不大:15 个完整社交网络、73 个写入片段、452 条监督动作。\n\n## 跑分领先,但离\"好用\"还有距离\n\n三个多方对话基准上,SpeakerMem-R1 比评测中最强的记忆\u002F检索基线分别高出 3.3、12.4、9.4 个百分点;在 EverMind-AI 公开报告的 EverMemBench 榜单上拿到 62.33%,领先 EverOS(60.08%)与 RippleMem(54.75%),论文称这是该榜已报告的最新框架中的最好成绩。在作为两人对话边界测试的 LoCoMo 全部 1,986 题上,它拿到 70.85%。但底色要看清:GroupMemBench 绝对准确率只有 47.9%,连一半都不到——多方归因对当前所有系统仍是硬骨头。\n\n## 为什么值得盯\n\n这篇论文的价值不在跑分,而在把\"说话人归因\"从检索问题升级成记忆结构问题:逐字轨道保真、结构轨道保关系,查询时再做证据组合。对做 Agent 记忆、客服机器人、会议助手的团队,这套双轨设计加\"3B 写入器 + RL\"的配方是可以直接复用的——代码以 MIT 协议开源,仓库包含推理包、基准运行器和写入器训练配方,论文也拿下 HuggingFace Daily Papers 9 月 24 日日榜第一(76 upvotes)。当所有人都在卷上下文长度时,记住\"谁在什么时候对谁说了什么\",可能才是长期记忆真正的瓶颈。\n\n论文: https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.26780\n代码: https:\u002F\u002Fgithub.com\u002F2022hpsk\u002FSpeakerMemR1","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.26780","b0060acd-5f64-427d-9673-5caea2940874",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"9e07bdec-2af7-4ca3-ac03-30170accddb2","en","SpeakerMem-R1: Dual-Track Memory Tells Who Said What","Zhejiang's SpeakerMem-R1: dual-track memory for multi-party dialogue. RL-trained 3B writer 57.38%→68.20%; 62.33% on EverMemBench. #1 HF Daily Papers.","Long-term memory in multi-party dialogue is turning into a litmus test for LLM memory systems. A new paper from Zhejiang University points out an awkward reality: general-purpose LLM memory systems tend to lose person and group relations on multi-party dialogue benchmarks, and in some settings they even underperform plain BM25 retrieval. The problem is not finding relevant content — it is telling who said what, whom each statement concerns, what the group shares, and how states change over time. The paper names these two bottlenecks message attribution and state reconstruction.\n\n## Dual-track memory: verbatim evidence plus structured state\n\nSpeakerMem-R1 splits memory into two complementary tracks. System 1 stores every message verbatim with speaker and time metadata, so questions that depend on exact wording can always return to the original utterance. System 2 has a writer distill conversations into four structured layers — per-speaker core memory, per-speaker profile, group interaction, and group insight — where each entry keeps owner, source, layer, timestamp, and provenance links, with only ADD, UPDATE, and NOOP actions and non-destructive updates. At query time, evidence from both tracks is combined by entity, event, and time before being handed to the answerer.\n\n## A 3B writer, trained with RL\n\nThe hard part of structured memory is that writing is error-prone: misattributed speakers, corrupted states. Instead of relying on a large model, the team trained Writer-R1 on Qwen2.5-3B using SpeakerLevenshtein rewards and speaker-conditioned GRPO, so the memory-writing component can be deployed locally. In a controlled evaluation of 305 held-out questions, RL lifts the SFT writer's mean accuracy from 57.38% to 68.20%; an external LLM writer reaches 71.48% under the same setting — a small model with a dedicated reward already closes most of the gap. The training data is modest too: 15 complete social networks, 73 writer segments, and 452 supervised actions.\n\n## Leading scores, but far from \"usable\"\n\nAcross three multi-party benchmarks, SpeakerMem-R1 beats the strongest evaluated memory\u002Fretrieval baselines by 3.3, 12.4, and 9.4 percentage points. On EverMind-AI's publicly reported EverMemBench leaderboard it reaches 62.33%, ahead of EverOS (60.08%) and RippleMem (54.75%) — which the paper calls the best reported result among the latest state-of-the-art frameworks. On all 1,986 LoCoMo questions, used as a two-person boundary test, it scores 70.85%. But keep the baseline in view: absolute accuracy on GroupMemBench is only 47.9%, below half — multi-party attribution remains a hard problem for every current system.\n\n## Why it deserves attention\n\nThe value of this paper is not the scores but the reframing: speaker attribution is a memory-structure problem, not a retrieval problem. The verbatim track preserves fidelity, the structured track preserves relations, and query-time composition merges both. For teams building agent memory, customer-service bots, or meeting assistants, the dual-track design plus the \"3B writer + RL\" recipe is directly reusable — the code is open-sourced under MIT with an inference package, benchmark runners, and writer-training recipes, and the paper took #1 Paper of the Day on HuggingFace Daily Papers for Sep 24 with 76 upvotes. While everyone races on context length, remembering who said what to whom and when may be the real bottleneck of long-term memory.\n\nPaper: https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.26780\nCode: https:\u002F\u002Fgithub.com\u002F2022hpsk\u002FSpeakerMemR1","speakermem-r1-multi-party-memory","2026-09-24T19:05:00Z","2026-09-24T19:09:43.332156Z","2026-09-24T19:09:43.332165Z",true,"agent",644,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"089195fb-7fe5-4ba9-a4bc-8e356fe5e923","BAAI把1000个GitHub仓库蒸馏成5000个技能,科研agent奖牌率31%冲到73%","baai-disco-repo-to-skill-library","2026-09-03T17:07:35+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"90af7ff5-b985-42d5-97c7-63a9579b7527","VitaBench 2.0：给 LLM Agent 出「长期用户建模」考卷，SOTA 也不及格","vitabench-2-0-long-term-user-modeling","2026-06-25T14:01:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"73e29aee-b368-423f-be27-7653f65b4775","DeepSeek DSec 公开:300 万沙盒日撑 V4.1 训练","deepseek-dsec-v4-1-sandbox-rl-training","2026-09-28T00:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"44740b4d-8c2c-44fc-8fff-fd89f3fb54ed","12 万美元 token 把 Copilot 运行时从 TypeScript 搬到 Rust","github-copilot-rust-migration-stephen-toub","2026-09-27T11:00:00+00:00"]