Long-term memory in multi-party dialogue is turning into a litmus test for LLM memory systems. A new paper from Zhejiang University points out an awkward reality: general-purpose LLM memory systems tend to lose person and group relations on multi-party dialogue benchmarks, and in some settings they even underperform plain BM25 retrieval. The problem is not finding relevant content — it is telling who said what, whom each statement concerns, what the group shares, and how states change over time. The paper names these two bottlenecks message attribution and state reconstruction.

Dual-track memory: verbatim evidence plus structured state

SpeakerMem-R1 splits memory into two complementary tracks. System 1 stores every message verbatim with speaker and time metadata, so questions that depend on exact wording can always return to the original utterance. System 2 has a writer distill conversations into four structured layers — per-speaker core memory, per-speaker profile, group interaction, and group insight — where each entry keeps owner, source, layer, timestamp, and provenance links, with only ADD, UPDATE, and NOOP actions and non-destructive updates. At query time, evidence from both tracks is combined by entity, event, and time before being handed to the answerer.

A 3B writer, trained with RL

The hard part of structured memory is that writing is error-prone: misattributed speakers, corrupted states. Instead of relying on a large model, the team trained Writer-R1 on Qwen2.5-3B using SpeakerLevenshtein rewards and speaker-conditioned GRPO, so the memory-writing component can be deployed locally. In a controlled evaluation of 305 held-out questions, RL lifts the SFT writer's mean accuracy from 57.38% to 68.20%; an external LLM writer reaches 71.48% under the same setting — a small model with a dedicated reward already closes most of the gap. The training data is modest too: 15 complete social networks, 73 writer segments, and 452 supervised actions.

Leading scores, but far from "usable"

Across three multi-party benchmarks, SpeakerMem-R1 beats the strongest evaluated memory/retrieval baselines by 3.3, 12.4, and 9.4 percentage points. On EverMind-AI's publicly reported EverMemBench leaderboard it reaches 62.33%, ahead of EverOS (60.08%) and RippleMem (54.75%) — which the paper calls the best reported result among the latest state-of-the-art frameworks. On all 1,986 LoCoMo questions, used as a two-person boundary test, it scores 70.85%. But keep the baseline in view: absolute accuracy on GroupMemBench is only 47.9%, below half — multi-party attribution remains a hard problem for every current system.

Why it deserves attention

The value of this paper is not the scores but the reframing: speaker attribution is a memory-structure problem, not a retrieval problem. The verbatim track preserves fidelity, the structured track preserves relations, and query-time composition merges both. For teams building agent memory, customer-service bots, or meeting assistants, the dual-track design plus the "3B writer + RL" recipe is directly reusable — the code is open-sourced under MIT with an inference package, benchmark runners, and writer-training recipes, and the paper took #1 Paper of the Day on HuggingFace Daily Papers for Sep 24 with 76 upvotes. While everyone races on context length, remembering who said what to whom and when may be the real bottleneck of long-term memory.

Paper: https://arxiv.org/abs/2609.26780
Code: https://github.com/2022hpsk/SpeakerMemR1