[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-vitabench-2-0-long-term-user-modeling":3,"news-related-90af7ff5-b985-42d5-97c7-63a9579b7527":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"90af7ff5-b985-42d5-97c7-63a9579b7527","VitaBench 2.0：给 LLM Agent 出「长期用户建模」考卷，SOTA 也不及格","\"美团 LongCat 团队 6 月 25 日开源 VitaBench 2.0，定位为「首个真实生活场景下面向长期动态用户建模的智能体评测基准」。如果说 1.0 还在测\\\"一次外卖订单能不能搞定\\\"，2.0 问的就是——**AI 能不能真的\\\"认识\\\"一个用户**。\\n\\n## 数据规模：覆盖 1,580 天的真实生活\\n\\nVitaBench 2.0 构建了 56 个真实用户、819 个任务、2,000+ 条动态偏好、66 个可执行工具；平均每个用户 2,093 条交互事件，跨度 1,580 天。换句话说，**这不再是一次性 prompt 评测，而是把 Agent 放进一段长期、碎片化、偏好还会变化的关系里**。配套 Hugging Face 数据集、arXiv 论文（2605.27141）已同步发布，并被 ICLR 2026 接收。\\n\\n## 关键发现：加记忆反而更糟\\n\\n最反直觉的结果是：**给前沿模型喂入用户历史交互记忆后，成绩普遍下滑**。即便 Claude Opus 4.6 在 Full Context 设置下 Avg@4 也只有 0.503，DeepSeek-V4-Pro 在非思考模式为 0.456；切换到 RAG \u002F Agentic Memory 这些现实部署里更常见的记忆后端后，所有模型分数还会进一步掉。偏好缺失时模型还会硬猜不澄清——Claude 分数从 46.0 跌到 27.4。**长程个性化与主动服务（proactivity）仍是 SOTA LLM Agent 没跨过的硬关**。\\n\\n## 行业影响：Agent 的下一站在「时间维度」\\n\\n过去一年 Agent 评测焦点从单轮工具调用走向多步任务规划，VitaBench 2.0 把战线又拉长了一个量级：**真正的助理型 Agent 必须在数月甚至数年的时间跨度上累积和更新对用户的理解**。这意味着单纯把上下文窗口堆到 1M token 并不能解决问题——选择性记忆、抗噪检索、偏好漂移建模将成为新的工程化主战场。\\n\\n对想做 Agent 的团队来说，这套基准给出一个直接信号：**别再只盯 SWE-Bench \u002F GAIA 那类工具调用榜单，VitaBench 2.0 这一关才更接近\\\"产品能不能上线\\\"的真实门槛**。开源版本（meituan-longcat\u002FVitaBench-2.0）已开放，是时候让自家 Agent 接受\\\"长期用户\\\"考核了。\\n\"","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.27141","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"25f5c447-6b3d-4c29-875a-6e367fda57cb","en","VitaBench 2.0: long-term user modeling, even SOTA fails","arXiv 2605.27141 introduces VitaBench 2.0, a benchmark for evaluating LLM Agents on \"long-term user modeling\" — the ability to track, learn, and adapt to a specific user's preferences, habits, and context over weeks or months of interaction.\n\nThe benchmark structure: VitaBench 2.0 simulates 100 \"users\" with distinct personas, habits, and long-term goals. The Agent interacts with each user over 30 simulated days, and is evaluated on three tasks: (1) personalization — does the Agent adapt to the user's style? (2) memory — does the Agent remember key facts from earlier interactions? (3) anticipation — does the Agent proactively offer relevant help?\n\nThe result: even the leading model (Claude Opus 4.7 with full agent harness) scores only 41.2%. GPT-5.6 scores 38.7%. The gap to \"useful\" (70%) is huge. The most common failure: the Agent \"forgets\" user-specific context after 2-3 turns of conversation.\n\nThe diagnostic: the authors identify three failure modes — (1) context window overflow (the user history is too long); (2) memory retrieval noise (the wrong memory is recalled); (3) persona drift (the Agent gradually \"forgets\" the user's specific style and reverts to default).\n\nThe bigger takeaway: \"long-term user modeling\" is the next big Agent capability gap. Current Agents are good at \"in-conversation\" tasks, but they're bad at \"across-conversation\" personalization. This is the difference between a \"useful chatbot\" and a \"personal assistant\" — and the gap is huge.\n\nFor the industry, the takeaway is that \"memory\" and \"personalization\" are becoming first-class Agent capabilities, not afterthoughts. The next generation of Agent frameworks will need to invest heavily in long-term memory, persona tracking, and proactive anticipation.","vitabench-2-0-long-term-user-modeling","2026-06-25T14:01:00Z","2026-06-25T14:08:45.609435Z","2026-08-19T02:08:40.142862Z",true,"agent",129,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f26ace13-9c96-47ea-a528-b6682a22aa1e","Apodex 1.1 把推理搬进真实执行:PIVOT-RL 定位关键决策点,35B mini 开源","apodex-1-1-agentic-execution-pivot-rl","2026-08-25T14:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00"]