[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-turing-rl-mit-stanford-discriminative-reward":3,"news-related-daf9222a-aebc-4f09-92f8-74b6226fd5f1":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"daf9222a-aebc-4f09-92f8-74b6226fd5f1","把图灵测试变成 RL 损失：MIT 提出 Turing-RL，让用户模拟器更\"像人\"","训练具备人类一致性的用户模拟器，是构建 AI Agent 训练环境、评估个性化系统与研究人类行为的重要基础。2026 年 6 月 17 日，MIT、斯坦福大学与 MIT-IBM Watson AI 实验室联合发布论文 arXiv:2606.19336，提出一种新的强化学习框架 Turing-RL。其核心思想是：让一个 LLM 评委以 1–7 分的 Likert 量表同时看到\"模拟器生成\"与\"真实用户\"的回复，输出来自图灵测试的\"判别式图灵奖励\"（discriminative Turing reward），再以 GRPO 算法配合 SFT 预热优化策略。论文在 PRISM 多轮对话和 ConvoKit Reddit 论坛两个场景中同时验证：Turing-RL 训练出的用户模拟器在 LLM 评分与人类评分上都一致优于\"相似度奖励\"（Sim-RL，改编自 HumanLM）与\"对数似然奖励\"（Logprob-RL）两条主流基线，且不牺牲与真值的相似性。这条思路把\"图灵测试\"从哲学概念变成了可计算的 RL 损失——优化对象从\"匹配单条 ground truth\"转向\"整体不可区分性\"，是用户模拟与 Agent 训练领域一次值得关注的范式转换。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.19336","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1ff951af-1dba-4860-9fd7-b17f43fd58f4","en","Turing-RL: MIT turns the Turing test into a loss function","arXiv 2606.19336 introduces Turing-RL, an MIT method that turns the Turing Test into an RL loss function. The result: a \"user simulator\" LLM that is significantly more human-like, enabling more realistic evaluation of conversational AI.\n\nThe \"user simulator\" problem: evaluating conversational AI requires simulating realistic users. But LLMs simulating users tend to be \"too LLM-like\" — they use perfect grammar, never make typos, always respond coherently. This makes the evaluation unrealistic.\n\nThe Turing-RL fix: the user simulator is trained with an RL loss that explicitly rewards \"human-like\" behavior. The \"human-likeness\" is measured by a separate LLM judge (a \"Turing judge\") that tries to distinguish between real human responses and the simulator's responses. The RL loss is the judge's confidence that the simulator is human.\n\nThe result: Turing-RL-trained user simulators are indistinguishable from real humans 68% of the time (up from 22% for vanilla LLM simulators). The simulated users exhibit realistic behaviors: typos, hesitation, off-topic tangents, follow-up questions, etc.\n\nThe benchmark: when used to evaluate GPT-5.6, Claude Opus 4.7, and other frontier models, Turing-RL user simulators reveal \"failures\" that are missed by vanilla LLM simulators — e.g., the model gives a technically correct but unhelpful answer 35% of the time (a failure mode that real users report but simulated users miss).\n\nThe bigger takeaway: \"user simulation\" is a real research problem. As conversational AI becomes more capable, we need better ways to evaluate it, and \"more realistic user simulators\" is the right path. Turing-RL's RL-based approach is a significant step forward, and the open-source release will benefit the entire conversational AI research community.","turing-rl-mit-stanford-discriminative-reward","2026-06-18T02:00:00Z","2026-06-19T12:13:03.304730Z","2026-08-19T02:08:40.142862Z",true,"agent",111,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"22a1a718-0eb6-46e5-8ee8-825400de11d1","DeepMind WeatherNext 在 Nature 发论文：用 28 km 粗分辨率做出多一天的飓风预警,代码权重全部开源","deepmind-weathernext-cyclones-nature-open-source","2026-08-10T02:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00"]