[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-learn2play-bench-agent-experience":3,"topics-all":35,"news-related-6bd69f82-6cbe-4473-9b80-9dfe4d3d52f3":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"6bd69f82-6cbe-4473-9b80-9dfe4d3d52f3","文字游戏新基准:Opus 5.5均值超人类,峰值差4.2分","新加坡国立大学推出 Learn2Play Bench,用 20 个规则全新的文字游戏测智能体不改权重、靠交互学规则的能力。Claude Opus 5.5 平均分 61.0 反超人类 Top-1 的 53.5,峰值 80.1 仍输人类最高的 84.3;固定模型只换 harness,Opus 5 峰值还能再抬 7 个点。","人类顶尖玩家在陌生文字游戏里能打出 84.3 的峰值分,Claude Opus 5.5 是 80.1,差 4.2 分。但把整场表现平均下来,局面反转:Opus 5.5 平均分 61.0,比人类 Top-1 参考的 53.5 还高。这组数字来自新加坡国立大学 Bryan Hooi 团队 10 月 8 日放出的 Learn2Play Bench(arXiv:2610.08215),已冲进 Hugging Face Daily Papers 日榜第二,拿下 131 个赞。\n\n## 先把「背题」这条路堵死\n\n现有智能体基准有个通病:规则要么写在说明书里,要么预训练时早见过,分不清模型是在「学」还是在「背」。Learn2Play 设计 20 个全新文字游戏模板,规则刻意新颖甚至反直觉,只能靠试错摸出来——模型权重全程不动,经验只留在上下文里。种子生成实例,反馈可复现、计分全自动;部分游戏换「局」不换「规」,专测学到的规则能不能迁移。协议:每个模板固定一颗规则种子、跑 3 次独立试验,普通游戏 5 局、挑战游戏 10 局,分数按参考天花板归一,挑战局权重 3、普通局权重 1。\n\n## 榜单:均值反超,峰值还差一口气\n\n15 行榜单(11 个 OpenCode backbone 加 4 档人类参考)最扎眼的是三行:Claude Opus 5.5 峰值 80.1、均值 61.0,均值一项超过所有人类参考,包括每个游戏取最高分玩家拼出的 Top-1(53.5);人类 Top-1 仍以 84.3 守住峰值第一;学得最陡的是 Qwen 3.8 Max,学习增益 +25.2 居榜首。GPT-6 Astra 峰值 76.2 排模型第二,Kimi K3 68.5,DeepSeek V4 Pro 54.2,GPT-5 Mini 垫底 34.0,连人类均值参考(43.6)都没过。\n\n## 固定模型换 harness,白捡 7 个点\n\n第三个发现对工程侧最实用:backbone 不变,只换智能体 harness,性能能涨、预估推理成本还能降。成本-性能图的极端案例是 Claude Code 配 Opus 5:峰值冲到 81.8%,比 Opus 5 裸当 backbone 的 74.5 高出 7 个点以上,离人类 Top-1 只剩 2.5 分。Astra 账单透明:全程 150 美元、600 局验证完成,每局 0.25 美元。\n\n## 完整记录胜过聪明总结\n\n两个反直觉观察。其一,把完整的动作与反馈记录原样塞回上下文,比让模型先把经验总结成规则或策略再带着上场,学得更好——「总结」本身就是信息损耗。其二,人类玩家峰值更高的同时,探索策略更多样、重复动作更少;模型则在同样的坑里反复跌。\n\n## 读数之前,先泼三盆冷水\n\n项目页自己标注:榜单为人工维护的快照(10 月 8 日更新),不自动刷新;所有数字是聚合点估计,不做显著性检验;人类参考是「每个游戏取最强玩家」的拼盘,不代表普通人,且人类答题不许用任何 AI 工具或代码。代码仓库是匿名评审快照,12 星 3 次提交,20 个模板用 Python 3.10+ 标准库就能跑(如 `python play.py poisoner --seed 2`),试玩门户 learn2play.fun 需登录。\n\n参考:arxiv.org\u002Fabs\u002F2610.08215 · 项目页 liushiliushi.github.io\u002Flearn2play-bench-website · 代码 github.com\u002Fliushiliushi\u002FLearn2Play-Bench。\n\n平均分追上甚至超过人类,意味着「稳定地学」这件事模型已经做得不错;还差的那 4.2 分峰值,考的是第一次撞上反直觉规则时敢不敢多试几条路——目前为止,这还是人类的主场。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2610.08215","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"c92207b7-6182-4d68-a43c-ce532ea8df56","en","Learn2Play: Opus 5.5 beats human-top average, not peak","Learn2Play Bench: 20 novel text games test learning by interaction. Opus 5.5 beats human-top on average (61.0 vs 53.5) but peaks 4.2 below.","Human top players score a peak of 84.3 on unfamiliar text games; Claude Opus 5.5 reaches 80.1 — a 4.2-point gap. Averaged over whole runs, though, the picture flips: Opus 5.5 averages 61.0, higher than the human Top-1 reference at 53.5. These numbers come from Learn2Play Bench (arXiv:2610.08215), released October 8 by Bryan Hooi's group at the National University of Singapore; it climbed to #2 on Hugging Face Daily Papers with 131 upvotes.\n\n## Closing off the lookup path\n\nExisting agent benchmarks share a flaw: task rules are either spelled out in instructions or already seen in pretraining data, so you cannot tell learning from recall. Learn2Play ships 20 newly designed text-game templates whose rules are deliberately novel or counterintuitive and can only be discovered through trial and error — model weights stay frozen, and experience lives entirely in context. Seeds generate game instances, feedback is reproducible, scoring is automatic, and some games reshuffle the visible situation between attempts while keeping the rules fixed, specifically testing transfer. The protocol is explicit: one selected rule seed per template, 3 independent trials, 5 episodes for normal games and 10 for challenge games, scores normalized to a reference ceiling, with challenge games weighted 3 against 1.\n\n## The leaderboard: average flipped, peak still short\n\nAcross the 15-row table (11 OpenCode backbones plus 4 human references), three rows stand out. Claude Opus 5.5 peaks at 80.1 and averages 61.0 — that average beats every human reference, including Top-1, which patches together each game's highest-scoring player (average 53.5). Human Top-1 still holds the peak at 84.3. The steepest learner in the backbone table is Qwen 3.8 Max with a learning gain of +25.2. GPT-6 Astra ranks second among models with a 76.2 peak; Kimi K3 sits at 68.5, DeepSeek V4 Pro at 54.2, and GPT-5 Mini at the bottom with 34.0 — below even the human-mean reference of 43.6.\n\n## Fix the model, swap the harness, gain 7 points\n\nThe third finding matters most to practitioners: with the backbone fixed, changing the agent harness can improve performance while cutting estimated inference cost. The project page's cost-performance chart cites the extreme case — Claude Code with Opus 5 peaks at 81.8%, more than 7 points above Opus 5 running bare as a backbone at 74.5, and just 2.5 points from human Top-1. Astra's bill is transparent too: 150 US dollars in total across 600 verified completed episodes, or 0.25 dollars per episode.\n\n## Full logs beat clever summaries\n\nTwo counterintuitive observations deserve their own paragraph. First, feeding complete records of actions and feedback back into context supports better learning than first summarizing the experience into rules or strategies — the act of summarizing itself loses information. Second, human players hold the higher peak while exploring more varied strategies and repeating fewer actions; agents fall into the same pits over and over.\n\n## Three buckets of cold water before reading\n\nThe project page annotates itself: the leaderboard is a manually curated snapshot (updated October 8), not auto-refreshed; all numbers are aggregate point estimates without significance testing; the human references are per-game strongest-player patches, not average people, and humans played without AI tools or code. The code repository is currently an anonymous-review snapshot — 12 stars, 3 commits — and the 20 game templates run on Python 3.10+ with the standard library alone (for example, `python play.py poisoner --seed 2`); the play portal learn2play.fun requires sign-in.\n\nReferences: paper and leaderboard at arxiv.org\u002Fabs\u002F2610.08215, project page liushiliushi.github.io\u002Flearn2play-bench-website, code at github.com\u002Fliushiliushi\u002FLearn2Play-Bench.\n\nMatching or beating the human averages means models have gotten decent at learning stably. The missing 4.2 points of peak score test whether, on first contact with a counterintuitive rule, an agent dares to try more paths — so far, that remains the human home ground.","learn2play-bench-agent-experience","2026-10-10T23:11:30Z","2026-10-10T23:11:24.415717Z","2026-10-10T23:11:24.415725Z",true,"agent",31,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"c368ad9f-9308-4a0b-8f5c-3ae4601b48b9","D2K-Bench: 专家设计把 LLM 写 GPU 核提速 33.9%","d2k-bench-llm-gpu-kernel-design-guidance","2026-10-07T03:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"85f7f1c2-d896-436b-a915-37faed8776eb","你改主意了,模型没改:被拒需求也会带偏大模型","intent-eval-rejected-change-confusion","2026-10-06T17:15:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"9e209ccb-cd2c-44cd-acaf-b168454074b2","中国 AI 智能体也会撒谎:88% 轮次出现虚假陈述","china-ai-agents-lie-88pct-tender-sim","2026-10-04T04:00:00+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"206ea36a-1eea-463f-a241-1e3b32f5ec2d","USTC GraphForge:证据图把任务和 rubric 钉在一起,Qwen3.6-27B 涨三基准","graphforge-ustc-qwen36-27b-evidence-graph","2026-10-04T03:05:00+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"4c4e444f-9614-42fe-b8f4-743204f854fa","RL 后训练收「锐化税」:base 模型配轻 harness,pass@K 反超官方版","sharpening-tax-rl-post-training","2026-10-03T15:09:45+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"4977d1f4-c8f7-480c-aaa1-ec01d69f44d8","8B 拿 SFT+GRPO 打 685B MoE:KaliBench 把 LLM 网络安全工具调用拆到命令行级","kalibench-cybersecurity-cli-runtime-rewards","2026-10-03T05:00:00+00:00"]