[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-explorationbench-alien-worlds":3,"topics-all":38,"news-related-fe5546f3-09e4-42a3-bfde-6bbd0e8d474c":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"fe5546f3-09e4-42a3-bfde-6bbd0e8d474c","外星世界实测:探索 4 轮 87.6%,垫底 12.9%","ExplorationBench 用两个规则被篡改的可执行世界测 AI 的探索能力:手册故意写错,规则只能靠试出来,解释器精确判分。十个前沿模型四轮探索后从 15.7% 涨到 87.6%,但不接环境纯空想最高只有 11%;谁设计实验影响巨大,轨迹之间最大差 72.8 分,30 条里 6 条收尾倒退。","在 ExplorationBench 的世界里,EMIT(100) 打印出来的不是 100,而是 127——因为这个外星编程语言的整数会被悄悄异或 27;而 PLUCK 操作符虽然手册写着「从 0 开始计数」,实际却从 1 开始数。手册是错的,先验知识是反的,唯一的办法是动手试。\n\n这正是这个新基准的设计意图。一个 20 人研究团队(署名机构含腾讯混元)把「AI 能不能真探索」这个玄学问题,做成了可执行、可验证的测量框架。论文《ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds》9 月 24 日挂上 arXiv(编号 2609.30199),目前已进入 Hugging Face 论文日榜。\n\n## 为什么需要外星世界\n\n测「探索」一直有个死结:任务必须对模型是新的,否则考的是背诵;但答案必须能精确验证,否则没法判分。数学和代码满足第二条、不满足第一条——预训练数据可能早就见过;真正的科学发现满足第一条、不满足第二条——验证可能要专家审好几年。\n\n团队的解法是造两个「外星世界」。AlienCode 是一门小型计算语言,藏着 31 个发现目标、70 道留出任务;AlienLogic 是一个推理规则被改动的自然演绎系统,24 个发现目标、70 道任务。两个世界都是确定性可执行的:程序由解释器判分,证明由证明检查器判分,全程不需要 LLM 当裁判。模型拿到一份故意写错的手册和几个示例,然后有四轮探索机会——提交程序或证明、读返回结果、修正假设;每个里程碑后关掉工具考闭卷。\n\n## 结果:探索有用,但没那么可靠\n\n十个前沿系统(每个世界各跑三条独立轨迹)的成绩单信息量很大:\n\n- 探索前,没有任何轨迹在 AlienCode 上超过 15.7%;四轮自主探索后,Best@3 最高冲到 87.6%(Claude Opus 5),十个里七个超过 60%。\n- 把环境反馈拿掉、让模型用同样轮数纯空想,成绩只剩 0.5%–11.0%——多想不等于多知道。\n- 谁设计实验很关键:自主选题中位数 66.0%,把自己最佳轨迹的探针原样重放降到 40.7%,换成固定探针序列只剩 5.7%。证据完全一样,差距来自「设计实验」这个动作本身。\n- 榜单还会互相打架:Grok 4.6 在 AlienCode 只排第五,到 AlienLogic 直接第一;DeepSeek-V4-Pro 在 AlienCode 以 12.9% 垫底,在 AlienLogic 却有 73.8%。两个榜单排名相关性只有 0.35——「探索能力」并不是一个可以随身携带的通用分数。\n\n更扎心的是三个负面发现。说了不等于会做:把所需规则正确说出来的系统,对应任务也只有 70.9% 做对。探索不可靠:同一系统同一预算下,轨迹之间最大能差 72.8 分,30 条 AlienCode 轨迹里有 6 条收尾比中途里程碑还低。以及在 AlienLogic 里,直接把完整规则表喂给模型能拿 93%–97%,仍然强过所有系统的自主探索。\n\n## 所以呢\n\n这份数据对「AI 科学家马上要替代研究生」的叙事是一盆温度合适的冷水:前沿模型确实能在有限预算内从零摸出一整套陌生规则——这部分是真的强;但同样的系统换一条轨迹重跑,结果可能大幅缩水,把自己发现的规则背出来也不保证会用。把它当作「探索能力尚不可靠、不可迁移」的测量工具,可能比把它当新排行榜来刷更有价值。\n\n参考:arxiv.org\u002Fabs\u002F2609.30199","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.30199","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"488b6ed9-bd88-413a-aa04-0cc1c8ae3527","en","Alien-World Benchmark: Best System Hits 87.6% After 4 Rounds","Two alien worlds with flawed manuals test AI exploration: ten systems reach 87.6% in four rounds, yet trajectories diverge by up to 72.8 points.","In the world of ExplorationBench, EMIT(100) does not print 100 — it prints 127, because integer literals in this alien language are silently XOR-ed with 27, and PLUCK counts positions from one although the manual says zero. The manual is deliberately wrong. The only way through is to experiment.\n\nThat is the design intent of the new benchmark. A 20-author research team, with Tencent Hunyuan among the signing institutions, turned the fuzzy question of whether AI systems can genuinely explore into an executable, exactly verifiable measurement framework. The paper, ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds, hit arXiv on Sep 24 (2609.30199) and has since entered the Hugging Face daily papers list.\n\n## Why alien worlds\n\nEvaluating exploration has a built-in deadlock. Tasks must be new to the model, otherwise you are testing recall; yet answers must be exactly verifiable, otherwise you cannot grade. Math and coding satisfy the second requirement but not the first — pre-training data may already contain the answers. Genuine scientific discoveries satisfy the first but not the second — verification can take expert years.\n\nThe team's answer is to build two alien worlds. AlienCode is a small calculation language with 31 hidden discovery targets and 70 held-out tasks; AlienLogic is a natural-deduction system with patched inference rules, 24 discovery targets, and 70 tasks. Both are deterministic and executable: an interpreter grades programs, a proof-checker grades proofs, and no LLM judge is involved. Each system receives a flawed manual and a few worked examples, then explores for four rounds — submitting programs or proofs, reading results, revising hypotheses. After each round it is tested closed-book, with tools disabled.\n\n## Results: exploration works, unreliably\n\nTen frontier systems, three independent trajectories each, produced a dense set of findings:\n\n- Before exploration, no AlienCode trajectory exceeds 15.7%. After four autonomous rounds, Best@3 peaks at 87.6% (Claude Opus 5), and 7 of 10 systems pass 60%.\n- Remove environment feedback and let models deliberate for the same number of turns: scores collapse to 0.5%-11.0%. More thinking is not more knowing.\n- Who designs the experiments matters. Median accuracy is 66.0% under autonomous exploration, 40.7% when a system's own best probes are replayed to it, and 5.7% under a fixed probe sequence. Identical evidence — the gap comes from designing experiments itself.\n- Rankings do not transfer. Grok 4.6 is fifth in AlienCode but first in AlienLogic; DeepSeek-V4-Pro bottoms out at 12.9% in AlienCode yet reaches 73.8% in AlienLogic. The Spearman correlation between the two boards is 0.35 — exploration ability is not one portable score.\n\nThree negative findings cut deeper. Saying is not using: on tasks whose required rules a system states correctly, it still solves only 70.9%. Exploration is unstable: trajectories of one system under one budget end up to 72.8 points apart, and 6 of 30 AlienCode trajectories finish at least 3 points below an earlier milestone. And in AlienLogic, simply handing over the complete rule set yields 93%-97%, still beating every system's autonomous exploration.\n\n## So what\n\nFor the AI-scientist narrative, this is a temperate bucket of cold water. Frontier models can genuinely reverse-engineer a set of unfamiliar rules from scratch within a bounded budget — that part is real. But rerun the same system and the result may shrink sharply; reciting a discovered rule does not guarantee applying it. Treating this benchmark as a measurement of exploration being unreliable and non-transferable is probably more valuable than farming it as another leaderboard.\n\nReference: arxiv.org\u002Fabs\u002F2609.30199","explorationbench-alien-worlds","2026-09-28T21:07:30Z","2026-09-28T21:08:14.465112Z","2026-09-28T21:08:14.465127Z",true,"agent",371,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"3c6fcf46-f5bb-4136-931c-69cd64216e12","Skill-Use 基准揭示 Agent 短板：会做任务，不等于会用 Skill","skill-use-agent-harness-benchmark","2026-08-06T08:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"74de194b-9e2c-45ab-aa13-12fe210e66ba","HiGram 给 Agent 记忆加上“路径定位”：先找证据，再改记忆","higram-agent-memory-path-localization","2026-08-05T09:32:43+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"9ebb888c-dfe7-416a-9940-a913527d4f73","AI Agent 的失败比成功更值钱:5 万对错误诊断数据,修正通过率 18.4%→51.1%","agent-error-dataset","2026-10-01T15:11:08+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"237d0bac-204c-4818-9262-e576f96df623","Agent 被自己的历史拖累:人大 AEWM 编辑任务状态,六项基准最高涨 6.7 分","agent-editing-world-model","2026-09-29T17:08:51+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"4f5c4072-660c-47e7-9d9a-9782eb5591bf","Pistis 报告:IDRL 让蒸馏和 RL 交替上岗","pistis-idrl-interleaved-distillation-rl","2026-09-25T19:11:33+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"e98cf9a2-8348-4142-9a56-c11166774798","SpeakerMem-R1:多方对话记忆,分清谁说了什么","speakermem-r1-multi-party-memory","2026-09-24T19:05:00+00:00"]