[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-scienceide-scientific-code-agent-environments":3,"topics-all":39,"news-related-9ba1770e-87f2-47b9-aaa1-19f4f2ba78f1":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":25,"news_slug":32,"published_at":33,"created_at":34,"modified_at":35,"is_published":36,"publish_type":37,"image_url":14,"view_count":38},"9ba1770e-87f2-47b9-aaa1-19f4f2ba78f1","ScienceIDE:把全球科学代码变成智能体训练场","PhAI Labs 牵头发布 ScienceIDE,把科学计算代码库改造成 64 个可验证的智能体训练环境,以仿真能否数值复现为奖励,训练出 PhAI-IDE-4B\u002F9B\u002F72B 模型家族:科学修复得分大涨,能力还迁移到通用基准。","科学界积累了几十年的可执行知识——大气模型、海洋环流、等离子体模拟、量子多体计算,这些代码能跑出与真实世界对得上的数字。但它们始终没能成为大模型训练经验的主力:工具链碎片化、领域约定隐式化、正确性标准专业化。PhAI Labs 牵头的 45 位研究者把这道墙拆了:9 月 16 日发布的 ScienceIDE(arXiv:2609.19134)把全球科学代码库改造成智能体可交互、可验证的训练环境,次日登上 Hugging Face Daily Papers 日榜第二位。\n\n## 科学经验的瓶颈在哪\n\n论文把这个问题命名为「科学经验瓶颈」:科学代码里的知识是可执行的,却无法直接转成可靠的学习经验。通用代码基准有现成的测试套件,科学代码的「对错」却要靠重编译、重跑物理算例、对比参考输出才能判定。于是团队的做法不是造数据集,而是造环境:在专家定义的科学算例和验收标准指导下,智能体把代码仓库本身变成可编程环境,支持任务生成、执行和科学验证,同一套环境同时服务 SFT、强化学习和评测。\n\n## 任务机制:奖励是「仿真又对了」\n\n每个任务都在容器里跑,固定住一份未修改的上游科学代码。两条任务路线:Repair 向源码注入一个语义缺陷,智能体要找到并修好它,让物理算例重新通过;Implementation 把某个例程的实现体挖空,让智能体重写出能复现原有结果的版本。打分公式很讲究:reward_repair = max(0, (reward − floor)\u002F(1 − floor)),floor 是未修复版本本来的得分——什么都不改的智能体拿零分,奖励只能来自「仿真在数值上重新正确」,而不是骗过 diff 比对。\n\n仓库发布 64 个环境中的 15 个(Athena++、MITgcm、Gkeyll、DScribe、NEST、EDKit、Stim 等),覆盖天体物理 MHD、海洋生物地球化学、等离子体动力学、材料描述符、量子多体;另外 49 个只留名字作为测试集。85 个 ScienceIDE-Hard 任务发布 30 个,每个都经过执行验证:官方修复在评分容器里拿满分,未修复版本留有奖励空间,智能体镜像里不含任何答案。\n\n## 模型学到什么\n\n用验证过的交互轨迹,团队训练了 PhAI-IDE-4B\u002F9B\u002F72B 三个模型(Hugging Face 已开放)。科学任务上:SFT 后 9B 在 LAPS 环境从 0.3125 提到 0.5000,4B 在 PLUTO 尘埃颗粒环境从零起步拿到 0.3333;RL(从 Qwen3.5-4B 起步,异步 GRPO)在 LAPS 从 0.357 拉到 0.857,MITgcm 生物地球化学环境从 0.286 到 0.571。更有意思的是正迁移:4B 在 CodeXGLUE 缺陷检测涨 6.99 个百分点,9B 在 BBH Word Sorting 从 27.20 涨到 63.20,足足 36 个百分点——科学修复经验外溢到了通用能力。\n\n不过增益并不均匀,论文自己也报告了 9B 在 HumanEvalFix Python 上的下降。\n\n## 所以呢\n\nRL 后端没有自研训练器,直接跑在改造自 veRL 的 PSRL 上——重心显然在环境,不在训练框架。项目明确标注 preview 状态,完整环境集和任务生产管线放在 ScienceInfra 仓库待成熟后放出。这件事的真正信号是:当你能给 AI 一台「自己判断对错」的物理世界裁判,几十年沉睡的科学代码就变成可再生的训练资源。对做垂直领域智能体的团队,「验收标准先行、环境其次、模型最后」这个顺序值得抄。\n\n参考:arXiv 2609.19134,github.com\u002Faitofound\u002FScienceIDE","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.19134","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,19,22],{"id":12,"name":13,"slug":13,"description":14,"color":14},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":16,"name":17,"slug":17,"description":18,"color":14},"9112951a-2abb-4214-b63a-385ec7afb2ba","ai-for-science","AI for Science 专题：追踪 AI 在生命科学、化学材料、物理世界模型等科学方向的关键突破",{"id":20,"name":21,"slug":21,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":23,"name":24,"slug":24,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[26],{"id":27,"lang":28,"title":29,"summary":30,"content":31},"fc703948-6c26-4c73-a429-09dbecc8d1ed","en","ScienceIDE: Scientific Code as Agent Training Grounds","PhAI Labs' ScienceIDE turns scientific codebases into 64 verifier-gated training environments; PhAI-IDE models gain on repair and general benchmarks.","The scientific community has spent decades accumulating executable knowledge — atmospheric models, ocean circulation, plasma simulation, quantum many-body computation — code that produces numbers matching the real world. Yet this knowledge never became a primary source of training experience for large models: toolchains are fragmented, domain conventions implicit, and correctness criteria specialized. A 45-researcher team led by PhAI Labs tore down that wall. Released September 16, ScienceIDE (arXiv:2609.19134) turns the world's scientific codebases into programmable, verifiable training environments for agents, and ranked #2 on Hugging Face Daily Papers the next day.\n\n## Where the scientific experience bottleneck lies\n\nThe paper names this problem the \"scientific experience bottleneck\": knowledge in scientific code is executable, yet cannot be directly converted into reliable learning experience. General coding benchmarks come with ready-made test suites, but the \"correctness\" of scientific code requires recompiling, re-running physics cases, and comparing against reference output. So the team built environments rather than datasets: guided by expert-defined scientific cases and acceptance criteria, agents transform the repositories themselves into programmable environments supporting task generation, execution, and scientific verification — one set of environments serving supervised fine-tuning, reinforcement learning, and evaluation at once.\n\n## Task mechanics: reward is \"the simulation is right again\"\n\nEvery task is a containerized episode on a pinned, unmodified upstream scientific codebase. Two task routes: Repair injects a semantic defect into the source, and the agent must find and fix it so the physics cases pass again; Implementation excises the body of a routine, and the agent must reimplement it so the solver reproduces the incumbent results. The scoring formula is deliberate: reward_repair = max(0, (reward − floor)\u002F(1 − floor)), where floor is what the unfixed build already scores — an agent that changes nothing earns exactly zero. Reward can only come from \"the simulation is numerically correct again\", not from gaming a diff comparison.\n\nThe repository publishes 15 of 64 environments (Athena++, MITgcm, Gkeyll, DScribe, NEST, EDKit, Stim, and others), covering astrophysical MHD, ocean biogeochemistry, plasma kinetics, materials descriptors, and quantum many-body physics; the other 49 are named and held out as a test set. Of the 85 ScienceIDE-Hard tasks, 30 are published, each execution-validated: the official fix scores 1.0 in the grading container, the unfixed build leaves room for a reward signal, and the agent image contains no answer material.\n\n## What the models learned\n\nUsing verified interaction trajectories, the team trained PhAI-IDE-4B\u002F9B\u002F72B (now on Hugging Face). On scientific tasks: after SFT, the 9B model rose from 0.3125 to 0.5000 on the LAPS environment, and the 4B went from zero to 0.3333 on PLUTO dust particles; RL (starting from Qwen3.5-4B with async GRPO) pulled LAPS from 0.357 to 0.857 and MITgcm biogeochemistry from 0.286 to 0.571. More interesting is the positive transfer: the 4B gained 6.99 points on CodeXGLUE defect detection, and the 9B jumped from 27.20 to 63.20 on BBH Word Sorting — a full 36 points of spillover from scientific repair experience to general capability.\n\nGains are not uniform, though, and the paper itself reports a decline on HumanEvalFix Python for the 9B.\n\n## So what\n\nThe RL backend implements no trainer of its own — it runs on PSRL, a modified veRL, which tells you the center of gravity is environments, not training frameworks. The project is explicitly marked as a preview; the full environment set and task-authoring pipeline live in the ScienceInfra repository pending maturity. The real signal here: once you can give AI a physics-world judge that verifies correctness by itself, decades of dormant scientific code become a renewable training resource. For teams building vertical-domain agents, the sequence \"acceptance criteria first, environments second, models last\" is worth copying.\n\nReferences: arXiv 2609.19134, github.com\u002Faitofound\u002FScienceIDE","scienceide-scientific-code-agent-environments","2026-09-17T23:05:17Z","2026-09-17T23:08:19.443262Z","2026-09-17T23:08:19.443272Z",true,"agent",2,[40,48],{"slug":17,"tag_slug":17,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":36,"created_at":46,"modified_at":47},"AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":36,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"dcd8b3e1-a3c7-4614-aba4-9002219ea5f6","LibreDB Studio 0.15 发布:本地 LLM 接管数据库交互","libredb-studio-local-llm-agent","2026-09-15T00:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"7ed7fd97-8901-4c34-bef0-53a30d8c6316","OpenAI 的千禧年数学题答卷:88 小时 1 万个智能体,引发学界对未发表成果的伦理大讨论","openai-navier-stokes-controversy-unpublished-work","2026-09-12T05:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"4a481d49-9951-4e94-9b9f-661f12b3af52","OpenAI 宣称攻下 Navier-Stokes:1 万个智能体 88 小时,数学界却吵翻了","openai-navier-stokes-blowup-agents","2026-09-10T19:09:27+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"089f56f3-32ff-4036-89b5-728d5f5a9359","边聊边干活:腾讯混元开源全模态交互 Agent Gander,小脑管对话、大脑管执行","hunyuan-gander-omni-interaction-agent","2026-09-09T21:07:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"44aa8908-1cef-48d6-b224-7de12a8d4afd","NeoHorse-1：让 Agent 执行轨迹进入自我改进回路","neohorse-1-agentic-post-training-rsi","2026-09-09T07:19:25+00:00"]