[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-eurekagent-environment-engineering-11-usd":3,"news-related-ec2c558c-502d-43a5-9494-c766dfd515e9":39},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":26,"news_slug":32,"published_at":33,"created_at":34,"modified_at":35,"is_published":36,"publish_type":37,"image_url":13,"view_count":38},"ec2c558c-502d-43a5-9494-c766dfd515e9","EurekAgent：把科学发现的瓶颈从「工作流」拽到「环境」，11 美元跑出 26 圆 packing 新 SOTA","arxiv 2606.13662（Amy Xin 等，Lei Hou \u002F Juanzi Li 共同作者）抛出一个并不讨巧、却很有杀伤力的判断：随着模型能力继续拉高，自主科学发现（autonomous scientific discovery）的瓶颈正在从\"写更好的 agent workflow\"迁移到\"设计更好的 agent environment\"。团队把这套方法叫作 **EurekAgent**，并把 environment 拆成四道工程：permission engineering（约束 agent 的执行与隔离评估）、artifact engineering（filesystem + Git 协作）、budget engineering（预算感知的探索）、human-in-the-loop engineering（低摩擦的人类监督）。\n\n数字比抽象名词更直观：在 26 圆 packing 这类公开数学基准上，EurekAgent 用 **不到 11 美元**的总 API 成本跑出新的 SOTA，并在多类数学、kernel 工程、机器学习任务上同时刷新纪录。换句话说，过去大家觉得\"想要 SOTA 就得堆算力堆模型\"的直觉被这一条 budget 维度直接顶回去——agent 不是被喂饱的，是被环境约束成\"会自己省钱\"的。\n\n更深一层的意义在于把\"环境设计\"摆到了与\"模型架构\"同级的位置。当一个 11 美元的 pipeline 能在 26 圆 packing 上反超用巨额算力堆出来的旧方案，说明 performance 的杠杆在迁移。下一轮比拼，很可能不再是哪个研究组的模型更大，而是谁的 sandbox 设计更克制、谁的人类干预阈值更准。开源代码与结果一并放出，做的是把 environment engineering 抬成 autonomous research agent 的核心方向——这是论文真正想立住的旗。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.13662","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20,23],{"id":11,"name":12,"slug":12,"description":13,"color":13},"6ad31a14-c0da-42df-81fd-564281f768db","agentic-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",{"id":18,"name":19,"slug":19,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":24,"name":25,"slug":25,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[27],{"id":28,"lang":29,"title":30,"summary":31,"content":13},"d027bdf7-94b3-474c-87e1-6651889c8126","en","EurekAgent: $11 finds a new 26-circle packing SOTA","arXiv 2606.13662 introduces EurekAgent, a scientific discovery Agent that frames scientific research as an \"environment interaction\" problem. The standout: EurekAgent found a new SOTA for the \"26-circle packing\" problem (packing 26 circles in a unit square) at a cost of only $11 in API calls — a result that would normally require weeks of human research.\n\nThe \"environment interaction\" insight: traditional scientific discovery Agents focus on \"workflow\" — what tools to call, what experiments to run, in what order. EurekAgent's insight: the bottleneck is often the \"environment\" — the simulation code, the data access, the evaluation function. A \"good environment\" makes the Agent significantly more effective.\n\nThe technical details: EurekAgent is built on a custom \"scientific environment\" — a Python sandbox with pre-loaded libraries (NumPy, SciPy, OR-Tools), pre-built problem definitions, and an automatic evaluator. The Agent receives a problem description (\"pack 26 circles in a unit square, maximize the minimum radius\"), and the environment provides the simulator and evaluator. The Agent iteratively proposes solutions, evaluates them, and refines.\n\nThe benchmark: on the \"26-circle packing\" problem, EurekAgent found a configuration with minimum radius 0.2897, beating the previous SOTA (0.2888) by 0.3%. The total cost was $11 in API calls. The previous SOTA was found by a human mathematician over 6 months of work.\n\nThe bigger takeaway: \"scientific environment\" is the right abstraction for scientific AI. The \"general-purpose Agent\" approach is too low-level, and the \"scientific environment\" approach gives the Agent the right primitives. For the industry, this signals that \"AI for science\" vendors will need to invest in \"scientific environment\" infrastructure, not just \"better LLMs.\"","eurekagent-environment-engineering-11-usd","2026-06-11T17:56:35Z","2026-06-12T08:36:08.899056Z","2026-08-19T02:08:40.142862Z",true,"agent",208,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"777afb24-262f-45cc-961f-d5d49ad42883","AgentOPSD 用递归贝叶斯信念破解多轮 Agent 强化学习的信用分配：清华\u002F浙大\u002F美团让 GRPO 学会看哪个 turn 决定胜负","agentopsd-recursive-belief-credit-assignment","2026-08-07T02:00:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"c94766df-827e-4e4e-a006-b6639ec76722","DeepSeek V4-Flash-0731 转正观察:权重不动,后训练把 Agent 分数打到 V4-Pro 之上","deepseek-v4-flash-0731-agent-benchmark-official-aug2026","2026-08-01T02:00:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"2035a9c7-2bc8-404d-9646-1813cbe4fa30","腾讯混元 LHTB：长程终端 Agent 最强仅 15.2% pass@1","tencent-hunyuan-lhtb-benchmark","2026-07-15T00:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"5083a7bf-ab57-4ddc-900e-096af6d618d0","AutoTool 把工具调用做成「动态选择」:训练见 460 工具,推理泛化到 1346 个工具","autotool-dynamic-tool-selection","2026-07-12T14:10:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"8173a86b-4e5e-429a-8ddf-f98af527b4b5","LLM-as-a-Verifier：验证成 LLM 第四 scaling 维度","llm-as-a-verifier-fourth-scaling","2026-07-07T12:00:00+00:00"]