[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gpt-6-astra-robot-arm-benchmark":3,"topics-all":38,"news-related-e12d2e7d-35b7-42d2-b02f-bdcc0a547878":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"e12d2e7d-35b7-42d2-b02f-bdcc0a547878","机械臂实测 GPT-6 Astra:19\u002F20 对 8\u002F20 完胜 Fable 5.1,精细插入却全员卡壳","独立评测机构 Robocurve 把 GPT-6 Astra 与 Claude Fable 5\u002F5.1 放上同一套双臂机器人:积木进碗任务 Astra 完成 19\u002F20,远超 Fable 5.1 的 8\u002F20,单轮成本约一半;但更难的拼图插入任务,双方都只完成 2\u002F20。","把 GPT-6 Astra、Claude Fable 5.1 和 Fable 5 放上同一套双臂机械臂,只给两个任务:把桌上的红色积木捡起来放进碗里;捏住圆形拼图件中央的旋钮,把它插进板上匹配的圆形凹槽。独立评测机构 Robocurve 9 月 4 日发布的这份报告,给前沿模型的「物理之手」划出了一条少见的清晰能力线,也是他们对 Fable 5 与 5.1 对比测试的后续。\n\n## 19\u002F20 对 8\u002F20:一边倒的积木任务\n\n每个模型每个任务各跑 20 次,全程共 120 次有效试跑。积木进碗任务上,GPT-6 Astra 完成 19 次,Claude Fable 5.1 完成 8 次,上一代 Fable 5 只成功 1 次。单轮耗时 Astra 约 2.5 分钟,Fable 5.1 要 6.8 分钟;按牌价估算单轮成本 0.94 美元对 2.12 美元,Robocurve 的图表把它标注为「完成率高 2.4 倍、便宜 2.3 倍」。\n\n评分不是非黑即白:人类评分员按「最高到达阶段」打 0 到 4 分,从「没有意图性接近」到「放到目标位置」,失败的试跑也会记录走到了哪一步。积木任务上 Astra 的平均阶段达到 3.95,几乎每次都走到终点;Fable 5.1 是 2.40,Fable 5 只有 1.30。\n\n## 输出 token 的差距比成功率更夸张\n\nAstra 单轮只输出约 2100 个 token,Fable 5.1 是 1.29 万,Fable 5 高达 1.92 万——用对手约六分之一的输出量,拿到了两倍以上的完成数。拼图任务上同样如此:2700 对 1.05 万,单轮成本 1.36 美元对 2.18 美元。\n\n## 精细插入:三个模型卡在同一步\n\n真正有意思的是更难的拼图插入:Astra 完成 2\u002F20,Fable 5.1 同样只有 2\u002F20,Fable 5 则一次未成。报告写道,Astra 能到达凹槽上方,然后卡在与 Fable 相同的最后一步。抓取、移动、对位都通过了,卡住的是把件按进去的动作。对一套只靠三个机位摄像头和本体状态做反馈的控制器来说,差的恐怕不是推理,而是纯视觉观测给不了的东西。\n\n## 这份评测该怎么读\n\nRobocurve 自己列出的局限值得原文引用:Astra 的试跑比 Fable 晚两天进行,没有交错安排;积木任务不在同一台机架上;评分员知道跑的是哪个模型,存在无意识偏袒的可能;Anthropic 的请求未启用 prompt 缓存,而 OpenAI 自动缓存了 Astra 约五分之一的输入且未折价——报告原话是,Astra 的成本「只会被高估」。\n\n所以结论要克制:在视觉抓取加粗放置这类任务上,Astra 拿到了又快、又便宜、又准的全面优势;但在精细插入这类操作上,前沿模型仍然集体不及格。这份评测的价值不在排名,而在它把下一个工程难题定位到了具体的一步:最后一段插入,还轮不到语言模型靠「多想」来解决。\n\n参考:Robocurve《GPT-6 Astra on robotic manipulation》 openai.robocurve.org\u002Fgpt-6-astra\u002F","https:\u002F\u002Fopenai.robocurve.org\u002Fgpt-6-astra\u002F","8f632b3d-848f-4235-9719-c766c15a7c9a",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"baf131c1-687a-49f4-87f6-4dd87c1c692f","gpt",{"id":19,"name":20,"slug":20,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":22,"name":23,"slug":23,"description":14,"color":14},"42e59a88-7795-47dc-a334-ef1e72c24347","openai",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"76b3be9e-11ff-41e1-8d36-d9c987d4052c","en","GPT-6 Astra on Robot Arms: 19\u002F20 vs 8\u002F20, All Stall on Insertion","Robocurve ran GPT-6 Astra vs Claude Fable 5\u002F5.1: 19\u002F20 vs 8\u002F20 on block-in-bowl at half the cost; both new models stalled at 2\u002F20 on insertion.","GPT-6 Astra, Claude Fable 5.1 and Fable 5 were each handed the same pair of robot arms and just two jobs: pick up the red block from the table and place it inside a bowl; grip a round puzzle piece by the knob at its center and insert it into the matching circular groove on the board. The independent evaluation shop Robocurve published the results on September 4, drawing an unusually clean capability line for the \"physical hands\" of frontier models, as a follow-up to their earlier Fable 5 versus 5.1 comparison.\n\n## 19\u002F20 vs 8\u002F20: a lopsided block task\n\nEach model ran 20 trials per task, 120 counted runs in total. On block-into-bowl, GPT-6 Astra completed 19 of 20, Claude Fable 5.1 managed 8, and the older Fable 5 just 1. Astra took about 2.5 minutes per trial to Fable 5.1's 6.8; estimated cost at list price was 0.94 USD per run against 2.12 USD, which Robocurve's chart labels as a 2.4x higher completion rate at 2.3x lower cost.\n\nScoring was not binary: a human grader scored each trial 0 to 4 by the highest stage reached, from \"no purposeful approach\" up to \"placed in its final position\", so failed runs still record how far they got. On the bowl task Astra's mean stage was 3.95, nearly always reaching the end; Fable 5.1 sat at 2.40 and Fable 5 at 1.30.\n\n## The output-token gap is more dramatic than the success rate\n\nAstra emitted roughly 2,100 output tokens per run, against 12,900 for Fable 5.1 and 19,200 for Fable 5 — about one-sixth the output for more than twice the completions. The puzzle task shows the same pattern: 2,700 vs 10,500 tokens, at 1.36 USD vs 2.18 USD per run.\n\n## Precision insertion: all three models stall at the same step\n\nThe harder puzzle insertion is where the report gets interesting: Astra completed 2 of 20, Fable 5.1 also 2 of 20, and Fable 5 none at all. Astra reaches the groove and then stalls at the same final step Fable does. Grasping, moving and positioning all pass; what stalls is pressing the piece home. For a controller working only from three camera views and proprioceptive state, the missing ingredient is probably not more reasoning but something pure visual observation cannot supply.\n\n## How to read this evaluation\n\nThe limitations Robocurve itself lists deserve quoting: Astra's trials ran two days after the Fable trials, not interleaved; the bowl comparison was not run on the same rig; the grader knew which model was running, leaving room for unconscious bias; Anthropic requests went out without prompt caching while OpenAI auto-cached about a fifth of Astra's input without applying a discount — in the report's own words, Astra's cost is \"if anything, overstated\".\n\nSo keep the conclusion measured: on vision-guided grasping and coarse placement, Astra holds a simultaneous faster, cheaper and more accurate advantage; on fine insertion, frontier models still collectively fail. The value of this evaluation is not the ranking but that it pinpoints the next engineering problem to one specific step: the final insertion is not something a language model solves by thinking harder.\n\nReference: Robocurve, \"GPT-6 Astra on robotic manipulation\" — openai.robocurve.org\u002Fgpt-6-astra\u002F","gpt-6-astra-robot-arm-benchmark","2026-09-07T19:13:54Z","2026-09-07T19:13:58.405143Z","2026-09-07T19:13:58.405153Z",true,"agent",179,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"0fd9ee7a-5b8f-49d2-9032-57f763de40e3","OpenAI 下一代模型 Astra 一口气破解 10 个数学难题:从 27 年未决的非 sofic 群到 46 年未动的高维球体堆积","openai-astra-ten-math-proofs-2026","2026-08-01T10:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"390c2437-4e4f-45ec-8270-67c5bfa4fa47","ChatGPT、Claude、Grok、Gemini 罕见同时下线,周四早晨全球 AI 集体失声","chatgpt-claude-grok-gemini-thursday-outage","2026-09-05T06:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"b5be4ce8-4a41-461c-9202-148e64fab329","GPT-6 Astra 系统卡:零日自用、对齐升 53%,CoT 可监控性反向下滑","gpt-6-astra-system-card-2026-monitorability","2026-09-04T03:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"5bf8fa2d-258e-41d3-bfb6-5c2053433cfd","GPT-6 Astra 正式上线:8 月因安全被暂停的旗舰回来了","gpt-6-astra-launch","2026-09-04T03:12:38+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"9e58d587-3c1b-44c5-ad36-daf23aeb42a2","微软叫停 tokenmaxxing:GitHub Copilot 默认切回 GPT-5.6 Sol,Parikh 设 token 预算","microsoft-token-budget-gpt-5-6-default","2026-09-03T00:30:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"6c5f1bc4-d877-483c-a0e9-70aff0e30dbe","微软内部 Ramp 账单:一名工程师 28 天烧掉 2.8 万美元 AI 费","microsoft-internal-ramp-ai-spending-28000-28-days","2026-08-31T03:00:00+00:00"]