[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-physbrain-1-5-open-embodied-base":3,"topics-all":38,"news-related-d055ddb8-4d82-4523-99b7-39c5f77e2ff7":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"d055ddb8-4d82-4523-99b7-39c5f77e2ff7","PhysBrain 1.5 开源：8B 具身基座 28 项评测均分 72.5，官方称追平 GPT-6-Astra","DeepCybo 联合中关村团队开源 PhysBrain 1.5：基于 Qwen3-VL 的 8B 具身基座模型，把环境理解、动作生成、未来状态预测统一进一套自回归 token 体系，28 项具身基准均分 72.5，官方称 14 项居被评测开源模型首位。","9 月 14 日，arXiv 上出现了一篇 54 位作者的技术报告：DeepCybo 联合中关村学院、中关村人工智能研究院的团队，发布了具身智能基座模型 **PhysBrain 1.5**，模型权重（2B 与 8B 两个规格）与评测工具链同步开源在 Hugging Face 与 GitHub 上。论文标题把野心写得很直白——\"从视觉语言模型到物理基础模型\"。\n\n## 一个模型，三种能力\n\n具身智能目前的常见做法，是给 VLM 外挂各种专用头：理解用一套模型，动作生成再接一个策略网络，世界模型又是另一个系统。PhysBrain 1.5 的路线恰好相反——理解、动作、预测共享同一个自回归骨架，用统一的 next-token prediction 目标联合优化，不做任务特定的输出头。\n\n具体拆开看三层能力：\n\n- **具身理解**：视觉空间感知、3D 与多视角推理、具身规划、指向与可供性定位、视觉轨迹推理；\n- **动作生成**：用 ActionPiece token 预测末端执行器轨迹块，一套动作码本跨控制配置、跨机器人本体复用；\n- **未来状态预测**：输出空间对齐的 RGB 图像、深度图与机器人掩码，直接\"想象\"动作执行后的场景。\n\n这个设计对应论文里强调的\"物理交互回路\"：观察驱动推理与动作，动作改变世界，更新后的观察再喂回下一轮交互。语言、动作、视觉三类 token 被编码为离散序列，在同一个 backbone 里联合训练。\n\n## 数据与训练：人类交互视频当预训练监督\n\n训练配方分两段。预训练阶段的具身监督**完全来自人类交互视频**——以任务为中心切片，把语义与空间上下文、恢复出的运动轨迹、后续观察配成对。之后通过监督微调适配，混合了人类演示、真实机器人轨迹与仿真经验三类数据。\n\n底座选择上，PhysBrain 1.5 构建于 Qwen3-VL 之上。从通用 VLM 出发而不是从零训练，也解释了论文强调的另一点：模型在具身能力之外**保留了通用多模态能力**，没有为了刷具身分数而丢掉底座的基本盘。\n\n## 跑分：28 项基准均分 72.5\n\n技术报告在 28 项具身理解基准上做了评测，覆盖五大类：基础视觉空间感知、空间与多视角理解、具身认知推理与规划、空间定位与可供性、视觉轨迹推理。\n\n- PhysBrain 1.5-8B 综合均分 **72.5**（0-100 分制，28 项无加权平均）；\n- 在被评测的开源模型中，**14 项基准排名第一、10 项排名第二**（含并列）；\n- 论文同时给出对比：表现与 GPT-6-Astra、Gemini 3.6 Flash 等领先闭源模型**相当**。\n\n需要说明的是，这些成绩是团队技术报告的自报结果，评测虽配套开源了 PhysBrainEvalKit 工具链，第三方独立复现还有待时间验证；\"追平闭源前沿\"的对比同样出自官方评测框架内。2B 版本均分 66.6，官方仅作参考、未参与排名。\n\n## 所以呢\n\n具身基座赛道正处在\"谁的 token 体系能统一感知-决策-预测\"的竞赛里。PhysBrain 1.5 的答案是把三类输出压进同一套离散 token，用人类视频当免费监督——这条路线如果被独立复现坐实，\"给 VLM 外挂专用头\"的拼装式架构可能会加速退场。对研究者而言，2B\u002F8B 权重、评测工具链、技术报告全套开放，复现门槛已经压到很低；下一个值得盯的问题，是 ActionPiece 这套跨本体动作码本能不能在更多构型的真实机器人上站住。\n\n参考：arXiv:2609.14973（arxiv.org\u002Fabs\u002F2609.14973）· GitHub: DeepCybo-PhysAI\u002FPhysBrain-1.5","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.14973","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"164b44d7-ecdb-4b0d-8367-6e3fd5e58c66","en","PhysBrain 1.5: Open 8B Embodied Base, 72.5 Avg on 28 Benchmarks","DeepCybo open-sources PhysBrain 1.5: an 8B embodied base on Qwen3-VL, 72.5 avg across 28 embodied benchmarks and 14 firsts among open models.","A 54-author technical report landed on arXiv on September 14: DeepCybo, together with the Zhongguancun Academy and the Zhongguancun Institute of Artificial Intelligence, released **PhysBrain 1.5**, an embodied foundation model whose weights (2B and 8B) and evaluation toolkit were open-sourced in the same move on Hugging Face and GitHub. The paper's title states the ambition plainly — \"From Vision-Language Models to Physical Foundation Models.\"\n\n## One Model, Three Capabilities\n\nThe common pattern in embodied AI today bolts specialized heads onto a VLM: one model stack for understanding, a policy network for action generation, yet another system as the world model. PhysBrain 1.5 takes the opposite route — understanding, action, and prediction share a single autoregressive backbone, jointly optimized under one next-token prediction objective, with no task-specific output heads.\n\nThe three capability layers:\n\n- **Embodied understanding**: visual-spatial perception, 3D and multi-view reasoning, embodied planning, pointing and affordance grounding, visual-trace reasoning;\n- **Action generation**: end-effector trajectory chunks predicted via ActionPiece tokens, with a unified action codebook reused across control configurations and robot setups;\n- **Future-state prediction**: spatially aligned RGB imagery, depth maps, and robot masks that \"imagine\" the scene after an action executes.\n\nThe design maps to the \"physical interaction loop\" the paper keeps returning to: observations guide reasoning and action, actions change the world, and updated observations feed the next round. Language, action, and visual tokens are encoded as discrete sequences and trained together inside one backbone.\n\n## Data and Training: Human Interaction Videos as Pre-training Supervision\n\nThe recipe has two stages. Pre-training draws its embodied supervision **entirely from human interaction videos** — task-centered episodes pairing semantic and spatial context with recovered motion and subsequent observations. The model is then adapted through supervised fine-tuning over a mixture of human demonstrations, real robot trajectories, and simulated experience.\n\nAs for the base: PhysBrain 1.5 is built on Qwen3-VL. Starting from a general VLM rather than from scratch also explains another claim the paper stresses — the model **retains general multimodal capabilities** alongside its embodied skills, rather than trading them away for benchmark points.\n\n## Scores: 72.5 Average Across 28 Benchmarks\n\nThe technical report evaluates embodied understanding across 28 benchmarks in five categories: foundational visual-spatial perception; spatial and multi-view understanding; embodied cognition, reasoning, and planning; spatial grounding, pointing, and affordance; and visual-trace and trajectory reasoning.\n\n- PhysBrain 1.5-8B scores **72.5 overall** (0–100 scale, unweighted mean across the 28 benchmarks);\n- Among the evaluated open-source models it ranks **first on 14 benchmarks and second on 10** (ties included);\n- The paper also draws the comparison: performance on par with leading proprietary models such as **GPT-6-Astra and Gemini 3.6 Flash**.\n\nOne caveat worth stating: the \"72.5 average, 14 open-source firsts\" figures are self-reported results from the team's technical report. The PhysBrainEvalKit evaluation toolkit is open-sourced alongside, but independent third-party reproduction remains to be seen; the \"on par with closed-source frontier\" comparison likewise sits inside the official evaluation frame, with no independent cross-vendor benchmark yet. The 2B variant scores 66.6 and is listed for reference only, excluded from ranking.\n\n## So What\n\nThe embodied-base race right now is about whose token system can unify perception, decision, and prediction. PhysBrain 1.5's answer compresses all three output types into one discrete token vocabulary and uses human videos as free supervision — if this route holds up under independent reproduction, the bolt-on architecture of \"VLM plus specialized heads\" could fade faster than expected. For researchers, the 2B\u002F8B weights, the evaluation toolkit, and the full technical report are all open, pushing reproduction costs to a low; the next question worth watching is whether the ActionPiece cross-embodiment action codebook stands up on real robots of more morphologies.\n\nReferences: [arXiv:2609.14973](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.14973) · [GitHub: DeepCybo-PhysAI\u002FPhysBrain-1.5](https:\u002F\u002Fgithub.com\u002FDeepCybo-PhysAI\u002FPhysBrain-1.5)","physbrain-1-5-open-embodied-base","2026-09-16T21:07:24Z","2026-09-16T21:07:29.686616Z","2026-09-16T21:07:29.686629Z",true,"agent",28,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"17006864-46a5-405c-a8cc-24507bbc5e37","YuE2-3B 开源:乐谱可编辑的音乐生成,官方基准反超 Suno v5","yue2-3b-editable-music-generation","2026-09-10T13:20:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"b6b9f5c8-0d71-4288-8782-0284fccfca8f","商汤 SenseNova-Vision：把「检测\u002F分割\u002F深度估计」统统塞进同一个生成式多模态基座","sensetime-sensenova-vision","2026-07-08T10:15:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"a36d9d97-42de-4c87-88e9-cdc173b9ab4b","VLX-Seek 1.5 把端侧具身感知切成 0.6B\u002F3B\u002F10B 三档：用 None 输出压住目标幻觉","vlx-seek-1-5","2026-07-06T02:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"f9bf6e21-2e8a-4571-ab7d-a4dba727b72a","ViiTorVoice-NAR：把 TTS 的「改一句重录」变成「改一词局部合成」","viitor-voice-nar-local-tts","2026-07-02T14:15:00+00:00"]