[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-gigabrain-0-7-embodied-vla-open-source":3,"news-related-e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","GigaAI 团队开源具身基座模型 GigaBrain-0.7:三系统架构统一理解、预测与动作,在 3.73 万小时异构数据上预训练,放出 3.5B Base 权重、训练代码与样例数据,Apache 2.0 协议;作者报告对比 π0.5 在零样本与任务成功率上显著提升。","具身智能的竞争正在从「谁的模型更大」转向「谁的数据更厚」。8 月 25 日,GigaAI 团队放出具身基座模型 GigaBrain-0.7 的训练代码、3.5B Base 权重和样例数据,全部以 Apache 2.0 开源;技术报告 8 月 17 日先行发布,GitHub 仓库已攒下 2.6k star,论文登上 Hugging Face Daily Papers 榜首、拿下 90 个 upvote。\n\n## 三系统:把世界模型搬进决策回路\n\nVLA(vision-language-action)模型已是通用具身智能体的主流范式,在结构化场景能完成长程任务,但架构红利与数据异构性能走多远仍是开放问题。GigaBrain-0.7 的答案是三系统架构:把理解、预测、动作统一进一个模型。官方亮点里最关键的一条是 System-3——把世界模型接入机器人的实时决策回路,先模拟评估、再执行动作;配合双金字塔框架,单个预训练模型可直接执行多类任务,即 One Model, Many Tasks。\n\n## 3.73 万小时异构数据 + 一阶段对齐\n\n预训练规模是另一个锚点:超过 37,000 小时(README 口径 37.3k hours)异构具身数据,覆盖真机、UMI、第一人称、仿真、世界模型生成五类来源。对齐阶段采用一阶段对齐训练,联合优化视觉-语言理解与多本体动作生成。对比对象包括自家 GigaBrain-0 系列与此前 SOTA 模型 π0.5,作者报告在零样本能力、语言条件指令遵循、后训练任务成功率上显著提升;在自建 Maker H01 平台与主流本体上覆盖家庭和工业场景。\n\n## 开源的不只是权重\n\nGitHub 仓库给出一条完整的后训练路径:数据用 LeRobot 格式,提供 norm stats 脚本与 delta mask 机制,内置 8 类本体 ID(AgileX、Agibot G1 及灵巧手、UMI、EgoDex、H01 等),训练栈钉死在 PaliGemma2 环境(三个 giga-* 组件全锁 1.1.0)。部署侧给了 AgileX Cobot Magic 与 Maker H01 两套 profile,H01 客户端可直接跑在 Jetson 上。仓库还披露了受控十步冒烟测试损失:pick-and-place 下降 51.27%,push-buttons 下降 31.97%——冒烟数据写进 README,在具身开源项目里不算常见。\n\n## 自报成绩与待验证清单\n\n官方亮点称其在 RoboColiseum 四项评测全部第一、Maker H01 任务成功率大幅领先,并放出 20 分钟以上、跨越 10 多个任务的一镜到底实机演示。但要泼冷水:这些数字全部出自团队自报,RoboColiseum、RoboTwin2.0、EBench 的 benchmark 代码和 VLM 评测代码还在 TODO 列表里,第三方复现暂无门路。作为参照,该团队 2026 年 2 月的 GigaBrain-0.1 曾登上 RoboChallenge 榜单第一,还将在 CVPR 2026 主办 GigaBrain Challenge 三条赛道(RoboTwin 仿真、GigaWorld 世界模型、RoboChallenge 真机)。\n\n## 所以呢\n\n具身基座的卷法已经清晰:拼数据小时数、拼本体覆盖、拼世界模型在环。GigaBrain-0.7 三件事一次做完并全部开源,对想在自有机型上做后训练的团队是现成起点;但 benchmark 代码落地之前,「四项全一」这类成绩不妨让子弹再飞一会儿。\n\n原始出处:arxiv.org\u002Fabs\u002F2608.15875;代码:github.com\u002Fopen-gigaai\u002Fgiga-brain-0","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.15875","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"9789edea-e75c-45bd-b2ad-52e2ce2c569c","en","GigaBrain-0.7 open-sources a 3.5B embodied VLA foundation","GigaAI open-sources GigaBrain-0.7: a three-system embodied VLA pretrained on 37.3k hours, shipping 3.5B weights, code and data under Apache 2.0.","Embodied AI competition is shifting from model size to data depth. On August 25, the GigaAI team released training code, 3.5B Base weights and sample data for GigaBrain-0.7, all under Apache 2.0; the technical report landed on August 17, the GitHub repo has already collected 2.6k stars, and the paper topped Hugging Face Daily Papers with 90 upvotes the same day.\n\n## Three systems: a world model inside the decision loop\n\nVision-language-action (VLA) models are now the dominant paradigm for generalist embodied agents, completing complex long-horizon tasks in structured settings. Whether architecture still has headroom, how heterogeneous data can scale, and how far generalization reaches remain open questions. GigaBrain-0.7's answer is a three-system architecture that unifies understanding, prediction and action in one model. The most notable entry in the official highlights is System-3: it plugs a world model into the robot's real-time decision loop, letting it simulate and evaluate before acting — different from the common practice of using world models merely as offline data generators. Together with a dual-pyramid framework, a single pretrained model handles many task types out of the box: One Model, Many Tasks.\n\n## 37.3k hours of heterogeneous data plus one-stage alignment\n\nPretraining scale is another anchor of this release: over 37,000 hours (37.3k per the README) of heterogeneous embodied data, spanning real-robot, UMI, egocentric, simulation, and world-model-generated sources. The alignment stage uses one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation. Compared with the team's own GigaBrain-0 series and prior state-of-the-art models including π0.5, the authors report substantial gains in zero-shot capabilities, language-conditioned instruction following, and post-training task success rates; on the in-house Maker H01 platform and mainstream embodiments, coverage spans both home and industrial scenarios.\n\n## What gets open-sourced goes beyond weights\n\nThe GitHub repo lays out a complete post-training path: data in LeRobot format, norm-stats computation scripts with a delta-mask mechanism, eight built-in embodiment IDs (from AgileX and Agibot G1\u002Fdexterous hands to UMI, EgoDex and H01), and a training stack pinned to the PaliGemma2 environment (Python 3.11.10, with giga-datasets, giga-train and giga-models all locked at 1.1.0). On the deployment side, two profiles are provided for AgileX Cobot Magic and Maker H01, and the H01 client runs directly on Jetson. The repo even discloses loss numbers from a controlled ten-step smoke test: pick-and-place down 51.27%, push-buttons down 31.97%. Publishing smoke-test numbers in a README is not a common practice among embodied open-source projects.\n\n## Self-reported results and the to-be-verified list\n\nThe official highlights claim first place in all four evaluations on RoboColiseum (a benchmark with strong sim-to-real alignment), a wide lead in task success rates on Maker H01, and a single-take demo of over 20 minutes spanning more than 10 tasks, described as an industry first. A dose of skepticism: all these numbers are self-reported, and the benchmark code for RoboColiseum, RoboTwin2.0 and EBench, plus the VLM evaluation code, still sit in the TODO list — third-party replication has no entry point yet. For context, the team's GigaBrain-0.1 took first place on the RoboChallenge leaderboard in February 2026, and the team will host the GigaBrain Challenge at CVPR 2026 with three tracks: RoboTwin (simulation), GigaWorld (world model) and RoboChallenge (real robot).\n\n## So what\n\nThe rules of embodied foundation-model competition are now clear: compete on data hours, embodiment coverage, and world models in the loop. GigaBrain-0.7 does all three at once and open-sources everything — a ready starting point for teams wanting post-training on their own robots. But until the benchmark code lands, let the \"first in all four\" claims fly for a while before you buy them.\n\nSource: arxiv.org\u002Fabs\u002F2608.15875; code: github.com\u002Fopen-gigaai\u002Fgiga-brain-0","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00Z","2026-08-26T23:09:46.211281Z","2026-08-26T23:09:46.211290Z",true,"agent",12,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"dccdd14b-babe-4883-b0ec-0cf5b1d85018","MiniMax Music 3.0 把「5 分钟完整歌曲」开源:Hybrid-LM + Flow-VAE 让音乐生成跨过录音室门槛","minimax-music-3-5min-song-open-source","2026-08-15T00:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"dd011592-f0aa-4d45-9229-56311232f9f0","OpenMOSS 开源 MOSS-VL-Realtime：11B 实时流视频 VLM","openmoss-vl-realtime","2026-07-19T03:55:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"d8c62859-54c8-4069-b776-8e623ca03029","Cohere Transcribe Arabic：2B 开源 ASR 登顶，WER 低 Whisper 11 点","cohere-transcribe-arabic","2026-07-16T04:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"c677680a-a08f-420c-8107-7816827707a2","小米开源 Xiaomi-Robotics-U0：38B 具身生成统一 Tokenizer","xiaomi-robotics-u0","2026-07-14T22:10:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"413b7c1f-e12b-4571-8076-8b5511360bbd","AlayaWorld开源:用3D缓存+DMD蒸馏破解长时视频世界模型一致性难题","alayaworld-long-video","2026-07-14T10:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"bfa4db7e-f3a2-4da6-9823-faa6ccef2274","高德 ABot-World Studio 把世界模型压进消费级 GPU：单卡可跑 + 全开源的另一种解法","gaode-abot-world-studio","2026-07-14T04:01:00+00:00"]