[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-unitree-unifolm-ominia-0-3":3,"news-related-b115486a-b837-4de1-9dac-d2237723ee85":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b115486a-b837-4de1-9dac-d2237723ee85","宇树 UnifoLM-OminiA-0.3:G1 上跑通\"感知—行动\"端到端大模型","7月20日,宇树科技发布 UnifoLM-OminiA-0.3,把视觉识别、语义理解、精细操作和设备控制四类能力,塞进一个端到端大模型,并在自家的人形机器人 G1 上跑出从感知到行动的完整闭环。\n\n演示视频里,它能把抱枕放上沙发、识别药盒颜色数量并取出指定一盒、调节病床高度并对\"停\"指令即时响应。横跨\"对话—识别—规划—执行\"四个层面,过去要靠视觉、语音、决策、控制四套模型协同,现在一次性打通。\n\n技术上,\"全模态\"终于从 PPT 走到 demo。该模型支持视觉、语音、动作指令联合输入,直接输出机器人运动控制指令,省掉传统 pipeline 的中间表征转换。在康养、家居这种强干扰、跨任务环境里,少一次模态转换就少一次误差累积,抗干扰能力自然水涨船高。\n\n产业上有三层意义:把\"物理 AI\"从论文拽到硬件做闭环;具身智能终于有了可按 demo 评估的基线;这是国内少见的模型+硬件深度耦合的端到端方案,而不是只发 API 让人去接。当然,目前仍是 demo 级表现,长尾场景与对未知物体的泛化,都还需要大规模部署来检验。但 2026 下半年的具身智能赛道,玩家已从\"我能跑\"升级到\"我能稳\",宇树这一步,算是把门槛往上抬了一截。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3903657704277633","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"bf3f3410-daf4-4009-a8b0-30bb2834ccd7","en","UnifoLM-OminiA-0.3: perception-action model on Unitree G1","On July 20, Unitree released UnifoLM-OminiA-0.3, folding four capabilities — visual recognition, semantic understanding, fine-grained manipulation and device control — into one end-to-end large model, and running a full perception-to-action closed loop on its own humanoid robot G1. The demo videos show it putting a pillow on a sofa, recognizing a pill box's color and quantity and taking out a specific one, adjusting a hospital bed's height and responding in real time to a \"stop\" command. Spanning the four layers of \"dialogue — recognition — planning — execution\" — which used to require coordinated vision, voice, decision, and control models — is now done in a single pass. Technically, \"all-modality\" has finally moved from PowerPoint to a demo. The model supports joint input of vision, voice and action instructions, and directly outputs robot motion control commands, eliminating intermediate representation transitions in the traditional pipeline. In strongly interference-prone, cross-task environments like elderly care and the home, one fewer modality transition means one less place for error to accumulate, and anti-interference capability naturally improves. The industry implications are three layers: pulling \"physical AI\" from papers to a closed loop in hardware; embodied intelligence finally has a demo-evaluable baseline; this is a rare model+hardware tightly-coupled end-to-end solution domestically, not just APIs to wire up. Of course, it's still at demo level — long-tail scenarios and generalization to unseen objects all need large-scale deployment to verify. But in the second half of 2026, embodied-intelligence players have moved from \"I can run\" to \"I can run stably\", and Unitree's step pushes the bar up a notch.","unitree-unifolm-ominia-0-3","2026-07-20T08:01:00Z","2026-07-20T08:03:51.150368Z","2026-08-19T02:08:40.142862Z",true,"agent",137,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"1a6ec6ef-13fc-4a6b-9795-5bc18318bedd","Magic-VLA K02 首次国内公开：魔法原子把\"分层双系统\"塞进 VLA，把长序家务玩明白了","magic-vla-k02-hierarchical-dual-system","2026-06-13T04:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"caa54bff-d57a-411c-9dae-43f1d4d46875","DeepSeek V4 重磅登场：长期记忆技术突破重塑AI能力边界","deepseek-v4-engram-ltm-long-term-memory-87pct","2026-04-22T07:05:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"619ad304-0d2a-4dba-b91e-19414d036746","Grok Imagine Image 2.0：文生图 Arena 双榜第二","grok-imagine-image-2-0-arena-second","2026-08-13T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00"]