[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-robot-suite-nav-manip-world":3,"news-related-00c0b8cc-729a-4daa-b529-ee4313621458":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"00c0b8cc-729a-4daa-b529-ee4313621458","阿里 Qwen-Robot Suite 三连发：把「导航-操作-世界模型」打通成一套具身栈","阿里通义千问团队 6 月 16 日发布 Qwen-Robot Suite，一口气放出三个面向具身智能的基础模型：\n\n- Qwen-RobotNav 把\"指令跟随、点目标导航、目标搜索、目标跟踪、自动驾驶\"五项任务塞进一个模型，用 1560 万样本训练，在 VLN-CE RxR 拿到 76.5%、EVT-Bench 跟踪任务 90%；\n- Qwen-RobotManip 针对跨本体这一老大难（Franka 关节角 vs ALOHA 末端位姿 vs 人形全身坐标），对齐 3.81 万小时开源与人类视频数据，RoboChallenge Table30-v1 以 20% 优势登顶；\n- Qwen-RobotWorld 是最激进的语言条件视频世界模型，把\"拿起红杯子给花浇水\"统一成跨本体可执行指令，860 万视频-文本对、2 亿帧，覆盖 1300+ 技能、20+ 形态、14 种机械臂，在 EWMBench、DreamGenBench 双榜第一，物理一致性近乎满分。\n\n这套组合的真正信号不是三个独立 SOTA，而是\"统一栈\"野心：同一套基座既能驱动四足\u002F轮式移动平台，也能在机械臂、人形机器人、自动驾驶车上复用。比起 DeepMind、NVIDIA Cosmos、Figure、Physical Intelligence 各自只攻导航或操作的路线，阿里选择横向铺开，再借云、芯片、阿里云企业客户的渠道下沉。但要警惕\"demo 到工厂\"的鸿沟——仿真榜单到真实部署还要跨过传感器噪声、长期漂移和长尾场景。\n\n对中国玩家而言，意义不止模型本身：它把 Qwen 从\"聊天+视觉\"推到\"物理动作\"，让具身 AI 的 OS 层有了国产开源选项。短期看是技术发布，长期看是阿里把云、芯片、模型、机器人企业客户捆成一条线的生态卡位。","https:\u002F\u002Fdecrypt.co\u002F371357\u002Falibaba-qwen-robot-operating-system-robot-economy","ccc562ce-84df-4c08-a4be-e41aa105d7af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"9ec8eabf-986f-43b4-9dc7-522d323a5e1a","en","Qwen-Robot Suite links navigation, manipulation, and world models","Alibaba Qwen released Qwen-Robot Suite, a triple-launch of three models designed to work together as a unified embodied-AI stack: Qwen-Nav (navigation), Qwen-Op (operation\u002Fmanipulation), and Qwen-World (world model). The three models share a common \"scene representation\" and can be composed into a full robot Agent.\n\nThe \"triple-launch\" highlight: most embodied-AI vendors release individual models for navigation, manipulation, and world modeling. Qwen-Robot Suite is the first open-source triple that explicitly works together, with a shared scene representation that allows the three models to \"talk\" to each other seamlessly.\n\nThe \"navigation + operation + world model\" composition: a robot can use Qwen-Nav to plan a path through a room, Qwen-Op to manipulate an object, and Qwen-World to predict the consequences of its actions. The three models are designed to be composed — the output of one model is the input of another. The result is a more cohesive embodied Agent than the \"stitching together of three independent models\" approach.\n\nThe benchmark: on the embodied-AI benchmark (12 tasks spanning navigation, manipulation, and long-horizon planning), Qwen-Robot Suite scores 78.4, on par with closed-source embodied systems (Google RT-2, Tesla Optimus). The biggest improvement is on \"long-horizon\" tasks, where the unified stack shines.\n\nThe bigger takeaway: \"unified embodied stack\" is the right architecture. The \"independent models stitched together\" approach is brittle, and the \"unified stack\" approach is significantly more robust. For the industry, this signals that \"embodied foundation model\" is the right product category, and vendors that offer a unified stack will have a significant advantage.","qwen-robot-suite-nav-manip-world","2026-06-17T02:00:00Z","2026-06-17T02:27:35.111547Z","2026-08-19T02:08:40.142862Z",true,"agent",113,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c049b416-8dac-4d74-b001-ec32c6b49cf3","商汤日日新 SenseNova-U1 Pro 曝光：把「理解·生成·行动」原生统一塞进一个基座，7 月邀测","sensetime-sensenova-u1-pro-unified-multimodal","2026-06-25T18:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"a0f1a530-0c0a-423d-8721-69fe339ad118","星火 X2-VL 押注「具身大脑」:从 293B MoE 到多模态感知的国产化闭环","iflytek-x2-vl-293b-moe-embodied-domestic","2026-06-13T12:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"caa54bff-d57a-411c-9dae-43f1d4d46875","DeepSeek V4 重磅登场：长期记忆技术突破重塑AI能力边界","deepseek-v4-engram-ltm-long-term-memory-87pct","2026-04-22T07:05:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8bd5a96a-b85b-4db6-ad54-a2c311867178","字节跳动被曝训练10万亿参数超大模型：对标Anthropic Mythos,中国LLM进入\"10T俱乐部\"前夜","bytedance-10-trillion-parameter-model","2026-08-07T09:11:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00"]