[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mistral-robostral-navigate-8b":3,"news-related-aebb8a81-713a-40b7-84dd-03213a6a808c":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"aebb8a81-713a-40b7-84dd-03213a6a808c","Mistral Robostral Navigate:8B 视觉语言模型只靠单目 RGB 在 R2R-CE 反超多传感器基线","Mistral 在 7 月 8 日发布 Robostral Navigate——他们的第一个具身导航模型,定位是 8B 视觉语言模型,输入只有一路普通 RGB 相机画面加一句自然语言指令,要把'看着摄像头走回家'做明白。过去把导航做好的方案几乎都靠 LiDAR、深度相机或多机位,而 Mistral 选择只喂一个普通摄像头。基准成绩最有说服力:在 R2R-CE 验证 unseen 上 Robostral Navigate 拿到 76.6% 成功率,比最强单相机基线高 9.7 个点,比依赖深度或多机位的最强系统还高 4.5 个点。Mistral 没有套壳现成开源 VLM,而是从一个自家为 pointing、counting、object localization 微调的视觉语言基座出发,把'知道东西在哪'自然延伸到'知道下一步往哪走',训练数据全部在仿真中造,共 400K 轨迹、6K 场景。工程上两个数字值得拎出来:prefix-caching 加 tree-based attention masking 把训练 token 量压缩 22 倍,原本几个月的训练被压到几天;CISPO 在线强化学习在监督训练之上又带来 3.2 个百分点独立增益,目前还看不到天花板。8B 模型不挑机器人形态,轮式、腿式、飞行都能跑,同一段指令可以穿过正常运转的办公区走完整条长链路任务。放在 Mistral 整体路线上,Robostral Navigate 是继 6 月 Physics AI 之后又一次'AI 进入物理世界'的延伸。导航被普遍视为通用机器人的基础能力,模型小、传感器依赖少、训练经济,意味着研究阶段产物向工厂、配送、酒店等真实场景迁移的门槛被显著压低了。","https:\u002F\u002Fmistral.ai\u002Fnews\u002Frobostral-navigate\u002F","2436174c-644b-4a65-9a98-e7a3b705569a",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"1d8cd467-f418-4564-a819-b77afba1c02b","en","Robostral Navigate: 8B VLM beats multi-sensor on monocular RGB","Mistral on July 8 released Robostral Navigate — their first embodied-navigation model, positioned as an 8B vision-language model with input of only a regular RGB camera feed plus a natural-language instruction, to \"watch the camera and walk home\". Past solutions for good navigation almost all relied on LiDAR, depth cameras, or multiple viewpoints, and Mistral chose to feed in only a single ordinary camera. The benchmark numbers are most convincing: on the R2R-CE validation unseen split, Robostral Navigate gets 76.6% success rate, 9.7 points above the strongest single-camera baseline, and 4.5 points above the strongest system that depends on depth or multiple viewpoints. Mistral didn't wrap an existing open-source VLM, but started from a vision-language base it had itself fine-tuned for pointing, counting, and object localization, naturally extending \"knowing where things are\" to \"knowing where to go next\", with all training data synthesized in simulation, totaling 400K trajectories and 6K scenes. Two engineering numbers are worth flagging: prefix-caching plus tree-based attention masking compresses training token volume by 22×, compressing what would have been several months of training into a few days; CISPO online reinforcement learning adds another independent 3.2 percentage point gain on top of supervised training, with no clear ceiling yet. The 8B model isn't picky about robot form, running on wheeled, legged, and aerial platforms, with the same instruction set able to traverse a normally-functioning office area and complete a full long-chain task. In Mistral's overall roadmap, Robostral Navigate is another extension of \"AI enters the physical world\" after the June Physics AI. Navigation is widely seen as the foundational capability of general robots, with small models, low sensor dependence, and economical training meaning the threshold for migration from research-stage products to real scenarios like factories, delivery, and hotels is significantly lowered.","mistral-robostral-navigate-8b","2026-07-09T14:15:00Z","2026-07-09T14:13:20.459778Z","2026-08-19T02:08:40.142862Z",true,"agent",97,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"cec093b3-47fe-490f-b0d7-57c07ba19758","Wan-Streamer v0.1：单模型端到端 550ms 实时交互","wan-streamer-v0-1-550ms-realtime","2026-06-23T18:01:03+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"bbf1a404-1f46-45f9-a61e-b6e210d28878","智元罗剑岚：把「部署-数据-迭代」打成飞轮，比堆参数更像具身智能的 Scaling Law","zhiyuan-luo-jianlan-flywheel-embodied","2026-06-17T06:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7fae0753-96df-4d4b-9ebd-cf0509c08b37","LLM架构演进：从规模竞赛到效率优化的范式转变","llm-architecture-evolution-2026-moe-multimodal-turboquant","2026-04-25T04:12:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"c9ba6037-e8c9-4007-98e5-32af59d92839","百度一镜 WAIC 首发数字人视频播客方案，文心多模态能力再突破","baidu-yijing-waic-digital-podcast","2026-07-19T08:02:00+00:00"]