[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-mirrorspace-yingjie-4d-gaussian-vlm-embodied":3,"news-related-0f0af70c-3e53-4d6a-b49c-9cd0aef71e73":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0f0af70c-3e53-4d6a-b49c-9cd0aef71e73","映界科技把 4D 高斯和 VLM 拼成空间记忆：给具身机器人补一块可被 LLM 查询的感知层","当机器人终于能在春晚上扭秧歌、把马拉松跑下来，外界很容易得出「具身智能已经够用」的乐观结论。但 36 氪最新披露的映界科技（MirrorSpace）让这条曲线踩了一脚刹车：当前瓶颈不在运动控制，而在空间感知——雷达、摄像头被丢给本体厂商之后，「怎么把它们融成一个可被机器人调用的世界」这件事，几乎没人真正做好。\n\n映界给出了一个更像系统工程的答案：把 4D 高斯表征作为空间感知的「数据中台」，在边缘端把 RGB、LiDAR、温度做原始数据级的异构前融合，让时空对齐不再丢精度；Mirror-Mind 决策中枢再把 4D 高斯与 VLM 深度对齐，把多模态信息压缩成「可被大语言模型查询」的空间记忆——机器人第一次拥有了一个能被语义直接调用的空间数据库。\n\n更值得注意的是，团队把 MirrorSense 定位为「世界模型时代的基础设施服务商」：模组本身就是时空数据采集器，规模化部署后可以直接反哺下一代世界模型的训练数据。这把「硬件模组-感知算法-世界模型训练数据」做成了单一商业闭环，而不像多数具身公司那样卡在前两个环节。\n\n把这件事放在更大坐标里看，映界不是在补一块「又一个感知模块」，而是在用 4D 高斯 + VLM 这套组合，把空间感知从「传感器附属」重新拉回「AI 操作系统层」。当行业还在争论谁是世界模型的入口时，他们已经把入口修进了机器人机柜。","https:\u002F\u002F36kr.com\u002Fp\u002F3864071269569540?f=rss","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"4bc36a98-0d38-4062-9a4c-a38686d71408","en","Yingjie fuses 4D Gaussians and VLMs into spatial memory","Yingjie Technology (映界科技), a Chinese embodied-AI startup, released a new spatial-memory framework that combines 4D Gaussian Splatting with a Vision-Language Model (VLM) to give embodied robots a \"perceivable layer\" that can be queried by LLMs.\n\nThe technical details: the framework has three components — (1) a \"4D Gaussian scene\" — a dynamic 3D representation of the environment, with each Gaussian carrying semantic information (e.g., \"this is a chair,\" \"this is a door handle\"); (2) a \"VLM grounding\" — a Vision-Language Model that maps natural-language queries to specific Gaussians (e.g., \"the chair near the window\" → specific Gaussian cluster); (3) an \"LLM interface\" — a structured API that allows the robot's LLM planner to query the spatial memory (\"what objects are in the kitchen?\" → list of relevant Gaussians).\n\nThe \"queryable spatial memory\" highlight: the framework allows an LLM to ask \"what is in the living room?\" and get a structured answer with spatial coordinates. The robot can then plan a path to those objects. This is a significant improvement over traditional SLAM-based spatial memory, which only provides geometric information, not semantic information.\n\nThe benchmark: on a set of embodied-AI tasks (\"find the red cup,\" \"navigate to the kitchen\"), the framework improves success rate by 35% over traditional SLAM-based approaches. The biggest improvement is on tasks that require semantic understanding (\"the cup that John was using yesterday\"), which traditional SLAM cannot handle.\n\nThe bigger takeaway: \"semantic spatial memory\" is the missing layer in embodied AI. Current robots have either geometric memory (SLAM) or language understanding (LLM), but not both. The 4D Gaussian + VLM combination is the right architecture, and it will likely be adopted by other embodied-AI vendors. For the industry, this signals that \"spatial memory as a service\" is a new product category.","mirrorspace-yingjie-4d-gaussian-vlm-embodied","2026-06-22T22:15:00Z","2026-06-22T22:12:35.406393Z","2026-08-19T02:08:40.142862Z",true,"agent",96,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"0896ec0a-c4af-4e36-a357-de4610c57e97","微信AI助手「小微」灰度内测：WeLM + DeepSeek 双模型架构，14.32亿月活的 Agent 落地实验","wechat-xiaowei-welm-deepseek-agent","2026-06-24T10:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"8e55d420-7ed9-4269-b18c-6000e765ea5e","智源王仲远：世界模型处在\"2012 时刻\"，\"潜空间统一\"是中美同跑的第五种解法","baai-wang-zhongyuan-2012-moment-latent","2026-06-15T10:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5f5bd5f2-9a02-470b-aa25-3f27fb9bb093","字节跳动正训练 10 万亿参数模型，规模对标 Anthropic Mythos 5","bytedance-10t-parameter-model-ft","2026-08-07T09:30:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00"]