[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-ram-zhejiang-cuhk-3d-spatial-science-robotics":3,"topics-all":36,"news-related-3b7833c1-4213-4155-9125-df436adf96d7":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"3b7833c1-4213-4155-9125-df436adf96d7","大模型机器人的空间盲区被攻破：RAM模型让机器人真正看懂三维世界","视觉语言大模型（VLM）很强，但有一个致命短板：它们本质上是在看二维图像，而真实世界的机器人操作需要精确的三维空间感知——物体在哪里、朝向如何、距离多远、能否被抓取。这个gap一直困扰着具身智能的落地。\n\n浙江人形机器人创新中心联合香港中文大学、浙江大学等机构，在国际顶刊《Science Robotics》上发表了RAM（Retrieval-Augmented Manipulation）三维空间理解与操作模型，首次系统性地解决了VLM的三维空间感知缺陷。\n\nRAM的核心思路是检索增强：不再让模型硬记所有三维知识，而是构建一个外部三维知识库，运行时动态检索与当前任务相关的空间信息。这相当于给VLM装了一个外接大脑，专门处理空间推理。对比端到端重新训练VLMoE的方式，这种方案成本低、迁移快，且不需要破坏原模型能力。\n\n实机验证结果令人印象深刻。在人形机器人平台上，语言指令驱动操作平均成功率达89.17%，图像引导操作成功率达92%。RAM还支持GPT、Qwen-VL等多款主流VLM，具备良好的模型兼容性——这意味着现有模型几乎不需要重新训练就能获得空间智能。\n\n为什么这值得关注？过去，机器人要实现可靠的抓取和操作，要么需要昂贵的端到端训练，要么依赖规则引擎但精度有限。RAM代表了一条中间路线：借助检索机制，在不完全重训的情况下，让基础模型获得接近专训的空间能力。这种外接知识库的范式有潜力迁移到其他需要精确空间推理的场景——比如自动驾驶的环境感知、工业机器人的精密装配。\n\n具身智能走到今天，缺的不是通用语言能力，而是物理世界的空间常识。RAM迈出了有趣的一步。","https:\u002F\u002Fwww.science.org\u002Fdoi\u002F10.1126\u002Fscirobotics.aea2092","e347a4b2-3269-4cc5-b792-e8b15d3a3bca",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"16d19d0f-4e04-4ebd-ba8e-e8dcb25e0758","en","RAM gives robots real 3D spatial understanding","Vision Language large Models (VLMs) are powerful, but have a fatal shortcoming: they essentially view 2D images, while real-world robot manipulation requires precise 3D spatial perception — where objects are, how they're oriented, how far away, whether they can be grasped. This gap has long plagued embodied-intelligence landing.\n\nZhejiang Humanoid Robot Innovation Center, jointly with the Chinese University of Hong Kong, Zhejiang University, and other institutions, published the RAM (Retrieval-Augmented Manipulation) 3D spatial understanding and manipulation model in the international top journal Science Robotics, systematically solving VLM's 3D spatial perception deficiencies for the first time.\n\nRAM's core idea is retrieval augmentation: instead of having the model hard-code all 3D knowledge, it constructs an external 3D knowledge base that dynamically retrieves task-relevant spatial information at runtime. This is equivalent to adding an external brain to VLM, specifically for spatial reasoning. Compared to end-to-end retraining of VLMoE, this approach has lower cost, faster migration, and doesn't need to break the original model's capabilities.\n\nReal-machine verification results are impressive. On a humanoid robot platform, language-instruction-driven manipulation achieves an average success rate of 89.17%, image-guided manipulation 92%. RAM also supports multiple mainstream VLMs including GPT and Qwen-VL, with good model compatibility — meaning existing models can gain spatial intelligence with virtually no retraining.\n\nWhy is this noteworthy? In the past, for robots to achieve reliable grasping and manipulation, they needed either expensive end-to-end training or rule engines with limited precision. RAM represents a middle path: via retrieval mechanisms, foundation models can gain near-specialized spatial capabilities without complete retraining. This external-knowledge-base paradigm has the potential to migrate to other scenarios requiring precise spatial reasoning — such as autonomous driving environment perception, industrial robot precision assembly.\n\nEmbodied intelligence has reached today, lacking not general language ability, but physical-world spatial commonsense. RAM has taken an interesting step.","ram-zhejiang-cuhk-3d-spatial-science-robotics","2026-05-06T13:00:00Z","2026-05-06T13:06:43.412347Z","2026-08-19T02:08:40.142862Z",true,"agent",157,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"4c1a894b-97fc-4d96-890c-8888167e23d8","元戎启行全面押注大模型：自动驾驶路线彻底转向","deeproute-large-model-autonomous-driving-ruanchong","2026-04-26T22:01:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"ad10985b-425c-4af1-9495-c63792a2b593","腾讯混元把语音识别打到 3% WER：Hy ASR 3.0 preview 让 ASR 从“逐字”走向“读语境”","tencent-hunyuan-hy-asr-3-0-preview-context-aware","2026-08-05T00:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"bd050dd6-4d85-4616-a004-55c23c533a24","腾讯混元合并大语言模型与多模态团队，成立基础模型部探索全模态统一","tencent-hunyuan-foundation-model-dept","2026-07-24T03:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"2ada2e69-25c2-4951-9e49-2b24a043393e","腾讯 Marvis 把 Agent 拽到端侧:混元要做 PC 集群,应用宝做了「系统级」分诊","tencent-marvis-on-device-agent","2026-07-23T20:30:00+00:00"]