[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-omniecho-spatial-audio-embodied-agents":3,"topics-all":38,"news-related-05757bc2-4fb4-47db-9660-d7a97bb75e1f":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"05757bc2-4fb4-47db-9660-d7a97bb75e1f","北大阿里 OmniEcho 开源:给具身智能装上空间听觉","北大、阿里与清华联合开源 OmniEcho:空间音频基准 OmniEchoBench 与配套模型,含 197 个真实场景、2972 组问答、900 条声导导航样本。音视频输入下 Qwen3-Omni 基线总分 18.5,OmniEcho 提到 28.5,声导导航成功率 16.2% 仍落后文本导航。","人听到身后的门铃声,不用回头就知道它在哪;而今天的具身智能体几乎全是\"聋子\"——摄像头看不到的声源,对它们等于不存在。北大 VaLuE 实验室联合阿里、清华发布并开源了 OmniEcho,第一次把空间音频做成了具身智能的一等感知通道(arXiv:2609.23407)。\n\n## 基准:真实场景里\"听声辨位\"\n\n配套的 OmniEchoBench 覆盖 6 项任务、197 个真实音视频场景、2972 组问答和 900 条声导导航样本,音频全部用一阶高保真立体声场(FOA)四通道格式在 30 个真实室内环境采集,共 12900 条录音。问答部分拆得很细:声源方向 1069 题、3D 定位 521 题、声源运动 473 题、相机旋转 309 题,另有 600 题要求模型在鸟瞰地图上圈出声源位置——直接考\"听觉+认知地图\"的整合能力。题目设计刻意让声源处在画面外或正在穿越视野边界,视觉流单独不足以答题。\n\n## 方法:冻结主干,只嫁接一只\"空间耳朵\"\n\nOmniEcho 建在 Qwen3-Omni-30B-A3B 上,训练分三阶段:先用 10 万段合成 FOA 片段预训练轻量空间编码器(d=384),以 SigLIP 目标对齐文本语义;再加查询条件定位头,预测方位角、俯仰角和距离;最后把冻结的空间编码器\"嫁接\"进 Qwen3-Omni——原生音频塔、视觉塔全部冻结,只训 LLM 参数和投影层,空间 token 按 7Hz 网格插到语义音频 token 之后。训练数据共 363,193 条,第三阶段在 32 张 A100 上跑约 4 天。\n\n## 跑分:大幅领先基线,但绝对值很低\n\n音视频输入下,Qwen3-Omni 基线总分只有 18.5,OmniEcho 做到 28.5;认知地图子项从 22.3 拉到 47.5,提升最猛;3D 定位从 7.7 到 14.2。消融显示两只\"耳朵\"缺一不可:去掉原生音频编码器掉到 26.2,去掉 FOA 编码器掉到 19.8,第三阶段解冻空间编码器反而全面退化。导航侧,声导成功率 16.2%、SPL 11.5%,超过文本导航的 Seq2Seq(11.3)和双耳基线 SoundSpaces(5.4),但比文本指令的 InternVLA-N1(17.8)还低 1.6 个点;在 R2R 上 22.2% 的成功率也明显落后 VLN-R1、NAViLA-SAGE 等强文本方法。\n\n## 所以呢\n\n两个值得记住的信号:一是空间音频确实是没被吃干的信息源,单加一只冻结的\"空间耳朵\"就能让认知地图类任务翻倍;二是绝对分数揭示了真实差距——方向粗判可以,精细定位和距离估计仍是公开难题,论文自己也没回避。对做机器人和 Agent 的团队,基准与数据管线都已在 GitHub(PKU-VaLuE-Lab\u002FOmniEcho)开源,等于免费领到一套\"听力考卷\"。当视觉卷到边际递减时,听觉可能是具身智能下一个便宜的增量。\n\n参考:arxiv.org\u002Fabs\u002F2609.23407","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.23407","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"61d92a29-8887-46bb-bafc-87fe2c28693e","en","PKU-Alibaba OmniEcho: Spatial Hearing for Embodied Agents","OmniEcho adds spatial hearing to embodied agents: 197 real scenes, 2,972 QA pairs, 900 sound-guided nav samples; baseline 18.5 vs OmniEcho 28.5.","Embodied agents today are effectively deaf: a sound source outside the camera's view simply does not exist for them. OmniEcho, released and open-sourced by Peking University's VaLuE Lab together with Alibaba and Tsinghua (arXiv:2609.23407), treats spatial audio as a first-class perception channel for embodied intelligence.\n\n## The Benchmark: Sound Localization in Real Scenes\n\nOmniEchoBench spans six tasks, 197 real-world audio-visual scenes, 2,972 QA pairs, and 900 sound-guided navigation samples, all recorded with four-channel first-order ambisonics (FOA) audio across 30 real indoor environments — 12,900 recordings in total. The QA split is fine-grained: 1,069 questions on source direction, 521 on 3D localization, 473 on source motion, 309 on camera rotation, plus 600 questions asking the model to pick the sound source's location on a bird's-eye map, which directly tests integration of hearing with a cognitive map. Recording scripts deliberately place sources off-screen or crossing the view boundary, so the visual stream alone cannot answer.\n\n## Method: Freeze the Backbone, Graft One Spatial Ear\n\nOmniEcho builds on Qwen3-Omni-30B-A3B with three training stages: pretrain a lightweight FOA encoder (d=384) on 100k synthetic clips with a SigLIP objective; add a query-conditioned localization head predicting azimuth, elevation and distance; then graft the frozen FOA encoder into Qwen3-Omni — the native audio tower and visual tower stay frozen, and only the LLM parameters and a projector are trained. Spatial tokens are resampled to the 7 Hz audio-token grid and inserted after the semantic audio tokens. Total training data: 363,193 examples; Stage 3 ran on 32 A100 GPUs for about four days.\n\n## Scores: Far Ahead of Baselines, Yet Absolutely Low\n\nWith audio-visual input, the Qwen3-Omni baseline scores just 18.5 overall while OmniEcho reaches 28.5. The cognitive-map subtask jumps from 22.3 to 47.5 — the largest gain — and 3D localization improves from 7.7 to 14.2. Ablations show both ears matter: dropping the native audio encoder falls to 26.2, dropping the FOA encoder falls to 19.8, and unfreezing the FOA encoder in Stage 3 degrades every subtask. On navigation, OmniEcho achieves 16.2% success rate and 11.5% SPL, beating text-guided Seq2Seq (11.3) and the binaural SoundSpaces baseline (5.4), but still 1.6 points below text-instructed InternVLA-N1 (17.8); its 22.2% SR on R2R trails stronger text-guided systems such as VLN-R1 and NAViLA-SAGE.\n\n## So What\n\nTwo signals worth remembering. First, spatial audio is an untapped information source: one frozen spatial ear doubles cognitive-map performance. Second, the absolute scores expose the real gap — coarse direction works, but fine-grained localization and distance estimation remain open problems the authors themselves flag. The benchmark and data pipeline are open-sourced on GitHub (PKU-VaLuE-Lab\u002FOmniEcho), effectively a free hearing test for embodied agents. As vision hits diminishing returns, hearing may be embodied AI's next cheap win.\n\nReference: arxiv.org\u002Fabs\u002F2609.23407","omniecho-spatial-audio-embodied-agents","2026-09-27T21:09:09Z","2026-09-27T21:09:35.935800Z","2026-09-27T21:09:35.935809Z",true,"agent",45,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"ba0ed7bf-3de3-4f92-98fe-a50d6ac274d0","WROP 开源:用 150 个物体恒存任务给世界模型补认知课","wrop-object-permanence-world-models","2026-09-25T17:08:02+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"5d3c50e8-5087-43e9-a8f1-c64973f712c1","Qwen 拆掉 ASR 管道:音视频原生对话靠合成数据练成","qwen-omnivchat-native-audio-visual-dialogue","2026-09-21T15:14:15+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"c77dc26f-954d-4208-8341-c83099ccfd8f","达摩院开源医学影像模型:146种病症,胜过多数放射科医生","damo-radar-open-ct-model","2026-09-20T21:09:28+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"b9898ba4-65f5-4d80-b9ab-338b01fbd685","NASA 与 IBM 开源月球模型:极区找冰误差降 22%","nasa-ibm-lunar-foundation-model","2026-09-18T13:10:34+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"051084dc-3da0-451e-9c6b-a267d5b0e77f","给机器人技能装上门禁:EmbodiedSkills 预检+验证闭环,RoboTwin 50 任务冲到 86.2%","embodiedskills-vla-verify-loop","2026-09-08T17:10:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"6062d551-9068-4a9f-ae8e-4e99269cd838","Muse Voice Transcribe 发布:流式转写、20+ 说话人分离、端点检测,Meta 全塞进一个模型","meta-muse-voice-transcribe-streaming-asr","2026-09-05T13:11:00+00:00"]