Embodied agents today are effectively deaf: a sound source outside the camera's view simply does not exist for them. OmniEcho, released and open-sourced by Peking University's VaLuE Lab together with Alibaba and Tsinghua (arXiv:2609.23407), treats spatial audio as a first-class perception channel for embodied intelligence.

The Benchmark: Sound Localization in Real Scenes

OmniEchoBench spans six tasks, 197 real-world audio-visual scenes, 2,972 QA pairs, and 900 sound-guided navigation samples, all recorded with four-channel first-order ambisonics (FOA) audio across 30 real indoor environments — 12,900 recordings in total. The QA split is fine-grained: 1,069 questions on source direction, 521 on 3D localization, 473 on source motion, 309 on camera rotation, plus 600 questions asking the model to pick the sound source's location on a bird's-eye map, which directly tests integration of hearing with a cognitive map. Recording scripts deliberately place sources off-screen or crossing the view boundary, so the visual stream alone cannot answer.

Method: Freeze the Backbone, Graft One Spatial Ear

OmniEcho builds on Qwen3-Omni-30B-A3B with three training stages: pretrain a lightweight FOA encoder (d=384) on 100k synthetic clips with a SigLIP objective; add a query-conditioned localization head predicting azimuth, elevation and distance; then graft the frozen FOA encoder into Qwen3-Omni — the native audio tower and visual tower stay frozen, and only the LLM parameters and a projector are trained. Spatial tokens are resampled to the 7 Hz audio-token grid and inserted after the semantic audio tokens. Total training data: 363,193 examples; Stage 3 ran on 32 A100 GPUs for about four days.

Scores: Far Ahead of Baselines, Yet Absolutely Low

With audio-visual input, the Qwen3-Omni baseline scores just 18.5 overall while OmniEcho reaches 28.5. The cognitive-map subtask jumps from 22.3 to 47.5 — the largest gain — and 3D localization improves from 7.7 to 14.2. Ablations show both ears matter: dropping the native audio encoder falls to 26.2, dropping the FOA encoder falls to 19.8, and unfreezing the FOA encoder in Stage 3 degrades every subtask. On navigation, OmniEcho achieves 16.2% success rate and 11.5% SPL, beating text-guided Seq2Seq (11.3) and the binaural SoundSpaces baseline (5.4), but still 1.6 points below text-instructed InternVLA-N1 (17.8); its 22.2% SR on R2R trails stronger text-guided systems such as VLN-R1 and NAViLA-SAGE.

So What

Two signals worth remembering. First, spatial audio is an untapped information source: one frozen spatial ear doubles cognitive-map performance. Second, the absolute scores expose the real gap — coarse direction works, but fine-grained localization and distance estimation remain open problems the authors themselves flag. The benchmark and data pipeline are open-sourced on GitHub (PKU-VaLuE-Lab/OmniEcho), effectively a free hearing test for embodied agents. As vision hits diminishing returns, hearing may be embodied AI's next cheap win.

Reference: arxiv.org/abs/2609.23407