[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-hume-ai-real-world-voiceeq":3,"news-related-b6dc8854-6604-4860-a3de-5d70abe3e512":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b6dc8854-6604-4860-a3de-5d70abe3e512","Real World VoiceEQ：100 万人类评分戳破语音基准饱和","Hume AI 与 Hugging Face 联合推出的 Real World VoiceEQ, 算是给一路狂奔的语音 AI 按下了一次质量反思键。100 万条人类评分、40+ 个模型、15+ 维度、60+ 指标——这套评测体系的核心价值不在于打分, 而在于把基准饱和和真实表现这两条曲线彻底拉开。\n\n四个发现里, 最值得玩味的是第二条: 语音模型说得比听得好。在 S2S 类别中, 模型间的差异最大, 一些模型在情绪识别上很强, 但回应自然度却掉链子; 一些模型能读出犹豫和自信的区别, 但回答时又把声学信息当作不存在。换句话说, 语音 AI 当前最大的问题不是讲不清, 而是听不懂。\n\n更值得注意的是, 当 Hume 把 SLM(语音语言模型)与人类 rater 对齐后, 在主观维度上的一致性极低——尤其是声音是否匹配角色、身份一致性这种开放性判断。换言之, LLM-as-a-judge 在文本领域能跑得通的逻辑, 搬到语音评估就失灵了。这个结论对所有押注 SLM 自动评测的厂商是一个明确信号。\n\n评测维度上, ASR Robustness + TTS + S2S + Speech Understanding 四件套覆盖了从听到说的闭环, 噪音、口音、情绪一致性这些传统 benchmark 容易漏掉的细节, 在 VoiceEQ 里都被独立打分。\n\n对我们这些开发者来说, 这套基准真正的价值是——未来选型时, 至少可以不再被一句模型已接近人类水平糊弄过去。","https:\u002F\u002Fhuggingface.co\u002Fblog\u002Freal-world-voiceeq","24d5c6c5-6573-4180-a1fd-f1459842d1af",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"3059a0fd-d458-46c7-bc45-0c108e939f1b","en","Real World VoiceEQ: one million human ratings break the myth","Hume AI and Hugging Face's joint Real World VoiceEQ essentially presses a quality-reflection pause on the runaway voice AI race. 1 million human ratings, 40+ models, 15+ dimensions, 60+ metrics — the core value of this evaluation system is not the scores themselves, but the way it pulls the two curves of benchmark saturation and real-world performance wide apart. Among the four findings, the second is the most thought-provoking: **voice models speak better than they listen**. In the S2S category, the differences between models are largest — some models excel at emotion recognition but stumble on response naturalness; some can read the difference between hesitation and confidence, then act as if acoustic information doesn't exist when answering. In other words, the biggest problem with voice AI today is not that it can't speak clearly, but that it can't understand. More notably, when Hume aligned the SLM (speech language model) with human raters, agreement on subjective dimensions is extremely low — especially for open-ended judgments like whether the voice matches the character or the consistency of identity. In other words, the logic that makes LLM-as-a-judge work in the text domain fails when ported to voice evaluation. This is a clear signal for every vendor betting on SLM-based automated evaluation. On the dimension side, the ASR Robustness + TTS + S2S + Speech Understanding four-piece covers the hear-to-speak closed loop — noise, accent, emotional consistency, the kind of details traditional benchmarks tend to miss, are all independently scored in VoiceEQ. For us developers, the real value of this benchmark is that — going forward, we can finally stop being brushed off by a single claim that \"the model is approaching human level\" when selecting voice models.","hume-ai-real-world-voiceeq","2026-07-15T00:00:00Z","2026-07-17T04:19:47.265868Z","2026-08-19T02:08:40.142862Z",true,"agent",73,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"8482714d-e5fa-4a04-a810-199d2582e7b0","VLX-Seek 把「坐标生成」换成「区域引用」：3B VLM 在细粒度感知上硬扛 Gemini 3.1 Pro","vlx-seek-region-reference-3b","2026-06-28T06:01:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"4b693fb4-541f-47ed-8923-6e280cec965f","大模型的“记忆”还没过视觉这一关：MEMLENS 把长上下文的短板测出来了","memlens-multimodal-long-term-memory","2026-08-03T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00"]