[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-qwen-fun-asr-nano-0-8b":3,"news-related-a2cd999e-e74f-43a9-a24b-3898cf5e9582":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a2cd999e-e74f-43a9-a24b-3898cf5e9582","「0.8B 干到 16.72% WER」:Fun-ASR-Nano 用「端到端+RAG」把工业 ASR 卷出新尺度","千问大模型今日正式升级 Fun-ASR-Realtime,把通义实验室的 Fun-ASR-Nano-2512(0.8B 参数)推到工业生产管线。这不是又一次普通迭代——它把 ASR 从「声学模型+后处理」两段式范式,推进到「一体化 LLM-ASR」的新阶段。\n\n最值得关注的是它把 RAG 技术塞进了流式识别:热词、命名实体、行业术语通过检索拼接到解码上下文,让同一份声学模型在教育、金融、医疗等领域都自带「专业词表」。端到端架构让声学特征直接映射到 token,ITN、标点、敏感词过滤在同一个网络里完成,显著减少传统 pipeline 的误差累积。\n\nGitHub FunAudioLLM\u002FFun-ASR(1.3k Star)README 公开的工业测试集(覆盖近场、远场、复杂背景、方言、口音、歌词、说唱 7 大场景)显示:Fun-ASR-Nano 平均 WER 16.72%,比参数更大的 GLM-ASR-Nano(1.5B, 26.13%)低近 10 个百分点,比 Whisper-large-v3(1.6B, 33.39%)好一倍以上。闭源旗舰 Fun-ASR(7.7B)虽然 WER 还能压到 12.70%,代价是近 10 倍体量。\n\n值得注意三点:其一,「小钢炮」逻辑再现——0.8B 跑赢 1.5-1.6B 同类,ASR 已进入「架构创新 > 堆参数」阶段;其二,工程能力是真正的护城河——原生 vLLM 引擎(3-5 倍批量加速)、llama.cpp\u002FGGUF 单二进制边缘部署、WebSocket VAD+Manual 双交互,这些活儿比模型架构更难抄;其三,31 语言 + 7 大方言 + 26 地区口音是千问对 Whisper 的差异化筹码。\n\n但也要警惕:闭源与开源的精度差距正在拉大(方言 15.21% vs 28.18%),ASR 进入「开源够用、闭源领先」的结构,对部署成本敏感的产品是双刃剑。","https:\u002F\u002Fgithub.com\u002FFunAudioLLM\u002FFun-ASR","998df6db-96e6-4b8e-8be1-cfa00a6cd177",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"80c7d1bb-c840-44c9-bd6c-bb506c2078b7","en","Fun-ASR-Nano: 0.8B hits 16.72% WER with end-to-end plus RAG","Qwen LLM today officially upgraded Fun-ASR-Realtime, pushing Tongyi Lab's Fun-ASR-Nano-2512 (0.8B parameters) to the industrial production pipeline. This isn't another routine iteration — it pushes ASR from the \"acoustic model + post-processing\" two-stage paradigm into the \"unified LLM-ASR\" new stage. The most noteworthy is that it stuffs RAG technology into streaming recognition: hotwords, named entities, and industry terms are spliced into the decoding context through retrieval, letting the same acoustic model in education, finance, healthcare and other domains all come with its own \"professional vocabulary\". The end-to-end architecture lets acoustic features map directly to tokens, with ITN, punctuation, and sensitive-word filtering all done in the same network, significantly reducing the error accumulation of traditional pipelines. The GitHub FunAudioLLM\u002FFun-ASR (1.3k Star) README's public industrial test set (covering 7 major scenarios: near-field, far-field, complex background, dialect, accent, lyrics, rap) shows: Fun-ASR-Nano averages WER 16.72%, nearly 10 percentage points lower than the larger GLM-ASR-Nano (1.5B, 26.13%), more than double that of Whisper-large-v3 (1.6B, 33.39%). Although the closed-source flagship Fun-ASR (7.7B) can still press WER to 12.70%, the cost is nearly 10× the size. Three points worth noting: first, the \"small steel cannon\" logic appears again — 0.8B beats 1.5-1.6B of the same kind, and ASR has entered the \"architectural innovation > parameter stacking\" stage; second, engineering capability is the real moat — native vLLM engine (3-5× batch speedup), llama.cpp\u002FGGUF single-binary edge deployment, WebSocket VAD+Manual dual interaction, these jobs are harder to copy than model architecture; third, 31 languages + 7 major dialects + 26 regional accents is Qwen's differentiation card against Whisper. But be wary: the accuracy gap between closed-source and open-source is widening (dialect 15.21% vs 28.18%), ASR is entering a \"open-source good enough, closed-source leading\" structure, which is a double-edged sword for deployment-cost-sensitive products.","qwen-fun-asr-nano-0-8b","2026-07-06T14:01:00Z","2026-07-06T14:14:50.681973Z","2026-08-19T02:08:40.142862Z",true,"agent",150,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"4bb31ede-b9c4-4762-86ae-9d3b008557ca","Hugging Face Summer 2026 报告:Qwen 拿下 15 万衍生模型, GGUF 仓库一年涨 464%","hugging-face-state-of-open-models-summer-2026","2026-08-18T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"1311adb6-dc19-41a7-a188-6760d9e53672","HF Summer 2026 报告:13 个下载量 Top 25 模型是 2022 年的老面孔","hugging-face-summer-2026-attention-adoption","2026-08-24T08:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"9389d1ed-dd2d-41cb-bbc5-9a543e2b2f71","开源报告里的「参数天花板」分水岭:中国实验室把上限拉到2.78T,美国还在130B徘徊","hf-summer-2026-china-open-weight-parameter-ceiling","2026-08-20T06:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"aa4d3e55-383d-4855-9965-cc6a4d2e38a7","Qwen 下载量 6 个月破 30 亿:开源模型的「默认底座」第一次换成了中国厂商","qwen-3-billion-downloads-open-weights","2026-08-15T23:20:00+00:00"]