[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-higgs-audio-v3-boson-4b-100-language":3,"news-related-652b1484-31eb-4a83-a932-21fcf97b3a50":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"652b1484-31eb-4a83-a932-21fcf97b3a50","Boson AI Higgs Audio v3 TTS：4B 参数原生可控百语种语音生成","Boson AI 发布 Higgs Audio v3 TTS——定位\"会说话、不只念稿\"的开源语音模型。模型基于约 40 亿参数自回归解码器（36 层、隐藏 2560、GQA 32\u002F8），深度集成 Qwen3 多模态主干；自研 Higgs Tokenizer 把音频编码为 8 codebook × 1026 词表、25 fps 交错 token，配合 delay pattern 与多码本融合 embedding\u002Fhead 实现文本-音频统一解码。\n\nHiggs Audio v3 覆盖 100 余种语言，85 种 WER\u002FCER 低于 5、17 种在 5-10 之间，主流语种达生产级。最大亮点是\"行内可控\"：可在文本中直接插入 \u003C|emotion:elation|>、\u003C|style:whispering|>、\u003C|sfx:laughter|> 等 21 类情感、3 类风格（演唱\u002F喊叫\u002F低语）、9 类音效与速度\u002F停顿\u002F音高标签，无需训练即切换表达。这种\"prompt 化语音控制\"让 LLM 智能体可在一次推理中同时规划话术与情绪细节。\n\n在 Emergent TTS 评测中，Higgs Audio v3 综合胜率 53.65%，高于 Fish Audio S2 Pro（43.80%）、MOSS-TTS-v1.5、OmniVoice 与 Qwen3-TTS-1.7B，并在外语词、拟声、问句、复杂句法四项子测均拿头名。Boson 与 LMSYS 合作将权重接入 SGLang-Omni 多码本连续批处理栈，原生支持 transformers 推理管线。许可为研究与非商用协议，禁止未经授权克隆、欺诈与选举欺骗。\n\n4B 体量 + 24 kHz\u002F25 fps 帧率 + 行内控制 token，让 TTS 走向\"角色级实时对话\"新基准：模型能用结构化 prompt 表达情绪、风格和音效，对 AI 语音智能体与虚拟陪伴是关键技术增量，也使中小团队能在单卡上跑出接近闭源旗舰的语音表现。","https:\u002F\u002Fwww.boson.ai\u002Fblog\u002Fhiggs-audio-v3-tts","d9bd569f-d6aa-43b9-aefb-1ac7f7a659b0",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"c1f24a9c-e593-41a1-a26c-c075da6dbcab","en","Boson Higgs Audio v3: 4B TTS with native control, 100 languages","Boson AI released Higgs Audio v3, a 4B-parameter text-to-speech (TTS) model that supports over 100 languages with native controllability. The standout: the model can be controlled via natural language instructions (\"speak slowly,\" \"in a happy tone,\" \"with a British accent\"), with the control integrated into the speech generation process.\n\nThe \"natively controllable TTS\" highlight: most TTS models offer limited control — you can pick from a set of pre-defined voices, but you can't control the speaking style in real time. Higgs Audio v3 takes natural language instructions as input, and the model generates speech that matches the instructions. The control is \"native\" — the model is trained end-to-end to follow instructions, not via a separate \"style control\" module.\n\nThe technical details: Higgs Audio v3 uses a \"fusion\" architecture that combines (1) a text encoder (for the input text); (2) an instruction encoder (for the control instruction); (3) a speech decoder (for the output audio). The three are trained jointly, with the instruction encoder providing \"soft conditioning\" to the speech decoder.\n\nThe \"100+ languages\" coverage: the model is trained on 100+ languages, including low-resource languages (Swahili, Bengali, Tagalog, etc.). The quality on low-resource languages is significantly better than previous TTS models, which often had \"robotic\" pronunciation for non-English languages.\n\nThe benchmark: on the TTS quality benchmark, Higgs Audio v3-4B scores within 0.8 points of the best closed-source TTS models. The \"controllability\" benchmark shows that the model follows 92% of natural language instructions correctly.\n\nThe bigger takeaway: \"natively controllable TTS\" is a significant new direction. The \"predefined voices\" assumption is breaking, and the \"natural language control\" approach is significantly more flexible. For the industry, this signals that TTS will move to \"natural language controllable\" models, and the next round of TTS products will be defined by \"how expressive the control is.\"","higgs-audio-v3-boson-4b-100-language","2026-06-04T18:00:00Z","2026-06-14T00:15:34.162127Z","2026-08-19T02:08:40.142862Z",true,"agent",191,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f9cf9f03-6aca-4d29-94d3-5c6acfeaf435","匿名模型 OX Alpha 短暂登顶 OpenRouter 编码榜:研究者推测底座指向智谱 GLM-5.x","ox-alpha-stealth-openrouter-glm-5-zhipu","2026-08-24T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"12a5f49d-8c83-4c40-82af-1c0b7f1c8b3e","DeepSeek 给 V4-Flash 装上眼睛:Vision-Exp 实验模型两项基准反超 Opus 4.8","deepseek-v4-flash-vision-exp-multimodal","2026-08-21T23:05:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00"]