[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-breeze-tts-2-open-source-realtime":3,"topics-all":38,"news-related-b1400260-ba9f-4e84-b658-ce53abba9304":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"b1400260-ba9f-4e84-b658-ce53abba9304","BreezeBlue 开源 Breeze TTS 2:3B 参数实时语音,五语种、可控设计、首包 133 毫秒","BreezeBlue 把 3B 参数 Breeze TTS 2 模型权重和 Apache 2.0 推理代码开源,在自建的 voice design、voice direction、Latency 三项 TTS 基准上同时跑到第一,首音频延迟在 H100 上稳定在 133 毫秒级,支持中英等 50 语种自然语音生成。","**BreezeBlue 是个从做\"虚拟人声\"起家的实时语音 AI 团队。**\n\n8 月 25 日,BreezeBlue 把第二代旗舰 TTS 模型 **Breeze TTS 2** 的权重和 PyTorch 推理代码同时开源。模型仅 3B 参数,一份权重覆盖中英等 50 种语言,可以做基于参考音频的声音克隆、纯自然语言描述的声音设计,也可以在保持音色一致的同时用指令切换情绪、节奏和语气。最关键的一项指标是把语音 AI 真正塞进实时对话场景的首音频延迟(Time To First Audio, TTFA)中位数压到 133.6 毫秒,加上 40ms 以下的\"加速模式\"已经能跟 ElevenLabs Flash v2.5、Fish Audio S2.1 Pro 这类顶级付费服务正面打。\n\n**\"三冠王\"全部出现在 BreezeBlue 自己发布的基准上。**\n\nBreezeBlue 这次同时开源了三套 TTS 评测基准——voice design、voice direction、Latency——用真实竞品做了完整对比。在 voice design benchmark 里,Breeze TTS 2 以 Role Fit 78.02、声线多样性 708 排名第一,比第二名 MiMo-v2.5-TTS 的 Role Fit 72.78 高出 5.24 分,声线多样性多 39%。在 voice direction 上,它以 4.25 分领先第二名 13%,同时将说话人相似度(SPK_SIM)维持在 0.67。在 latency 这一项,它的 TTFB p50 119.4ms、TTFA p50 133.6ms、TTFA p95 163.3ms,均低于 ElevenLabs Flash v2.5、Fish Audio S2.1 Pro、Cartesia Sonic 3.5、xAI TTS、Speechify Simba 3.2、Async Flash 1.5 等头部闭源对手。\n\n需要明确说明:这三套基准是 BreezeBlue 团队自己开源的,目前三项冠军的成绩均来自该团队而非交叉独立评测。这一点读者在做横向对比时务必留意。\n\n**一个 3B 模型为什么能同时做\"用文字写声音\"和\"实时出声音\"。**\n\nBreeze TTS 2 把四类输入约定塞进同一份权重:\n- Voice Clone:传一段参考音频和它的逐字稿,模型保留音色、节奏和情绪;\n- Voice Design:不传参考音频,直接用自然语言指令\"中年男性、磁性低音、节奏铿锵\"写出全新声线;\n- Voice Direction:克隆一个声音后,再用自然语言控制每一句的情绪、强度和节奏;\n- Vocal Events:文本内联\"(laugh) \u002F (sigh) \u002F (cough)……\"或中文\"[笑] \u002F [叹气]\"等标记,在合成结果里直接插入对应声音事件。\n\n它的推理栈是 Apache 2.0 协议的 PyTorch 实现,在 NVIDIA H100 上 eager 模式显存 7.7 GiB;启用 `--fast-all` 全流程 CUDA Graph 编译后,体积升至 14.4 GiB,但能把首包延迟再削掉一段。中文和英文写在同一份权重里,日、西、法、德、韩、葡、印地等 50 种语言同模型生成,主页示例覆盖动画、游戏、电影电视、文学舞台、神话民俗、新闻播报六种语气走向。\n\n**权重不是 Apache 2.0,商业落地要先谈授权。**\n\n代码 Apache 2.0 没问题,但模型权重、衍生模型和自托管输出走 BreezeBlue 自己的\"研究 + 非商用\"许可,商业部署需要 RESONIA 公司书面授权。这个安排跟 MiniMax-Music3、IndexTTS-2.5 这类近期国产开源 TTS 类似——研究社群可直接下载、企业集成还得谈授权——但和 ElevenLabs、Cartesia 这种纯闭源商用相比,开源 TTS 正在被压缩到一个\"免费研究 + 付费商业\"的二元结构,平台侧的预算未来会进一步往前者倾斜。\n\n模型卡也把语音克隆的合规红线直接写进 License:未授权的声纹复制、冒充、欺诈和其他非法用途一律不许碰。这一点对国内做 AI 客服 \u002F 数字人 \u002F 短视频配音的团队尤其重要——国内监管现在对\"逼真克隆他人声音\"的态度比去年更严,集成之前必须把合规走通,不能只看许可证字面。\n\n**实时语音 AI 这条赛道正在被卷回到\"低延迟\"。**\n\n这一波 TTS 发布潮(MiniMax-Music3、IndexTTS-2.5、Qwen-Audio-3.0-TTS、Breeze TTS 2)有一个共同方向——把 TTFA \u002F TTFB 从一秒以上压进 200ms 以内。原因不复杂:实时语音 agent、游戏 NPC、互动故事这些场景,自然对话节奏要求首包延迟 ≤ 300ms,否则用户会感到明显的\"顿挫\"。各家都在拼三条线:声线多样性(让同一个模型支持\"上千种角色\")、情绪可控(让声音真的能\"演\")、流式协议(WebSocket 单会话多轮、PCM 流式回灌)。\n\nBreeze TTS 2 把 TTFA p50 做到 133.6ms 这个数字,客观上把实时语音 AI 的工程门槛又拉高了一档。下一步可以关注两个数字:一是第三方独立基准(LMSYS、Artificial Analysis 之外的)能不能复现它的\"三冠\";二是它的商用授权价格、是否对硬件 \u002F 语言 \u002F 声纹数做限制——这两个变量会直接决定它在企业 SDK 选型表里的真实排位。\n\n原始发布:https:\u002F\u002Fbreezeblue.ai\u002Fbreeze-tts-2\n","https:\u002F\u002Fbreezeblue.ai\u002Fbreeze-tts-2","6e803912-6cd7-4472-a19f-08efe581e893",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"3ea57c7a-fbc6-4f92-8cef-a971c57ce913","en","BreezeBlue open-sources Breeze TTS 2: a 3B real-time speech model with 50-language support, controllable voice design, and 133 ms first-audio latency","BreezeBlue open-sourced the 3B-parameter Breeze TTS 2 weights and Apache 2.0 inference code. On its self-published voice design, voice direction and latency benchmarks the model takes first place across all three; TTFA on an NVIDIA H100 sits at 133 ms, with 50-language support including Chinese and English.","**BreezeBlue is a real-time voice AI team that grew out of building virtual-character voices.**\n\nOn August 25, BreezeBlue open-sourced both the weights and the PyTorch inference code for **Breeze TTS 2**, its second-generation flagship TTS model. The model is just 3B parameters, and a single checkpoint covers 50 languages including Chinese and English. It can do reference-audio-based voice cloning, zero-shot voice design from a natural-language description, and it can switch emotion, pace and delivery on demand while keeping the speaker identity stable. The key headline number is that the median Time To First Audio (TTFA) sits at 133.6 ms — and with the warmed-up \"fast\" path it has gone under 40 ms — putting it in the same range as top paid services such as ElevenLabs Flash v2.5 and Fish Audio S2.1 Pro.\n\n**The \"three crowns\" all come from BreezeBlue's own benchmarks.**\n\nAlongside the model, BreezeBlue open-sourced three TTS evaluation suites — voice design, voice direction, and latency — and ran a head-to-head comparison against real competitors. On the voice design benchmark, Breeze TTS 2 ranks first with a Role Fit score of 78.02 and 708 distinct voices, beating the second-place MiMo-v2.5-TTS by 5.24 Role Fit points and 39% more voice diversity. On voice direction it leads with 4.25, 13% ahead of the next entry, while keeping speaker similarity (SPK_SIM) at 0.67. On latency, the model posts TTFB p50 119.4 ms, TTFA p50 133.6 ms, and TTFA p95 163.3 ms — all lower than ElevenLabs Flash v2.5, Fish Audio S2.1 Pro, Cartesia Sonic 3.5, xAI TTS, Speechify Simba 3.2, and Async Flash 1.5.\n\nOne important caveat: these three benchmarks were open-sourced by the BreezeBlue team itself, so all three first-place scores come from that team rather than from any cross-independent evaluation. Readers doing a horizontal comparison should keep that in mind.\n\n**How can a 3B model do \"write a voice from text\" and \"play it in real time\" at the same time?**\n\nBreeze TTS 2 packs four input conventions into one checkpoint:\n- Voice Clone: feed a reference audio clip with its exact transcript and the model preserves timbre, rhythm, and emotion.\n- Voice Design: no reference audio needed; describe the speaker in natural language — \"middle-aged male, magnetic bass, staccato cadence\" — and generate a brand-new voice.\n- Voice Direction: clone a voice, then steer each line's emotion, intensity, and pacing with natural-language instructions.\n- Vocal Events: inline tags such as `(laugh) \u002F (sigh) \u002F (cough)` in English or `[笑] \u002F [叹气]` in Chinese become actual sound events inside the synthesized audio.\n\nThe inference stack is Apache 2.0 PyTorch. On NVIDIA H100, the eager path uses roughly 7.7 GiB of GPU memory; enabling `--fast-all` (full CUDA-Graph compilation) raises this to 14.4 GiB while shaving more off the first-audio latency. Chinese and English share the same checkpoint, and the same model generates Japanese, Spanish, French, German, Korean, Portuguese, Hindi and other languages up to 50 in total. The launch page demonstrates voice personas across animation, games, film and TV, literature and stage, mythology and folklore, and news broadcasting.\n\n**The weights are not Apache 2.0 — commercial deployment still requires a license.**\n\nThe code is Apache 2.0, but the model weights, derivative models, and self-hosted outputs sit under BreezeBlue's own \"research and non-commercial\" license, and commercial deployment requires written authorization from RESONIA. The structure is the same one we have seen on MiniMax-Music3 and IndexTTS-2.5 recently — research users can download freely, enterprises still have to negotiate. Compared with fully closed commercial offerings from ElevenLabs and Cartesia, the open-weights TTS space is settling into a \"free for research, paid for commercial\" two-tier structure, and platform-side budgets will increasingly tilt toward the former.\n\nThe model card also writes voice-cloning compliance boundaries directly into the license: unauthorized voice cloning, impersonation, fraud, and other unlawful or harmful uses are prohibited. This matters for any Chinese team building AI customer service, digital humans, or short-video voice-overs — regulators in China have been visibly stricter on \"realistic cloning of another person's voice\" over the past year, and you have to clear the compliance question before shipping, not just glance at the license text.\n\n**Real-time voice AI is being driven back to \"low latency\".**\n\nThe current wave of TTS releases (MiniMax-Music3, IndexTTS-2.5, Qwen-Audio-3.0-TTS, Breeze TTS 2) all share one direction — push TTFA \u002F TTFB from \"over one second\" to \"under 200 ms\". The reason is simple: in real-time voice agents, game NPCs, and interactive stories, the natural conversational rhythm needs a first-audio latency of ≤ 300 ms, otherwise users feel an obvious \"stutter\". Every team is racing along three axes: voice diversity (thousands of personas from one model), emotional controllability (let the voice actually \"act\"), and streaming protocols (single WebSocket session across many turns, raw PCM flowing back while the model is still generating).\n\nBreeze TTS 2 setting TTFA p50 at 133.6 ms effectively raises the engineering bar for real-time voice AI by another notch. Two numbers to watch next: first, whether third-party independent benchmarks (outside LMSYS and Artificial Analysis) can reproduce the \"three crowns\"; second, how the commercial license treats hardware, language, and voice-clone-count caps. Those two variables will decide its real place on the enterprise SDK shortlist.\n\nOriginal announcement: https:\u002F\u002Fbreezeblue.ai\u002Fbreeze-tts-2\n","breeze-tts-2-open-source-realtime","2026-08-29T10:00:00Z","2026-08-29T09:08:02.732376Z","2026-08-29T09:08:02.732389Z",true,"agent",556,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"dc2f4ead-963c-4a8e-bd41-400bebf83bb4","物理、几何、外观一个模型全包:Puffin-World 开源,相机 roll 误差低至 0.26°","puffin-world-native-3d-world-states","2026-09-06T19:09:41+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"630ed9ae-699e-4115-8e75-33196ea6db28","MiniMax Music 3 开源:8B+0.6B 双 LLM 写五分钟完整歌,8GB 显存能跑","minimax-music3-open-weights-architecture","2026-08-29T13:30:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"d8bc7b5e-9eb0-475e-91b7-5a3390d2c6a6","2026年开源LLM爆发：Meta、阿里、Google竞相发布新一代模型","open-source-llm-boom-2026-q1-meta-alibaba-google","2026-04-24T04:06:08+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"51c13e24-8072-404c-a8d4-75c40cff05ee","Ling-3.0-flash-VL 开源：124B MoE 只激活 5.5B，视觉塞进 Agent 闭环","ling-3-0-flash-vl-open-weights","2026-09-15T13:18:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"550cee5e-18e8-4236-9304-7207ebc221a8","Agnes 3.0 Flash 开源:72 层仅 18 层带 KV 缓存,33B 单卡跑 262k 上下文","agnes-3-0-flash-preview-open-weights","2026-09-13T15:20:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"108af093-b226-4372-9cf0-77323ffc5456","小鹏 X-AuT 给语音大模型剪枝:音频塔砍 4 层,车载推理提速 21.4%","xpeng-x-aut-audio-encoder-pruning","2026-09-12T19:06:47+00:00"]