[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-hunyuan-auk-speech-editing":3,"topics-all":38,"news-related-cd49f913-cde7-4cf3-8d93-24508653180e":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"cd49f913-cde7-4cf3-8d93-24508653180e","腾讯混元开源AuK:1.5B语音模型统一生成与编辑,4步推理快4.5倍","腾讯混元开源AuK语音基础模型:1.5B参数,30.3亿指令-音频实例、195万小时监督训练,把TTS、内容编辑、增强分离、副语言编辑收进同一套自然语言指令接口;蒸馏版AuK-Flash 4步推理、加速4.5倍,权重以MIT协议放出。","语音处理的工具箱向来很碎：合成用一个模型，降噪用一个模型，人声分离又一个模型，改个音高、换个情绪还得再找工具。9 月 9 日，腾讯混元团队把 AuK 开源——一个 1.5B 参数的语音基础模型，试图用「自然语言指令 + 音频上下文」这一个接口，把语音生成和编辑的全部环节收进同一个模型。\n\n## 五大任务族，一套指令接口\n\n按照技术报告，AuK 的训练覆盖五个任务族：语音生成、内容编辑、增强与分离、副语言编辑、声学编辑，对应到实际能力是 16 类任务——零样本 TTS、无参考音频的指令 TTS、改写说词的内容编辑、保留旋律改歌词的歌词编辑、按半音调整音高、调速、调音量、换情绪、换音色、去口音、增删呼吸笑声等非语言声、气声与正常语音互转，再加上降噪增强、人声分离、音乐分离和按内容定位的目标说话人提取。\n\n支撑这个能力面的数据规模不小：约 30.3 亿条指令-音频实例、195 万小时有效监督。所有任务共用同一个消息式调用接口，一条指令说清楚要什么，音频按需传入。\n\n## 三件套架构：MLLM + 联合 VAE + 混合整流流\n\n架构上 AuK 是三件套：一个多模态大语言模型负责语义条件（仓库披露当前实现用 Qwen2.5-Omni-3B 做编码器）；一个在语音、通用音频、音乐上联合训练的 VAE 负责声学条件；生成主体是混合整流流 Transformer——先走双流 MMDiT 块，再进统一单流 DiT 块。\n\n训练分三步走：先生成热身，再进生成-编辑联合预训练，后训练阶段按任务分化——开放式编辑用人类反馈偏好优化，语音生成走基于奖励的强化学习。\n\n## AuK-Flash：4 步推理，4.5 倍加速\n\n推理成本用蒸馏压：官方用一致性初始化加任务路由的 Decoupled DMD 蒸馏出 AuK-Flash，固定 4 步推理、免分类器自由引导，相同条件下比完整模型快 4.5 倍。官方报告称其在零样本与指令控制的语音生成、通用指令编辑评测中处于领先水平，在信号级修复任务上保持竞争力——这是官方评测口径，独立复现还有待社区验证。\n\n## 开源交付：MIT 协议 + 全套工具链\n\n交付面相当完整：以 MIT 协议放出代码和权重（AuK 与 AuK-Flash 两个变体），HuggingFace 与 ModelScope 双平台提供权重下载和在线 Demo，自带 ComfyUI 节点、命令行与 Python API，微调管线用 JSONL 数据格式加动态 batching。开源当天 GitHub 仓库已有 73 个 star。\n\n论文与代码：[arXiv:2609.08936](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.08936) · [GitHub 仓库](https:\u002F\u002Fgithub.com\u002FTencent-Hunyuan\u002FAuK)\n\n## 所以呢\n\nAuK 的思路值得注意：它没有在某个单一语音任务上卷指标，而是把 LLM 领域验证过的「指令统一 + 大规模多任务预训练」范式搬到语音模态，跟图像、视频领域近一年的统一模型趋势同构。1.5B 的体量加 4 步推理的蒸馏版，意味着这条路线已经摸到本地部署的门槛。\n\n对做语音应用的开发者，MIT 协议加现成微调管线基本消除了试用阻力；对垂直语音工具厂商，问题则尖锐得多——当生成、编辑、分离被一个 1.5B 模型统一，单点功能的护城河还剩多宽？","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.08936","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"0492fb2d-205e-41e9-9c08-67a011794705","en","AuK: Tencent's Open 1.5B Speech Generation and Editing Model","Tencent AuK: open 1.5B speech model unifying TTS, editing and separation in one instruction interface; Flash does 4-step inference, 4.5x faster.","Speech tooling has always been fragmented: one model for synthesis, another for denoising, another for separation, and pitch or emotion edits require yet more tools. On September 9, Tencent's Hunyuan team open-sourced AuK, a 1.5B-parameter speech foundation model that pulls generation and editing into a single interface built on natural-language instructions plus audio context.\n\n## Five Task Families, One Instruction Interface\n\nAccording to the technical report, AuK was trained across five task families — speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing — mapping to 16 concrete tasks: zero-shot TTS, reference-free instruct TTS, speech content editing, melody-preserving lyric editing, semitone-level pitch control, speed and volume adjustment, emotion and timbre changes, de-accenting, adding or removing nonverbal sounds like breaths and laughs, whisper conversion, speech enhancement, speech separation, music separation, and content-based target speaker extraction.\n\nThe data scale behind this is substantial: roughly 3.03 billion instruction-audio instances and 1.95 million hours of effective supervision. Every task shares the same message-based call — one instruction string, optional audio input.\n\n## Architecture: MLLM + Joint VAE + Hybrid Rectified Flow\n\nAuK combines three components: a multimodal LLM for semantic conditioning (the repo discloses the current implementation uses Qwen2.5-Omni-3B as encoder), a VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that runs dual-stream MMDiT blocks followed by unified single-stream DiT blocks.\n\nTraining proceeds in stages: generation-only warm-up, then joint generation-editing pre-training, then diverging post-training — human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation.\n\n## AuK-Flash: 4-Step Inference, 4.5x Faster\n\nInference cost is compressed via distillation: consistency initialization plus task-routed Decoupled DMD yields AuK-Flash, which performs fixed 4-step inference without classifier-free guidance and achieves a 4.5x wall-clock speedup over the full model under matched conditions. The report claims leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration — an official-eval caveat pending independent replication.\n\n## Open Delivery: MIT License and a Full Toolchain\n\nThe release is unusually complete: code and weights for both AuK and AuK-Flash under the MIT license, weight downloads and live demos on HuggingFace and ModelScope, ComfyUI nodes (ComfyUI-AuK), CLI and Python APIs, and a fine-tuning pipeline built on JSONL data with dynamic batching. The repo picked up 73 stars on release day.\n\nPaper and code: [arXiv:2609.08936](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.08936) · [GitHub repo](https:\u002F\u002Fgithub.com\u002FTencent-Hunyuan\u002FAuK)\n\n## So What\n\nAuK's approach is notable: rather than chasing benchmarks on a single speech task, it ports the instruction-unified, large-scale multi-task pretraining playbook — proven in LLMs — to the speech modality, mirroring this year's unified-model trend in image and video. At 1.5B parameters with a 4-step distilled variant, the route is already brushing against local-deployment thresholds.\n\nFor speech-application developers, MIT licensing plus a ready fine-tuning pipeline removes nearly all friction to trying it. For vendors of vertical speech tools, the question is sharper: when generation, editing, and separation collapse into one 1.5B model, how wide is the moat around any single-point feature?","tencent-hunyuan-auk-speech-editing","2026-09-09T09:12:00Z","2026-09-09T09:07:25.714256Z","2026-09-09T09:07:25.714271Z",true,"agent",82,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"40095b51-97b0-4fd4-9b1d-f636c970572e","阿里 Qwen 团队发布 Qwen3.8-Max:2.4 万亿参数 MoE 模型首度开放权重","qwen3-8-max-2-4t-moe-open-weights","2026-08-07T02:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"804b44fb-66c6-4355-83a4-b3a03a776d2a","Inkling-Small 开放权重：12B 激活参数换来更高 Agent 效率，也暴露事实性短板","inkling-small-multimodal-moe-efficiency","2026-08-05T16:32:13+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"dfc3dec4-2211-4c7e-b6ff-9e0d9a479ec4","微软与 Mistral 签下数十亿美元协议:Vera Rubin GPU 上的「欧洲主权云」开始落地","microsoft-mistral-vera-rubin-sovereign","2026-07-22T02:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"cf01282f-8a64-49a8-a608-9b806ccfbea3","Mira Murati 实验室 Inkling 开源：975B MoE 不卷\"最强\"，押注\"可定制\"","thinking-machines-inkling","2026-07-15T22:00:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"1a27bedc-012d-4d63-85e1-ddc57aabd8bf","ByteDance UniVR 让模型「在视觉空间里思考」：34B 参数逼近 Gemini 3 Pro + Nano Banana 2","bytedance-univr-34b","2026-07-14T12:10:00+00:00"]