[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-yue2-3b-editable-music-generation":3,"topics-all":38,"news-related-17006864-46a5-405c-a8cc-24507bbc5e37":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"17006864-46a5-405c-a8cc-24507bbc5e37","YuE2-3B 开源:乐谱可编辑的音乐生成,官方基准反超 Suno v5","OpenMOSS 放出音乐生成模型 YuE2-3B:3B 参数支持中英双语,先用 ABC 乐谱做符号规划再合成音频,24GB 显卡 71 秒出一首带人声的完整歌曲,官方 WildSongBench 基准 best-of-8 平均 6.9632 微超 Suno v5,CC BY-NC 4.0 研究许可。","开源音乐生成一直有个尴尬:文本和图像模型早已进入「生成—修改—再生成」的工作流,音乐却还是按下按钮、整首抽卡。9 月 10 日,音乐生成模型 YuE2-3B 上架 Hugging Face([模型卡](https:\u002F\u002Fhuggingface.co\u002Fmrfakename\u002FYuE2-3B)),据开源模型追踪站 The Open Weights 报道,它出自 OpenMOSS,思路正是冲着这个痛点:先写乐谱,再生成歌曲。\n\n## 乐谱变成一等公民\n\n主干是一套 AR–NAR Mixture-of-Transformers:模型先输出 ABC 记谱法写成的乐谱与语义 token,再经 flow matching 生成声学隐向量,最后由 VAE 解码成 48 kHz 立体声。规划与合成在接口层也是分开的——`pipe.plan()` 先出乐谱,你可以手动改旋律、改和弦,再交给后续管线渲染音频。三档控制粒度:旋律加和弦、只锁旋律、完全放开。翻唱走同一条路:先用 SheetSage2 把原曲转成 ABC 乐谱,再用 Qwen3-ASR 等工具对齐歌词,换风格重新渲染。\n\n模型卡还展示了一个 9 步 14 版本的 agentic editing 案例:把「换成爵士、加段萨克斯 solo」这类反馈交给 agent,由它改乐谱、改风格、改歌词,YuE2 渲染下一版——音乐生成从一次性抽卡变成可对话的协作。\n\n## 71 秒一首,家用显卡跑得动\n\n参数量 3B 级,部署门槛随之下降:官方实测 RTX 4090 上一首 3.6 分钟的歌 71 秒生成,峰值显存约 11.2 GiB,无需量化;一张 24GB 显卡加 24GB 内存即可跑通创作、翻唱、编辑全流程。服务器侧配 vLLM 0.19,32 路并发时吞吐 3231 tokens\u002F秒、每小时约 373 首。\n\n## 官方基准上的成绩单\n\n在团队自建的 WildSongBench(192 条 prompt)上,YuE2 best-of-8 的 SongBench 平均分 6.9632,高于表内全部开源与闭源对手——Suno v5 为 6.8721,开源阵营的 LeVo 2(6.3247)、MiniMax Music 3(6.2830)被明显拉开;零样本翻唱任务 SHS100K 上 CLEWS Hit@1 达 71.3%。需要说明:这是官方基准、官方口径,技术报告标注 coming soon,数字宜等独立复测再下结论([追踪站报道](https:\u002F\u002Ftheopenweights.com\u002Fnews\u002Fyue2-3b-m3dl))。\n\n权重以 CC BY-NC 4.0 许可放出——可研究、可个人使用,商用不行;模型卡显示近 30 天下载 19 次,它才刚进入社区视野。\n\n对做音乐的人,这次发布真正值得记的不是跑分,而是它把「乐谱」——音乐行业的源代码——还给了用户:不满意改两小节重新渲染,比整首重抽划算得多。3B 级参数就把完整歌曲生成塞进家用显卡,也再次说明垂直模态未必需要通用大模型的体量。接下来看正式技术报告与社区翻玩。","https:\u002F\u002Ftheopenweights.com\u002Fnews\u002Fyue2-3b-m3dl","d67f395d-c387-4cbc-b571-5c79814a7bda",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"88ad6dca-b520-46d5-ac47-47b7cc351fcb","en","YuE2-3B Brings Editable Scores to Open Music Generation","YuE2-3B writes editable ABC scores, then renders full vocal songs in 71s on a 24GB GPU; official WildSongBench best-of-8 edges Suno v5.","Open music generation has long had an awkward gap: text and image models already live in a generate-revise-regenerate loop, while song models still work like pulling a slot-machine lever. On September 10, the music generation model YuE2-3B landed on Hugging Face ([model card](https:\u002F\u002Fhuggingface.co\u002Fmrfakename\u002FYuE2-3B)); according to the open-weights tracker The Open Weights, it comes from OpenMOSS, and its whole design attacks that pain point: write the score first, then sing it.\n\n## The score becomes a first-class citizen\n\nThe backbone is an AR–NAR Mixture-of-Transformers: the model first emits a score written in ABC notation plus semantic tokens, then produces acoustic latents through flow matching, and a VAE decodes them into 48 kHz stereo audio. Planning and synthesis are also separated at the API level — `pipe.plan()` returns the score, you can edit melody and chords by hand, then hand it back to the pipeline for rendering. Three control levels are offered: melody plus chords, melody-only, and fully free generation. Covers follow the same road: transcribe the original song into an ABC score with SheetSage2, align lyrics with tools like Qwen3-ASR, then re-render in a new style.\n\nThe model card also demos an agentic-editing case of 9 steps and 14 versions: feedback like \"make it jazz, add a sax solo\" goes to an agent that revises the score, style and lyrics while YuE2 renders the next version — music generation turns from a one-shot gamble into a conversational collaboration.\n\n## A song in 71 seconds on a consumer GPU\n\nWith a 3B-class parameter count, the deployment bar drops accordingly: official tests show a 3.6-minute song generated in 71 seconds on an RTX 4090, peak VRAM around 11.2 GiB, no quantization needed; a single 24GB GPU plus 24GB of host RAM covers the full create-cover-edit workflow. On the server side, with vLLM 0.19 at 32-way concurrency the system sustains 3,231 tokens\u002Fs and roughly 373 songs per hour.\n\n## The report card on an official benchmark\n\nOn the team's own WildSongBench (192 prompts), YuE2 best-of-8 reaches a SongBench average of 6.9632, above every open and proprietary system in the table — Suno v5 sits at 6.8721, and open-source rivals LeVo 2 (6.3247) and MiniMax Music 3 (6.2830) trail well behind; on the zero-shot cover task SHS100K, CLEWS Hit@1 hits 71.3%. One caveat: these are official numbers on a self-built benchmark, the technical report is marked \"coming soon,\" and the figures deserve independent replication before being treated as settled ([tracker coverage](https:\u002F\u002Ftheopenweights.com\u002Fnews\u002Fyue2-3b-m3dl)).\n\nThe weights ship under CC BY-NC 4.0 — fine for research and personal use, off limits for commercial deployment; the model card shows 19 downloads in the last 30 days, so it has only just entered the community's field of view.\n\nFor people who make music, the real takeaway is not the benchmark score but the fact that the score itself — the source code of the music industry — is handed back to the user: editing two bars and re-rendering beats re-rolling an entire song. And a 3B-class model squeezing full song generation onto a consumer GPU is one more sign that vertical modalities may not need frontier-model scale. What to watch next: the formal technical report and what the community does with it.","yue2-3b-editable-music-generation","2026-09-10T13:20:00Z","2026-09-10T13:07:34.625149Z","2026-09-10T13:07:34.625160Z",true,"agent",199,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"b6b9f5c8-0d71-4288-8782-0284fccfca8f","商汤 SenseNova-Vision：把「检测\u002F分割\u002F深度估计」统统塞进同一个生成式多模态基座","sensetime-sensenova-vision","2026-07-08T10:15:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"a36d9d97-42de-4c87-88e9-cdc173b9ab4b","VLX-Seek 1.5 把端侧具身感知切成 0.6B\u002F3B\u002F10B 三档：用 None 输出压住目标幻觉","vlx-seek-1-5","2026-07-06T02:00:00+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"f9bf6e21-2e8a-4571-ab7d-a4dba727b72a","ViiTorVoice-NAR：把 TTS 的「改一句重录」变成「改一词局部合成」","viitor-voice-nar-local-tts","2026-07-02T14:15:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"51c13e24-8072-404c-a8d4-75c40cff05ee","Ling-3.0-flash-VL 开源：124B MoE 只激活 5.5B，视觉塞进 Agent 闭环","ling-3-0-flash-vl-open-weights","2026-09-15T13:18:00+00:00"]