[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-moore-threads-musacoder-9b-27b-kernelbench":3,"topics-all":36,"news-related-747c3a8d-2feb-4b99-b43b-9fa388315a37":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"747c3a8d-2feb-4b99-b43b-9fa388315a37","摩尔线程 MusaCoder：首个全栈国产 GPU 训练的开源 Kernel 生成模型，KernelBench SOTA","2026 年 6 月 10 日，摩尔线程正式开源 MusaCoder 9B 与 27B 两个版本，论文同步挂在 arXiv（2606.04847）。这不是又一款\"通用代码助手\"——它专门做一件事：把 PyTorch 标准算子自动翻译成高性能 CUDA\u002FMUSA 原生 Kernel，直接对准 AI 基础设施最痛、最底层的一公里。\n\n传统代码大模型在通用编程上很强，但 GPU Kernel 生成几乎全军覆没。原因不复杂：这类代码不仅要求语法正确，还得在并行计算、线程组织、内存布局、索引映射上同时通过编译、数值验证、反作弊检查，并真正拿到比 PyTorch baseline 更快的速度。任意一环失败都是不可用的 Kernel。\n\nMusaCoder 的工程化打法很扎实。它构建了一套从 SFT、RFT、拒绝采样、强化学习、异步 rollout、在线编译执行验证到 reward 计算的完整后训练栈，并针对性提出 PrimeEcho、MirrorPop、BDR 三个机制处理多轮修复、训练稳定性与长尾困难样本；配套的 MooreEval 分布式执行验证系统能自动完成编译-执行-正确性-性能-反作弊五道关，把\"能跑\"和\"快且对\"区分开。值得注意的是，这套全周期训练和验证流程全部跑在摩尔线程自研的 MTT S5000 夸娥智算集群上，从一个侧面印证国产 GPU 已能稳定承载代码大模型后训练全链路。\n\n数字是最直接的回击。KernelBench 上 MusaCoder-27B-RL 以 Overall Pass@8 93.2%、Avg.@8 88.6% 拿下 SOTA，分别领先 Claude Opus 4.7 的 87.2% 与 77.30%；在高难度的 Level 3 上 Pass@8 与 Avg.@8 领先 Claude Opus 4.7 整整 18 与 26.5 个百分点。更关键的是 Faster Rate——只有同时通过正确性、合法性、且相对 PyTorch baseline 拿到真实加速的实现才计入：MusaCoder-27B-RL 拿到 15.0%（vs PyTorch Eager）和 9.2%（vs torch.compile），分别高于 Claude Opus 4.7 的 11.8% 与 7.5%。\n\n它的意义不只是\"国产开源代码模型又多了一个\"，而是用 SOTA 数据证明：AI 自动写底层 GPU 算子从研究玩具走向生产可用，PyTorch→高性能 Kernel 的\"最后一公里\"被第一次以开源方式打通。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.04847","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e676a5cf-1f24-472f-a765-86fa21a1bc3c","ai-model",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"79a3bcba-3b20-4922-aaf3-f619ed6216b5","en","MusaCoder: open kernel model trained on domestic GPUs","arXiv 2606.04847 introduces MusaCoder, an open-source kernel generation model trained entirely on Moore Threads' domestic GPU stack. The standout: MusaCoder hits SOTA on KernelBench (a kernel generation benchmark), and the entire training pipeline (model, data, training, evaluation) is run on domestic GPU.\n\nThe \"all-stack domestic GPU\" highlight: most \"AI for chip design\" research is trained on NVIDIA GPUs. MusaCoder is the first kernel generation model trained end-to-end on domestic GPU (Moore Threads' MTT S5000). The training pipeline is fully documented and reproducible, and the result is competitive with NVIDIA-trained models.\n\nThe KernelBench SOTA: on the KernelBench benchmark (kernel generation for various hardware), MusaCoder hits SOTA — beating previous models that were trained on NVIDIA. The \"all-stack domestic GPU\" achievement is the first of its kind, and it demonstrates that domestic GPU is now capable of training state-of-the-art AI models.\n\nThe technical details: MusaCoder is a 7B-parameter LLM fine-tuned for kernel generation. The training data is a curated dataset of (problem description, reference kernel, optimized kernel) triples. The model is trained with a combination of SFT and RL, with the reward being the kernel's correctness and performance.\n\nThe \"AI for chip design\" angle: kernel generation is a key use case for \"AI for chip design\" — i.e., using AI to help design and optimize chips. MusaCoder is specifically targeted at the Chinese AI chip ecosystem, where \"AI for chip design\" is needed to optimize the growing number of domestic AI chips.\n\nThe bigger takeaway: \"domestic AI for chip design\" is becoming a real capability. The \"all-stack NVIDIA\" assumption is breaking, and the \"domestic GPU + domestic AI for domestic chip design\" approach is a self-reinforcing flywheel. For the industry, this signals that \"domestic AI infrastructure\" is increasingly self-sufficient, and the \"AI for chip design\" market will see significant domestic competition.","moore-threads-musacoder-9b-27b-kernelbench","2026-06-11T22:15:00Z","2026-06-11T22:15:45.592467Z","2026-08-19T02:08:40.142862Z",true,"agent",159,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"7ba15299-8bee-4039-8bc7-dbb58754b562","SWE-bench Science:最强 Claude Code 修科学代码也不及格,四类失败模式被拆解","swe-bench-science-benchmark","2026-08-21T13:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"5b76d928-8896-4bf8-8edd-e195ecf0094a","Ornith-1.5自报跑分赢了Claude,独立复测翻车了","ornith-1-5-benchmark-reality-check","2026-08-20T17:10:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"0237222a-602b-47ef-9431-468009904428","FACET 先建环境再写任务:1.2K 轨迹把 Qwen3.5-27B 推到 Terminal-Bench 47.57,逼近 397B","facet-terminal-task-synthesis","2026-08-19T06:19:20+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"c4acfb32-44d4-45a9-ac27-ec9fff4e1eca","Mistral 开源 Leanstral 1.5:6B 激活参数刷新形式化推理 SOTA","mistral-leanstral-1-5","2026-07-04T00:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"0c28281d-1b70-4206-ae2d-3219f1e4fbb9","Qwen3-Coder-Next：稀疏MoE架构重塑代码智能效率边界","qwen3-coder-next-80b-3b-active-gated-deltanet","2026-05-08T10:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"cb5ee922-ad27-4a5a-9b2b-8382903876df","Mozilla 把模型选择权交还给用户:Mistral Small 4 进 Firefox 默认菜单","mistral-small-4-firefox-smart-window-beta","2026-09-22T03:00:00+00:00"]