[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kimi-k3-minitriton-gpu-compiler":3,"news-related-a151db0c-d832-4df2-ac03-2d4e58b26e99":45},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":32,"news_slug":38,"published_at":39,"created_at":40,"modified_at":41,"is_published":42,"publish_type":43,"image_url":13,"view_count":44},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","Moonshot 在 Kimi K3 技术博客里把 LLM 的 agentic 能力直接推到了系统软件层级:模型不只在 benchmark 上和 Fable 5、GPT 5.6 Sol 贴身竞争,它真的在 sandbox 里连续 24 小时 profile、重写并 benchmark 了四个 GPU kernel 任务,在 AttnRes、KDA 和一个 512 维头部的 MLA kernel 上和 Fable 5 持平、显著压过 Opus 4.8、GPT 5.5 和 5.6 Sol。更有杀伤力的是另一个 case:K3 自己写出了一个 Triton-like 编译器 MiniTriton,带上自研的 tile-level IR、MLIR 之上的优化通道和 PTX 代码生成 pipeline,在多组 roofline benchmark 上和 Triton、torch.compile 互有胜负,还能稳定跑完 nanoGPT 端到端训练,loss 曲线几乎贴着参考实现走。K3 的架构骨架是 Kimi Delta Attention(线性注意力)+ Attention Residuals(跨层残差)+ Stable LatentMoE(896 选 16),搭配 SiTU 激活和 Gated MLA,官方给出的 scaling efficiency 比 K2 提升约 2.5 倍,2.8 万亿参数全部由 MXFP4 权重 + MXFP8 激活训出来,服务侧则依赖 64 卡以上的 supernode 配置才跑得动。技术报告还没发,论文里承诺的细节要等 7 月 27 日权重开源一起落地。问题是这些 case 都跑在 Moonshot 自家 sandbox 里,能不能复现还得等社区下场——但把「自己写编译器」写进 model card 这一步,确实把 agentic benchmark 的天花板往上抬了一截。","https:\u002F\u002Fwww.kimi.com\u002Fblog\u002Fkimi-k3","0ec8f614-42c7-4256-8591-209e1e39eb6b",[10,14,17,20,23,26,29],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":24,"name":25,"slug":25,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":27,"name":28,"slug":28,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":30,"name":31,"slug":31,"description":13,"color":13},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",[33],{"id":34,"lang":35,"title":36,"summary":37,"content":37},"4f80a8cc-e0ff-4a9a-b291-ed61a19e02b2","en","Kimi K3 builds MiniTriton: LLMs compile their own GPU kernels","In the Kimi K3 technical blog, Moonshot pushes its LLM's agentic capabilities straight to the system-software layer: the model isn't just benchmark-competitive with Fable 5 and GPT-5.6 Sol — it actually profiled, rewrote, and benchmarked four GPU kernel tasks for 24 straight hours inside a sandbox, matching Fable 5 on AttnRes, KDA, and a 512-dim-head MLA kernel, and significantly beating Opus 4.8, GPT 5.5 and 5.6 Sol. Even more striking is the other case: K3 itself wrote out a Triton-like compiler called MiniTriton, complete with a self-designed tile-level IR, an optimization pass over MLIR, and a PTX code generation pipeline — it trades wins and losses with Triton and torch.compile on multiple roofline benchmarks, and can stably run end-to-end nanoGPT training, with the loss curve almost glued to the reference. K3's architectural skeleton is Kimi Delta Attention (linear attention) + Attention Residuals (cross-layer residual) + Stable LatentMoE (16-of-896), with SiTU activation and Gated MLA. The official scaling efficiency is ~2.5x over K2; all 2.8T parameters are trained with MXFP4 weights + MXFP8 activations, and the serving side requires a 64+ GPU supernode to run. The technical report hasn't been published yet — the paper's promised details have to land together with the July 27 weight release. The open question is whether all of these cases are reproducible on community infrastructure, since they all ran inside Moonshot's own sandbox — but the act of writing \"we built our own compiler\" into the model card does push the ceiling of agentic benchmarks up another notch.","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00Z","2026-07-26T10:03:43.050611Z","2026-08-19T02:08:40.142862Z",true,"agent",208,{"items":46},[47,52,57,62,67,72],{"id":48,"title":49,"news_slug":50,"published_at":51},"3d8b9b1a-e038-466f-9b6b-304f911e35a7","Kimi K3 开源三件套 MoonEP\u002FFlashKDA\u002FAgentEnv:Moonshot 把 2.8T MoE 训练栈完整交底","kimi-k3-moonep-flashkda-agentenv","2026-07-28T04:30:00+00:00",{"id":53,"title":54,"news_slug":55,"published_at":56},"4d436945-18e9-4d69-a4c8-c1e3e975ab33","MiniMax M3发布：稀疏注意力打通百万token上下文，开源模型编程能力逼近闭源前沿","MiniMax-m3-sparse-attn-million-token-msa","2026-06-04T01:00:00+00:00",{"id":58,"title":59,"news_slug":60,"published_at":61},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00"]