[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kimi-k3-moonep-flashkda-agentenv":3,"news-related-3d8b9b1a-e038-466f-9b6b-304f911e35a7":39},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":26,"news_slug":32,"published_at":33,"created_at":34,"modified_at":35,"is_published":36,"publish_type":37,"image_url":13,"view_count":38},"3d8b9b1a-e038-466f-9b6b-304f911e35a7","Kimi K3 开源三件套 MoonEP\u002FFlashKDA\u002FAgentEnv:Moonshot 把 2.8T MoE 训练栈完整交底","7 月 27 日,月之暗面在 Kimi K3 开放日把模型权重、技术报告之外的三项核心 Infra 技术一并开源——高性能通信库 MoonEP、KDA 线性注意力算子 FlashKDA,以及与 KVCache.ai 合作的分布式 RL 沙箱 AgentEnv。Kimi K3 是一个 2.8 万亿参数 MoE 模型,每次推理仅激活 16\u002F896 专家,工程上的最大挑战是「高稀疏度下如何稳住训练与通信」: MoonEP 把细粒度专家并行的通信做到不均衡负载下仍接近线性扩展,让超大 EP 域不再因少数专家过热而卡死;FlashKDA 给出 KDA 的 Triton\u002FCUDA 级实现,在 H20 上相比 flash-linear-attention 基线 prefill 速度提升 1.72–2.22 倍,可直接作为后端替换;AgentEnv 则提供高保真、快速 fork 的沙箱,支撑 Kimi K3 后训练里大规模并行 Agent 任务。三件套覆盖「通信—算子—沙箱」三个最常被开源社区忽略的工程角,让任何人都能在同等规格下复现 K3 级别的训练流水线。结合 KDA + AttnRes 3:1 混合、Stable LatentMoE 的 Quantile Balancing、Per-Head Muon 等架构创新,Kimi K3 在算力受限的前提下把规模化效率相对 K2 提升 2.5 倍——开源的不仅是模型,更是中国大模型第一次把「2.8T 怎么训」的全部工程细节摊在桌面上。","https:\u002F\u002Fgithub.com\u002FMoonshotAI\u002FMoonEP","0ec8f614-42c7-4256-8591-209e1e39eb6b",[10,14,17,20,23],{"id":11,"name":12,"slug":12,"description":13,"color":13},"7ac06d8e-b074-4147-abfc-ffaa4c6b8744","ai-efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":24,"name":25,"slug":25,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[27],{"id":28,"lang":29,"title":30,"summary":31,"content":31},"88b586ef-20e3-48d3-bb79-06a9024381c7","en","Kimi K3 open-sources MoonEP, FlashKDA, AgentEnv training stack","On July 27, at the Kimi K3 open day, Moonshot AI released not only the model weights and technical report but also three core infrastructure technologies — the high-performance communication library MoonEP, the KDA linear-attention kernel FlashKDA, and the distributed RL sandbox AgentEnv co-developed with KVCache.ai. Kimi K3 is a 2.8-trillion-parameter MoE model that activates only 16 of 896 experts per inference, so the central engineering challenge is: how do you keep training and communication stable at extreme sparsity? MoonEP delivers near-linear scaling for fine-grained expert-parallel (EP) communication even under imbalanced loads, so very large EP domains no longer stall when a few hot experts create a bottleneck. FlashKDA provides a Triton\u002FCUDA-level implementation of KDA, achieving a 1.72x–2.22x prefill speedup over the flash-linear-attention baseline on H20, and can be dropped in as a backend replacement. AgentEnv offers a high-fidelity, fast-fork sandbox that supports the large-scale parallel agent workloads in K3's post-training pipeline. Together, the three pieces cover \"communication — kernels — sandbox\", the three engineering corners most often overlooked by the open-source community, allowing anyone with comparable hardware to reproduce a K3-class training pipeline. Combined with architectural innovations such as a 3:1 KDA + AttnRes hybrid, Quantile Balancing in Stable LatentMoE, and Per-Head Muon, Kimi K3 improves scaling efficiency by 2.5x over K2 under constrained compute — and open-sourcing the model is only part of the story. For the first time, a Chinese frontier model is laying out every engineering detail of \"how to train a 2.8T model\" on the table.","kimi-k3-moonep-flashkda-agentenv","2026-07-28T04:30:00Z","2026-07-28T04:06:17.036638Z","2026-08-19T02:08:40.142862Z",true,"agent",291,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"f6e4aab0-7693-4c2c-bb66-c1641fc2cc3e","Ox Alpha 谜底揭晓:智谱 GLM-5.3-Flash,MIT 开源 320B MoE","ox-alpha-glm-5-3-flash-reveal","2026-08-27T13:30:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"491f4904-c854-4925-b3e3-e34b8afd5e50","KDA+MLA 混合栈下沉到 1.3B 激活:Ling-3.0-tiny 把 MoE 端侧化,INT4 跑出 115 tok\u002Fs","ling-3-tiny-kda-mla-edge-deployment","2026-08-18T00:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"b0183d10-bcfd-44ed-a178-a2c813f10b69","国家超算互联网AI社区上线Kimi K3:2.8万亿参数MoE一键调用,开源大模型有了国产算力底座","kimi-k3-cnsc-internet-launch","2026-07-28T09:30:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"a16a4f36-cb00-4f14-a394-80d49075a323","上海AI Lab 发布 Intern-S2-Preview-397B：把「记忆」与「思考」拆开，397B 跑出万亿模型效果","shai-lab-intern-s2-preview-397b","2026-07-18T02:00:00+00:00"]