[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kimi-k3-48h-saturate-chinese-compute-supernode":3,"news-related-bce0fe8f-14de-4ffc-8c22-2d798e711e73":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"bce0fe8f-14de-4ffc-8c22-2d798e711e73","Kimi K3 上线 48 小时打满集群:开源旗舰正在把推理算力拖进新一轮\"卖方周期\"","Kimi K3 上线仅 48 小时就打满现有推理集群,触发月之暗面暂停 C 端新用户订阅。银河证券最新研报指出,这一事件把开源旗舰的真实算力饥渴暴露在卖方视野:国产开源模型能力逼近全球顶级闭源模型,倒逼头部厂商以更大规模训练和更快迭代维持差异化,同时持续引爆推理算力需求。Kimi K3 已与多家国产算力平台完成 Day-0 极速适配,正在拉动服务器、交换机、光模块、液冷到电源整条国产超节点产业链的配套需求。","## Kimi K3 上线 48 小时打满集群:开源旗舰正在把推理算力拖进新一轮\"卖方周期\"——而国产超节点卡在了最紧的那一环\n\n**2026 年 8 月 2 日**,一则来自卖方的研报把最近 LLM 圈最戏剧化的数字扒了出来:月之暗面 Kimi K3 上线后 48 小时,请求量就把现有推理集群打满,公司不得不宣布暂停 C 端新用户会员订阅。银河证券在最新报告中把这件事定性成\"**Kimi K3 并非抑制算力需求**\",并由此拉开了对国产超节点产业链的全面推荐。这不是普通的\"国产替代\"逻辑,而是首次由头部开源旗舰的真实生产流量触发的、对国产算力链的**自下而上**的卖方推荐。\n\n### 48 小时打满,意味着什么\n\n\"打满\"这两个字在 LLM 圈并不稀罕,但这次格外扎眼:Kimi K3 是 2.8 万亿参数 MoE,6 月公开推理栈、7 月底权重进一步扩散,围绕它的调用方既包括 ToC 用户,也包括 ToB 集成方。理论上,MoE 的稀疏激活把单次请求的算力开销压到了接近稠密 7B 的水平——但所有这些优化都被**两件事**轻松抵消:**请求量**与**长上下文**。前者在 K3 这类\"敢开源且敢放开给 C 端\"的产品上几乎以指数级暴涨,后者把每条请求的单次 FLOPs 又往上拉了一大截。\n\n更关键的是,48 小时打满这件事发生在**模型开源后不久**。这意味着:不再是闭源模型的 API 容量紧张问题,而是**整个开源生态的算力分配问题**。任何一家想跑 K3 全尺寸权重或蒸馏版的厂商,都得提前把机柜准备好——而机柜不是按月可以扩的。\n\n### 卖方为什么把\"Day-0 适配\"看作真信号\n\n银河证券研报里反复强调的一个词是**\"Day-0 极速适配\"**。过去两年,国产算力平台(无论是 Ascend、寒武纪、海光,还是各类互联交换方案)和开源大模型的搭配常常**滞后 1-3 个月**——等芯片厂商拿到模型权重、做完算子 mapping、走通编译器栈、再到推理框架的端到端测试,机会窗口早就过了。\n\n而这次 K3 是**还没正式开源**——仅仅是预热版本,各国产算力平台就已经把 Day-0 适配跑通了。这背后至少有两件事值得记一笔:一是 K3 在工程层面对编译、kernel 调度做了较深度的开放(配合 MiniTriton 这类自举式 GPU 编译器栈);二是国产算力平台把\"开源旗舰 Day-0\"作为**标准动作**写进了内部排期,而不是可选项。这两点合起来,意味着国产超节点不再仅仅是\"替代者\",而是\"**第一个吃到生态红利的人**\"。\n\n### 不是一锤子需求,是结构性的\n\n卖方推荐国产超节点产业链的逻辑里,有一个常被忽略的回路:**开源旗舰每前进一步,都要反过来要求头部闭源厂商加速训练与迭代**。也就是说,Kimi K3 拉爆的不只是**自己的**推理算力,还顺手抬高了 GPT、Claude、Gemini 这一档次的训练成本——因为训练侧必须做更大规模、更快节奏的实验,才能维持差异化。这条传导链是双向的:\n\n- 开源旗舰 → 推理算力需求激增 → 国产超节点现货吃紧\n- 开源旗舰 → 闭源厂商训练投入加码 → 训练侧算力水位被撑高\n\n两个方向同时把水位线往上抬,卖方当然愿意喊\"产业链配套需求\"。\n\n### 但国产超节点要接住的挑战并不小\n\n1. **互联与拓扑**:千亿到万亿级 MoE 对 NVLink\u002F自研互联的拓扑、跨节点带宽、collective 库的优化提出了远超传统训练的要求。不是\"插上卡就能跑\"。\n2. **液冷与功耗密度**:超节点形态下单机柜功耗密度正在突破 100kW,液冷不是\"加分项\"而是\"门槛\"。冷却系统的国产化与机房改造的协同,会影响真正放量节奏。\n3. **编译器栈与一致性**:开源模型升级飞快,Day-0 适配的代价是**持续维护**——一次适配不等于永远适配,这对国产算力平台的工程团队规模和反应速度提出了很高要求。\n4. **光模块与配套**:当集群规模上量后,光模块、交换机、电源的瓶颈并不会被\"算力\"自己掩盖——它们会一个个浮上水面。\n\n### 怎么看这件事\n\nKimi K3 的 48 小时\"打满-暂停\"不是一次偶发事件,而是**一个信号**:开源旗舰已经正式进入了\"自下而上拉动算力链\"的阶段。这意味着:\n- 对**国产超节点及其上游(光模块、液冷、电源、交换机)**来说,接下来 3-6 个月最值得跟踪的不是哪家模型更聪明,而是**哪家算力平台能把 Day-0 适配做成流水线**\n- 对**闭源头部厂商**来说,差异化必须从\"更大\"转向\"更快+更专\",否则开源旗舰的算力马拉松会把他们拖进比拼烧钱速度的困局\n- 对**整个生态**来说,LLM 行业第一次出现了\"**算力卖方周期**\"驱动开源模型迭代的现象——这与早几年\"算法领先驱动算力投入\"的方向正好反过来\n\n> 一句话:Kimi K3 把国产超节点从\"备选项\"拉成了\"必选项\"。接下来要看的是配套链条谁能跟上、跟上多快。\n\n—— *本文综合自银河证券 2026 年 8 月研报要点与 36 氪、界面新闻相关报道,数据时点为 2026-08-02。*","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3921888283848068","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":19,"name":20,"slug":20,"description":14,"color":14},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":22,"name":23,"slug":23,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"32e32174-9f39-45d9-bac8-fafb519165a6","en","Kimi K3 saturates its cluster in 48 hours of launch","Just 48 hours after launch, Kimi K3's request volume saturated the existing inference cluster, forcing Moonshot AI to pause new C-end subscriber sign-ups. In a fresh research note, Galaxy Securities frames the event not as compute demand being suppressed, but as the moment open-source flagships exposed their true hunger for inference compute: domestic open-source models are now approaching the capability ceiling of the world's top closed-source models, forcing frontier vendors to keep training larger and iterating faster; meanwhile Kimi K3 has already achieved Day-0 adaptation with multiple domestic compute platforms, pulling through demand across servers, switches, optical modules, liquid cooling, and power — the full domestic super-node supply chain.","## Kimi K3 Saturated Its Cluster in 48 Hours: How an Open-Source Flagship Just Pulled Inference Compute Into a New Sell-Side Cycle — and Why Domestic Super-Nodes Are the Bottleneck\n\n**August 2, 2026.** A sell-side research note just pulled the most dramatic number from this week's LLM cycle into the open: 48 hours after launch, Kimi K3's request volume saturated Moonshot AI's existing inference cluster, forcing the company to pause new C-end subscriber sign-ups. In its latest note, Galaxy Securities characterizes this not as compute demand being \"suppressed\" by open-source releases, but as the moment when an open-source flagship exposed its true compute hunger — and, for the first time, generated a **bottom-up** sell-side recommendation across the full domestic super-node supply chain.\n\nThis is not the usual \"domestic substitution\" narrative. This is the first time real production traffic from a top-tier open-source model has triggered a sell-side recommendation for the domestic compute stack.\n\n### What \"48 hours to saturation\" actually means\n\n\"Saturated\" is not rare in the LLM world, but this one cuts deeper. Kimi K3 is a 2.8-trillion-parameter MoE. The inference stack was opened up in June, weights spread further in late July, and the call surface now includes both ToC users and ToB integrators. In theory, MoE's sparse activation should keep per-request compute close to a dense 7B baseline — but two forces swat that assumption aside: **request volume** and **long context**. Request volume on a model as bold as K3 (open weights, real C-end availability) explodes almost exponentially; long context pushes the per-request FLOPs back up.\n\nThe more important point: saturation happened **shortly after open-sourcing**. This is no longer a closed-source API capacity story — it is a **whole open ecosystem compute-allocation story**. Any vendor that wants to run K3 full weights or a distilled variant has to have racks ready in advance. And racks are not something you can add month by month.\n\n### Why sell-side treats \"Day-0 adaptation\" as the real signal\n\nGalaxy's note hammers on a phrase: **\"Day-0 rapid adaptation\"**. Over the past two years, the gap between a domestic compute platform (Ascend, Cambricon, Hygon, plus a roster of interconnect and switching vendors) and an open-source flagship has often been **1–3 months**: by the time a chip vendor ingests the weights, builds the operator mappings, lights up the compiler stack, and runs end-to-end inference framework tests, the window is gone.\n\nThis time, K3 wasn't even formally open-sourced — just a pre-release — and domestic compute platforms had **already shipped Day-0 adaptation**. Two things made that possible: first, K3's engineering stack opened up the compiler and kernel scheduling layers in some depth (alongside the MiniTriton-style self-bootstrap GPU compiler work); second, domestic compute platforms have elevated \"open-source flagship Day-0\" into a **default line item** in their planning, not an option. Together, these moves re-position domestic super-nodes from \"substitutes\" to \"**first to harvest the ecosystem dividend**\".\n\n### Not a one-shot demand spike — a structural shift\n\nThere is a loop in the sell-side logic that often goes unnoticed: **every step an open-source flagship takes forward pressures closed-source leaders to accelerate training and iteration**. In other words, Kimi K3's saturation does not just light up **its own** inference compute — it also raises the training-cost bar for GPT, Claude, and Gemini tiers, because training has to scale further and iterate faster to defend differentiation. The transmission is bidirectional:\n\n- Open-source flagship → inference demand surges → domestic super-node spot capacity tightens\n- Open-source flagship → closed-source training spend ramps → training-side compute waterline rises\n\nBoth directions push the waterline up, and sell-side is content to call that \"supply-chain pull-through\".\n\n### But the challenges the domestic super-node chain must catch are not small\n\n1. **Interconnect and topology.** Trillion-parameter MoE pushes NVLink and domestic interconnect far beyond classical training topologies: cross-node bandwidth and collective-library optimization become gating items. You cannot just \"plug in cards and run\".\n2. **Liquid cooling and power density.** At super-node scale, single-rack power density is breaking past 100 kW. Liquid cooling is no longer a bonus — it is the entry ticket. The choreography between domestic cooling vendors and data-center retrofit work will dictate true ramp timing.\n3. **Compiler stack and continuity.** Open-source models iterate fast; Day-0 adaptation is a **maintenance burden** — one adaptation does not equal forever-adaptation. That puts real pressure on engineering team size and reaction speed at every domestic compute platform.\n4. **Optical modules and supporting gear.** Once cluster scale goes up, the bottlenecks hidden behind \"compute\" — optics, switches, power — surface one after another.\n\n### How to read this\n\nThe 48-hour \"saturated-paused\" event for Kimi K3 is not an incident; it is a **signal**: open-source flagships have formally entered the phase of \"bottom-up pulling on the compute chain\". What that means:\n\n- For **domestic super-nodes and their upstream (optics, liquid cooling, power, switches)**, the next 3–6 months are not about which model is smarter — they are about **which compute platform can turn Day-0 adaptation into a pipeline**.\n- For **closed-source frontier vendors**, differentiation has to shift from \"bigger\" to \"faster + more specialized\" — otherwise the open-source flagship compute marathon will drag them into a burn-rate arms race.\n- For **the ecosystem overall**, this is the first time the LLM industry has seen a **compute sell-side cycle** drive open-source iteration — almost the inverse of the earlier era when algorithmic breakthroughs drove compute investment.\n\n> One line: Kimi K3 just promoted domestic super-nodes from \"an option\" to \"the option\". The next thing to watch is who in the supporting chain can keep up, and how fast.\n\n— *Compiled from Galaxy Securities' August 2026 research note and related coverage on 36Kr and Jiemian, with data as of 2026-08-02.*","kimi-k3-48h-saturate-chinese-compute-supernode","2026-08-02T06:04:11Z","2026-08-02T06:04:47.642599Z","2026-08-02T06:04:47.642612Z",true,"agent",138,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"1ab2732b-519c-45fd-acd7-27af601b4cb9","摩尔线程 MTT AICUBE 开启预售：自研\"长江\"SoC 50 TOPS 异构算力，本地大模型终于跑进家庭","moore-threads-aicube-yangtze-50-tops-home","2026-06-18T10:51:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"4aa9534a-778e-4cd7-8194-fdf3097249b8","OpenAI Jalapeño Hot Chips 实测:峰值每瓦 1.9×,延迟压到 1 秒","openai-jalapeno-hot-chips-benchmark-2026","2026-08-26T02:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"bb1d01e1-c61f-4428-b91a-41000a26bec5","从「堆硬件」到「卖Token」:10余家上市公司押注Token工厂,算力行业 TaaS 模式浮出水面","china-ai-token-factory-taas-shift-2026","2026-07-31T04:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"b0183d10-bcfd-44ed-a178-a2c813f10b69","国家超算互联网AI社区上线Kimi K3:2.8万亿参数MoE一键调用,开源大模型有了国产算力底座","kimi-k3-cnsc-internet-launch","2026-07-28T09:30:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"b86fa77c-b9cb-406f-af26-7c79329c6d43","真武M890超节点跑通Qwen3.8：国内首个2万亿参数模型推理集群落地阿里云百炼","alibaba-zhenwu-m890-qwen-3-8","2026-07-23T10:05:00+00:00"]