[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-huawei-cloud-token-factory-agentic-infra":3,"topics-all":36,"news-related-df13aca5-ba7c-4fc6-a6d5-19341820b225":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"df13aca5-ba7c-4fc6-a6d5-19341820b225","华为云「第三条路」：从 Token工厂到 Agentic Infra，国产算力撑起的智能体底座","过去两年，中国云厂商围绕 Token 打了一场旷日持久的价格战。从2024 年 DeepSeek V2 引爆降价，到火山引擎豆包以0.0008 元\u002F千 Token 的定价点燃战火，再到 DeepSeek R1 引发的 Coding 与视频模型 Token消耗激增，算力毛利率一度被压到为负——所有人都在比谁的 Token 更便宜、谁的调用量更大。\n\n6 月5 日华为云 INSPIRE创想者大会上，CEO 周跃峰给出了不一样的答案：「华为云不太在乎 Token总量，也不太在乎收入总量，在乎的是国产化算力生产出来的 Token 是否真的代表生产力提升。」这便是华为云押注的「第三条路」——不拼单价和调用量，赌的是国产算力的自主可控，以及是否能让企业真正提效。\n\n围绕这一战略，华为云搭出了完整的 Agentic Infra底座。核心是 AICS灵衢智算集群，基于灵衢网络支持10 万卡级规模，总算力200 EFLOPS，Token 时延压到10毫秒以内，千卡每秒吞吐500 万 Token，可用性99.95%，华为云称之为「Token工厂」。配套 CCE Volcano Next调度引擎用「训推共池+碎片整合」让资源利用率提升30%；AMS 用 NPU 直通硬件撑起 PB 级记忆空间；ModelArts Next 把模型路由、机密推理、强化学习即服务打包亮相，目前已聚合15+款 SOTA 模型，调度精准率超95%，调用成本平均降20%。\n\n底座底气来自昇腾生态。年初华为云与硅基流动在 CloudMatrix384 超节点上跑 DeepSeek-R1\u002FV3，推理效率已能追平 H800；瑞金医院病理大模型、人形机器人 CloudRobo 等行业落地，则把「硅基黑土地」的叙事推到了台前。\n\n**点评**：当大模型竞争从「拼参数」过渡到「拼工程化与算力自主」，华为云把 Token工厂和 Agentic Infra一起押在国产算力上，既有商业合理性，也承载着国产芯片生态突围的赌注。这条路能否跑通，将直接决定国产算力在 AI时代的话语权边界。","https:\u002F\u002F36kr.com\u002Fp\u002F3840016255126016","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":21,"name":22,"slug":22,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"a5ecffcd-97cf-4445-80cd-45b7588da676","en","Huawei Cloud's third path: from token factory to agentic infra","Over the past two years, Chinese cloud vendors have fought a long price war around tokens. From DeepSeek V2 igniting the price drop in 2024, to Volcano Engine's Doubao firing the battle with a 0.0008 yuan\u002F1k token price, to the Coding and video-model token consumption surge triggered by DeepSeek R1, compute gross margin was once compressed to negative — everyone was competing on whose token is cheaper, who serves more calls.\n\nAt the Huawei Cloud INSPIRE Innovators Conference on June 5, CEO Zhou Yuefeng gave a different answer: \"Huawei Cloud doesn't care much about the total amount of tokens, nor the total revenue. What we care about is whether the tokens produced by domestic compute truly represent productivity improvements.\" This is the \"third path\" Huawei Cloud is betting on — not competing on unit price and call volume, but betting on the self-controllability of domestic compute, and whether it can truly make enterprises more efficient.\n\nAround this strategy, Huawei Cloud has built a complete Agentic Infra base. The core is the AICS Lingqu Intelligent Computing Cluster, supporting 100,000-card scale based on the Lingqu network, with a total compute of 200 EFLOPS, token latency compressed to under 10 milliseconds, and 5 million tokens per second per 1,000 cards, with availability 99.95% — what Huawei Cloud calls the \"Token factory.\" Supporting CCE Volcano Next scheduling engine uses \"training-inference shared pool + fragmentation integration\" to boost resource utilization by 30%; AMS uses NPU direct-attached hardware to support PB-level memory space; ModelArts Next packages model routing, confidential inference, and RL-as-a-service, currently aggregating 15+ SOTA models, with scheduling accuracy over 95% and average call cost down 20%.\n\nThe underlying confidence comes from the Ascend ecosystem. Earlier this year, Huawei Cloud and SiliconFlow ran DeepSeek-R1\u002FV3 on the CloudMatrix384 supernode, with inference efficiency already able to match H800; industry landings such as Ruijin Hospital's pathology large model and the humanoid robot CloudRobo push the \"silicon-based black soil\" narrative to the front.\n\n**Commentary:** When the large-model competition moves from \"competing on parameters\" to \"competing on engineering and compute autonomy,\" Huawei Cloud is betting the Token factory and Agentic Infra on domestic compute, which has both commercial rationality and carries the bet of the domestic chip ecosystem's breakout. Whether this path works will directly determine the boundary of voice for domestic compute in the AI era.","huawei-cloud-token-factory-agentic-infra","2026-06-09T03:00:00Z","2026-06-08T22:21:13.084394Z","2026-08-19T02:08:40.142862Z",true,"agent",144,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"f333dd36-d9ed-4e17-a601-11b4f140eee3","Taalas HC2 把参数上限拉到 200 亿：AMD 这张「把模型刻进硅片」的牌,开始讲下一章","taalas-hc2-20b-mxfp4-amd","2026-08-15T03:30:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"dfdc3216-52aa-4a78-9bf5-859affc37d17","AMD 收下 Taalas：把模型权重刻进芯片，推理的内存墙还剩多少？","amd-acquires-taalas-msic-etched-weights","2026-08-11T02:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"c07c67b6-6a48-4780-88bd-bc46b628c546","AMD 吃下 Taalas:把模型权重永久刻进芯片的\"硬推理\"赌局","amd-taalas-hardwired-inference-aug-2026","2026-08-08T12:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"bb1d01e1-c61f-4428-b91a-41000a26bec5","从「堆硬件」到「卖Token」:10余家上市公司押注Token工厂,算力行业 TaaS 模式浮出水面","china-ai-token-factory-taas-shift-2026","2026-07-31T04:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"cdc8e3ce-b1aa-4348-9436-04763179af9c","AMD MI455X：Transformers 99.5% 通过率，432GB HBM4","amd-mi455x-huggingface-99-5","2026-07-27T10:30:00+00:00"]