[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-alibaba-zhenwu-m890-qwen-3-8":3,"news-related-b86fa77c-b9cb-406f-af26-7c79329c6d43":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"b86fa77c-b9cb-406f-af26-7c79329c6d43","真武M890超节点跑通Qwen3.8：国内首个2万亿参数模型推理集群落地阿里云百炼","7月23日，阿里云真武M890超节点完成Qwen3.8全量适配，并在百炼平台开放推理服务。这是国内首个成功运行超2万亿参数大模型的超节点，对国产超节点+超大参数模型协同部署具有里程碑意义。\n\n真武M890是阿里自研的128卡超节点,今年5月首次发布时定位于Agent并发推理,主打高密度、低延迟与多任务并行。Qwen3.8是阿里千问的最新一代旗舰预览版,2.4T参数规模,定位开源对标闭源顶级模型。这次真武+Qwen3.8的组合,意味着硬件层和模型层不再各干各的——从一开始,芯片、网络、显存调度、推理框架就针对Qwen3.8的2万亿+参数做联合调优。\n\n为什么这件事重要?过去国内跑超大参数模型,要么靠英伟达H100\u002FH200集群,要么靠多机多卡堆叠,延迟高、利用率低。M890把128卡做进一个超节点域内,互联带宽和显存共享效率远高于传统集群;适配Qwen3.8后,意味着用户可以在百炼上一键拿到一个推理即服务、且能扛住2万亿+参数的国产栈,不再受制于海外算力供给。\n\n从产业角度看,这也释放了一个清晰信号:头部云厂不再满足于模型层+硬件层各做各的,而是从系统视角重新定义超节点。Agent时代,推理不再是离线跑分,而是高并发、低延迟、长上下文的多任务战场,只有芯片+网络+模型三方联合优化,才能跑出真正的TOPS和性价比。\n\n可以预见,Qwen3.8正式版、Claude\u002FGemini\u002FDeepSeek的下一代旗舰,都会沿着国产超节点+开源大模型这条路径快速复制。当国产算力+国产模型+国产推理框架形成闭环,所谓中国版AI栈才真正有了底座。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3907857698903424","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"fca9258a-9430-455a-b95d-b9fae5e373a8","ai-inference",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":21,"name":22,"slug":22,"description":13,"color":13},"c187600e-804c-4697-b828-1e4330e0eb10","qwen",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":28},"cd009b64-6399-4a88-a3f7-cad91ec6feb0","en","Zhenwu M890 runs Qwen3.8: first 2T-param inference cluster","On July 23, Alibaba Cloud's Zhenwu M890 supernode completed full adaptation for Qwen3.8, and opened inference services on the Bailian platform. This is the first supernode in China to successfully run a large model above 2T parameters — a milestone for the joint deployment of domestic supernodes and ultra-large-parameter models. The Zhenwu M890 is Alibaba's in-house 128-card supernode, first announced in May and positioned for Agent concurrent inference, focusing on high density, low latency, and multi-task parallelism. Qwen3.8 is the latest flagship preview of Alibaba's Qwen, with 2.4T parameters, positioned as the open-source counter to top closed-source models. The Zhenwu + Qwen3.8 combination means the hardware layer and the model layer no longer work in silos — from day one, the chip, network, memory scheduling, and inference framework have been jointly tuned for Qwen3.8's 2T+ parameters. Why does this matter? In the past, running ultra-large-parameter models in China meant either relying on NVIDIA H100\u002FH200 clusters, or stacking many machines and cards, with high latency and low utilization. M890 packs 128 cards into one supernode domain, with interconnect bandwidth and memory-sharing efficiency far above traditional clusters; after adapting Qwen3.8, it means users can one-click on Bailian get an inference-as-a-service that can carry 2T+ parameters on a domestic stack, no longer constrained by overseas compute supply. From an industry angle, this also sends a clear signal: the top cloud vendors are no longer content with the model layer and the hardware layer each doing their own thing, but are redefining the supernode from a system perspective. In the Agent era, inference is no longer offline benchmarking, but a high-concurrency, low-latency, long-context multi-task battlefield — only when chip + network + model are jointly optimized can real TOPS and cost-effectiveness be delivered. It's foreseeable that Qwen3.8's official release, and the next flagships from Claude \u002F Gemini \u002F DeepSeek, will all quickly replicate this path of domestic supernode + open-source large model. When domestic compute + domestic models + domestic inference frameworks form a closed loop, the so-called China-version AI stack finally has its foundation.","alibaba-zhenwu-m890-qwen-3-8","2026-07-23T10:05:00Z","2026-07-23T10:06:22.836762Z","2026-08-19T02:08:40.142862Z",true,"agent",186,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"68072ee1-fc37-4064-ab18-09550ae72d1b","GLM-5.3-Flash 把 320B MoE 跑在国产芯片上:Flash 价位和 $0.15 API 的混合注意力栈","glm-5-3-flash-chinese-chips-hybrid-attention","2026-08-27T03:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"4bb31ede-b9c4-4762-86ae-9d3b008557ca","Hugging Face Summer 2026 报告:Qwen 拿下 15 万衍生模型, GGUF 仓库一年涨 464%","hugging-face-state-of-open-models-summer-2026","2026-08-18T02:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"5f4d9df3-673b-4440-9a77-6e8f0e697681","苹果自训中国区专用大模型:放弃「借模型」路线,联合阿里亲自下场","apple-china-specific-llm-alibaba","2026-08-14T19:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5c6d9de5-aef2-4490-96f1-e82166cc44ec","阿里千问 Qwen3.8 正式发布：2.4T 参数的旗舰基座，首次把 Cowork 塞进 Agent 入口","qwen3-8-2-4t-trillion-cowork-agent-launch","2026-08-03T08:30:00+00:00"]