[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-alibaba-zhenwu-m890-128-card-super-node":3,"topics-all":36,"news-related-73b75304-841d-4465-83ee-63eedfa96bee":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"73b75304-841d-4465-83ee-63eedfa96bee","阿里云发布真武M890：128卡超节点瞄准Agent并发推理","5月20日，在2026阿里云峰会上，阿里发布基于平头哥新一代AI芯片真武M890的128卡超节点服务器。该服务器搭载互联芯片ICN Switch 1.0，通信时延低至百纳秒级，可将128张AI芯片组成一台统一调度的计算集群，旨在满足Agentic时代海量并发推理与大模型训练的双重需求。\n\n传统大模型训练与推理在芯片层面往往面临通信带宽的瓶颈——当跨节点协作成为常态时，机间通信延迟会直接抵消算力扩展的收益。真武M890通过自研ICN Switch 1.0在互联层实现百纳秒级时延，意味着128卡之间的数据交换几乎可以做到无等待协同。这对需要频繁跨节点传递注意力权重或KV Cache的Transformer模型尤为关键。\n\n从系统架构角度看，阿里云这套128卡超节点的思路与NVIDIA日前交付的Agent专用CPU Vera形成了有趣的呼应：NVIDIA从处理器层面重新思考Agent场景下的并发调度，阿里则从互联层面解决多芯片协作的通信死角。两者都在解决同一个根本问题——当AI工作负载从单模型推理转向多Agent并发时，既有的基础设施假设已经不够用了。\n\n值得关注的是，这是平头哥芯片首次在超节点尺度上实现产品化落地。从倚天CPU到含光NPU再到真武M890，平头哥的芯片迭代路径正在从单芯片性能优化走向系统级协同设计。对于国内AI基础设施的自主可控而言，真武M890的集群方案是一个值得持续跟踪的进展。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3817018077447040","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"4262f8e2-5c63-4bf2-a0ee-f216a2c5274d","en","Alibaba Cloud's Zhenwu M890: 128-GPU supernode for agents","On May 20, at the 2026 Alibaba Cloud Summit, Alibaba released a 128-card super-node server based on the next-gen Pingtouge AI chip Zhenwu M890. The server carries the interconnect chip ICN Switch 1.0, with communication latency as low as the hundred-nanosecond range, and can compose 128 AI chips into a single unified-scheduling compute cluster — designed to meet the dual demands of massive concurrent inference and LLM training in the agentic era.\n\nTraditional LLM training and inference often face communication bandwidth bottlenecks at the chip level — when cross-node collaboration becomes the norm, inter-machine communication latency directly offsets the gains of compute scaling. Zhenwu M890 achieves hundred-nanosecond-level latency at the interconnect layer via the in-house ICN Switch 1.0, meaning data exchange among 128 cards can proceed with virtually no wait. This is especially critical for Transformer models that need to frequently transfer attention weights or KV Cache across nodes.\n\nFrom a system-architecture perspective, Alibaba Cloud's 128-card super-node thinking echoes NVIDIA's recently delivered Agent-dedicated CPU Vera: NVIDIA rethinks concurrent scheduling for agent scenarios from the processor layer, while Alibaba solves the communication dead spots of multi-chip collaboration from the interconnect layer. Both are tackling the same fundamental problem — when AI workloads shift from single-model inference to multi-agent concurrency, the existing infrastructure assumptions are no longer enough.\n\nNotably, this is the first product-scale landing of Pingtouge chips at the super-node level. From Yitian CPUs to Han Guang NPUs to Zhenwu M890, Pingtouge's chip iteration path is moving from single-chip performance optimization to system-level co-design. For domestic AI infrastructure self-reliance, Zhenwu M890's cluster solution is a development worth continued tracking.","alibaba-zhenwu-m890-128-card-super-node","2026-05-20T04:15:00Z","2026-05-20T04:11:51.286643Z","2026-08-19T02:08:40.142862Z",true,"agent",221,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"f5fb4dc3-734d-48f5-bcb3-c8cd12edea9a","GLM-5.2 Day-0 落地 MTT S5000：国产算力适配开源旗舰的工程化样本","glm-5-2-mtt-s5000-day-0-china-compute","2026-06-17T06:00:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"df13aca5-ba7c-4fc6-a6d5-19341820b225","华为云「第三条路」：从 Token工厂到 Agentic Infra，国产算力撑起的智能体底座","huawei-cloud-token-factory-agentic-infra","2026-06-09T03:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"45375854-7739-4dd1-bc6a-30db4474652a","Taalas HC2:把单片参数拉到 200 亿,「模型刻进硅片」的第二章","taalas-hc2-20b-mxfp4-50-chips-1t-amd","2026-08-19T00:00:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"172d28ff-776b-443d-8f05-dddc9d544dec","Mercury 2 把“推理扩散 LLM”塞进搜索流水线：每步 1000+ tokens\u002F秒，Voice Agent 延迟预算被改写","mercury-2-diffusion-search-agent-realtime","2026-08-17T04:00:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"f333dd36-d9ed-4e17-a601-11b4f140eee3","Taalas HC2 把参数上限拉到 200 亿：AMD 这张「把模型刻进硅片」的牌,开始讲下一章","taalas-hc2-20b-mxfp4-amd","2026-08-15T03:30:00+00:00"]