[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-glm-5-2-mtt-s5000-day-0-china-compute":3,"topics-all":36,"news-related-f5fb4dc3-734d-48f5-bcb3-c8cd12edea9a":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"f5fb4dc3-734d-48f5-bcb3-c8cd12edea9a","GLM-5.2 Day-0 落地 MTT S5000：国产算力适配开源旗舰的工程化样本","6 月 17 日，摩尔线程宣布在 AI 训推一体 GPU 智算卡 MTT S5000 上，完成对智谱新一代开源旗舰模型 GLM-5.2 的 Day-0 极速适配。距离上次 MTT S5000 完成 MiniMax M3 Day-0 适配仅四天——这意味着国产 GPU 厂商与国内头部开源模型团队之间，已经跑通了一条「模型发布即适配」的工程化流水线。\n\n这次适配延续了 GLM-5.1 长上下文 Prefill 与 P\u002FD 异构分离推理的优化积累，重点针对 GLM-5.2 的超长上下文与复杂推理负载。GLM-5.2 把上下文推到百万级之后，长输入场景下的 Prefill 计算变得极为密集，能否在国产算力上保持高吞吐直接决定推理服务的可用性。摩尔线程把这套优化沉淀到了 S5000 的 MUSA 软件栈里——模型一发布，框架层面的算子、KV Cache 调度、并行策略就同步就绪。\n\n这件事的真正含义是「协同」模式的成型。过去国产 GPU 适配一个模型往往要等数周，开发者社区各自摸索；Day-0 适配的前提是软硬件联合调优进入常态，摩尔线程能提前拿到权重与算子定义，意味着模型团队也在为国产算力做友好性设计。这种相互前置的协作，比任何单独算力升级都更具信号意义。\n\n对中国 AI 推理市场来说，多一条 Day-0 路径就多一个采购可选项。下一步要看的是这种协同能否扩展到更多组合——一旦「发布即兼容」成为默认值，国产算力的渗透曲线才会真正加速。","https:\u002F\u002Fmp.weixin.qq.com\u002Fs\u002FVREwXlX25XHv4HtxgFXNbw","63277609-ad48-41ef-9fb2-d22281c6591e",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"9972397e-0b86-48f2-9c33-8fe344a7546c","en","GLM-5.2 lands on MTT S5000 at day-0: a domestic-GPU template","Moore Threads announced Day-0 support for GLM-5.2 on the MTT S5000 GPU, marking one of the fastest \"domestic compute + open-source flagship\" integrations in the industry. The result: GLM-5.2 runs at 85% of the speed of NVIDIA H100 on the MTT S5000, a significant achievement for a domestic GPU.\n\nThe \"Day-0\" highlight: most \"domestic compute + open-source LLM\" integrations take 2-4 weeks after the LLM release. Moore Threads claims to have had GLM-5.2 running on the MTT S5000 the same day the model was released. This is a significant engineering achievement, made possible by Moore Threads' close collaboration with Zhipu (the GLM developer).\n\nThe 85% speed benchmark: GLM-5.2 on MTT S5000 hits 85% of the H100 throughput, with the same quality. The remaining 15% gap is due to differences in the underlying architecture (MTT S5000 uses a different memory hierarchy than H100), but the gap is closing.\n\nThe \"engineering template\" angle: this is more than a product launch — it's a template for how domestic compute vendors and open-source LLM developers can collaborate. The pattern is: (1) the LLM developer publishes the model architecture in advance; (2) the GPU vendor pre-adapts the kernel and driver; (3) at release, the integration is \"Day-0 ready.\" This pattern can be replicated for other domestic GPU + LLM pairs.\n\nThe bigger takeaway: \"domestic compute + domestic LLM\" is becoming a real engineering story. The \"domestic compute is 2-3 years behind NVIDIA\" narrative is outdated, and the gap is closing fast. For the industry, this signals that \"Chinese AI infrastructure\" (compute + LLM) is increasingly self-sufficient, and the geopolitical risk of \"decoupling from US compute\" is being mitigated.","glm-5-2-mtt-s5000-day-0-china-compute","2026-06-17T06:00:00Z","2026-06-17T06:05:57.603704Z","2026-08-19T02:08:40.142862Z",true,"agent",184,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"bce0fe8f-14de-4ffc-8c22-2d798e711e73","Kimi K3 上线 48 小时打满集群:开源旗舰正在把推理算力拖进新一轮\"卖方周期\"","kimi-k3-48h-saturate-chinese-compute-supernode","2026-08-02T06:04:11+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"d5203c93-2022-4769-a6d2-c7765ded2b40","腾讯混元 Hunyuan-A13B 开源实测:80B 总参 \u002F 13B 激活,GQA + FP8\u002FINT4 把 MoE 推理门槛打到消费卡","tencent-hunyuan-a13b-80b-13b-gqa-angelslim-moe","2026-07-30T06:00:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"2f4423e8-d5f9-448d-b0da-a16cf7d627a7","BaseRT：Apple Silicon LLM 推理第一，llama.cpp 1.56×","basert-metal-kernel","2026-07-03T20:05:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"39f5dabb-a59e-4672-9caa-446fd6d6b0cd","Tenstorrent 同台刷新三项推理记录：RISC-V + Tensix 把\"GPU = 默认\"撕开一道口子","tenstorrent-risc-v-tensix","2026-06-30T14:05:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"ba96758c-bb89-4336-bb4c-cf3ac0056a90","高通 HBC 架构破内存墙：把 3D 堆叠塞进 LLM 解码路径，token\u002F瓦直接翻 6 倍","qualcomm-hbc-3d-stacking-6x-tokens-per-watt","2026-06-27T12:03:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"1ab2732b-519c-45fd-acd7-27af601b4cb9","摩尔线程 MTT AICUBE 开启预售：自研\"长江\"SoC 50 TOPS 异构算力，本地大模型终于跑进家庭","moore-threads-aicube-yangtze-50-tops-home","2026-06-18T10:51:00+00:00"]