[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-minimax-h3-mtt-s5000-day-zero-stack":3,"news-related-83bf0960-2a51-4519-8326-6977527a68d9":39},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":37,"view_count":38},"83bf0960-2a51-4519-8326-6977527a68d9","MiniMax H3 三小时跑上 MTT S5000：Day-0 适配真正比拼的是软件栈","MiniMax H3 开源当天，摩尔线程称仅用 3 小时便在单机八卡 MTT S5000 节点完成稳定部署，串起 SGLang-MUSA、MATE、muDNN 与 MUSA Runtime。本文拆解这次 Day-0 适配的工程含义：国产 GPU 的竞争焦点，正从“能不能跑”转向框架兼容、算子优化和持续跟进模型架构的速度。","# MiniMax H3 三小时跑上 MTT S5000：Day-0 适配真正比拼的是软件栈\n\n模型开源当天，另一套非 CUDA 计算栈能多快把它跑起来？摩尔线程给出的答案是 **3 小时**。\n\n8 月 3 日，MiniMax 开源多模态生成模型 H3。摩尔线程随后宣布，团队已在单机八卡 MTT S5000 节点完成快速部署与稳定运行，并打通从推理框架到算子库、编译器和运行时的适配链路。相比“又一款模型支持国产 GPU”，这次更值得看的，其实是 Day-0 支持背后的软件工程。\n\n## H3 为什么比普通语言模型更难适配\n\nH3 同时接收文本、图片、音频和视频，可生成最高 2K、最长 15 秒且带原生音频的视频。摩尔线程披露，多模态上下文让 H3 的序列长度方差扩大约 3 倍，理解与生成两个阶段的负载也明显上升。\n\n这意味着推理系统面对的不只是更大的矩阵乘法，还包括长度高度不规则的输入、跨模态数据流和视频生成链路。硬件峰值算力再高，如果框架无法调度、关键算子没有高效实现、运行时频繁搬运数据，模型依然只能“勉强启动”，谈不上稳定服务。\n\n## 三小时不是魔法，而是提前铺好的四层能力\n\n据摩尔线程介绍，团队先拆解模型架构、分析核心技术并梳理典型算子，随后贯通了完整软件路径：\n\n- **SGLang-MUSA** 承接模型与推理框架，且已于今年 4 月合入 SGLang 主线；\n- **SGLang-Diffusion 与 sgl-kernel** 覆盖多模态生成子系统和高性能算子接入；\n- **MATE 与 muDNN** 负责能力检测、算子选择、后端调度，以及 Attention、GEMM 等核心计算的优化实现；\n- **MTCC 与 MUSA Runtime** 完成编译和运行时执行，形成从上层框架到底层 GPU 的闭环。\n\n所以，“3 小时完成适配”并不等于工程师临时写完了一套后端。更准确的说法是：此前对主流框架、算子接口和运行时兼容性的投入，在新模型发布时被快速复用。Day-0 能力本质上是一种软件资产的复利。\n\n## 这次结果证明了什么，又没有证明什么\n\n它首先证明，国产 GPU 的竞争正在从“模型能不能跑”进入“新架构多久能稳定接入”的阶段。模型迭代按周计算，硬件若每次都要等待数月适配，即使参数漂亮，也很难进入开发者的真实工作流。\n\n但也要保持克制。摩尔线程目前披露的是单机八卡节点的运行状态与适配时间，没有公布端到端生成延迟、吞吐、显存占用、能耗，也没有给出与其他硬件的同配置对比。因此，这是一项重要的**兼容性和工程效率里程碑**，还不是性能胜负的最终答案。下一步真正有说服力的材料，应是可复现的基准、并发服务表现和长期稳定性数据。\n\n## 国产算力的护城河，会越来越像软件公司\n\nH3 的开源让企业可以本地部署并结合自有数据定制；而非 CUDA 平台若能在发布当天跟进，就能降低用户迁移和试用的时间成本。对国产 GPU 厂商而言，未来比拼的不只是芯片面积与理论算力，更是能否持续进入 PyTorch、SGLang 等上游生态，能否让模型代码少改甚至不改，能否把一次算子优化复用到下一代模型。\n\n**芯片决定性能上限，软件栈决定这块芯片能否及时进入生产。Day-0 适配的真正价值，不是“跑起来”三个字，而是把等待新模型支持的时间从月缩短到小时。**","https:\u002F\u002Fmp.weixin.qq.com\u002Fs?__biz=Mzg3MTU3Mjc4OQ==&mid=2247492753&idx=1&sn=62bb2ee5d25a71ee13a78d6f52267f32&chksm=cfb7a97c927339ad29befafd9cf2032c3e10ce56493bfcb26d40fdb1d0b4cb19f57525208e3e#rd","63277609-ad48-41ef-9fb2-d22281c6591e",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e0d31e94-ce47-4c8f-831c-d3d2926d42f3","hardware",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"8ca4f22f-58aa-47ad-8e7b-fbb2c239fa34","en","MiniMax H3 runs on MTT S5000 in three hours: software stack wins","On the day MiniMax open-sourced H3, Moore Threads said it achieved stable deployment on a single eight-card MTT S5000 node in only three hours, linking SGLang-MUSA, MATE, muDNN, and MUSA Runtime. This article examines the engineering significance of that day-zero result and argues that competition among non-CUDA GPU platforms is shifting from basic compatibility toward framework integration, operator optimization, and the speed of supporting new model architectures.","# MiniMax H3 Runs on MTT S5000 in Three Hours: Day-Zero Support Is Really a Software-Stack Test\n\nHow quickly can a newly open-sourced model run on a computing stack outside the CUDA ecosystem? Moore Threads says its answer for MiniMax H3 was **three hours**.\n\nOn August 3, MiniMax released H3 as an open multimodal generation model. Moore Threads then reported that it had deployed the model and achieved stable operation on a single eight-card MTT S5000 node. The team connected the inference framework, operator libraries, compiler, and runtime into a complete execution path. The important part is not simply that another model can run on a domestic GPU. It is what this day-zero result reveals about the maturity of the surrounding software.\n\n## Why H3 Is Harder to Port Than a Conventional Language Model\n\nH3 accepts text, images, audio, and video as inputs. It can generate videos at up to 2K resolution, as long as 15 seconds, with native audio. According to Moore Threads, multimodal context increased the variance in H3's sequence lengths by roughly three times, while both the understanding and generation phases imposed substantially heavier compute loads.\n\nThat creates challenges beyond large matrix multiplications. The inference system must handle irregular input lengths, cross-modal data flows, and a video-generation pipeline. High theoretical throughput alone is insufficient: if the framework cannot schedule workloads efficiently, critical kernels are poorly implemented, or the runtime moves data too often, a model may technically start but still be unsuitable for a stable service.\n\n## Three Hours Was Not Magic; It Was the Payoff From Four Prepared Layers\n\nMoore Threads says its engineers decomposed the architecture, analyzed its core techniques, and identified representative operators before linking the following stack:\n\n- **SGLang-MUSA** serves as the model and inference-framework layer and was merged into the main SGLang project in April;\n- **SGLang-Diffusion and sgl-kernel** provide the multimodal generation subsystem and high-performance operator integration;\n- **MATE and muDNN** handle capability detection, kernel selection, backend scheduling, and optimized implementations of core operations such as Attention and GEMM;\n- **MTCC and MUSA Runtime** cover compilation and execution, completing the path from the upper-level framework to the GPU.\n\nIn other words, “three-hour adaptation” does not mean engineers built an entire backend after H3 appeared. It means previous investments in framework compatibility, operator interfaces, and runtime infrastructure could be reused immediately. Day-zero support is the compounding return on accumulated software assets.\n\n## What the Result Proves—and What It Does Not\n\nThe result shows that competition among Chinese GPU platforms is moving beyond the question of whether a model can run at all. The more relevant question is how quickly a new architecture can be integrated reliably. Model families now change on a weekly cadence. A hardware platform that requires months of custom work for each release will struggle to enter developers' real workflows, regardless of attractive peak specifications.\n\nHowever, the claim should be interpreted carefully. Moore Threads disclosed the adaptation time and operation on an eight-card node, but it did not publish end-to-end generation latency, throughput, memory use, power consumption, or a like-for-like comparison with other accelerators. This is therefore a meaningful **compatibility and engineering-efficiency milestone**, not a final verdict on performance. Reproducible benchmarks, concurrent-serving results, and long-duration stability data would be the stronger next evidence.\n\n## The Moat for Domestic Compute Will Look Increasingly Like a Software Company\n\nBecause H3 is open, enterprises can deploy it locally and customize it with proprietary data. A non-CUDA platform that supports the model on release day reduces the time cost of evaluation and migration. For GPU vendors, the contest will increasingly depend not only on chip area and theoretical compute, but on sustained participation in upstream ecosystems such as PyTorch and SGLang, minimal code changes for users, and the ability to reuse operator optimizations across successive models.\n\n**The chip defines the performance ceiling, but the software stack determines whether that chip reaches production in time. The real value of day-zero support is not merely that a model “runs”; it is that the wait for support can shrink from months to hours.**","minimax-h3-mtt-s5000-day-zero-stack","2026-08-02T19:43:00Z","2026-08-03T06:10:17.711815Z","2026-08-03T06:10:17.711824Z",true,"agent","https:\u002F\u002Fmmbiz.qpic.cn\u002Fmmbiz_jpg\u002FrnU4nRsEjiayMWULbicdEOrfcwGvS2icaVHYj3XgvOINriaKy1epb2UYV2ee1q628F4F0WS8IkVUGTvW73ovVBN5GH6l8uVoCzarZlCwl8CWOUo\u002F0?wx_fmt=jpeg",148,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"8f4e0907-919f-410c-b543-7b52260659c2","「Thinking with Video」把推理拉出文本：Sora-2 在 MATH 跑到 92%，多模态统一架构有了新候选","thinking-with-video-sora-2-fudan-92-math","2026-06-16T12:00:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"f4c705fd-47c9-481a-807f-8001820070f8","InfinityEdit:三注意力轻量适配器,把视频编辑推进无界流时代","infinityedit-infinite-video-editing-adapter","2026-08-25T13:00:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00"]