[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-roboneo-minimax-h3-multimodal-editing":3,"news-related-6e3002da-c1fd-4a6d-b903-4f65b976dd04":39},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":37,"view_count":38},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","8 月 3 日,美图旗下 RoboNeo 宣布接入 MiniMax 刚开源的多模态生成模型 MiniMax H3,主打多模态理解与精细化视频局部编辑。同日,摩尔线程基于 MTT S5000 完成 H3 的 Day-0 适配。从开源到商用、从训练到推理,H3 正在被打磨成一款真正可用的多模态生产力工具。","## RoboNeo 接住了 MiniMax H3 落地的第一棒\n\n8 月 3 日,美图旗下的 RoboNeo 宣布正式接入 MiniMax H3,强化多模态理解与视频局部编辑能力(来源:36氪美图旗下 RoboNeo 接入 MiniMax H3 报道)。在多模态生成模型还在卷参数、卷分辨率的当口,一个真实的视频编辑产品选择用 H3 作为底座,这本身是个值得展开的信号。\n\n按 36氪披露的口径,接入 H3 之后 RoboNeo 能做到的事包括:\n\n- **多模态统一理解**:文本、图片、视频、音频同框处理,模型在内部把它们当成一条统一的生成输入。\n- **精细化视频局部编辑**:人物替换、物体增减、背景修改、特效调整、音色迁移、台词修改。\n\n这两条能力都直接对应 MiniMax 在官方博客里列出的 H3 核心卖点 —— \"general-purpose omni-modal generation model\",一个模型同时处理文本、图片、视频、音频的生成与编辑(来源:MiniMax Blog · MiniMax H3, 2026-08-03)。\n\n换句话说:H3 不是又一个只能\"生成 8 秒漂亮片段\"的视频模型,它想要成为能跑进真实工作流的\"生成 + 编辑\"基础设施。\n\n---\n\n## H3 的设计哲学,为什么决定了 RoboNeo 这种用法\n\nMiniMax 在官方博客里把 H3 的设计原则写得很直白:过去两代(Hailuo 01、Hailuo 02)是把任务拆开做 —— 图像、视频、音频各一套专家模型;H3 反过来,把任务**收敛**到一个统一的生成框架里。\n\n具体落点有三个:\n\n1. **H3-Omni Transformer** 替代了 Hailuo 02 的\"专用架构\",作者明确说,因为 H3 要做任务泛化,过去那些架构上的特殊优化反而成了\"不必要的复杂度\"。\n2. **训练侧把多种模态、多种任务尽可能早地融合**,通过专用的 captioning pipeline 把大约 100K token 的原始素材压缩到平均 4K token 的\"上下文全模态表征\"(`Contextual Omni Representation`)。\n3. **2K 输出用 in-context 再生完成**,不用外挂超分模块:H3 base model 自己把低分辨率结果在多模态上下文里再升一次,小型文字和细节恢复明显好于传统超分。\n\n这套结构放在今天的多模态竞争里看,**最直接的副产品就是可控编辑**。RoboNeo 接进来用的人物替换、台词修改、背景更换 —— 都是\"在已有视频上、听自然语言指令、改某一块\"的场景,这恰恰是 H3 这条\"任务统一 + 语言作为通用桥梁\"路线最擅长的活。\n\n---\n\n## 同一天:摩尔线程跑通 Day-0 适配\n\n发布同一天的另一件事是摩尔线程基于 AI 训推一体智算卡 MTT S5000 和自家 MUSA 软件栈,完成了 H3 的 Day-0 适配与运行(来源:36氪报道,2026-08-03)。\n\n这件事容易被忽视,但意义不小:\n\n- **国内推理栈开始把\"多模态大模型\"当一类基础设施做适配**,不再是每个应用层从零搞兼容。\n- MTT S5000 + MUSA 一天之内能跑起来,意味着 H3 在硬件兼容性上是**主动设计**过的 —— MiniMax 在博客里也明说\"hardware compatibility has been a key consideration since the earliest stages of H3's design\"。\n\n这两件事叠加在一起看,H3 正在形成一个**\"模型开源 + 推理栈 Day-0 + 应用层 Day-1 集成\"** 的闭环节奏。模影像样的发布见过不少,但能同时在开源、硬件、应用三层都\"零日响应\"的,目前中文圈并不多。\n\n---\n\n## 商用反馈会变成 H3 的照妖镜\n\n从训练范式到产品落地,中间隔着一条很长的沟。RoboNeo 这种 P 端产品在用 H3 之后必然会撞上几类问题,这些反馈会反过来定义 H3 后续版本的优先级:\n\n- **可控 vs 风格化**:商用场景对\"改一处不能动其他\"的稳定性要求非常高。H3 现在的\"自然语言指令 + in-context 再生\"路线能不能撑得住电商素材、广告口播这类对一致性极端敏感的场景,是关键考验。\n- **2K 默认输出 vs 推理成本**:官方说 H3 在 2K 的每秒价格低于主流模型的三分之一,在 768p 是主流模型 720p 的一半 —— 这是相对值不是绝对值。一旦商用流量上来,推理 TCO 会是第二个被反复拉出来算的账。\n- **多模态上下文长度**:处理一段 15 秒视频时上下文接近 100K token 被压到 4K,这个压缩比效果如何,在 RoboNeo 这类带历史素材的产品里很容易被量化。\n\n所以这次接入,对 MiniMax 来说不是\"又一家客户用了我们\"这种面子工程,而是真正把 H3 的\"任务泛化\"主张放到工业压力测试里。\n\n---\n\n## 所以呢:多模态生成的下半场,看谁先被产品用顺手\n\n2026 年的多模态生成赛道,各家要么拼原生分辨率,要么拼上下文长度,要么拼与硬件 \u002F 模型的兼容性。H3 走了一条更显眼但也更难走的路 —— **\"一个模型做所有任务 + 开源 + 2K 默认\"**,商业落点直接押在\"可编辑\"上。\n\nRoboNeo、ComfyUI Day-0 支持、Moore Threads Day-0 适配 —— 这一串节奏让 H3 在一周内就被放进了\"开源模型 + 推理栈 + 编辑产品\"的实际工作流。下一阶段值得盯的不是 MiniMax 又发了几张 demo,而是 RoboNeo 这类真实产品用 H3 跑一两个月之后,**用户能否真的用它替代传统视频编辑流水线里的某一节**。\n\n如果能,H3 才是真正把多模态生成从\"演示\"推进到\"生产力\"的那一个;如果不能,后面排队的 Veo、Kling、Sora 们,会在同一个痛点上迅速补位。\n\n**参考来源:**\n- 美图旗下 RoboNeo 接入 MiniMax H3,36氪,2026-08-03 (https:\u002F\u002Fwww.36kr.com\u002Fnewsflashes\u002F3923674190999173)\n- 摩尔线程完成 MiniMax H3 多模态生成模型适配,36氪,2026-08-03 (https:\u002F\u002Fwww.36kr.com\u002Fnewsflashes\u002F3923276079263113)\n- MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities,MiniMax Blog,2026-08-03 (https:\u002F\u002Fwww.minimax.io\u002Fblog\u002Fminimax-h3)","https:\u002F\u002Fwww.36kr.com\u002Fnewsflashes\u002F3923674190999173","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"f8ff89df-abfd-4456-9bc0-91edab3399e6","en","MiniMax H3's first landing: Meitu RoboNeo tests editability","On August 3, Meitu's RoboNeo announced its integration with MiniMax's newly open-sourced multimodal generation model MiniMax H3, focusing on multimodal understanding and fine-grained video editing. The same day, Moore Threads completed Day-0 adaptation of H3 on its MTT S5000 card. From open source to commercial use, from training to inference, H3 is being polished into a truly usable multimodal productivity tool.","## RoboNeo Catches the First Commercial Pitch from MiniMax H3\n\nOn August 3, Meitu's RoboNeo announced a formal integration with MiniMax H3, strengthening multimodal understanding and fine-grained local video editing capabilities (source: 36Kr report on Meitu's RoboNeo integrating MiniMax H3). At a time when multimodal generation models are still racing on parameter counts and resolution, a real video editing product has chosen to use H3 as its backbone — a signal worth digging into.\n\nAccording to 36Kr's disclosure, after integrating H3, RoboNeo can do the following:\n\n- **Unified multimodal understanding**: text, image, video, and audio processed in a single frame; the model treats them as one unified generation input internally.\n- **Fine-grained local video editing**: character swap, object add\u002Fremove, background change, effect tweak, voice tone transfer, line rewriting.\n\nBoth capabilities directly map to MiniMax's H3 core selling points listed in its official blog — \"a general-purpose omni-modal generation model,\" one model handling both understanding and generation across text, images, video, and audio (source: MiniMax Blog · MiniMax H3, 2026-08-03).\n\nIn other words, H3 is not just another \"generate 8 seconds of pretty clips\" video model — it wants to become generative-plus-editing infrastructure that can run inside real workflows.\n\n---\n\n## Why H3's Design Philosophy Determines This Kind of Usage\n\nMiniMax spelled out H3's design principle pretty bluntly on the official blog: the previous two generations (Hailuo 01, Hailuo 02) split tasks apart — separate expert models for image, video, audio. H3 reverses that and **converges** tasks into a single generative framework.\n\nThree specific anchors:\n\n1. **H3-Omni Transformer** replaces Hailuo 02's specialized architecture. The authors explicitly say that because H3 does task generalization, those prior architectural tricks became \"unnecessary complexity.\"\n2. **Early training-stage fusion** of modalities and tasks, with a dedicated captioning pipeline that compresses roughly 100K tokens of raw material down to an average of 4K tokens of \"contextual omni representation.\"\n3. **2K output via in-context regeneration** — no external super-resolution module. The H3 base model upscales its own low-resolution output within the multimodal context, recovering small text and fine detail noticeably better than traditional upscaling.\n\nThis architecture, viewed against today's multimodal competition, has a **most-direct byproduct: controllable editing.** RoboNeo's character swap, line rewrite, background change — all are \"on existing video, take a natural-language instruction, modify one piece\" — exactly the kind of task H3's \"unified tasks + language as the universal bridge\" line was designed for.\n\n---\n\n## Same Day: Moore Threads Completes Day-0 Adaptation\n\nAnother thing happening on the same day: Moore Threads, based on its AI training-inference integrated card MTT S5000 and its MUSA software stack, completed Day-0 adaptation and runtime for H3 (source: 36Kr report, 2026-08-03).\n\nThis is easy to overlook but not insignificant:\n\n- **Domestic inference stacks are starting to treat \"multimodal large models\" as a category of infrastructure to adapt**, not something each application layer reinvents from scratch.\n- MTT S5000 + MUSA being runnable within a day means H3 was actively designed for hardware compatibility — MiniMax stated on the blog, \"hardware compatibility has been a key consideration since the earliest stages of H3's design.\"\n\nStacked together, these events show H3 building a cadence of **\"model open-source + inference stack Day-0 + application-layer Day-1 integration.\"** Fancy model launches are common, but having \"zero-day response\" simultaneously on open source, hardware, and application layers — in the Chinese AI community today that is rare.\n\n---\n\n## Commercial Feedback Becomes H3's Litmus Test\n\nFrom training paradigm to product deployment sits a long ditch. Once RoboNeo and similar P-end products start using H3, certain issues will inevitably surface, and those issues will redefine the priorities of H3's next versions:\n\n- **Controllability vs. stylization**: commercial scenarios demand extreme stability around \"change one piece, don't touch the rest.\" Whether H3's \"natural-language instructions + in-context regeneration\" line holds up under the consistency-sensitive demands of e-commerce assets and ad voice-overs is the key test.\n- **2K default vs. inference cost**: MiniMax says H3's per-second price at 2K is less than a third of mainstream models, and at 768p is less than half of the mainstream 720p price — that is relative, not absolute. Once commercial traffic ramps, inference TCO will be the second number repeatedly pulled up in spreadsheets.\n- **Multimodal context length**: with 15-second video the context is roughly 100K tokens compressed to 4K — how that compression ratio plays out will be easy to quantify in RoboNeo's historical-asset style workflow.\n\nSo this integration is not, for MiniMax, \"another customer used us\" prestige — it is putting H3's task-generalization claim under real industrial stress test.\n\n---\n\n## So What: The Second Half of Multimodal Generation Hinges on Who Gets Production-Used First\n\nIn 2026's multimodal generation race, everyone is pushing on native resolution, context length, or hardware \u002F model compatibility. H3 has taken a more visible but harder path — **\"one model for all tasks + open source + 2K default** — with the commercial bet squarely on \"editability.\"\n\nRoboNeo, ComfyUI Day-0 support, Moore Threads Day-0 adaptation — that streak puts H3 inside real workflows of \"open-source model + inference stack + editing product\" within a single week. What to watch next is not how many more demos MiniMax releases, but whether real products like RoboNeo can **actually replace one specific segment of the traditional video editing pipeline** after one or two months of use.\n\nIf yes, H3 is the one that genuinely moves multimodal generation from \"demo\" to \"productivity.\" If not, Veo, Kling, Sora, and the queue behind them will quickly fill the same gap.\n\n**References:**\n- Meitu's RoboNeo integrates MiniMax H3, 36Kr, 2026-08-03 (https:\u002F\u002Fwww.36kr.com\u002Fnewsflashes\u002F3923674190999173)\n- Moore Threads completes MiniMax H3 multimodal model adaptation, 36Kr, 2026-08-03 (https:\u002F\u002Fwww.36kr.com\u002Fnewsflashes\u002F3923276079263113)\n- MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities, MiniMax Blog, 2026-08-03 (https:\u002F\u002Fwww.minimax.io\u002Fblog\u002Fminimax-h3)","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02Z","2026-08-03T18:02:17.425224Z","2026-08-03T18:02:17.425251Z",true,"agent","https:\u002F\u002Ffile.cdn.minimax.io\u002Fpublic\u002F3df321d9-42bd-4be0-ac58-71f22377a11f.png",102,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"9f82c248-0592-421f-9fd8-ebd2100dcaf5","VideoChat3 全开源 4B 视频 MLLM 一次打通四种能力,I3D-ViT 把时空 token 砍掉 16×","videochat3-4b-mllm","2026-07-15T02:00:00+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"3851a096-37d6-4a45-bfb7-57b2fd65d992","京东开源 JoyAI-Echo：5 分钟长视频生成首次解决「跨镜头一致性」难题，DMD 蒸馏跑出 7.5× 加速","joyai-echo-jd-5-min-cross-shot-dmd-7-5x","2026-06-12T02:01:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00"]