[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-minimax-h3-omnimodal-video-unified-generation-editing":3,"news-related-6f375936-79af-4622-a75e-d802ade563e0":39},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":37,"view_count":38},"6f375936-79af-4622-a75e-d802ade563e0","MiniMax H3 不只是 2K 视频：它想把生成、参考和编辑收回一个模型","MiniMax 发布通用全模态视频模型 H3，把文本、图像、视频与音频统一为上下文，可生成最高 2K、最长 15 秒并带原生立体声的视频。它真正值得关注的不是规格，而是用 H3-VAE、Omni Transformer 与上下文再生成，把分裂的生成、参考和编辑任务收回一个模型。","## 视频生成终于开始告别‘一个任务一个模型’\n\n过去两年，视频模型的进步主要体现在清晰度、时长和运动稳定性上，但工作流依然很碎：文生视频、图生视频、首尾帧、主体参考、动作迁移、配音、音效和编辑，往往需要不同模型或不同管线。MiniMax 新发布的 H3 试图改变的正是这件事。它把文本、图像、视频和音频放进统一上下文，可生成最高 2K、最长 15 秒、带原生立体声的视频。\n\n这不只是‘又一个 2K 视频模型’。H3 的核心判断是：**任务边界本身正在成为生成模型的瓶颈。** 当创作者想表达‘参考视频一的镜头运动，让图片二中的人物按音频三演唱’时，传统工作流要先拆任务，再逐段调用工具；H3 则希望让用户直接用自然语言描述素材之间的关系，由同一模型完成理解、生成和编辑。\n\n## 三项技术选择，比演示视频更值得看\n\nMiniMax 公布了四个关键组件，其中三项直接指向成本与泛化能力。\n\n第一是 **Contextual Omni Representation**。团队不只给目标视频写描述，还让专用理解管线描述输入素材之间、输入与目标之间的关系。官方称，单份源素材通常需要约 10 万 token 的推理，再蒸馏成平均约 4000 token 的表示。语言在这里不是简单提示词，而是把不同模态和任务统一起来的中间层。\n\n第二是重新设计的 **H3-VAE**。更高压缩率让有效序列长度提升 4 倍，既降低训练和推理成本，也成为原生 2K 输出的基础。视频模型最昂贵的部分之一，就是时空 token 数量随分辨率和时长迅速膨胀；压缩器如果只追求还原而缺乏可学习性，后面的 Transformer 仍会被长序列拖垮。H3 把 tokenizer 当成核心架构，而不是前处理模块。\n\n第三是 **H3-Omni Transformer** 与异构训练调度。引入全模态上下文后，样本序列长度方差扩大到三倍，理解与生成的计算形态也不一样。MiniMax 将两类负载分开优化，再做样本间负载均衡，称端到端训练吞吐提升接近 30%。这类工程改进不够吸睛，却直接决定一个视频模型能否持续迭代。\n\n## ‘上下文再生成’比传统超分更聪明\n\nH3 没有用独立超分模型把低清结果放大到 2K，而是让基础模型带着原始多模态上下文重新生成。传统超分只能根据已有像素猜细节，遇到小字、品牌标识和复杂纹理时容易编造；上下文再生成则能重新读取原始文字、图片、音频和参考视频，理论上更有机会恢复正确内容。\n\n官方还宣称，H3 在 2K 下每秒价格低于主流模型的三分之一，768p 下也不到主流 720p 模型的一半。不过这仍是厂商口径，技术报告尚未发布，模型权重也只是承诺在未来几天开放。没有第三方基准、显存需求和真实开源许可证之前，‘开源’还不能算完全落地。\n\n## 真正的行业信号：视频模型正在变成创作系统\n\nH3 的意义不在于单次生成是否压过某个闭源模型，而在于它把视频模型的竞争从‘更漂亮的短片’推向‘更少的任务切换’。如果统一预训练能覆盖生成、参考、编辑和音画联合建模，创作工具的护城河将从拼接多个专家模型，转向如何组织上下文、数据和反馈。\n\n所以，接下来最值得盯的不是 H3 的样片，而是权重何时真的开放、第三方硬件能否跑起来，以及统一模型在复杂编辑中是否仍能保持角色和品牌一致性。视频生成的下一阶段，可能不是更长的片段，而是一个模型真正接管整条创作链。","https:\u002F\u002Fwww.minimax.io\u002Fblog\u002Fminimax-h3","70524a06-fc44-487c-ac6b-4a0186f66a45",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"4f214978-cac1-4f39-aa4b-f92a0d0934b7","transformer",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"18d12800-c096-4cc6-a3c5-d46e5bb46217","en","MiniMax H3 unifies generation, reference, and editing","MiniMax has introduced H3, a general-purpose omni-modal video model that accepts text, images, video and audio in one context and generates up to 15-second 2K video with native stereo sound. The more important shift is architectural: H3 uses a redesigned VAE, an Omni Transformer and in-context regeneration to unify generation, reference and editing tasks that were previously split across specialized pipelines.","## Video generation is starting to move beyond one model per task\n\nFor the past two years, progress in video generation has been measured mainly by resolution, duration and motion quality. Yet production workflows remain fragmented: text-to-video, image-to-video, first-and-last-frame generation, subject reference, motion transfer, dubbing, sound effects and editing are often handled by separate models or pipelines. MiniMax H3 is designed to challenge that fragmentation. It accepts text, images, video and audio in a unified context and generates video at up to 2K resolution, 15 seconds in length, with native stereo sound.\n\nThat makes H3 more than another model with a 2K headline. Its central thesis is that **task boundaries have become a bottleneck**. A request such as ‘use the camera move from video one, make the person in image two sing, and match audio three’ normally has to be decomposed into multiple tools. H3 aims to understand the relationships among those inputs and carry out understanding, generation and editing inside one model.\n\n## Three technical choices matter more than the demos\n\nThe first is **Contextual Omni Representation**. MiniMax says its understanding pipeline describes not only the target video, but also the relationships among source assets and between the context and the target. A source item can require about 100,000 tokens of inference before being distilled into roughly 4,000 tokens on average. Language therefore acts as a general interface that represents tasks and cross-modal relationships, rather than merely serving as a prompt.\n\nThe second is a redesigned **H3-VAE**. Its higher compression ratio produces a fourfold gain in effective sequence length, lowering training and inference costs while enabling native 2K output. This is important because spatiotemporal token counts expand rapidly with resolution and duration. A video tokenizer that reconstructs well but is hard for the downstream model to learn still leaves the Transformer buried under long sequences. H3 treats the tokenizer as a core architectural component.\n\nThe third is the **H3-Omni Transformer** and its heterogeneous training system. Adding omni-modal context tripled the variance in sequence length, while understanding and generation created different compute profiles. MiniMax separated those workloads for hardware optimization and balanced heterogeneous compute across samples, reporting an end-to-end training throughput gain of nearly 30%. This is less spectacular than a demo reel, but it determines whether a video model can iterate economically.\n\n## In-context regeneration is a smarter approach to 2K\n\nInstead of attaching a conventional super-resolution network, H3 asks the base model to regenerate its own low-resolution result while seeing the original multimodal context again. Standard super-resolution can only infer details from existing pixels and often invents small text, logos or fine textures. Regeneration can revisit the original text, images, audio and reference video, giving the model a better chance to restore semantically correct details.\n\nMiniMax also claims that H3 costs less than one-third as much per second as mainstream models at 2K, and that its 768p output costs less than half as much as mainstream 720p generation. Those figures remain vendor claims. The full technical report has not yet been released, and the weights are promised for the coming days rather than already available. Until independent benchmarks, memory requirements and a concrete license appear, the open-model claim should be treated as incomplete.\n\n## The larger signal: video models are becoming production systems\n\nH3 matters less as a possible winner in a single quality comparison than as an attempt to reduce task switching. If unified pretraining can cover generation, reference, editing and synchronized audio-visual modeling, the competitive advantage of creative tools will shift away from chaining specialist models and toward organizing context, data and feedback around a general model.\n\nThe next things to watch are therefore not the showcase clips. They are whether the weights are actually released, whether third-party hardware can run the model efficiently, and whether character, text and brand consistency survive complex editing. The next stage of video generation may not be a longer clip; it may be one model taking responsibility for the entire creative pipeline.","minimax-h3-omnimodal-video-unified-generation-editing","2026-08-03T04:08:31Z","2026-08-03T04:14:53.550355Z","2026-08-03T04:14:53.550367Z",true,"agent","https:\u002F\u002Ffile.cdn.minimax.io\u002Fpublic\u002F3df321d9-42bd-4be0-ac58-71f22377a11f.png",136,{"items":40},[41,46,51,56,61,66],{"id":42,"title":43,"news_slug":44,"published_at":45},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":47,"title":48,"news_slug":49,"published_at":50},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":52,"title":53,"news_slug":54,"published_at":55},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":57,"title":58,"news_slug":59,"published_at":60},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"2fbfa6c5-3bf5-4279-a353-6324396b2d36","字节 Seedance 2.5 把单段视频拉到 30 秒：视频生成终于\"能用\"了？","bytedance-seedance-2-5-30s-video-model","2026-07-31T06:00:00+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"d3e01f3d-745b-4c98-9289-38081a3f5f06","FLUX 3：图像\u002F视频\u002F音频统一进 flow matching","bfl-flux-3-flow-matching","2026-07-27T10:00:00+00:00"]