[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-fal-h3-max-post-trained-video":3,"topics-all":42,"news-related-1bf979b5-2b52-4a19-8e2d-e0a549bb24ba":60},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":28,"news_slug":35,"published_at":36,"created_at":37,"modified_at":38,"is_published":39,"publish_type":40,"image_url":15,"view_count":41},"1bf979b5-2b52-4a19-8e2d-e0a549bb24ba","fal 后训练版 MiniMax H3:5 秒视频约 3 秒生成,吞吐 35 倍","fal 在开源 MiniMax H3 权重上后训练出 H3 Max:5 秒视频约 3 秒生成,吞吐约为官方端点 35 倍,人工偏好评测三项居首,Design Arena 独立评测也称其速度超 50 倍;但分辨率止步 768p,仅支持文生视频与图生视频。","视频生成行业长期默认一条铁律:质量更高的模型,推理必然更慢。8 月 27 日,推理平台 fal 正面挑战了这条假设——基于 MiniMax 开源 H3 权重后训练出的 H3 Max,生成 5 秒视频约需 3 秒,吞吐约为官方端点的 35 倍,同时拿下人工偏好评测整体质量、提示理解、美学三项第一(见 [fal 官方博客](https:\u002F\u002Fblog.fal.ai\u002Fintroducing-h3-max-by-fal\u002F))。\n\n## 不是蒸馏复刻,是「后训练 + 推理栈」协同设计\n\nfal 团队在开源的 MiniMax H3 权重上引入大量新数据做后训练,重点瞄准提示遵循与视觉质量;fal 的 X 官方帖还提到,后训练算力的相当大一部分花在了可验证的 RL 任务上。与传统步进蒸馏「逼近教师模型」不同,fal 声称要的是更快的更好模型,且基座核心能力(统一多模态上下文、原生音画同步)保持完好。\n\n推理侧同样关键。H3 Max 架构围绕 fal 自研推理引擎设计——团队过去四年专门优化扩散与生成媒体负载,训练与服务全部跑在 NVIDIA GB200 NVL72 系统上,fal 称其单芯片性能最高达上一代加速器的 2 倍。fal 的纪律是:降精度、砍采样步数这类常见加速手段,只有模型仍守住内部质量评测排名时才保留——优化与质量是同一个问题。\n\n## 榜单:官方评测之外,还有两个独立第三方\n\nfal 把 H3 Max 与 12 个主流视频模型做头对头人工偏好对比,名单包括官方 MiniMax H3 端点、Gemini Omni Flash、Wan 3.0、Seedance 2.5、Kling 3 与 Veo 3.1,用贝叶斯 Elo 加 95% 置信区间聚合。H3 Max 三项维度全部第一,且对每一个对比模型都赢得多数对局。\n\n这一结论不只在 fal 自己的评测里。Artificial Analysis 与 Design Arena 两个独立评测机构的榜单上,H3 Max 同样排名第一;Design Arena 称其「以超过 50 倍的速度交付 MiniMax H3 的质量,建立了新的速度-偏好帕累托前沿」。fal 的帕累托图显示,H3 Max 生成耗时约 2.4 秒,而 MiniMax H3 与 Wan 3.0 都要超过一分钟。MiniMax H3 团队也在 fal 博客中背书,称双方「从第一天起就紧密合作」。\n\n## 读细一点:768p 上限,和一个容易被误读的名字\n\n速度有代价。第三方页面 Morphic 的 FAQ 指出,H3 Max 为速度调优、分辨率止步 768p(原版 H3 可到 2K),且只覆盖文生视频与图生视频,原版混合输入的参考生成与编辑能力不在。名字也值得留意:H3 Max 容易被当成 MiniMax 官方新版本,实际上 MiniMax 从未发布过叫这个名字的模型——它是 fal 的改装版,只在 fal 平台提供。\n\n还有一个常被忽略的对照:fal 产品 FAQ 写道,原版标准 H3 在 fal 栈上只比 MiniMax 自家推理快约 15 倍——35 倍增益里相当部分来自后训练与协同设计,而非换个托管商。\n\n## 所以呢\n\nH3 Max 的意义超过单个产品发布。它验证了开源权重的「二阶价值」:MiniMax 开放 H3 权重,收获的不只是社区好感,还有 fal 这样的基础设施厂商在其之上做出官方端点做不到的速度-质量组合。「质量必然慢」的假设被实测数据打破;开发者选模型该看的不再是孤立跑分,而是质量、延迟、成本构成的前沿。首周五折的定价表明 fal 瞄准高吞吐生产负载。当视频生成走向交互式与批量生产,「把模型研究和推理优化当作同一个问题」可能就是下一代生成媒体公司的分水岭。","https:\u002F\u002Fblog.fal.ai\u002Fintroducing-h3-max-by-fal\u002F","9b3b9d87-0c6d-4253-95e8-426c121b96c0",[11,16,19,22,25],{"id":12,"name":13,"slug":13,"description":14,"color":15},"f3be854a-50f0-411c-893d-16d0df6def02","h3-series","MiniMax H3 专题：持续追踪 H3 的发布、开源、蒸馏与部署全链路",null,{"id":17,"name":18,"slug":18,"description":15,"color":15},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":20,"name":21,"slug":21,"description":15,"color":15},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":23,"name":24,"slug":24,"description":15,"color":15},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":26,"name":27,"slug":27,"description":15,"color":15},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[29],{"id":30,"lang":31,"title":32,"summary":33,"content":34},"e6cdd454-5306-462a-9769-534f04040177","en","fal's H3 Max: post-trained MiniMax H3 makes 5s clips in 3s","fal's H3 Max: post-trained MiniMax H3 makes 5s clips in ~3s, ~35x official throughput, #1 in preference and Design Arena evals; capped at 768p.","Generative video has long run on an assumed tradeoff: a higher-quality model must be slower at inference. On August 27, inference platform fal challenged that assumption head-on with H3 Max, a post-trained version of MiniMax's open-weight H3 model. It generates a 5-second video in roughly 3 seconds — about 35x the throughput of the official MiniMax H3 endpoint — while ranking #1 in overall quality, prompt understanding, and aesthetics in fal's human preference evaluations ([fal blog](https:\u002F\u002Fblog.fal.ai\u002Fintroducing-h3-max-by-fal\u002F)).\n\n## Not a distillation copy, but model-and-inference co-design\n\nThe fal team started from the open-weights MiniMax H3 and introduced substantial new data during post-training, focused on prompt adherence and visual quality. According to fal's official X posts, a large portion of that post-training compute went to verifiable RL tasks. Unlike traditional step-distillation, which aims to match the base model, fal says it aimed for a better model at much faster speed — with the base model's core capabilities (unified multimodal context, natively synchronized audio and video) intact.\n\nThe inference side matters just as much. H3 Max's architecture was designed around fal's in-house inference engine — a team that has spent four years optimizing diffusion and generative-media workloads. Training and serving ran entirely on NVIDIA GB200 NVL72 systems, which fal says deliver up to 2x the per-chip performance of the previous-generation accelerators. Crucially, fal set a hard rule: common speedups such as reduced precision or fewer sampling steps only survived if the model still held its position in internal quality evaluations. Optimization and quality were treated as one problem, not two.\n\n## The leaderboard: two independent third parties agree\n\nfal benchmarked H3 Max against twelve leading video models in head-to-head human preference studies — including the official MiniMax H3 endpoint, Gemini Omni Flash, Wan 3.0, Seedance 2.5, Kling 3, and Veo 3.1 — aggregated with Bayesian Elo ratings and 95% confidence intervals. H3 Max ranked #1 across all three dimensions and won the majority of matchups against every model tested.\n\nThe claim doesn't rest solely on fal's own evaluations. In independent benchmarks from Artificial Analysis and Design Arena, H3 Max also ranks #1 among video models. Design Arena's wording: \"MiniMax H3 Max by fal delivers the quality of MiniMax H3 at more than 50x the speed,\" establishing what it calls a new speed–preference Pareto frontier. fal's own Pareto chart shows H3 Max generating in about 2.4 seconds while MiniMax H3 and Wan 3.0 take over a minute. The MiniMax H3 team also endorsed the work in fal's post, saying the two companies had \"worked closely ... since day one.\"\n\n## Read the fine print: a 768p ceiling and a misleading name\n\nSpeed has a price. The third-party page Morphic notes that H3 Max is tuned for speed rather than maximum resolution — it stops at 768p where the original H3 reaches 2K — and covers only text-to-video and image-to-video, dropping the reference generation and editing endpoints that mix images, clips, and audio. The name itself deserves a caveat: it is easy to read \"H3 Max\" as an official MiniMax tier, but MiniMax has never announced a model by that name. It is fal's tuned edition, hosted exclusively on fal.\n\nOne more overlooked comparison: fal's own product FAQ states that the standard H3 on fal's stack runs about 15x faster than MiniMax's own inference. In other words, a large share of the 35x gain comes from post-training and model-system co-design — not merely from switching hosting providers.\n\n## So what\n\nH3 Max matters beyond a single product launch. It validates the \"second-order value\" of open weights: by open-sourcing H3, MiniMax gained more than goodwill — infrastructure vendors like fal built a speed-quality combination on top that the official endpoint doesn't offer. The \"quality must be slow\" assumption now has counter-evidence; the right lens for model selection is no longer an isolated benchmark score but the frontier across quality, latency, and cost. The 50%-off first week signals that fal is chasing high-volume production workloads, not demos. As video generation moves toward interactive and high-volume production, \"treat model research and inference optimization as the same problem\" may well be the dividing line for the next generation of generative-media companies.","fal-h3-max-post-trained-video","2026-08-29T23:02:12Z","2026-08-29T23:05:26.363369Z","2026-08-29T23:05:26.363378Z",true,"agent",387,[43,52],{"slug":44,"tag_slug":44,"title_zh":45,"title_en":46,"intro_zh":47,"intro_en":48,"id":49,"is_active":39,"created_at":50,"modified_at":51},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":13,"tag_slug":13,"title_zh":53,"title_en":54,"intro_zh":55,"intro_en":56,"id":57,"is_active":39,"created_at":58,"modified_at":59},"MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":61},[62,67,72,77,82,87],{"id":63,"title":64,"news_slug":65,"published_at":66},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"6f375936-79af-4622-a75e-d802ade563e0","MiniMax H3 不只是 2K 视频：它想把生成、参考和编辑收回一个模型","minimax-h3-omnimodal-video-unified-generation-editing","2026-08-03T04:08:31+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"83bf0960-2a51-4519-8326-6977527a68d9","MiniMax H3 三小时跑上 MTT S5000：Day-0 适配真正比拼的是软件栈","minimax-h3-mtt-s5000-day-zero-stack","2026-08-02T19:43:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"7ddc323f-fc52-406a-b6df-79b7393e121b","高德开源 DreamX-Creator:7B 原生音视频生成,2K 输出","dreamx-creator-7b-native-audio-video","2026-09-01T13:10:00+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"aad00b18-d354-48b5-ad21-62b53150b8c6","MiniMax H3 开源实测:你下载的权重,和 API 里跑的不是同一个模型","minimax-h3-local-vs-api-gap","2026-08-15T17:07:24+00:00",{"id":88,"title":89,"news_slug":90,"published_at":91},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00"]