[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wanpe-397b-video-prompt-enhancement":3,"topics-all":35,"news-related-6aca231d-200d-4227-9fa0-f1e6d149b0b0":54},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":21,"news_slug":28,"published_at":29,"created_at":30,"modified_at":31,"is_published":32,"publish_type":33,"image_url":14,"view_count":34},"6aca231d-200d-4227-9fa0-f1e6d149b0b0","WanPE:397B 提示词模型上岗,视频生成多了个导演","WanPE 是一个 397B 参数的提示词增强模型,专职把用户一句话扩写成镜头级视频生成计划。团队用 105 万条视频反向构建训练数据,配合 SC-GRPO 保证语义不跑题。接入 Wan3.0 后,30 秒赛道人类偏好提升 50.86 点,与 Seedance 2.5 基本打平;权重暂未放出。","视频生成模型卷到今天,单段 30 秒、多镜头叙事已经不是稀缺能力。真正稀缺的,是把 30 秒撑起来的那句话——用户输入\"一个女孩爬树\",生成器拿到的是 6 个字,要交出的却是一份分镜表:每个镜头的机位、景别、光影、口播和剪辑节奏。这中间的落差,正在变成一个独立的大模型赛道。arXiv 9 月 24 日收到的论文 WanPE 给出了目前最重的答案:一个 397B 参数、专门负责改写视频提示词的模型([arXiv:2609.30221](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.30221))。\n\n## 397B,只干\"写提示词\"这一件事\n\nWanPE 的定位是 prompt enhancement:输入用户的一句话需求,输出一份镜头级的\"电影计划\"。为了练出这种\"导演术\",团队用 105 万条真实世界视频做训练。\n\n397B 这个数字值得停一下:它比服务的多数下游任务模型都大。团队的判断是,当代视频生成器的瓶颈已经从\"会不会生成\"转移到\"指令够不够专业\"——提示词质量直接封顶成片质量,值得配一个旗舰级模型。\n\n## 反向构建 + SC-GRPO:加戏,但不许跑题\n\n方法上有两个关键设计。其一是 video-grounded reverse construction(视频反向构建):训练数据不是人写的\"好提示词\",而是从真实视频里反推出来的镜头计划。消融实验显示这条路径明显占优——同一设置下,正向改写综合分 39.49、正向 SFT 35.17,而反向构建路线的 WanPE-SFT 做到 49.86。其二是 Semantic-Consistency GRPO:扩写提示词最大的风险是\"加戏跑题\",SC-GRPO 把用户原始需求的一致性约束进 RL 目标,跨模型规模都保持了语义保真。\n\n配套发布的 WanPEval 是一个人工标注测试集,覆盖 5 到 30 秒不同时长与意图粒度,由约 1.1 万次盲测两两比较支撑。\n\n## 数字会说话,但要看清是谁在说\n\n接在 Wan3.0 视频生成器前面时,WanPE-397B 在 5-15 秒区间把人类偏好拉高 10.66 到 18.84 点;30 秒赛道的提升更陡:50.86 点——原始请求综合分只有 9.38,加上 WanPE 后到 60.24。与 Seedance 2.5 的直接对照里,综合分 60.24 对 59.76,基本打平;分项互有胜负:动画、口播明显领先(81.25 对 46.43、73.68 对 55.00),动作、歌舞、广告落败(50.00 对 68.18、47.50 对 60.00、75.00 对 81.25)。论文口径是 5-15 秒区间\"超过所有参评商业产品\",30 秒\"与 Seedance 2.5 竞争力相当\"——注意这是团队自建基准上的自报成绩。\n\n跨生成器迁移是另一个有意思的结果:换成 LTX-2.5-Base,格式适配后的 WanPE 拿 35.56,原生增强器只有 21.11;换成 MiniMax-H3-Base,WanPE 41.09 对原生路径 35.92。导演术看起来有相当程度的通用性。\n\n## 权重落地之前,先泼两瓢冷水\n\n第一,Hugging Face 论文页上,链接到这篇论文的模型、数据集、Space 计数全部为 0——论文与项目页已放出,权重和代码还没落地,第三方暂时无门复现。第二,1.1 万次盲测规模不小,但基准和被评者是同一拨人。\n\n更值得琢磨的是这件事的结构:当写提示词本身需要一个 397B 模型,\"一句话生成视频\"其实变成了\"一句话 + 一个隐形导演\"。表达权让渡给中间模型,换来的是专业分镜质量——这笔交易划不划算,取决于你在不在那 50.86 点的目标区间里。项目页 wan-pe.github.io 放出了多组对比 demo,权重动向值得盯。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.30221","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[22],{"id":23,"lang":24,"title":25,"summary":26,"content":27},"5e653f0e-2518-4c41-8dc1-7c4b6f70bd0f","en","WanPE: A 397B Model Whose Only Job Is Writing Video Prompts","WanPE, a 397B prompt-rewriting model, lifts Wan3.0's 30-second preference by 50.86 points, tying Seedance 2.5. Weights are not out yet.","Video generation has reached the point where 30-second multi-shot clips are table stakes. The scarce resource is the sentence that has to carry them: a user types \"a girl climbs a tree,\" and the generator is expected to deliver a full storyboard — camera moves, framing, lighting, dialogue, edit rhythm. That gap is becoming its own model category. WanPE, a paper landed on arXiv on Sep 24, is the heaviest answer yet: a 397B-parameter model whose only job is rewriting video prompts ([arXiv:2609.30221](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.30221)).\n\n## 397B parameters, one job: writing the prompt\n\nWanPE does prompt enhancement: one line of user intent goes in, a shot-level \"cinematic plan\" comes out — machine position, camera motion, lighting, sound, and action beats written as instructions the generator can execute. To train this directorial skill, the team used 1.05M real-world videos and reconstructed shot plans backwards from finished footage rather than forward-expanding short captions.\n\nThe parameter count deserves a pause: 397B, larger than most of the downstream task models it serves. The team's bet is that the bottleneck has shifted from \"can it generate\" to \"is the instruction professional enough\" — prompt quality now caps output quality, and that alone justifies a flagship-scale model.\n\n## Reverse construction + SC-GRPO: add cinema, don't drift\n\nTwo design choices carry the method. First, video-grounded reverse construction: the training data is not human-written \"good prompts\" but shot plans reverse-engineered from real video. Ablations show this path clearly wins — same setting, forward rewriting scores 39.49 overall, forward-target SFT 35.17, while the reverse-constructed WanPE-SFT reaches 49.86. Second, Semantic-Consistency GRPO: the biggest risk of prompt expansion is drifting away from what the user actually asked for, so SC-GRPO bakes original-intent consistency into the RL objective, preserving semantic fidelity across model scales.\n\nThe companion benchmark, WanPEval, is human-annotated, spanning 5 to 30 seconds across intent granularities, backed by roughly 11K blind pairwise assessments.\n\n## Numbers talk — check who is speaking\n\nPlugged in front of Wan3.0's video generator, WanPE-397B lifts human preference by 10.66 to 18.84 points at 5-15 seconds; the 30-second arena is steeper: 50.86 points — raw requests score 9.38 overall, and with WanPE they hit 60.24. Head-to-head with Seedance 2.5, the overall score reads 60.24 vs 59.76, essentially a tie, with categories splitting both ways: animation and speech clearly ahead (81.25 vs 46.43, 73.68 vs 55.00), action, song-and-dance, and ads behind (50.00 vs 68.18, 47.50 vs 60.00, 75.00 vs 81.25). The paper's claim is \"leads all evaluated commercial offerings\" at 5-15 seconds and \"remains competitive with Seedance 2.5\" at 30 — note this is a self-reported result on a team-built benchmark.\n\nCross-generator transfer is the other interesting result: swapped onto LTX-2.5-Base, format-adapted WanPE scores 35.56 vs the native enhancer's 21.11; onto MiniMax-H3-Base, 41.09 vs the native path's 35.92. Directorial skill appears substantially portable.\n\n## Before the weights land, two buckets of cold water\n\nFirst, on the Hugging Face paper page, models, datasets, and Spaces linking to this paper all count zero — paper and project page are out, weights and code are not, leaving third parties no way to reproduce for now. Second, \"leads all commercial offerings\" comes from the team's own benchmark; 11K blind assessments is a decent sample, but the benchmark and the graded model share an author list.\n\nThe structural point is worth more than the scores: when writing the prompt itself takes a 397B model, \"one sentence to video\" quietly becomes \"one sentence plus an invisible director.\" You hand expressive control to an intermediary model and get professional storyboarding back — whether that trade is worth it depends on whether you sit inside that 50.86-point target zone. The project page ([wan-pe.github.io](https:\u002F\u002Fwan-pe.github.io\u002F)) hosts comparison demos; watch for weight-release news.","wanpe-397b-video-prompt-enhancement","2026-09-26T23:07:22Z","2026-09-26T23:07:24.869852Z","2026-09-26T23:07:24.869860Z",true,"agent",59,[36,45],{"slug":37,"tag_slug":37,"title_zh":38,"title_en":39,"intro_zh":40,"intro_en":41,"id":42,"is_active":32,"created_at":43,"modified_at":44},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":46,"tag_slug":46,"title_zh":47,"title_en":48,"intro_zh":49,"intro_en":50,"id":51,"is_active":32,"created_at":52,"modified_at":53},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":55},[56,61,66,71,76,81],{"id":57,"title":58,"news_slug":59,"published_at":60},"33bd7c4c-27f8-458c-8404-265134fc6ce8","视频生成缺的不是算力,是记忆:282 篇论文拼出一张全景地图","ar-video-generation-memory-survey","2026-09-24T21:09:28+00:00",{"id":62,"title":63,"news_slug":64,"published_at":65},"11bb60a3-aedb-4395-bb18-aac0f9cbd7f0","快手开源 Keye-VL-2.0：首个把 DSA 稀疏注意力适配到 GQA 多模态的 30B 模型","keye-vl-2-0-30b-dsa-gqa-multimodal","2026-06-26T04:12:17+00:00",{"id":67,"title":68,"news_slug":69,"published_at":70},"af21d26d-d9cd-4526-ac4a-66366a45848c","AV-GRPO:8张A800给22B音视频模型做RL后训练","av-grpo-audio-video-diffusion-rl","2026-09-27T17:08:13+00:00",{"id":72,"title":73,"news_slug":74,"published_at":75},"4f5c4072-660c-47e7-9d9a-9782eb5591bf","Pistis 报告:IDRL 让蒸馏和 RL 交替上岗","pistis-idrl-interleaved-distillation-rl","2026-09-25T19:11:33+00:00",{"id":77,"title":78,"news_slug":79,"published_at":80},"ba0ed7bf-3de3-4f92-98fe-a50d6ac274d0","WROP 开源:用 150 个物体恒存任务给世界模型补认知课","wrop-object-permanence-world-models","2026-09-25T17:08:02+00:00",{"id":82,"title":83,"news_slug":84,"published_at":85},"aa9b279e-07c6-4d0c-86f0-403c4af321fa","蚂蚁 RULER:六维评分表给 SVG 生成重写奖励信号","ruler-rubric-rewards-svg-generation","2026-09-23T23:06:24+00:00"]