Video generation has reached the point where 30-second multi-shot clips are table stakes. The scarce resource is the sentence that has to carry them: a user types "a girl climbs a tree," and the generator is expected to deliver a full storyboard — camera moves, framing, lighting, dialogue, edit rhythm. That gap is becoming its own model category. WanPE, a paper landed on arXiv on Sep 24, is the heaviest answer yet: a 397B-parameter model whose only job is rewriting video prompts (arXiv:2609.30221).
397B parameters, one job: writing the prompt
WanPE does prompt enhancement: one line of user intent goes in, a shot-level "cinematic plan" comes out — machine position, camera motion, lighting, sound, and action beats written as instructions the generator can execute. To train this directorial skill, the team used 1.05M real-world videos and reconstructed shot plans backwards from finished footage rather than forward-expanding short captions.
The parameter count deserves a pause: 397B, larger than most of the downstream task models it serves. The team's bet is that the bottleneck has shifted from "can it generate" to "is the instruction professional enough" — prompt quality now caps output quality, and that alone justifies a flagship-scale model.
Reverse construction + SC-GRPO: add cinema, don't drift
Two design choices carry the method. First, video-grounded reverse construction: the training data is not human-written "good prompts" but shot plans reverse-engineered from real video. Ablations show this path clearly wins — same setting, forward rewriting scores 39.49 overall, forward-target SFT 35.17, while the reverse-constructed WanPE-SFT reaches 49.86. Second, Semantic-Consistency GRPO: the biggest risk of prompt expansion is drifting away from what the user actually asked for, so SC-GRPO bakes original-intent consistency into the RL objective, preserving semantic fidelity across model scales.
The companion benchmark, WanPEval, is human-annotated, spanning 5 to 30 seconds across intent granularities, backed by roughly 11K blind pairwise assessments.
Numbers talk — check who is speaking
Plugged in front of Wan3.0's video generator, WanPE-397B lifts human preference by 10.66 to 18.84 points at 5-15 seconds; the 30-second arena is steeper: 50.86 points — raw requests score 9.38 overall, and with WanPE they hit 60.24. Head-to-head with Seedance 2.5, the overall score reads 60.24 vs 59.76, essentially a tie, with categories splitting both ways: animation and speech clearly ahead (81.25 vs 46.43, 73.68 vs 55.00), action, song-and-dance, and ads behind (50.00 vs 68.18, 47.50 vs 60.00, 75.00 vs 81.25). The paper's claim is "leads all evaluated commercial offerings" at 5-15 seconds and "remains competitive with Seedance 2.5" at 30 — note this is a self-reported result on a team-built benchmark.
Cross-generator transfer is the other interesting result: swapped onto LTX-2.5-Base, format-adapted WanPE scores 35.56 vs the native enhancer's 21.11; onto MiniMax-H3-Base, 41.09 vs the native path's 35.92. Directorial skill appears substantially portable.
Before the weights land, two buckets of cold water
First, on the Hugging Face paper page, models, datasets, and Spaces linking to this paper all count zero — paper and project page are out, weights and code are not, leaving third parties no way to reproduce for now. Second, "leads all commercial offerings" comes from the team's own benchmark; 11K blind assessments is a decent sample, but the benchmark and the graded model share an author list.
The structural point is worth more than the scores: when writing the prompt itself takes a 397B model, "one sentence to video" quietly becomes "one sentence plus an invisible director." You hand expressive control to an intermediary model and get professional storyboarding back — whether that trade is worth it depends on whether you sit inside that 50.86-point target zone. The project page (wan-pe.github.io) hosts comparison demos; watch for weight-release news.