36Kr's recent AI evaluation published a horizontal review of 6 mainstream AI video generation tools, spanning US and China head vendors — ByteDance Seedance 2.0, Kuaishou Kling 3.0, Alibaba Wan 2.2, OpenAI Sora 2.0, Google Veo 3.1, and Runway Gen-4.5. On the surface it's a selection guide, but it actually reveals a deeper signal: this track has moved from "single-point racing" into the "all-scenario differentiation" stage. The technical routes are diverging. Seedance 2.0 emphasizes 4D multimodal mixed input of image, text, audio, and video plus director-style storyboard scheduling; Veo 3.1 makes audio-visual sync a native capability, directly generating voice and BGM, skipping the tedious post-process; Wan 2.2 bets on open source + private deployment. Almost all head players have already moved from UNet to DiT, to enjoy the Scaling Law dividend — this aligns with another 36Kr report on Seedance 2.0's step-jump effect after going from 100B to 200B+ parameters. Market positioning is also differentiating. ByteDance focuses on C-end zero threshold and short-drama ecosystem; Kuaishou locks down Chinese complex body movement; Alibaba takes the open-source enterprise-grade route; OpenAI excels at physical consistency and long video; Google emphasizes Gemini ecosystem integration; Runway focuses on deep integration of professional post-production toolchains. There's no longer the possibility of "one product fits all" — every player is using differentiated moats to avoid the pure price war. My judgment: AI video generation has entered the "specialized division of labor" stage. The moat isn't in model capability itself, but in whether the model can be embedded into high-value production scenarios. ByteDance has the most complete commercial closed loop domestically with the Hongguo+Douyin+Volcano Engine flywheel; globally Veo is regaining its position with Gemini ecosystem integration. This differentiation has no end point.