Alibaba's video generation model HappyHorse released version 1.1, with five major upgrades addressing the shortcomings of 1.0. The model is open-sourced and available via Alibaba Cloud's PAI platform.
The five upgrades:
- Motion quality: smoother motion, with reduced jitter and popping artifacts. The motion smoothness score improves from 7.2 to 8.7 (out of 10).
- Camera control: explicit camera control is now supported — the model can pan, zoom, track, and orbit based on user-specified camera paths.
- Multi-subject consistency: the model can now maintain consistent identity for 3+ subjects across a 30-second video, with 92% identity consistency.
- Audio-visual sync: the model can generate synchronized audio (sound effects, ambient music) along with the video, with sub-100ms lip-sync accuracy for speech.
- Long-video extension: 1.1 can extend a generated video beyond 30 seconds, with explicit "scene transition" markers to maintain coherence.
The benchmark: HappyHorse 1.1 scores 79.2 on the VBench long-video benchmark, on par with Kling 2.5 and slightly below Sora 2. The biggest improvement over 1.0 is in camera control and multi-subject consistency, which were the two biggest user complaints.
The commercial angle: HappyHorse 1.1 is available via Alibaba Cloud at $0.06 per second of 1080p video. The first batch of enterprise customers include Taobao (for product video generation) and Youku (for short-form content).
The bigger takeaway: "incremental improvement" is still the dominant mode in video generation. HappyHorse 1.1 is not a "breakthrough" — it's a "5 fixes + 1 polish" release. The video generation space is maturing, and vendors are focusing on fixing specific shortcomings rather than introducing entirely new capabilities. For the industry, this signals that "production-grade video generation" is the next milestone, and the focus is on reliability and controllability.