36Kr's feature on BAAI (Beijing Academy of Artificial Intelligence) researcher Wang Zhongyuan argues that world models are at a "2012 moment" — comparable to deep learning's breakthrough around 2012. The argument: the "latent space unification" approach is the fifth major paradigm in AI, and the US and China are racing toward it together.

The "2012 moment" framing: Wang draws a parallel to 2012, when AlexNet won ImageNet and triggered the deep learning revolution. World models in 2026 are at a similar inflection point — the recent breakthroughs (Genie 3, Sora 2, Wan-Streamer) suggest that "world model" is the next general-purpose AI paradigm.

The "latent space unification" insight: the "fifth answer" is the idea that all modalities (text, image, video, audio, action) can be unified in a single latent space, and a single model can operate over this unified space. This is a more ambitious version of the "multimodal" idea — instead of having separate encoders for each modality, a single latent space captures everything.

The "US-China race" angle: Wang argues that the US and China are both pursuing latent space unification, but with different approaches. The US labs (OpenAI, Anthropic, Google) are pursuing it via scaling (bigger models, more data). The Chinese labs (BAAI, Qwen, DeepSeek) are pursuing it via architecture (MoE, SSM, novel attention). The two approaches will likely converge.

The bigger takeaway: "latent space unification" is the next major AI paradigm. The "2012 moment" framing suggests we're on the cusp of a major transformation, and the "latent space unification" approach is the most promising direction. For the industry, this signals that "AI as a unified latent space" is the long-term vision, and the next 5-10 years will be defined by progress toward this goal.