A joint team from CUHK-Shenzhen, Huazhong University of Science and Technology and USTC released VideoRAE, demonstrating for the first time that frozen representations from "understanding" video foundation models such as V-JEPA 2 and VideoMAEv2 can be directly adapted into generation-friendly latents. On UCF-101 class-conditional video generation, the AR/DiT two-track gFVD scores reach 40 and 93 respectively. In a 2B text-to-video comparison, replacing LTX-VAE yields simultaneous improvements on all three VBench dimensions and ~5× faster convergence. Code is open-sourced.