ByteDance Seed and CUHK MMLab propose the Chain-of-Frame reasoning paradigm, based on OpenCoF-17K (11 task categories / 17,312 samples / 4 curation pipelines) doing LoRA fine-tuning on Wan2.2-I2V-A14B, no architecture change needed to comprehensively beat the baseline on four external video-reasoning benchmarks (VIPER, Gen-ViRe, MME-CoF, RULER-Bench); further introducing visual/text reasoning-token mechanisms, with the former leaning toward perception and the latter toward semantics, pushing the new path of "video generation model as a pure visual reasoner" from paper concept to engineering reproducibility.