Peking University + Tencent Hunyuan's joint team open-sources GEAR (Guided End-to-End AutoRegression for Image Synthesis), using a dual read-out architecture to truly end-to-end co-train the VQ tokenizer and AR generator, pushing gFID convergence speed on ImageNet to about 10× the LlamaGen-REPA baseline, and shifting the alignment cost from the tokenizer side to the AR side, providing a minimally changed, directly drop-in training paradigm for the next generation of open-source visual generation models.