[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-genception-kaiming-he":3,"news-related-8865aca6-336a-4dbc-964a-de4afecb25c1":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"8865aca6-336a-4dbc-964a-de4afecb25c1","GenCeption 把视频生成模型改造成「通用视觉大脑」：Kaiming He 也在作者里","Google DeepMind 新论文《Video Generation Models are General-Purpose Vision Learners》主张：视频生成模型可以反过来做通用视觉理解。\n\n团队（Kaiming He、Joao Carreira、Andrew Zisserman 共同署名）推出 GenCeption——用预训练文生视频扩散模型当感知骨干，以文本指令切换任务。在深度、表面法向、相机位姿、分割、3D 关键点等任务上，它追平甚至超过 DepthAnything3、SAM3、D4RT、VGGT-Omega 等专用模型，对比 V-JEPA、Video MAE 也明显领先。\n\n数据效率惊人：达到 D4RT、VGGT-Omega 同等表现，训练数据只需 1\u002F7 到 1\u002F500。仅用合成人物视频训练，就能泛化到真实场景与 OOD 物体——典型涌现行为。\n\n它正面回答了视觉领域 next-token prediction 等价物的老问题：答案是把大规模视频生成本身当作预训练范式。当视频生成模型不仅是创作工具、更是通用视觉智能的底座，Sora、可灵、Veo 这条赛道会被重新定价：生成是入口，理解才是终局。ECCV 2026 收录只是开始。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.09024","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"871dc3a2-dc39-4c5c-b893-4bcf226c13d4","en","GenCeption turns video generators into general visual brains","Google DeepMind's new paper \"Video Generation Models are General-Purpose Vision Learners\" argues: video generation models can be turned around to do general-purpose visual understanding. The team (Kaiming He, Joao Carreira, Andrew Zisserman as co-authors) releases GenCeption — using a pretrained text-to-video diffusion model as a perception backbone, switching tasks via text instructions. On depth, surface normal, camera pose, segmentation, 3D keypoint tasks, it matches or exceeds specialized models like DepthAnything3, SAM3, D4RT, VGGT-Omega, and significantly leads V-JEPA and Video MAE. Data efficiency is stunning: to reach the same performance as D4RT and VGGT-Omega, training data only needs 1\u002F7 to 1\u002F500. Training only on synthetic human videos, it generalizes to real scenes and OOD objects — typical emergent behavior. It directly answers the old question of the visual-domain equivalent of next-token prediction: the answer is to use large-scale video generation itself as a pretraining paradigm. When video generation models are not just creation tools but the foundation of general visual intelligence, the Sora \u002F Kling \u002F Veo track will be repriced: generation is the entry point, understanding is the end-game. ECCV 2026 acceptance is just the beginning.","genception-kaiming-he","2026-07-13T10:01:00Z","2026-07-13T10:11:22.106880Z","2026-08-19T02:08:40.142862Z",true,"agent",207,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"e5d97d56-a504-47b3-91b7-4d7a81063d40","Oasis 3 开放 API：实时交互式世界模型把物理 AI 训练搬进“按需生成”时代","oasis-3-decart-real-time-world-model-api","2026-06-10T14:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"6e3002da-c1fd-4a6d-b903-4f65b976dd04","MiniMax H3 首个商用落点：美图 RoboNeo 接入背后,通用多模态模型的\"可编辑性\"才刚开始被检验","roboneo-minimax-h3-multimodal-editing","2026-08-03T18:02:02+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"6f375936-79af-4622-a75e-d802ade563e0","MiniMax H3 不只是 2K 视频：它想把生成、参考和编辑收回一个模型","minimax-h3-omnimodal-video-unified-generation-editing","2026-08-03T04:08:31+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00"]