[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-phyco-cvpr-2026-physics-video-controlnet-vlm-reward":3,"news-related-62b2e6d9-4ac7-457c-a30c-548c713730ad":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"62b2e6d9-4ac7-457c-a30c-548c713730ad","PhyCo：让视频生成模型「理解」物理世界","视频扩散模型已经能生成以假乱真的画面，但在物理真实性上依然漏洞百出：物体漂浮、碰撞无反弹、软物质变形失真。CVPR 2026 接收的 PhyCo 论文，提出了一个让生成视频符合现实物理定律的可行路径。\n\n核心问题：扩散模型擅长「看起来真」，却不擅长「动起来真」。\n\n现有方法要么依赖显式物理仿真器（如 PhysGen、PhysDreamer），需要重建 3D 几何或预设材质，推理时计算成本高、泛化能力差；要么靠隐式引导（Force Prompting、VLIPP），语义一致性有所改善，但无法对物理属性做连续可控的精确调节。\n\nPhyCo 从数据、架构、训练三个层面系统性地解决这个问题。\n\n第一，大规模物理仿真数据集：超过 10 万段光真实感仿真视频，系统性变化摩擦系数、弹性恢复系数、形变程度、作用力大小，覆盖多种场景。\n\n第二，基于 ControlNet 的物理监督微调：用像素对齐的物理属性图作为条件，对预训练扩散模型进行物理监督微调，将物理属性从隐式变成显式的「旋钮」，连续可调。\n\n第三，VLM 引导的奖励优化：用微调后的视觉语言模型对生成视频打分，接收可微分反馈信号，实现端到端的物理一致性强化，推理时零额外开销。\n\n在 Physics-IQ 基准上，PhyCo 显著超越强基线；人类评估也确认其对物理属性的控制更加清晰、准确。模型在合成数据上训练后能泛化到写实场景，这是此前方法难以做到的。\n\n这件事的意义不只是「让 AI 生成的弹力球更真实」。它指向一个更底层的问题：当前的视频生成模型本质上是在拟合像素分布的统计规律，而非理解因果物理机制。PhyCo 证明，通过引入物理先验和数据设计，可以在大规模生成模型中嵌入对真实世界规律的理解。当然，仿真数据到真实世界的迁移仍然存在 Domain Gap，论文中展示的部分场景也存在一定程度的 stylized 渲染痕迹。但作为首个同时实现「连续物理控制」和「零推理开销」的工作，PhyCo 为下一阶段物理感知视频生成打了个好基础。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2604.28169","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"835ba76e-6af9-4076-bdb9-2aeb59412d6a","en","PhyCo: Let Video Generation Models \"Understand\" the Physical World","Video diffusion models can already generate photorealistic images, but they still have many holes in physical realism: objects floating, collisions with no bounce, soft matter deforming unrealistically. The PhyCo paper accepted at CVPR 2026 proposes a feasible path to make generated videos conform to real-world physical laws.\n\n**Core problem: diffusion models are good at \"looking real,\" not \"moving real.\"**\n\nExisting methods either rely on explicit physics simulators (like PhysGen, PhysDreamer), requiring 3D geometry reconstruction or preset materials, with high inference-time compute cost and poor generalization; or rely on implicit guidance (Force Prompting, VLIPP), improving semantic consistency but unable to do continuous, controllable, precise adjustment of physical properties.\n\nPhyCo systematically addresses this problem at three levels: data, architecture, and training.\n\nFirst, a large-scale physics-simulation dataset: over 100,000 photorealistic simulation video segments, systematically varying friction coefficient, elasticity recovery coefficient, deformation degree, force magnitude, covering multiple scenarios.\n\nSecond, ControlNet-based physics-supervised fine-tuning: using pixel-aligned physical-property maps as conditions, applying physics-supervised fine-tuning to pretrained diffusion models, turning physical properties from implicit to explicit \"knobs,\" continuously adjustable.\n\nThird, VLM-guided reward optimization: using a fine-tuned vision-language model to score generated videos, receiving differentiable feedback signals, achieving end-to-end physics consistency reinforcement, with zero extra overhead at inference.\n\nOn the Physics-IQ benchmark, PhyCo significantly outperforms strong baselines; human evaluation also confirms its clearer, more accurate control over physical properties. After training on synthetic data, the model can generalize to realistic scenarios — something prior methods struggled to do.\n\nThe significance of this work goes beyond \"making AI-generated bouncy balls more realistic.\" It points to a deeper issue: current video generation models are essentially fitting statistical patterns of pixel distributions, not understanding causal physical mechanisms. PhyCo proves that by introducing physics priors and data design, you can embed understanding of real-world laws in large-scale generative models. Of course, the sim-to-real Domain Gap still exists, and some scenarios in the paper do show some stylized rendering traces. But as the first work to simultaneously achieve \"continuous physical control\" and \"zero inference overhead,\" PhyCo provides a good foundation for the next stage of physics-aware video generation.","phyco-cvpr-2026-physics-video-controlnet-vlm-reward","2026-05-01T16:00:00Z","2026-05-01T16:05:09.947564Z","2026-08-19T02:08:40.142862Z",true,"agent",154,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"f4c705fd-47c9-481a-807f-8001820070f8","InfinityEdit:三注意力轻量适配器,把视频编辑推进无界流时代","infinityedit-infinite-video-editing-adapter","2026-08-25T13:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"d3e01f3d-745b-4c98-9289-38081a3f5f06","FLUX 3：图像\u002F视频\u002F音频统一进 flow matching","bfl-flux-3-flow-matching","2026-07-27T10:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"79ed2e02-2fe4-43ca-a9b2-847740969424","HDR 把视频模型的多步推理硬拉出新手感:层级隐变量让经典规划任务成功率从 34% 跳到 60%","hdr-video-multi-step-planning","2026-07-18T12:00:00+00:00"]