[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-4danyone-monocular-video-4d-human":3,"news-related-2874a2e5-beae-4627-8f6f-a34cf2cc8d7a":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","浙大、蚂蚁集团等机构发布4DAnyone:输入一段未标定单目视频,先生成数十路视角一致的目标视频,再提升为4D高斯泼溅人体。针对分组生成的结构漂移,用RCP把参考上下文压到固定预算、TCR在去噪中轮换分组共享信息,在DNA-Rendering与DyMVHumans上优于既有方法,代码权重已开源。","8 月 20 日提交到 arXiv 的一项研究(arXiv:2608.20335)把「随手拍一段视频」和「可自由转动视角的 4D 人体」直接连在了一起:4DAnyone 框架输入一段未经标定的单目视频,先生成数十路多视角一致的目标视频,再把这些视频提升为 4D 高斯泼溅(4DGS)人体模型。论文由浙江大学、蚂蚁集团、Robbyant、HKUST、CUHK 五家机构合作完成,项目页标注为 SIGGRAPH Asia 2026,代码与模型权重均已公开。\n\n## 用生成模型「补齐」相机阵列\n\n按照项目页的描述,照片级真实的 4D 人体重建通常依赖一组经过标定、时间同步的相机阵列,这严重限制了它在真实场景中的使用。4DAnyone 的思路是换个方向:重建需要的本来就是「那个阵列本来会拍到的几十路视频」,那就让视频扩散模型把这些视频生成出来,再用生成结果训练 4DGS。项目页的定位一句话就能讲完:单视频进,4D 人体出——不用 rig,不用标定,不用三脚架。\n\n## 一致性为什么会在几十个视角上崩掉\n\n直接把相机控制的视频扩散模型拉到这个规模,多视角一致性会失效。论文把根因归结为「有界注意力上下文」问题:当目标视角数量超过单次 DiT 前向的容量,生成就必须分组进行,而这暴露出两个耦合的瓶颈。参考侧,条件里包含所有已生成视角,上下文随视角数呈 O(N) 增长,跨视角的外观引导反而被稀释;目标侧,彼此不相交的分组无法直接交换信息,产生全局结构漂移。\n\n对应的两个设计相当工程化:\n\n- **Reference Context Packing(RCP)**:把持续增长的参考视角压缩为固定长度的混合分辨率上下文,参考侧复杂度从 O(N) 降到 O(1);\n- **Target Context Routing(TCR)**:在去噪过程中轮换目标视角的分组方式——高噪声步骤跨组共享上下文,低噪声步骤稳定细节。\n\n条件信号上还有一个关键取舍:用 3D 骨架提供「稀疏但准确」的几何线索,替代脆弱的稠密深度图与带噪声的相机参数。论文认为这是它在野外素材上能稳定泛化的原因。\n\n## 游戏引擎造数据,两个基准验证\n\n训练数据方面,团队用自研游戏引擎构建了 MVGameHuman 数据集,并与光场(light-stage)数据、野外视频混合训练。评测在 DNA-Rendering 与 DyMVHumans 两个基准上进行:论文报告其在新视角视频质量与下游 4DGS 重建两项上均优于既有方法,并表现出稳健的野外泛化;最终的 4DGS 模型训练用到了 FreeTimeGS。\n\n## 开源与社区热度\n\n代码托管在 GitHub(ant-research\u002F4DAnyone),权重发布在 Hugging Face(AntResearch\u002F4DAnyone)。8 月 21 日,这篇论文登上 Hugging Face Daily Papers 当日第三,拿到 64 个 upvote。对做自由视角体积视频、数字人内容的工作流来说,这是一个可以直接上手跑的管线,而不是只存在于论文里的方法。\n\n从更大的视角看,4DAnyone 代表的是一类新解法:当重建(4DGS)与生成(视频扩散)在「多视角一致性」这个接口上相遇,瓶颈不再只在重建侧,而在生成侧的上下文管理——RCP 和 TCR 本质上都是在给 DiT 的注意力上下文做「预算管理」。这个思路对其他需要大规模一致生成的任务同样有参考价值:上下文不够用时,压缩它,而不是无限堆叠它。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2608.20335","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"7b67033c-19e6-4052-a626-e681bba64c7a","diffusion",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"405796c8-e60e-4842-bf94-3e3959596868","en","4DAnyone: One Casual Monocular Video In, A 4D Human Out","4DAnyone reconstructs a 4D Gaussian Splatting human from one uncalibrated monocular video, generating tens of consistent views to bridge the gap.","A study submitted to arXiv on August 20 (arXiv:2608.20335) connects \"a casually shot video\" directly to \"a 4D human you can orbit freely\": the 4DAnyone framework takes an uncalibrated monocular video as input, first generates tens of multiview-consistent target videos, and then lifts them into a 4D Gaussian Splatting (4DGS) human model. The paper is a collaboration between five institutions - Zhejiang University, Ant Group, Robbyant, HKUST, and CUHK - with the project page listing SIGGRAPH Asia 2026 as the venue; both code and model weights are already public.\n\n## Filling In the Virtual Camera Array with a Generative Model\n\nAs the project page describes it, photorealistic 4D human reconstruction normally depends on a calibrated array of synchronised cameras, which severely limits real-world use. 4DAnyone flips the direction: what reconstruction needs is precisely \"the tens of videos that array would have recorded\", so let a video diffusion model generate them, and train 4DGS on the results. The project page's pitch is one line: single video in, 4D human out - no rig, no calibration, no tripod.\n\n## Why Consistency Breaks at Tens of Views\n\nScaling an off-the-shelf camera-controlled video diffusion model to this size breaks multiview consistency. The paper traces the failure to a \"bounded-attention-context\" problem: once the number of target views exceeds the capacity of a single DiT forward pass, generation must be split into groups, which exposes two coupled bottlenecks. On the reference side, conditioning on all previously generated views makes the context grow as O(N), weakening cross-view appearance guidance. On the target side, disjoint groups cannot directly exchange information, causing global structural drift.\n\nThe two corresponding designs are decidedly engineering-flavored:\n\n- **Reference Context Packing (RCP)** compresses the growing reference views into a fixed-length mixed-resolution context, dropping reference-side complexity from O(N) to O(1);\n- **Target Context Routing (TCR)** rotates the target-view groupings during denoising - sharing context across groups at high-noise steps and stabilizing details at low-noise steps.\n\nThere is one more key trade-off in the conditioning signal: a 3D skeleton supplies sparse-but-accurate geometric cues, replacing fragile dense depth and noisy camera parameters. The authors credit this for the model's robust in-the-wild generalization.\n\n## Game-Engine Data, Two Benchmarks\n\nFor training data, the team built the MVGameHuman dataset with an in-house game engine, mixing it with light-stage data and in-the-wild videos. Evaluation runs on DNA-Rendering and DyMVHumans: the paper reports it outperforms prior methods on both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization; the final 4DGS model is trained with FreeTimeGS.\n\n## Open Source and Community Traction\n\nCode is hosted on GitHub (ant-research\u002F4DAnyone) and weights on Hugging Face (AntResearch\u002F4DAnyone). On August 21 the paper reached #3 on Hugging Face Daily Papers with 64 upvotes. For workflows built on free-viewpoint volumetric video and digital humans, this is a pipeline you can actually run, not just a method that lives in a paper.\n\nSeen from a wider angle, 4DAnyone represents a new class of solutions: where reconstruction (4DGS) and generation (video diffusion) meet at the interface of \"multiview consistency\", the bottleneck is no longer only on the reconstruction side but in context management on the generation side - RCP and TCR are, in essence, budget management for the DiT attention context. The same lesson transfers to other tasks that need consistency at scale: when you run out of context, compress it - do not just keep stacking it.","4danyone-monocular-video-4d-human","2026-08-20T17:59:53Z","2026-08-22T19:10:08.098811Z","2026-08-22T19:10:08.098827Z",true,"agent",56,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"6f9e9f94-9dcc-4c6c-b254-6c5d0fe8ed37","京东开源 JoyAI-Video-Edit:16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-realtime-diffusion","2026-08-10T00:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"bcdc10bc-2f08-4c39-8ffa-e7e34041c112","京东开源 JoyAI-Video-Edit:用 16B 多模态扩散 Transformer 把视频编辑推进「边播边改」实时流时代","jd-joyai-video-edit-real-time-streaming","2026-08-05T03:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"f4c705fd-47c9-481a-807f-8001820070f8","InfinityEdit:三注意力轻量适配器,把视频编辑推进无界流时代","infinityedit-infinite-video-editing-adapter","2026-08-25T13:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"18d2aa73-7244-4b10-b611-46475e17327e","ForgeWM开源:一步去噪72FPS的可玩世界模型,8张卡复现全流程","forgewm-few-step-playable-world-model","2026-08-24T21:10:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00"]