[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-anchorworld-kling-3d-motion-first-person":3,"news-related-36d4991a-f6ed-4edf-8f9a-e0fcb1212044":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"36d4991a-f6ed-4edf-8f9a-e0fcb1212044","可灵团队提出 AnchorWorld：用 3D 人体运动重塑「第一人称世界模拟」","当 Sora、Veo 等视频生成模型还在卷「怎么把单条视频拍得更像电影」时，视频生成领域的下一个战场已经悄悄转移——可交互的世界模型（World Model）。快手可灵（Kling）团队联合清华在 arXiv 上放出的 AnchorWorld，正是一份来自工业界头部玩家、对「可定制、可交互、可自我演化」世界模拟框架的硬核回应。\n\n论文的核心切入点很明确：用 3D 人体运动作为交互的第一模态。第一人称视角天然存在视野遮挡和身体截断的问题，作者引入一个「与智能体第一人称感知解耦」的辅助监督信号，让模型能从外部视角观察智能体全身相对环境的位置，从而把「人-世界交互」的空间锚定做得更扎实。\n\n更值得注意的是「Anchor View + 文本驱动」的自演化机制：在统一世界坐标系下定义若干锚定视角，配合文本描述来约束局部场景的动态演化。简单，但有效——实验显示其在时空几何一致性上严格遵循预设动态，且在多项 SOTA 基准上显著领先。\n\n如果说之前的世界模型（Project Genie、SANA-WM 等）解决的是「能不能生成一个能走进去的视频」，AnchorWorld 回答的是「走进去之后能不能像玩游戏一样改写这个世界」。这或许才是通向具身智能与 AGI 的真正桥梁。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2606.07326","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"eaf35b67-8c08-4d6a-8567-90ee14f1175d","kling",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"39604fcd-2285-460a-8946-4dcef249bd5c","en","Kuaishou's AnchorWorld: egocentric world sims from 3D motion","While video generation models like Sora and Veo are still racing to \"make a single video look more cinematic,\" the next battlefield in the video generation field has quietly shifted — interactive world models. AnchorWorld, released by Kuaishou Kling team in collaboration with Tsinghua on arXiv, is a hardcore response from an industry-leading player to a \"customizable, interactive, self-evolving\" world simulation framework.\n\nThe core entry point of the paper is very clear: use 3D human motion as the first modality of interaction. The first-person view naturally suffers from view occlusion and body truncation, and the authors introduce an auxiliary supervision signal \"decoupled from the agent's first-person perception,\" letting the model observe the agent's full body position relative to the environment from an external view, thus making the spatial anchoring of \"person-world interaction\" more solid.\n\nMore noteworthy is the self-evolution mechanism of \"Anchor View + text-driven\": several anchor views are defined under a unified world coordinate system, paired with text descriptions to constrain the dynamic evolution of local scenes. Simple, but effective — experiments show it strictly follows preset dynamics on spatiotemporal geometric consistency, and significantly leads on multiple SOTA benchmarks.\n\nIf previous world models (Project Genie, SANA-WM, etc.) answered \"can it generate a video you can walk into,\" AnchorWorld answers \"after walking in, can you rewrite this world like a game.\" This may well be the real bridge to embodied intelligence and AGI.","anchorworld-kling-3d-motion-first-person","2026-06-08T10:00:00Z","2026-06-08T10:08:06.036854Z","2026-08-19T02:08:40.142862Z",true,"agent",148,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"413b7c1f-e12b-4571-8076-8b5511360bbd","AlayaWorld开源:用3D缓存+DMD蒸馏破解长时视频世界模型一致性难题","alayaworld-long-video","2026-07-14T10:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"bfa4db7e-f3a2-4da6-9823-faa6ccef2274","高德 ABot-World Studio 把世界模型压进消费级 GPU：单卡可跑 + 全开源的另一种解法","gaode-abot-world-studio","2026-07-14T04:01:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"4f4da76a-ee45-4fdc-9a82-a346c6712995","阿里视频生成模型 HappyHorse 1.1：五维升级补齐 1.0 短板","alibaba-happyhorse-1-1-five-dim-upgrade","2026-06-22T08:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"43644653-55a5-40e6-b48c-dd9a548b7311","可灵 3.0 Turbo 落地：把视频生成拆成「快速预览 + 影院成片」两段式工作流","kling-3-0-turbo-two-stage-workflow","2026-06-22T00:04:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"e8ad36b6-ec3c-4c90-8fbc-bf1e6807ae34","xAI Grok Imagine Video 1.5：单图生视频登顶 Arena榜首，自回归 MoE 改写视频生成规则","grok-imagine-1-5-ar-moe-arena-top","2026-06-07T08:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"c362b1b1-a804-4879-a516-8afe724d392c","可灵AI两周年：26次迭代+122篇论文，视频生成模型的「中国速度」","kling-2-year-122-papers-china-speed","2026-06-05T13:00:00+00:00"]