[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wrop-object-permanence-world-models":3,"topics-all":38,"news-related-ba0ed7bf-3de3-4f92-98fe-a50d6ac274d0":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"ba0ed7bf-3de3-4f92-98fe-a50d6ac274d0","WROP 开源:用 150 个物体恒存任务给世界模型补认知课","物体恒存——东西被挡住仍知道它存在——是人类认知先验,视频世界模型未必有。31 人团队开源 WROP:150 个认知任务、150 万条语料加 300 题考试,14 个视频模型同考,自训 16B 模型拿续写组第一;论文两天冲上 HF 日榜第一,数据权重训练栈全部放出。","给一块布盖住玩具,几个月大的婴儿会伸手去掀——东西看不见了,但知道它还在。这种能力叫物体恒存性(object permanence),连同\"固体不会互相穿过\"的固体性直觉,是人类空间认知的地基。那么问题来了:被大家叫作\"世界模型\"的视频生成模型,到底有没有这份直觉?一支 31 人的研究团队用一套全开源的\"认知考试\"给出了系统回答,论文 9 月 23 日挂上 arXiv,两天后冲到 Hugging Face 日榜第一。\n\n## 一套可以无限换皮的考题\n\nWROP(World Reasoning with Object Permanence)由 150 个手工设计的认知科学任务组成,分为六类。每个任务配一个 Blender 生成器:速度、光照、机位等干扰项随便变,任务自身的认知结构不动,单个任务就能量产一万条以上样本。团队最终放出 150 万条训练语料,外加一套 300 题的考试——考的不是画面清晰度,而是\"挡板后面那个球还在不在原处\"这类物理直觉。\n\n## 14 个模型同考,自训 16B 拿下小组第一\n\n14 个视频模型进了考场,分三类:3 个 reference-to-video、7 个编辑类、4 个续写(continuation)类。盲测两两对打排 Elo 的结果是:团队用这套语料微调出的 16B 世界模型 PWM-WROP,在续写类里排第一,总榜第三——总榜前两名,是两个 reference-to-video 模型的统计性平局。训练用它同时开源的 PWM 训练栈,原生 PyTorch,跑在 AWS Trainium2 上。\n\n## 全开源,以及该泼的冷水\n\n数据、考题、14 个模型的答卷与得分、模型权重、训练栈,全部放出。社区反应也快:9 月 25 日论文在 Hugging Face Daily Papers 日榜第一,拿到 153 个 upvote,提交论文的作者本人还在评论区现身讲解。冷水照旧要泼:盲测 Elo 由团队自己组织,\"续写组第一\"是论文自报口径,目前没有第三方复现;16B 的 PWM-WROP 终归是在自家考题上验证自家模型,离真实世界视频里的恒存性还有距离。\n\n世界模型这半年拼算力、拼时长、拼物理一致性,这份工作指了另一条路:认知科学里最基础的先验——物体恒存、固体性——可以直接变成可规模化的训练数据。连\"看不见的东西还在\"都要专门补课,恰恰说明视频模型离真正的\"世界\"还缺最底层的一块。下一个问题:补完恒存性,该补哪门认知课?\n\n参考:[arXiv:2609.28654](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28654) · 项目页:[object-permanence.world](https:\u002F\u002Fobject-permanence.world\u002F)","https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28654","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":19,"name":20,"slug":20,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":22,"name":23,"slug":23,"description":14,"color":14},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"4b38c2b1-e15b-4135-8df6-03def345d370","en","WROP: 150 Object-Permanence Tasks to School Video World Models","WROP open-sources 150 object-permanence tasks, a 1.5M corpus and a 300-question exam; 14 video models tested, a 16B fine-tune leads its class.","Drape a cloth over a toy and an infant will reach to lift it — out of sight is not gone. That ability, object permanence, together with the solidity intuition that objects cannot pass through each other, is foundational to human spatial cognition. So the awkward question: do video generation models — the systems everyone now calls \"world models\" — actually have it? A 31-author team delivers a systematic answer in the form of a fully open-source cognitive exam. The paper landed on arXiv on Sep 23 and topped Hugging Face's Daily Papers board two days later.\n\n## An exam that can re-skin itself indefinitely\n\nWROP (World Reasoning with Object Permanence) consists of 150 hand-designed cognitive-science tasks across six categories. Each task ships with a Blender generator: speed, lighting, camera angle and other nuisance parameters get randomized while the task's cognitive structure stays fixed, letting a single task scale to 10,000+ samples. The team released a 1.5M-sample training corpus plus a 300-question exam — testing not text-to-video fidelity but physical intuition of the \"is the ball still behind the occluder\" kind.\n\n## 14 models sat the exam; a fine-tuned 16B took its group\n\nFourteen video models entered, spanning three classes: 3 reference-to-video, 7 edit, and 4 continuation. In a blind pairwise Elo study, PWM-WROP — the team's 16B world model fine-tuned on the corpus — ranked first among continuation models and third overall, behind a statistical tie between two reference-to-video models. Training ran on PWM, their simultaneously released native-PyTorch stack on AWS Trainium2.\n\n## Fully open — and the cold water\n\nData, exam, all 14 model answers with scores, weights, and the training stack are all public. The community moved fast: the paper hit #1 on Hugging Face Daily Papers for Sep 25 with 153 upvotes, and the author who submitted it showed up in the comments to explain the work. The cold water, as ever: the blind Elo was organized by the team itself, \"first among continuation models\" is the paper's own framing with no independent replication yet, and a 16B model validated on its own benchmark still has distance to cover before real-world video.\n\nFor half a year the world-model race has been about compute, clip length and physical consistency. This work points at another lane: the most basic cognitive priors — object permanence, solidity — can be turned directly into scalable training data. That \"unseen objects persist\" needs remedial training at all shows how much of the \"world\" in world models is still missing. Next question: which cognitive course comes after permanence?\n\nReference: [arXiv:2609.28654](https:\u002F\u002Farxiv.org\u002Fabs\u002F2609.28654) · Project: [object-permanence.world](https:\u002F\u002Fobject-permanence.world\u002F)","wrop-object-permanence-world-models","2026-09-25T17:08:02Z","2026-09-25T17:08:29.903636Z","2026-09-25T17:08:29.903647Z",true,"agent",344,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"9fa15063-58b1-4b87-a0cb-0e19ffc4dc6c","4B 参数横扫四大具身基准：开悟世界模型让小模型重新定义 SOTA","kaiwu-4b-world-model-sensetime-72x","2026-06-12T04:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"05757bc2-4fb4-47db-9660-d7a97bb75e1f","北大阿里 OmniEcho 开源:给具身智能装上空间听觉","omniecho-spatial-audio-embodied-agents","2026-09-27T21:09:09+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"af21d26d-d9cd-4526-ac4a-66366a45848c","AV-GRPO:8张A800给22B音视频模型做RL后训练","av-grpo-audio-video-diffusion-rl","2026-09-27T17:08:13+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"33bd7c4c-27f8-458c-8404-265134fc6ce8","视频生成缺的不是算力,是记忆:282 篇论文拼出一张全景地图","ar-video-generation-memory-survey","2026-09-24T21:09:28+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"9210c86b-3e1d-434b-bbfd-78e62c698aed","WorldCrafter 开源:给视频世界模型装上可查询的 3D 记忆,转一圈回来还是那个房间","worldcrafter-video-world-model-3d-memory","2026-09-22T19:08:48+00:00"]