Drape a cloth over a toy and an infant will reach to lift it — out of sight is not gone. That ability, object permanence, together with the solidity intuition that objects cannot pass through each other, is foundational to human spatial cognition. So the awkward question: do video generation models — the systems everyone now calls "world models" — actually have it? A 31-author team delivers a systematic answer in the form of a fully open-source cognitive exam. The paper landed on arXiv on Sep 23 and topped Hugging Face's Daily Papers board two days later.

An exam that can re-skin itself indefinitely

WROP (World Reasoning with Object Permanence) consists of 150 hand-designed cognitive-science tasks across six categories. Each task ships with a Blender generator: speed, lighting, camera angle and other nuisance parameters get randomized while the task's cognitive structure stays fixed, letting a single task scale to 10,000+ samples. The team released a 1.5M-sample training corpus plus a 300-question exam — testing not text-to-video fidelity but physical intuition of the "is the ball still behind the occluder" kind.

14 models sat the exam; a fine-tuned 16B took its group

Fourteen video models entered, spanning three classes: 3 reference-to-video, 7 edit, and 4 continuation. In a blind pairwise Elo study, PWM-WROP — the team's 16B world model fine-tuned on the corpus — ranked first among continuation models and third overall, behind a statistical tie between two reference-to-video models. Training ran on PWM, their simultaneously released native-PyTorch stack on AWS Trainium2.

Fully open — and the cold water

Data, exam, all 14 model answers with scores, weights, and the training stack are all public. The community moved fast: the paper hit #1 on Hugging Face Daily Papers for Sep 25 with 153 upvotes, and the author who submitted it showed up in the comments to explain the work. The cold water, as ever: the blind Elo was organized by the team itself, "first among continuation models" is the paper's own framing with no independent replication yet, and a 16B model validated on its own benchmark still has distance to cover before real-world video.

For half a year the world-model race has been about compute, clip length and physical consistency. This work points at another lane: the most basic cognitive priors — object permanence, solidity — can be turned directly into scalable training data. That "unseen objects persist" needs remedial training at all shows how much of the "world" in world models is still missing. Next question: which cognitive course comes after permanence?

Reference: arXiv:2609.28654 · Project: object-permanence.world