World models are the hottest thing this year, but most evaluation still stays at "are the generated videos good-looking". Meituan LongCat's open-source WBench pulls the battlefield to "can you really walk into this world and control it": 289 test cases, 1058 interaction turns, covering four interaction types — navigation, subject action, event editing, viewpoint switching — with a unified interface letting text-driven models, camera-pose models, and keyboard-control models compete on the same field. The conclusions from the first test of 20 frontier models are quite hard-hitting: no omnipotent model exists; navigation capability is almost zero-correlated with video quality — the model "knows" what the world looks like, but doesn't know where it is in the world; under multi-turn interaction, navigation scores drop a sharp 33 points from the first to the fourth turn, exposing pose-error accumulation turn-by-turn as a structural hard injury of the iterative-generation paradigm; viewpoint switching is the universally recognized hardest item, with an average score of only 30.7. WBench's value isn't just the leaderboard, but as a starting point for the research-paradigm shift from "passive generation" to "active interaction".