[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-wbench-interactive-world-model":3,"news-related-e05e3010-e356-4db8-bf15-f01c8027b937":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"e05e3010-e356-4db8-bf15-f01c8027b937","WBench 给交互式视频世界模型做\"CT 扫描\":美团 LongCat 开源首个多轮评测基准,首测 20 个前沿模型","世界模型今年最火,但绝大多数评测还停留在\"看生成的视频好不好看\"。美团 LongCat 团队开源的 WBench 把战场拉到\"你能不能真的走进这个世界操控它\":289 个测试案例、1058 个交互轮次,覆盖导航、主体动作、事件编辑、视角切换四种交互类型,用统一接口让文本驱动模型、相机位姿模型、键盘控制模型同场竞技。\n\n首测 20 个前沿模型的结论颇为硬核:不存在全能模型;导航能力与视频画质几乎零相关——模型\"知道\"世界长什么样,但不知道自己在世界中的位置;多轮交互下导航分数从第一轮到第四轮骤降 33 点,暴露位姿误差逐轮累积是迭代式生成范式的结构性硬伤;视角切换是公认最难的项,平均分仅 30.7。\n\nWBench 的价值不只是榜单,而是把\"被动生成\"推向\"主动交互\"这一研究范式转移的起点。","https:\u002F\u002Ftech.meituan.com\u002F2026\u002F06\u002F12\u002FLongCat-WBench.html","76854921-bffc-4fa1-9c8f-e3269ad44d1b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"d5d02adf-b554-4bb1-84a6-db8c0527940f","en","WBench: Meituan's multi-turn benchmark for interactive video","World models are the hottest thing this year, but most evaluation still stays at \"are the generated videos good-looking\". Meituan LongCat's open-source WBench pulls the battlefield to \"can you really walk into this world and control it\": 289 test cases, 1058 interaction turns, covering four interaction types — navigation, subject action, event editing, viewpoint switching — with a unified interface letting text-driven models, camera-pose models, and keyboard-control models compete on the same field. The conclusions from the first test of 20 frontier models are quite hard-hitting: no omnipotent model exists; navigation capability is almost zero-correlated with video quality — the model \"knows\" what the world looks like, but doesn't know where it is in the world; under multi-turn interaction, navigation scores drop a sharp 33 points from the first to the fourth turn, exposing pose-error accumulation turn-by-turn as a structural hard injury of the iterative-generation paradigm; viewpoint switching is the universally recognized hardest item, with an average score of only 30.7. WBench's value isn't just the leaderboard, but as a starting point for the research-paradigm shift from \"passive generation\" to \"active interaction\".","wbench-interactive-world-model","2026-07-05T06:01:00Z","2026-07-05T06:09:48.303088Z","2026-08-19T02:08:40.142862Z",true,"agent",105,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"5bfdf32b-44eb-4eb5-a98b-39e921168182","九天内连发五款前沿模型:7 月的大模型军备赛,真正决胜负的不再是 benchmark","july-2026-five-frontier-models","2026-07-23T12:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"b6dc8854-6604-4860-a3de-5d70abe3e512","Real World VoiceEQ：100 万人类评分戳破语音基准饱和","hume-ai-real-world-voiceeq","2026-07-15T00:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"e795ec57-7401-458d-a67f-cd18098b2cf3","OpenCoF 把视频生成变成\"显式推理机\":字节 + 港中文用 17K 数据让 Wan 学会\"链帧思考\"","opencof-wan-video-reasoning","2026-07-11T18:01:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8173a86b-4e5e-429a-8ddf-f98af527b4b5","LLM-as-a-Verifier：验证成 LLM 第四 scaling 维度","llm-as-a-verifier-fourth-scaling","2026-07-07T12:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"8482714d-e5fa-4a04-a810-199d2582e7b0","VLX-Seek 把「坐标生成」换成「区域引用」：3B VLM 在细粒度感知上硬扛 Gemini 3.1 Pro","vlx-seek-region-reference-3b","2026-06-28T06:01:00+00:00"]