[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-kaiwu-4b-world-model-sensetime-72x":3,"news-related-9fa15063-58b1-4b87-a0cb-0e19ffc4dc6c":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"9fa15063-58b1-4b87-a0cb-0e19ffc4dc6c","4B 参数横扫四大具身基准：开悟世界模型让小模型重新定义 SOTA","大晓机器人（商汤系）旗下开悟世界模型（Kairos）最近同时在 RoboTwin 2.0、LIBERO-Plus、WorldModelBench Robot、DreamGen 四项全球权威具身智能评测中拿下第一，把 Cosmos 2.5-14B、Wan 2.2-5B、Lingbot 等百亿俱乐部选手压在身后。考虑到它只有 4B 参数、23.5GB 显存占用，这不仅是跑分赢，更像是在质疑大参数=高性能的行业惯性。技术差异在架构层面。开悟 3.0 走多模态理解—生成—预测原生一体路线，把物理因果链和思维链直接编进决策过程，而不是像多数同行那样在视频扩散模型外挂运动接口。其自研的混合时间线性注意力算子是真正放量点——A800 上 10 秒生成任务仅耗时 9.5 秒，对比 Cosmos 2.5 的 687.2 秒提速 72 倍，云侧 1:1 实时推理也因此首次成为可能。更值得关注的是端侧落地。它是行业首个在 Jetson Thor T5000 平台跑出 1:1.5（生成时间：视频时长）实时生成的具身世界模型，意味着机器人本体可以想到即可做到，省掉中间转译环节。一脑多形泛化也跑通了——同一权重可同时驱动单臂、双臂、灵巧手，覆盖智元 G1、松灵 PIPER、宇树 G1 等不同硬件。当视频生成赛道还在拼更长的上下文、更大的窗口时，开悟选了一条反方向的路：用 4B 模型、几 GB 显存、端侧实时，去撬动具身智能从仿真到真机的最后一公里。benchmark 第一只是个引子——真正值得跟踪的是它在机器人量产里能不能稳定交付。如果跑分之外它也能在工厂、产线上稳定干活，那世界模型这个词可能要从内容生成赛道，重新分类到具身基础设施。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3849612388570374","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"120fa59a-ff6f-4537-9bf5-f818df636a0e","benchmark",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":18,"name":19,"slug":19,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",{"id":21,"name":22,"slug":22,"description":13,"color":13},"ebe5dcd1-46b1-4298-b8c2-8e0e2f456e56","video-generation",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"dd25011a-5a19-473c-90f8-0a9702695d38","en","A 4B world model sweeps four embodied benchmarks","36Kr reports on a 4B-parameter \"world model\" from Kaiwu (开悟) that sweeps four embodied AI benchmarks, redefining SOTA. The standout: a 4B model matches or beats 70B+ models on embodied tasks, marking a major milestone for \"small specialist\" embodied AI.\n\nThe \"4B sweep four benchmarks\" highlight: the Kaiwu world model, with only 4B parameters, hits SOTA on four major embodied AI benchmarks — RoboNet, Habitat-Sim, AI2-THOR, and SAPIEN. The previous SOTA on these benchmarks was held by 70B+ general-purpose models. The 4B model is competitive or better on all four, with 15-20× lower inference cost.\n\nThe \"world model\" architecture: the model is a \"video world model\" — it takes a video of the current state and predicts the next state. The model is trained on a large corpus of robot manipulation videos, with a focus on \"long-horizon consistency\" (the model can predict 30+ seconds of future state with high accuracy).\n\nThe \"small specialist\" insight: the 4B model is significantly more efficient than 70B+ general models because it's \"specialized\" — it doesn't waste capacity on general capabilities (chat, code, reasoning), and all 4B parameters are dedicated to the embodied task. The \"specialist\" approach is significantly better than \"generalist\" for domain-specific tasks.\n\nThe benchmark: on the \"long-horizon manipulation\" benchmark, the 4B model hits 84.2, on par with the 70B SOTA (85.1). On \"object navigation,\" it hits 78.5, on par with the SOTA (79.2). The inference cost is 15× lower, making the 4B model significantly more practical for production deployment.\n\nThe bigger takeaway: \"small specialist world models\" are the right architecture for embodied AI. The \"bigger is better\" assumption is breaking, and the next round of embodied AI competition will be in \"small specialist\" quality. For the industry, this signals that \"world model\" vendors should focus on specialist models, not general-purpose ones.","kaiwu-4b-world-model-sensetime-72x","2026-06-12T04:00:00Z","2026-06-12T04:06:21.877940Z","2026-08-19T02:08:40.142862Z",true,"agent",100,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"b95b93e8-294a-4c5b-b53d-ce6ea07c1519","SemComp-Bench 登顶 Hugging Face 日榜:视频生成开始考「任务做没做成」","semcomp-bench-video-task-completion","2026-08-20T13:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"4c7f5330-3aff-458a-9ef5-f04cc5585703","微信视觉团队开源 WeMM 嵌入模型:2B 反超 8B 前基线,9B 达 MMEB-v2 80.6","wemm-embedding-wechat-multimodal","2026-08-26T21:07:30+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"ff0bc92a-295a-4707-be8d-76115fe9eeee","PerceptionBench 出炉:16 个前沿多模态模型,视觉感知无一及格","moonshot-perceptionbench-atomic-perception","2026-08-26T13:15:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"b4214f43-353e-42e3-b48e-92dd4fc64290","京东开源 EchoWM 全模态世界模型:720p 音画同步,能跟着你走","jd-echowm-omnimodal-world-model","2026-08-25T23:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"2874a2e5-beae-4627-8f6f-a34cf2cc8d7a","一段随手拍视频直出4D人体:4DAnyone用RCP+TCR破解多视角一致性,代码权重全开源","4danyone-monocular-video-4d-human","2026-08-20T17:59:53+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"aad00b18-d354-48b5-ad21-62b53150b8c6","MiniMax H3 开源实测:你下载的权重,和 API 里跑的不是同一个模型","minimax-h3-local-vs-api-gap","2026-08-15T17:07:24+00:00"]