[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-lingbot-va-2-embodied-video":3,"topics-all":36,"news-related-c7e500ff-b7bf-4f18-a9d5-0217c2925d6b":55},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"c7e500ff-b7bf-4f18-a9d5-0217c2925d6b","LingBot-VA 2.0:首个\"具身原生\"视频-动作世界模型,Robbyant 拒绝\"借壳\"路线","7月10日,蚂蚁集团旗下具身智能公司 Robbyant 推出 LingBot-VA 2.0,定位为行业首个\"具身原生\"视频-动作世界模型。它**不从视频生成模型微调而来,而从零自回归预训练**,目标单一:让机器人准确预测动作将如何改变环境,并据此决定下一步。\n\n主流路线普遍\"先用视频生成模型做世界模型,再微调给机器人\"——但内容创作追求视觉质量,机器人控制需要物理精度,这种\"借壳\"经常导致灾难性遗忘和泛化下降。VA 2.0 改走四件套:**Semantic Visual-Action Tokenizer** 在视觉压缩阶段对齐语义与动作信息;**Strict Causal Pre-training** 保证单向时序;**MoE** 扩容不损速度;**Enhanced Asynchronous Inference** 让机器人边执行边预测,形成闭环。落地数据直接命中痛点:**单 GPU 150 Hz 实时推理,20 段演示即可零参数更新的 in-context learning 泛化到新任务**。\n\nVA 2.0 是 Robbyant 本周\"6 模型连发\"的收官之作。此前发布的 LingBot-Depth 2.0、LingBot-Vision、LingBot-VLA 2.0、LingBot-World 2.0、LingBot-Video 覆盖感知、仿真、动作三层级,VA 2.0 把\"动作 + 仿真\"压成统一模型,完成具身原生全栈拼图。\n\n评论:这条路线最值得关注的不是某个 benchmark,而是**范式选择**——把\"借用数字内容模型\"换成\"为物理世界从头造一个\"。在 Genie 3、Veo 主宰数字世界的当下,Robbyant 给具身赛道提供了一类不同样本:不追最炫的视频生成,把物理一致性与实时性放第一位。短期不如 Sora-2 类模型在公众视野里显眼,但对工业、养老、医疗辅助这类\"必须跑得稳\"的场景,这才是真正的入场券。","https:\u002F\u002Fsecure.businesswire.com\u002Fnews\u002Fhome\u002F20260709654440\u002Fen\u002FRobbyant-Launches-LingBot-VA-2.0-Built-Natively-for-Embodied-AI-and-Physical-World-Control","1453aaad-7e99-4f0d-9db7-aeda513f128b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"471c51be-e620-49df-bd6c-0b5504f53f00","ant-group",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"4fceab98-3018-4765-b867-a0d0bcf88459","en","LingBot-VA 2.0: an embodied-native video-action world model","On July 10, Ant Group's embodied-intelligence subsidiary Robbyant released LingBot-VA 2.0, positioned as the industry's first \"embodied-native\" video-action world model. It is **not fine-tuned from a video generation model, but autoregressively pretrained from scratch**, with a single goal: let the robot accurately predict how actions will change the environment, and decide the next step accordingly. The mainstream route is generally \"first use a video generation model as a world model, then fine-tune it for robots\" — but content creation pursues visual quality, and robot control needs physical precision; this \"borrowed shell\" often causes catastrophic forgetting and generalization degradation. VA 2.0 takes a different path with a four-piece set: **Semantic Visual-Action Tokenizer** aligns semantic and action information at the visual-compression stage; **Strict Causal Pre-training** ensures unidirectional time sequence; **MoE** scales capacity without losing speed; **Enhanced Asynchronous Inference** lets the robot predict while executing, forming a closed loop. The landing data directly hits the pain point: **150 Hz real-time inference on a single GPU, just 20 demonstrations enable in-context learning with zero parameter updates generalizing to new tasks**. VA 2.0 is the closing piece of Robbyant's \"6 models in a week\" release. The previously released LingBot-Depth 2.0, LingBot-Vision, LingBot-VLA 2.0, LingBot-World 2.0, and LingBot-Video cover the perception, simulation, and action layers; VA 2.0 compresses \"action + simulation\" into a unified model, completing the embodied-native full-stack puzzle. Commentary: what's most worth noting about this path isn't a single benchmark, but the **paradigm choice** — replacing \"borrowing digital-content models\" with \"building from scratch for the physical world\". At a time when Genie 3 and Veo dominate the digital world, Robbyant provides a different kind of sample for the embodied track: not chasing the flashiest video generation, but putting physical consistency and real-time performance first. In the short term it's less eye-catching in public than Sora-2-class models, but for industrial, elderly-care, and medical-assistance scenarios that \"must run stably\", this is the real entry ticket.","lingbot-va-2-embodied-video","2026-07-10T12:00:00Z","2026-07-11T12:08:13.215276Z","2026-08-19T02:08:40.142862Z",true,"agent",130,[37,46],{"slug":38,"tag_slug":38,"title_zh":39,"title_en":40,"intro_zh":41,"intro_en":42,"id":43,"is_active":33,"created_at":44,"modified_at":45},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":47,"tag_slug":47,"title_zh":48,"title_en":49,"intro_zh":50,"intro_en":51,"id":52,"is_active":33,"created_at":53,"modified_at":54},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":56},[57,62,67,72,77,82],{"id":58,"title":59,"news_slug":60,"published_at":61},"76c05fca-5560-41d3-8cdb-311556f7e845","蚂蚁灵波 LingBot-Depth 2.0：把机器人深度估计从「看懂」推向「看准」","ant-lingbot-depth-2-0","2026-07-07T04:30:00+00:00",{"id":63,"title":64,"news_slug":65,"published_at":66},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":68,"title":69,"news_slug":70,"published_at":71},"aebb8a81-713a-40b7-84dd-03213a6a808c","Mistral Robostral Navigate:8B 视觉语言模型只靠单目 RGB 在 R2R-CE 反超多传感器基线","mistral-robostral-navigate-8b","2026-07-09T14:15:00+00:00",{"id":73,"title":74,"news_slug":75,"published_at":76},"5de8c559-2ab5-44a7-b1d6-97cc8e499b25","蚂蚁灵波开源 LingBot-VLA 2.0：6 万小时数据 + 17 个品牌,把具身基座卷向跨构型","ant-lingbot-vla-2-0","2026-07-08T06:30:00+00:00",{"id":78,"title":79,"news_slug":80,"published_at":81},"cec093b3-47fe-490f-b0d7-57c07ba19758","Wan-Streamer v0.1：单模型端到端 550ms 实时交互","wan-streamer-v0-1-550ms-realtime","2026-06-23T18:01:03+00:00",{"id":83,"title":84,"news_slug":85,"published_at":86},"bbf1a404-1f46-45f9-a61e-b6e210d28878","智元罗剑岚：把「部署-数据-迭代」打成飞轮，比堆参数更像具身智能的 Scaling Law","zhiyuan-luo-jianlan-flywheel-embodied","2026-06-17T06:30:00+00:00"]