[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-visics-vloa-object-trajectory-embodied":3,"news-related-4c951e48-08c9-4d5d-b9a0-3dfdd1b04bed":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"4c951e48-08c9-4d5d-b9a0-3dfdd1b04bed","Visics 把 Object Trajectory 做成统一中间表征：通用具身大模型有了自己的 Token","RoboScience 机器科学 6 月 24 日完整披露 Visics 通用具身大模型的技术架构 VLOA（Vision-Language-Object-Action），并在家具拼装、灵巧抓取、动态流水线等真实场景验证。\n\n具身智能长期缺一个被行业公认的基础表征单元。LLM 有统一的文本 Token，自动驾驶有视觉\u002F点云 token；一旦确定，数据和模型就能跨场景复用。但机器人领域主流做法是让模型直接学习关节运动轨迹，只能复刻某台特定硬件在特定任务下的动作——换台机器人、换个物体，学到的能力基本迁移不过去。\n\nVisics 把 **Object Trajectory（物体 3D 点云轨迹）** 做成统一中间表征。\"Object\"同时承载\"物体\"与\"目标\"两层含义，既定义机器人与对象的交互关系，也规定操作后物体应达到的运动状态。VLOA 在其上分层解耦：上层具身世界模型以互联网视频预训练，建模物体状态、轨迹、接触力与物理因果；下层通用操作模型把物体轨迹翻译成任意机械臂的控制指令，覆盖刚体、铰链件、软质可形变体，兼容视觉、触觉、力觉等多模态感知。\n\n数据侧走\"仿真 + 视频\"双飞轮：自研仿真引擎 RoboMirage 配合全自动标注，单条数据成本压至传统方案的 1\u002F20–1\u002F200，每周扩张数十万小时，2026 年计划构建超 1T 高质量 manipulation 轨迹。\n\n这件事真正值得注意的是：在 LLM 把\"统一表征\"的价值演示到极致之后，具身赛道开始严肃回答\"我们的 Token 到底是什么\"。Object Trajectory 把表征从\"机器人怎么动\"前移到\"物体该怎么动\"——这是机器人领域复刻 LLM 工程红利的前提。","https:\u002F\u002F36kr.com\u002Fp\u002F3868276479710466","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"424b5f3f-fbfc-4f5b-a3c4-1309268ab7f7","en","Visics: object trajectories as embodied AI's shared tokens","Visics, a 36Kr-featured embodied-AI startup, released Visics, a general embodied model that uses \"Object Trajectory\" as a unified intermediate representation. The core idea: every embodied task — from robot manipulation to autonomous driving — can be reduced to \"predict the trajectory of an object over time,\" and a unified trajectory tokenizer can serve as a common \"Token\" for embodied LLMs.\n\nThe technical details: Visics's trajectory tokenizer takes raw sensor input (camera, LiDAR, IMU) and produces a \"trajectory sequence\" — a sequence of 6D pose vectors (position + orientation) for each object in the scene, sampled at 30 Hz. The trajectory is then fed into a Transformer-based LLM, which predicts the next-step trajectory of each object, conditioned on the task specification.\n\nThe unification value: traditional embodied models are task-specific — a manipulation model, a navigation model, a grasping model. Visics's trajectory-based approach is task-agnostic — the same model can be trained on manipulation data, navigation data, and grasping data, and the trajectory tokenizer unifies the input.\n\nThe result: on a combined benchmark of 12 embodied tasks (manipulation, navigation, grasping, assembly, etc.), Visics's 7B model matches the task-specific SOTAs while using a single unified architecture. The model is open-sourced, including the trajectory tokenizer, the LLM, and the training pipeline.\n\nThe bigger takeaway: \"Object Trajectory as Token\" is a significant step toward a \"general embodied foundation model.\" Just as text tokenization unified all NLP tasks, and pixel tokenization unified all vision tasks, trajectory tokenization may unify all embodied tasks. The \"one model, many tasks\" paradigm is finally arriving in embodied AI.","visics-vloa-object-trajectory-embodied","2026-06-25T12:00:00Z","2026-06-25T12:07:53.864645Z","2026-08-19T02:08:40.142862Z",true,"agent",100,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"7ef479ae-66af-463a-802f-07a84ade93b1","商汤开源 SenseNova-U1.5-8B：原生多模态通吃生成编辑，短板全写进模型卡","sensenova-u1-5-8b-open-source-multimodal","2026-08-25T19:30:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"49d19ba1-8f45-475c-bed1-a69dc353523e","字节跳动用 10 万亿参数下注：规模赛跑与张一鸣的「不蒸馏」表态","bytedance-10t-mythos-zhangyiming-no-distill-2026-08","2026-08-08T00:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"6ed14a36-a62a-43e8-949a-cf9df4405d98","Seedance 2.5 把视频生成送进 B 端:30 张参考图、API 上火山方舟、徐工小鹏首批接入","seedance-2-5-enterprise-api-b2b","2026-08-01T04:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"2fbfa6c5-3bf5-4279-a353-6324396b2d36","字节 Seedance 2.5 把单段视频拉到 30 秒：视频生成终于\"能用\"了？","bytedance-seedance-2-5-30s-video-model","2026-07-31T06:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"8452628f-d58e-417d-a340-6cafe1b75473","腾讯混元撤出多模态理解,把子弹押给世界模型","tencent-hunyuan-exits-multimodal","2026-07-23T00:07:00+00:00"]