[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-embodied-cpp-cpp-runtime":3,"news-related-ea745180-90a1-4e71-a5ed-018b243359a7":37},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":30,"published_at":31,"created_at":32,"modified_at":33,"is_published":34,"publish_type":35,"image_url":14,"view_count":36},"ea745180-90a1-4e71-a5ed-018b243359a7","Embodied.cpp：C++ 统一 VLA 部署，显存砍到三分之一","【摘要】东南大学 SAIL Lab 推出 Embodied.cpp——一个面向异构机器人的 C++ 推理 runtime,采用五层模块化设计(input adapters \u002F sequence builders \u002F backbone execution \u002F head plugins \u002F deployment adapters),统一多速率闭环控制、batch-1 延迟优先推理和可扩展算子。在 HY-VLA、pi0.5 与 LingBot-VA Transformer block 上的实测表现硬核:VLA 闭环任务成功率 100% 与 91%,WAM Transformer block 显存从 312.2 MiB 砍到 88.1 MiB(不到三分之一)。GitHub 开源,可直接上车机器人与模拟器。\n\n【正文】过去两年,具身智能(VLA、WAM)在模型层不断刷出 SOTA,但「模型写得好,跑不起来」成了行业共同的尴尬——每个团队一套 Python 推理栈、对每种硬件写一份胶水代码、每个机器人一个独立 backend。东南大学 SAIL Lab 这次把视角从训练端转向部署端,推出 Embodied.cpp:一个面向异构机器人的 C++ 推理 runtime,直接登顶 Hugging Face 7 月 6 日 #2 trending。\n\n论文(arXiv 2607.02501)的核心思路是从 VLA\u002FWAM 模型架构里抽出\"共享执行路径\",分成五层:input adapters → sequence builders → backbone execution → head plugins → deployment adapters。这套抽象带来三件硬通货:统一支持多速率闭环控制、batch-1 延迟优先推理、可扩展算子与 I\u002FO,让同一份 runtime 跑在机器人、模拟器和不同加速器上,无需为每个模型家族重写一套胶水。\n\n实测数据也不含糊——HY-VLA 闭环任务成功率 100%、pi0.5 91%;WAM 基准把 Transformer block 的内存从 312.2 MiB 砍到 88.1 MiB,不到三分之一。代码开源在 github.com\u002FSEU-PAISys\u002FEmbodied.cpp,配套 Hugging Face 仓库同步发布,接口设计把模型侧的\"上肢动作\"和后端的\"硬件差异\"彻底解耦。\n\n在「模型月月新」的具身圈子里,真正把部署工程做扎实的项目反而稀缺——当各家还在比 demo 成功率,能稳定\"上车\"的运行时基础设施才是规模化落地的入场券。Embodied.cpp 提醒我们:具身智能的下半场,胜负不在参数大小,而在边缘设备能不能跑得稳、跑得快、跑得便宜。","【摘要】东南大学 SAIL Lab 推出 Embodied.cpp——一个面向异构机器人的 C++ 推理 runtime,采用五层模块化设计(input adapters \u002F sequence builders \u002F backbone execution \u002F head plugins \u002F deployment adapters),统一多速率闭环控制、batch-1 延迟优先推理和可扩展算子。在 HY-VLA、pi0.5 与 LingBot-VA Transformer block 上的实测表现硬核:VLA 闭环任务成功率 100% 与 91%,WAM Transformer block 显存从 312.2 MiB 砍到 88.1 MiB(不到三分之一)。GitHub 开源,可直接上车机器人与模拟器。\n\n【正文】过去两年,具身智能(VLA、WAM)在模型层不断刷出 SOTA,但「模型写得好,跑不起来」成了行业共同的尴尬——每个团队一套 Python 推理栈、对每种硬件写一份胶水代码、每个机器人一个独立 backend。东南大学 SAIL Lab 这次把视角从训练端转向部署端,推出 Embodied.cpp:一个面向异构机器人的 C++ 推理 runtime,直接登顶 Hugging Face 7 月 6 日 \n论文(arXiv 2607.02501)的核心思路是从 VLA\u002FWAM 模型架构里抽出\"共享执行路径\",分成五层:input adapters → sequence builders → backbone execution → head plugins → deployment adapters。这套抽象带来三件硬通货:统一支持多速率闭环控制、batch-1 延迟优先推理、可扩展算子与 I\u002FO,让同一份 runtime 跑在机器人、模拟器和不同加速器上,无需为每个模型家族重写一套胶水。\n\n实测数据也不含糊——HY-VLA 闭环任务成功率 100%、pi0.5 91%;WAM 基准把 Transformer block 的内存从 312.2 MiB 砍到 88.1 MiB,不到三分之一。代码开源在 github.com\u002FSEU-PAISys\u002FEmbodied.cpp,配套 Hugging Face 仓库同步发布,接口设计把模型侧的\"上肢动作\"和后端的\"硬件差异\"彻底解耦。\n\n在「模型月月新」的具身圈子里,真正把部署工程做扎实的项目反而稀缺——当各家还在比 demo 成功率,能稳定\"上车\"的运行时基础设施才是规模化落地的入场券。Embodied.cpp 提醒我们:具身智能的下半场,胜负不在参数大小,而在边缘设备能不能跑得稳、跑得快、跑得便宜。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.02501","7437aeb9-930c-4866-a2e9-48003c1a792b",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":19,"name":20,"slug":20,"description":14,"color":14},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":14},"da0955d1-ef78-48eb-a244-d00c485529b7","en","Embodied.cpp: one C++ runtime for VLA, a third of the memory","[Summary] Southeast University's SAIL Lab released Embodied.cpp — a C++ inference runtime for heterogeneous robots, using a five-layer modular design (input adapters \u002F sequence builders \u002F backbone execution \u002F head plugins \u002F deployment adapters), unifying multi-rate closed-loop control, batch-1 latency-priority inference, and extensible operators. Hardcore measurements on the HY-VLA, pi0.5, and LingBot-VA Transformer block: VLA closed-loop task success rates of 100% and 91%, WAM Transformer block memory cut from 312.2 MiB to 88.1 MiB (less than one-third). GitHub open-source, can be directly mounted on robots and simulators. [Body] Over the past two years, embodied intelligence (VLA, WAM) has continuously posted SOTA at the model layer, but \"the model is well-written but cannot run\" has become a shared embarrassment in the industry — each team has its own Python inference stack, writes a piece of glue code for every kind of hardware, and each robot has an independent backend. Southeast University's SAIL Lab this time shifts the perspective from the training end to the deployment end, releasing Embodied.cpp: a C++ inference runtime for heterogeneous robots, directly topping Hugging Face's #2 trending on July 6. The paper's (arXiv 2607.02501) core idea is to extract the \"shared execution path\" from the VLA\u002FWAM model architecture, divided into five layers: input adapters → sequence builders → backbone execution → head plugins → deployment adapters. This abstraction brings three hard currencies: unified support for multi-rate closed-loop control, batch-1 latency-priority inference, and extensible operators and I\u002FO, letting the same runtime run on robots, simulators, and different accelerators without rewriting a set of glue for each model family. The measurement data is also unambiguous — HY-VLA closed-loop task success rate 100%, pi0.5 91%; the WAM benchmark cuts the Transformer block's memory from 312.2 MiB to 88.1 MiB, less than one-third. Code is open-sourced at github.com\u002FSEU-PAISys\u002FEmbodied.cpp, with a companion Hugging Face repository published simultaneously, and the interface design completely decouples the \"upper-limb actions\" on the model side from the \"hardware differences\" on the backend. In the embodied circle where \"models come out monthly\", projects that truly solidify deployment engineering are instead rare — when each vendor is still comparing demo success rates, runtime infrastructure that can stably \"board the vehicle\" is the entry ticket for scaled deployment. Embodied.cpp reminds us: the second half of embodied intelligence is decided not by parameter size, but by whether edge devices can run stably, fast, and cheap.","embodied-cpp-cpp-runtime","2026-07-06T18:05:00Z","2026-07-06T18:08:23.111756Z","2026-08-19T02:08:40.142862Z",true,"agent",182,{"items":38},[39,44,49,54,59,64],{"id":40,"title":41,"news_slug":42,"published_at":43},"e3c0b314-d7b7-4901-b2b0-08ca5ef08ac7","GigaBrain-0.7开源:37k小时数据+三系统架构,世界模型进VLA决策回路","gigabrain-0-7-embodied-vla-open-source","2026-08-26T23:15:00+00:00",{"id":45,"title":46,"news_slug":47,"published_at":48},"ce70384a-990b-4994-bfb6-27775be45661","TensorRT Edge-LLM 0.10.0：边端第一个统一的 C++ 多模态推理栈","tensorrt-edge-llm-0-10-multimodal-runtime","2026-08-23T00:00:00+00:00",{"id":50,"title":51,"news_slug":52,"published_at":53},"a91067a3-4fa4-4e88-a25a-18ba3bea21ea","Google 把\"加密推理\"摆上桌面：HEIR 编译器让预训练模型在密文上直接跑","google-heir-compiler-encrypted-ai-inference","2026-08-14T14:00:00+00:00",{"id":55,"title":56,"news_slug":57,"published_at":58},"cdc8e3ce-b1aa-4348-9436-04763179af9c","AMD MI455X：Transformers 99.5% 通过率，432GB HBM4","amd-mi455x-huggingface-99-5","2026-07-27T10:30:00+00:00",{"id":60,"title":61,"news_slug":62,"published_at":63},"dd011592-f0aa-4d45-9229-56311232f9f0","OpenMOSS 开源 MOSS-VL-Realtime：11B 实时流视频 VLM","openmoss-vl-realtime","2026-07-19T03:55:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"b3a59736-4791-4834-8dc2-a851833a63ae","LightOn-rerank：2B 模型同时排文本和文档页，listwise 把 pointwise 打成过去式","lighton-rerank-listwise","2026-07-16T22:30:00+00:00"]