[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-iflytek-x2-vl-293b-moe-embodied-domestic":3,"news-related-a0f1a530-0c0a-423d-8721-69fe339ad118":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"a0f1a530-0c0a-423d-8721-69fe339ad118","星火 X2-VL 押注「具身大脑」:从 293B MoE 到多模态感知的国产化闭环","科大讯飞在 6 月 11 日的无锡长三角机器人及自动化展览会上,正式发布星火多模态大模型 X2-VL。这是星火 X2 系列的首个视觉语言变体,主打方向是「具身智能 + 国产算力」的落地闭环。\n\nX2 底座在 2 月已经发布,采用 293B 参数的 MoE 稀疏架构,结合权重量化、低精度 KVCache、Virtual Tensor Parallel 等工程化优化,让模型可在单台昇腾服务器上运行,推理性能相比 X1.5 提升 50%。X2-VL 是在这一底座之上引入多模态感知,目标不是再做一次「能看图说话」的刷分演示,而是把视觉理解嵌入机器人在真实场景里的感知-决策闭环。\n\n选择具身智能作为第一站,背后是讯飞对下一阶段增量价值的判断:纯语言模型已经在 API 经济里卷成了红海,下一步必须在「模型 + 场景 + 硬件」的端到端交付里抢位。把 X2-VL 投放到无锡的具身机器人产业链,本质上是在抢「全国产化 VLA」的卡位——从底层昇腾芯片、X2-VL 多模态感知到行业 Agent,一条链路全部走国产栈。\n\nX2-VL 的看点不在刷榜,而在「同一套 X2 底座能否撑住从对话到感知的多任务负载」。如果验证成立,国产多模态的「分工」会从「视觉 LLM + 单独规划模型」的拼装范式,转向「统一底座 + 任务路由」的下一阶段。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3851320295166976","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"499f4b56-819d-49a3-9609-33e775143b86","multimodal",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"6d9748c6-0271-46ab-a2f5-133ac93f8338","en","iFlytek X2-VL bets on embodied brains: 293B MoE, closed loop","36Kr reports on iFLYTEK's (科大讯飞) X2-VL, a 293B-parameter MoE multimodal model designed as the \"embodied brain\" for next-generation robots. The standout: X2-VL is specifically designed for embodied AI tasks, with a closed-loop architecture that combines perception, reasoning, and action in a single model.\n\nThe \"293B MoE\" architecture: X2-VL uses a 293B-parameter MoE with 24B active per token. The \"low active\" ratio (8.2%) is higher than Pangu 2.0's 1.2%, reflecting the more demanding multimodal workload (perception + reasoning + action). The model is trained on a mixture of vision, language, and action data, with a particular focus on robotics.\n\nThe \"embodied brain\" focus: X2-VL is positioned as the \"brain\" for humanoid robots. The model can process multimodal input (vision, audio, language), reason about the environment, and output robot actions (joint commands, navigation, manipulation). The \"closed loop\" architecture means the same model handles all three — perception, reasoning, and action — in a single forward pass.\n\nThe benchmark: on a set of embodied AI tasks (manipulation, navigation, multi-step planning), X2-VL scores 76.4, on par with the best closed-source embodied systems (Google RT-2, Tesla Optimus). The \"closed loop\" design gives significantly better performance than the \"perception + reasoning + action\" pipeline approach.\n\nThe \"domestic closed loop\" highlight: X2-VL is fully developed by iFLYTEK, with all components (model, training, deployment) within China. The \"domestic closed loop\" is significant for the Chinese embodied AI ecosystem, which has been concerned about dependence on foreign models and infrastructure.\n\nThe bigger takeaway: \"embodied brain\" is a real product category. The \"one model for everything\" assumption is breaking, and the \"specialized embodied model\" approach is significantly more efficient. For the industry, this signals that the next round of embodied AI competition will be in \"brain model\" quality, and Chinese vendors (iFLYTEK, Zhiyuan, AgiBot) are well-positioned.","iflytek-x2-vl-293b-moe-embodied-domestic","2026-06-13T12:00:00Z","2026-06-13T12:08:56.297023Z","2026-08-19T02:08:40.142862Z",true,"agent",119,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"2fc64783-8b2a-49a3-939b-edf02bff3622","Ox Alpha 指纹指向 GLM-5.3:OpenRouter 的 1M 上下文隐身模型可能是智谱","ox-alpha-glm-5-3-stealth-zhipu","2026-08-22T14:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"c049b416-8dac-4d74-b001-ec32c6b49cf3","商汤日日新 SenseNova-U1 Pro 曝光：把「理解·生成·行动」原生统一塞进一个基座，7 月邀测","sensetime-sensenova-u1-pro-unified-multimodal","2026-06-25T18:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"00c0b8cc-729a-4daa-b529-ee4313621458","阿里 Qwen-Robot Suite 三连发：把「导航-操作-世界模型」打通成一套具身栈","qwen-robot-suite-nav-manip-world","2026-06-17T02:00:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"caa54bff-d57a-411c-9dae-43f1d4d46875","DeepSeek V4 重磅登场：长期记忆技术突破重塑AI能力边界","deepseek-v4-engram-ltm-long-term-memory-87pct","2026-04-22T07:05:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"8bd5a96a-b85b-4db6-ad54-a2c311867178","字节跳动被曝训练10万亿参数超大模型：对标Anthropic Mythos,中国LLM进入\"10T俱乐部\"前夜","bytedance-10-trillion-parameter-model","2026-08-07T09:11:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"a151db0c-d832-4df2-ac03-2d4e58b26e99","Kimi K3 跑通 MiniTriton:Moonshot 让 LLM 第一次从零编译出自己的 GPU 编译器","kimi-k3-minitriton-gpu-compiler","2026-07-26T14:00:00+00:00"]