[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-xing4-29b-a4b-ascend-moe":3,"topics-all":38,"news-related-2bd44b6f-5688-471f-930e-17a93984e8e7":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"2bd44b6f-5688-471f-930e-17a93984e8e7","中国电信开源星辰 Xing4.0:昇腾全栈训练的 29B MoE","中国电信 AI 公司 9 月 17 日开源 Xing4.0-29B-A4B:29B 总参数、每 token 激活 4B 的 MoE,原生 256K 上下文,官方称是该规模首个全程用昇腾 NPU 与 MindSpore 训练的模型,训练吞吐提升约 96%,SWE-bench Verified 官方自报 75.0。","中国电信的 AI 子公司把 TeleChat 系列的下一代模型直接开源了:9 月 17 日,Xing4.0-29B-A4B 权重同步上架 Hugging Face、ModelScope 与 Modelers 三个平台,附带 FP8 和 GGUF 量化版本,Apache-2.0 许可证,组织名为 XingChen-AGI。\n\n## 29B 总参、4B 激活的高稀疏 MoE\n\n架构上这是一颗典型的高稀疏 MoE:29B 总参数、每 token 仅激活 4B;64 个路由专家每 token 取 4 个,外加 1 个共享专家;注意力用 MLA,40 层,hidden size 3584。上下文原生 256K,可扩展到 512K。官方将其定位为 agent 导向,架构组合写作 mHC + MLA + MTP,其中 MTP 在官方给出的 vLLM 启动命令里被用作投机解码方法。\n\n## 真正的看点:昇腾全栈训练\n\n模型卡里最重的一句声明是:Xing4.0-29B-A4B 是该规模上首个完全在昇腾 NPU 平台、用 MindSpore 框架训练的模型(官方口径,单一发布方声明)。训练在昇腾 910C 集群上进行,通过细粒度 MoE 通信优化、选择性重计算、DVM 自动图算融合与昇腾 C 融合算子的多层协同优化,整体训练吞吐比开箱配置提升约 96%。这组数字意味着国产算力栈进入了深度调优阶段——README 里还专门致谢 DeepSeek 团队的架构设计带来的稳定性与效率。\n\n## 跑分:agent 场景是主场\n\n官方提交的评测里,Xing4.0-29B-A4B 在 SWE-bench Verified 拿到 75.0,对比 Gemma4-26B-A4B 的 53.0 与 Qwen3.6-35B-A3B 的 76.0;Terminal-Bench 2.1 拿 57.5,明显超过两个对手的 30.0 和 51.5;SWE-bench Multilingual 66.0,Claw-Eval 76.55,DeepresearchBII 60.80。数学与工具调用上互有胜负:AIME2026 为 90.0,低于 Qwen3.6 的 92.7;Tau3-Bench 64.63,同样略低于 Qwen3.6 的 67.2。以上均为官方自报口径,评测配置(温度、上下文窗口、运行次数)在模型卡脚注里逐项列明。\n\n## 生态位:框架 PR 还没合入\n\n部署侧覆盖 vLLM、SGLang、KTransformers,微调支持 LLaMA-Factory 与 MindFormers,并对 OpenCode、Claude Code、OpenClaw、Hermes 等 agent 框架做了格式对齐;但要注意,SGLang、vLLM、TensorRT-LLM、llama.cpp、KTransformers 五个框架的支持 PR 截至开源时全部处于待审状态,短期部署需要走 PR 分支。社区反应尚在早期:上架两天,HF 显示近 30 天下载 7,278 次,GitHub 65 star,已出现 6 个微调版本与 7 个量化版本。\n\n对运营商背景的团队来说,这是把「自研大模型」从新闻稿推进到可复现工件的一步;对行业来说,昇腾栈上跑出这个规模的 MoE,比跑分本身更值得盯。模型卡与权重见 [Hugging Face](https:\u002F\u002Fhuggingface.co\u002FXingChen-AGI\u002FXing4.0-29B-A4B)。","https:\u002F\u002Fhuggingface.co\u002FXingChen-AGI\u002FXing4.0-29B-A4B","e96fe25f-0c49-47cd-bfa2-6dfe71aa784f",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"eac2ba5f-7c1c-424c-8bf9-8ec96a1a5362","en","China Telecom Open-Sources Xing4.0: 29B MoE Trained on Ascend NPUs","China Telecom open-sources Xing4.0-29B-A4B (Apache-2.0): 29B-total\u002F4B-active MoE, 256K context, SWE-bench Verified 75.0, trained entirely on Ascend NPUs.","China Telecom's AI subsidiary has open-sourced the next generation of its TeleChat lineage: on September 17, Xing4.0-29B-A4B weights landed simultaneously on Hugging Face, ModelScope and Modelers, joined by FP8 and GGUF variants, all under an Apache-2.0 license and published under the XingChen-AGI organization.\n\n## A 29B-total, 4B-active sparse MoE\n\nArchitecturally this is a highly sparse mixture-of-experts model: 29B total parameters with only 4B activated per token; 64 routed experts with 4 active per token plus 1 shared expert; MLA attention, 40 layers, hidden size 3584. Context is natively 256K, extensible to 512K. The vendor positions it as agent-oriented, describing the architecture as mHC + MLA + MTP, and MTP appears as the speculative-decoding method in the official vLLM launch command.\n\n## The real story: trained entirely on Ascend\n\nThe heaviest claim in the model card is that Xing4.0-29B-A4B is the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework (vendor's own wording, a single-publisher claim). Training ran on Ascend 910C clusters; through fine-grained MoE communication optimization, selective recomputation, DVM automatic graph-operator fusion and Ascend C fused operators, overall training throughput improved by roughly 96% over out-of-the-box settings. The README explicitly thanks the DeepSeek team for architecture design inspiration.\n\n## Benchmarks: agent workloads are home turf\n\nIn vendor-submitted evaluations, Xing4.0-29B-A4B scores 75.0 on SWE-bench Verified, versus 53.0 for Gemma4-26B-A4B and 76.0 for Qwen3.6-35B-A3B; 57.5 on Terminal-Bench 2.1, clearly ahead of 30.0 and 51.5; SWE-bench Multilingual at 66.0, Claw-Eval at 76.55 and DeepresearchBII at 60.80. It trades blows on math and tool use: AIME2026 is 90.0 against Qwen3.6's 92.7, and Tau3-Bench 64.63 against 67.2. All figures are vendor-reported, with evaluation configs (temperature, context window, run counts) spelled out in the model card footnotes.\n\n## Ecosystem: framework PRs still in flight\n\nDeployment covers vLLM, SGLang and KTransformers; fine-tuning supports LLaMA-Factory and MindFormers; format alignment is done for agent frameworks including OpenCode, Claude Code, OpenClaw and Hermes. Note, however, that support PRs for SGLang, vLLM, TensorRT-LLM, llama.cpp and KTransformers were all still pending review at release time, so short-term deployment means riding PR branches. Community traction is early: two days after listing, HF shows 7,278 downloads over the last month, the GitHub repo sits at 65 stars, and 6 finetunes plus 7 quantizations are already posted.\n\nFor a carrier-backed team, this moves the in-house LLM from press release to reproducible artifact; for the industry, a MoE at this scale trained on the Ascend stack matters more than the benchmark table. Model card and weights: [Hugging Face](https:\u002F\u002Fhuggingface.co\u002FXingChen-AGI\u002FXing4.0-29B-A4B).","xing4-29b-a4b-ascend-moe","2026-09-19T15:10:00Z","2026-09-19T15:11:08.876722Z","2026-09-19T15:11:08.876737Z",true,"agent",76,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"804ab59a-a8d6-4b61-bf74-8f6f2bdae83c","智谱把 Flash 做成一件正经事:一次说清 GLM-5.3-Flash 的架构和 benchmark 真相","glm-5-3-flash-hybrid-attention-architecture","2026-08-27T08:00:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"8ebbcd9c-31ee-4baa-b395-b104bd87c8e1","Kimi K2.8 Preview 把 K3 的百万上下文下放给免费档：月之暗面的「过日子」模型登场","kimi-k2-8-preview-coding","2026-09-17T03:00:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"21fe3c11-4ba4-4801-b6fc-60c4ae559dc1","Yandex 逆流开源:35B 参数的 T5 MoE,每个 token 只激活 0.6B","yandex-aliceai-t5-sparse-moe","2026-09-16T19:11:43+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"8e730a3d-439b-45cf-961d-f77cf01469fd","Cohere 开源 218B 翻译专用 MoE:25B 激活,自测评分超 DeepL,2×H100 可部署","cohere-north-small-translate","2026-09-11T19:07:20+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"2b37a19b-1dde-4238-bef5-39b1d19157f1","OpenBMB 开源 MiniCPM5-2B:2B 端侧模型平均分超对比集 4B 级","openbmb-minicpm5-2b-on-device","2026-09-07T17:02:00+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"d941056b-c2e7-42e5-965a-a982c20b1169","Qwen3.8-Flash-Next 架构细节:Gated Residual 多分支残差 + QSA micro-block 稀疏注意力","qwen3-8-flash-next-cost-efficiency-architecture","2026-09-02T02:00:00+00:00"]