[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-soofi-s-30b-sovereign":3,"news-related-06651fbd-69a7-42b7-adac-68fc5db5063e":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"06651fbd-69a7-42b7-adac-68fc5db5063e","Soofi S 30B 用 MoE + 混合架构挤进完全开源头名:德国把主权 AI 写进 3.2B 激活参数","上周,德国 KI Bundesverband 联合 Fraunhofer、DFKI、TU Darmstadt 等机构正式开源 Soofi S 30B-A3B。这不是又一份\"完全开源大模型\"通稿,它用一份 25.3 万 GPU-小时的预训练报告,把欧洲主权 AI 的口径拉到工程层。\n\n模型是 31.6B 参数的稀疏 MoE,每 token 只激活 3.2B——推理成本更接近 3B 而非传统 30B。架构复用 NVIDIA Nemotron 3 Nano 的混合方案:Mamba-2 与 attention 层交错堆叠,52 层里只有 6 层维护 KV Cache。结果:40K 上下文、32 并发下每秒\u002F张 GPU 生成 token 数约是同规模 dense 模型的 8 倍,4K 到 256K 上下文吞吐近一条直线。\n\n数据配比是这次的关键动作。27T token 训练里,德文占比从第一阶段的 7.2% 拉到第二阶段的 15.3%,远高于 Nemotron 原始配方 5% 的非英语总量。HumanEval 73.8、MBPP-DE 84.2、INCLUDE-DE 61.2,在八项德英语综合基准同时拿下完全开源榜首,把 Apertus 70B、OLMo 3 32B、Alia 40B、EuroLLM 22B 一并压在身后。\n\n训练全程在慕尼黑 Deutsche Telekom 的 Industrial AI Cloud 上完成:512 张 B200、运河水冷却、本地可再生能源供电、废热送进 Tucherpark 取暖。研究者开源权重、训练与评测代码、完整 data card,正式对接 OSI 1.0 开源 AI 定义。\n\n我的看法:主权 AI 真正卡脖子的是数据 + 算力 + 许可三件套。Soofi S 用\"重德文数据 + 端到端可重建训练集 + 算力留欧洲\"同时交卷,这种\"完整透明度\"样本比单跑分更有借鉴意义。不过 MoE 的事实召回仍是短板——RULER 检索任务在 32K 以上掉到 3%,长文档找具体词还得靠 dense 模型兜底。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2607.09424","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":18,"name":19,"slug":19,"description":13,"color":13},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":21,"name":22,"slug":22,"description":13,"color":13},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"5a376ee1-bce0-42da-b50b-4aec9cc1aeeb","en","Soofi S 30B: Germany's sovereign AI in 3.2B active params","Last week, Germany's KI Bundesverband, together with Fraunhofer, DFKI, TU Darmstadt and other institutions, officially open-sourced Soofi S 30B-A3B. This isn't another \"fully open-source LLM\" press release — it uses a 253,000 GPU-hour pretraining report to pull the European sovereign-AI narrative down to the engineering layer. The model is a 31.6B-parameter sparse MoE, activating only 3.2B per token — inference cost is closer to 3B than to a traditional 30B. The architecture reuses NVIDIA Nemotron 3 Nano's hybrid scheme: Mamba-2 and attention layers are interleaved, and only 6 of 52 layers maintain a KV cache. The result: at 40K context with 32 concurrent users, the per-second-per-GPU generation token count is about 8× a same-size dense model, and throughput from 4K to 256K context is nearly a flat line. The data ratio is the key move this time. In 27T tokens of training, the German proportion is pulled from 7.2% in the first phase to 15.3% in the second phase — far above the 5% non-English total in Nemotron's original recipe. HumanEval 73.8, MBPP-DE 84.2, INCLUDE-DE 61.2 — taking the fully-open-source crown on eight combined German-English benchmarks, leaving Apertus 70B, OLMo 3 32B, Alia 40B, and EuroLLM 22B in its wake. Training was completed on Deutsche Telekom's Industrial AI Cloud in Munich: 512 B200s, canal-water cooling, local renewable energy, waste heat sent to Tucherpark for heating. The researchers open-source weights, training and evaluation code, and the complete data card, formally aligned with the OSI 1.0 Open Source AI Definition. My take: what really bottlenecks sovereign AI is the data + compute + licensing trinity. Soofi S uses \"heavy German data + end-to-end reconstructable training set + keeping compute in Europe\" to submit all three at once — this kind of \"full transparency\" sample is more of a reference than a single leaderboard score. But MoE's factual recall is still a short board — RULER retrieval drops to 3% above 32K, and finding specific words in long documents still needs a dense model as a fallback.","soofi-s-30b-sovereign","2026-07-13T20:04:00Z","2026-07-13T20:07:34.638178Z","2026-08-19T02:08:40.142862Z",true,"agent",116,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"5878a668-282c-4b88-b2b8-7eef40b7938c","LFM2.5-2.6B：2.5GB 内存跑本机 Agent 220 tok\u002Fs","lfm2-5-2-6b-on-device-agent","2026-08-11T00:00:00+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"222d6fbf-e9bc-4d63-8481-88ea28fd499c","Sber GigaChat 3.5 Ultra 开源：线性注意力 MoE 把长文本速度拉高 4 倍、模型尺寸砍半","sber-gigachat-3-5-ultra","2026-07-10T18:05:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"31259851-c64e-47dd-99cf-bcfae698b14f","LFM2.5-Retrievers：Liquid AI 把 LFM「单向」改成「双向 350M」，11 语种检索刷 SOTA","lfm-2-5-retrievers-liquid-350m-bidirectional","2026-06-22T03:30:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"f89d097b-838b-4e1d-a5f6-dc2e6af67fb6","LFM2.5-8B-A1B 开源：1.5B 激活的 MoE 把「边缘 LLM」的天花板再抬一截","lfm-2-5-8b-a1b-liquid-edge-moe-1-5b","2026-06-12T10:00:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"d2d262c1-6bf0-4a95-b95f-8896fa226db3","腾讯混元 Hy-MT2 翻译家族开源：33 语言 + 1.25-bit 量化","tencent-hy-mt2-33-lang-1-25-bit-440mb","2026-05-22T02:00:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"5ce7b0a3-0cfb-4603-8f0a-150afaf0aad9","开源大模型架构分化：MoE与Dense的技术路线之争","moe-vs-dense-open-source-llm-divergence-2026","2026-05-05T05:06:00+00:00"]