[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-aleph-alpha-kolibri-1-open-moe":3,"topics-all":38,"news-related-0372db20-feaf-42b8-98bd-e42d9c550306":57},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"0372db20-feaf-42b8-98bd-e42d9c550306","德国Kolibri开源:78B参数只激活3.46B","德国 Aleph Alpha 统一日开源 Kolibri-1:78B 总参、每 token 仅激活 3.46B 的德英双语 MoE,Apache 2.0 许可,上下文最高 1M。官方跑分数学项对标 4 倍激活量模型,德语语料 80% 自建,分词器德语压缩率压过 GPT-5。","10 月 3 日德国统一日当天,德国 AI 公司 Aleph Alpha 开源了推理模型 Kolibri-1:总参数 78B、每 token 仅激活 3.46B 的德英双语 MoE,Apache 2.0 许可,上下文最高 1M token。在 MoE 卷总参数的当下,「78B 总量、3.46B 激活」的组合外加一份少见的德语数据工程账本,值得拆开看。\n\n## 为什么是 78B,不是 123B\n\n内部前身 Kolibri Origin 是 30.6B 总参 \u002F 3.27B 激活的验证模型,6 月完成预训练;9 月 11 日,Kolibri 完成 20T token 预训练。官方博客披露,团队试过放大到 123B:两张 H100 上,123B 只能同时服务 3 个 256k token 长上下文请求,78B 能扛 18 个,解码还快 28%。最终架构是 50 层 MoE,每层 384 个专家(1 共享 + 6 路由),40 层用 512 token 滑窗注意力、每 5 层保留一层全局注意力;FP8 权重,最低 2×H100 起步。支持四档推理力度与工具调用;1M 上下文已验证,生产环境建议控制在 262,144 token 内。\n\n## 跑分有硬点,短板也摆在明面上\n\n官方自报评测里,数学最亮:AIME 2025 拿 96.9,高于 12B 激活的 Nemotron 3 Super(91.7)和 Qwen3.6-35B-A3B(84.6);GPQA diamond 84.3;银行 τ³-bench 38.1,是同表 Qwen3.6(10.6)的三倍多。官方称可对标 4 倍激活量模型。但同一张表也写着另一面:英语总分 75.5,不敌 27B 稠密的 Qwen3.8(80.2);TerminalBench 2.1 只有 27.7,对比 Qwen3.8 的 76.8;SWE-Bench Verified 66.4,低于 Qwen3.6-35B 的 73.8。强项在 agentic RAG:Honeypot 80.8 领先全部对比模型。幻觉控制是差异化卖点:AA-Omniscience 上「不作答而非答错」的比例 44%(前代 15%),Merlin-Arthur 协议训练出的 M\u002FA grounding 分 0.23,同表多数为 0。\n\n## 德语才是护城河\n\n最值得记录的是德语数据账本:20T 预训练 token 中 21.3%(约 4.3T)是德语,背后是 2.4T 的去重德语池,80% 由官方自建——1.3T 有机德语网页,加约 1T 用 LLM 改写的德语,机器翻译只占约 6%。理由很直白:翻译文本携带源语言的文化指纹,模型说德语、看世界却是美国的。配套 UniBPE 分词器在官方对比中德语压缩率 4.90 字节\u002Ftoken,高于 GPT-5(4.35)、DeepSeek V4(3.72)、Kimi K3(3.28)。算力账同样透明:768 张 B200、预训练 21 天,38 次中断全部自动恢复,总能耗约 950 MWh。\n\n## 所以呢\n\nKolibri 的看点不在「又一个开源 MoE」,而在另一条路线:不追万能多语种,把单一语言的数据工程做穿——改写而非翻译、词法感知分词——再用 3.46B 激活的经济学换受监管行业的本地部署市场。当然,全部跑分都是官方自报口径,第三方复现尚未出现。对中文社区,它提了个镜像问题:我们做了这么多语种,有没有哪一个是把中文数据账本算到 token 级的?\n\n参考:官方发布博客 aleph-alpha.com\u002Fen\u002Fblog\u002Fkolibri-has-landed-a-sovereign-open-weight-model,HF 模型卡 huggingface.co\u002FAleph-Alpha\u002FKolibri-1。","https:\u002F\u002Faleph-alpha.com\u002Fen\u002Fblog\u002Fkolibri-has-landed-a-sovereign-open-weight-model\u002F","482bca33-35da-4346-a1f3-a241a532d801",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":19,"name":20,"slug":20,"description":14,"color":14},"d11f0044-8aef-487c-bebe-89ce4683a4a3","moe",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"632d02f9-b009-4b9d-82e3-8f35949cde82","en","Aleph Alpha Open-Sources Kolibri-1: 78B MoE, Only 3.46B Active","Aleph Alpha open-sources Kolibri-1: 78B total, 3.46B active. Official math evals match 4x-larger models; 80% of German training data is self-curated.","On Germany's Day of Reunification, October 3, German AI company Aleph Alpha open-sourced Kolibri-1, a German-English mixture-of-experts reasoning model with 78B total parameters and only 3.46B active per token. The weights are on Hugging Face under an Apache 2.0 license, with a maximum context of 1,048,576 tokens. At a moment when most MoE releases compete on total parameter count, this \"78B total, 3.46B active\" combination — plus an unusually detailed German data-engineering ledger — deserves a closer look.\n\n## Why 78B, not 123B\n\nKolibri's internal predecessor, Kolibri Origin, was a validation model with 30.6B total \u002F 3.27B active parameters and a 65k context window; it finished pre-training on June 11. Three months later, on September 11, Kolibri finished pre-training on 20T tokens. The official blog discloses that the team tried scaling the model to 123B: on two H100s, a 123B model could serve only 3 concurrent 256k-token long-context requests, while 78B handles 18 and decodes 28% faster — so 123B was cut. The final architecture is a 50-layer MoE with 384 experts per layer (1 shared + 6 routed); 40 layers use 512-token sliding-window attention and every 5th layer keeps full attention. Weights are stored in FP8 (128×128 blocks) with an FP8 KV cache, a ~78GB footprint, and a 2×H100 minimum. The model supports four reasoning-effort levels (none\u002Flow\u002Fmedium\u002Fhigh) and tool calling; the 1M context is validated, but the company recommends staying within 262,144 tokens in production.\n\n## Strong scores, visible weaknesses\n\nIn self-reported evaluations (run with the company's open eval-framework), math is the highlight: 96.9 on AIME 2025, above the 120B-total \u002F 12B-active Nemotron 3 Super (91.7) and Qwen3.6-35B-A3B (84.6); 84.3 on GPQA diamond; and 38.1 on the banking τ³-bench, more than triple the same table's Qwen3.6 (10.6). Aleph Alpha claims it matches models with up to four times its active parameters across math, code, grounding, and long context. But the same table shows the other side: an English overall score of 75.5, behind the 27B dense Qwen3.8 (80.2); TerminalBench 2.1 at just 27.7 versus Qwen3.8's 76.8; and SWE-Bench Verified at 66.4, below Qwen3.6-35B's 73.8. The real strength is agentic RAG: 80.8 on Honeypot, ahead of every compared model. Hallucination control is the differentiator: on AA-Omniscience, the \"abstains instead of answering wrong\" rate is 44% (predecessor: 15%), and abstention training via the in-house Merlin-Arthur protocol yields an M\u002FA grounding score of 0.23 where most compared models sit at 0.\n\n## German is the moat\n\nThe most noteworthy part of this release is the German data ledger: 21.3% of the 20T pre-training tokens (about 4.3T) are German, backed by a 2.4T deduplicated German pool that is 80% self-built — 1.3T of organic German web plus roughly 1T of LLM-rephrased German (encyclopedia-style, Q&A-style), with machine translation at only ~6%. The rationale is blunt: translated text carries the cultural fingerprint of its source language — a model that speaks German but sees an American world (chancellor, not president). The accompanying UniBPE tokenizer (BPE base + Unigram objective for merge selection) reaches 4.90 bytes\u002Ftoken German compression in the official comparison, ahead of GPT-5 (4.35), DeepSeek V4 (3.72), and Kimi K3 (3.28). The compute ledger is equally transparent: 768 B200 GPUs, 21 days and 392k GPU-hours of pre-training with 38 unplanned interruptions all auto-recovered, and an estimated ~950 MWh total energy for pre-training, mid-training, and long-context training combined.\n\n## So what\n\nThe point of Kolibri is not \"another open-source MoE\" but the alternative route it demonstrates: skip universal multilinguality, do the data engineering of one language all the way through — rephrasing over translation, morphology-aware tokenization — and use the economics of 3.46B active parameters to win on-premise deployment in regulated industries. The model card names public administration, industrials, and aerospace as targets, and the company's EU GPAI Code of Practice signatory status is stated up front. Caveat: every benchmark so far is self-reported; third-party replication has not appeared yet. For the Chinese open-source community, it raises a mirror question: across all the languages we ship, is there a single one whose data ledger is accounted down to the token level?\n\nReferences: official [launch blog](https:\u002F\u002Faleph-alpha.com\u002Fen\u002Fblog\u002Fkolibri-has-landed-a-sovereign-open-weight-model\u002F), [HF model card](https:\u002F\u002Fhuggingface.co\u002FAleph-Alpha\u002FKolibri-1), and [The Open Weights entry](https:\u002F\u002Ftheopenweights.com\u002Fnews\u002Fkolibri-1-v2uw).","aleph-alpha-kolibri-1-open-moe","2026-10-03T19:14:02Z","2026-10-03T19:14:04.940789Z","2026-10-03T19:14:04.940805Z",true,"agent",917,[39,48],{"slug":40,"tag_slug":40,"title_zh":41,"title_en":42,"intro_zh":43,"intro_en":44,"id":45,"is_active":35,"created_at":46,"modified_at":47},"ai-for-science","AI for Science 2026：从 UniPert 到 GPT-Rosalind 的硬核进化","AI for Science 2026: from UniPert to GPT-Rosalind","生命科学、化学材料、物理世界模型——AI 正在从\"语言工具\"变成\"实验伙伴\"。本专题收录 AI 在三大科学方向的关键节点：UniPert 统一基因与化学扰动空间、GPT-Rosalind 端到端生命科学推理、达摩院 AI 智能体 28 小时找到 4 种超导新材料、Anthropic Claude Science 把工作台做成标准品。","From language tool to lab partner — AI is reshaping life sciences, chemistry\u002Fmaterials, and physical world models. This topic covers the key milestones: UniPert unifying genetic-chemical perturbation spaces, GPT-Rosalind's end-to-end life-sciences reasoning, DAMO's AI agent discovering 4 superconducting materials in 28 hours, and Anthropic's Claude Science workbench going mainstream.","988a4300-5fab-41c4-b5d8-63711a2dc757","2026-09-10T01:34:15.296649Z","2026-09-10T01:34:15.296663Z",{"slug":49,"tag_slug":49,"title_zh":50,"title_en":51,"intro_zh":52,"intro_en":53,"id":54,"is_active":35,"created_at":55,"modified_at":56},"h3-series","MiniMax H3 系列：从开源权重到 35 倍吞吐","MiniMax H3 Series: from open weights to 35x throughput","MiniMax H3 自 2026 年 8 月开源以来节奏密集：官方把生成、参考与编辑收回一个模型；ComfyUI 当天压进 RTX 3060；摩尔线程 3 小时完成国产 GPU 适配；fal 后训练版把吞吐拉到 35 倍；FastH3 蒸馏再砍推理成本。本专题持续追踪 H3 的发布—开源—蒸馏—部署全链路。","Since MiniMax open-sourced H3 in August 2026 the pace has been relentless: one unified omni-modal model, same-day ComfyUI support down to an RTX 3060, a 3-hour Day-0 port to Moore Threads GPUs, fal's post-trained H3 Max at 35x throughput, and FastH3 distillation cutting inference cost further. This topic tracks the full H3 chain — release, open weights, distillation, deployment.","83ef0daa-3c31-4cb3-86ed-e5ee58654d5f","2026-09-08T07:33:19.942193Z","2026-09-08T07:33:19.942209Z",{"items":58},[59,64,69,74,79,84],{"id":60,"title":61,"news_slug":62,"published_at":63},"f3c43720-650b-4054-9349-a1386e06c8fc","Le Chonk 把法国拉回非美\u002F美头部:38 分的 Mistral Large 4","mistral-large-4-le-chonk-intelligence-index-38","2026-10-08T03:30:00+00:00",{"id":65,"title":66,"news_slug":67,"published_at":68},"5ddba2a9-781e-4763-b1e0-1e20c6480391","Mistral Large 4:1万亿参数MoE,月底开源","mistral-large-4-1t-moe","2026-10-06T23:10:00+00:00",{"id":70,"title":71,"news_slug":72,"published_at":73},"9fa3427c-cf69-46c7-9720-cd3b646a155b","B站开源35B翻译模型:3B激活,150种语言","bilibili-index-translate-35b-moe","2026-10-04T13:30:00+00:00",{"id":75,"title":76,"news_slug":77,"published_at":78},"0bc17892-3c47-4081-86a4-3d90afa0c54b","小米 MiMo-V2.6 开源:万亿 MoE 追平 Grok 4.7,Flash 三分之一价格保九成战力","xiaomi-mimo-v2-6-open-weights","2026-09-22T13:02:37+00:00",{"id":80,"title":81,"news_slug":82,"published_at":83},"21fe3c11-4ba4-4801-b6fc-60c4ae559dc1","Yandex 逆流开源:35B 参数的 T5 MoE,每个 token 只激活 0.6B","yandex-aliceai-t5-sparse-moe","2026-09-16T19:11:43+00:00",{"id":85,"title":86,"news_slug":87,"published_at":88},"8e730a3d-439b-45cf-961d-f77cf01469fd","Cohere 开源 218B 翻译专用 MoE:25B 激活,自测评分超 DeepL,2×H100 可部署","cohere-north-small-translate","2026-09-11T19:07:20+00:00"]