On Germany's Day of Reunification, October 3, German AI company Aleph Alpha open-sourced Kolibri-1, a German-English mixture-of-experts reasoning model with 78B total parameters and only 3.46B active per token. The weights are on Hugging Face under an Apache 2.0 license, with a maximum context of 1,048,576 tokens. At a moment when most MoE releases compete on total parameter count, this "78B total, 3.46B active" combination — plus an unusually detailed German data-engineering ledger — deserves a closer look.

Why 78B, not 123B

Kolibri's internal predecessor, Kolibri Origin, was a validation model with 30.6B total / 3.27B active parameters and a 65k context window; it finished pre-training on June 11. Three months later, on September 11, Kolibri finished pre-training on 20T tokens. The official blog discloses that the team tried scaling the model to 123B: on two H100s, a 123B model could serve only 3 concurrent 256k-token long-context requests, while 78B handles 18 and decodes 28% faster — so 123B was cut. The final architecture is a 50-layer MoE with 384 experts per layer (1 shared + 6 routed); 40 layers use 512-token sliding-window attention and every 5th layer keeps full attention. Weights are stored in FP8 (128×128 blocks) with an FP8 KV cache, a ~78GB footprint, and a 2×H100 minimum. The model supports four reasoning-effort levels (none/low/medium/high) and tool calling; the 1M context is validated, but the company recommends staying within 262,144 tokens in production.

Strong scores, visible weaknesses

In self-reported evaluations (run with the company's open eval-framework), math is the highlight: 96.9 on AIME 2025, above the 120B-total / 12B-active Nemotron 3 Super (91.7) and Qwen3.6-35B-A3B (84.6); 84.3 on GPQA diamond; and 38.1 on the banking τ³-bench, more than triple the same table's Qwen3.6 (10.6). Aleph Alpha claims it matches models with up to four times its active parameters across math, code, grounding, and long context. But the same table shows the other side: an English overall score of 75.5, behind the 27B dense Qwen3.8 (80.2); TerminalBench 2.1 at just 27.7 versus Qwen3.8's 76.8; and SWE-Bench Verified at 66.4, below Qwen3.6-35B's 73.8. The real strength is agentic RAG: 80.8 on Honeypot, ahead of every compared model. Hallucination control is the differentiator: on AA-Omniscience, the "abstains instead of answering wrong" rate is 44% (predecessor: 15%), and abstention training via the in-house Merlin-Arthur protocol yields an M/A grounding score of 0.23 where most compared models sit at 0.

German is the moat

The most noteworthy part of this release is the German data ledger: 21.3% of the 20T pre-training tokens (about 4.3T) are German, backed by a 2.4T deduplicated German pool that is 80% self-built — 1.3T of organic German web plus roughly 1T of LLM-rephrased German (encyclopedia-style, Q&A-style), with machine translation at only ~6%. The rationale is blunt: translated text carries the cultural fingerprint of its source language — a model that speaks German but sees an American world (chancellor, not president). The accompanying UniBPE tokenizer (BPE base + Unigram objective for merge selection) reaches 4.90 bytes/token German compression in the official comparison, ahead of GPT-5 (4.35), DeepSeek V4 (3.72), and Kimi K3 (3.28). The compute ledger is equally transparent: 768 B200 GPUs, 21 days and 392k GPU-hours of pre-training with 38 unplanned interruptions all auto-recovered, and an estimated ~950 MWh total energy for pre-training, mid-training, and long-context training combined.

So what

The point of Kolibri is not "another open-source MoE" but the alternative route it demonstrates: skip universal multilinguality, do the data engineering of one language all the way through — rephrasing over translation, morphology-aware tokenization — and use the economics of 3.46B active parameters to win on-premise deployment in regulated industries. The model card names public administration, industrials, and aerospace as targets, and the company's EU GPAI Code of Practice signatory status is stated up front. Caveat: every benchmark so far is self-reported; third-party replication has not appeared yet. For the Chinese open-source community, it raises a mirror question: across all the languages we ship, is there a single one whose data ledger is accounted down to the token level?

References: official launch blog, HF model card, and The Open Weights entry.