What happened?
On 3 October 2026, the Day of German Unity, Aleph Alpha released Kolibri-1. The weights are on Hugging Face under Apache 2.0. You can run them on hardware you control. There is no hosted inference provider on the card yet.
It is an English-German mixture-of-experts model. The card counts 78,103,074,560 parameters in total and 3,457,573,120 active per token. That is 78.1 billion and 3.46 billion. The blog rounds the active count to 3 billion. The tech report says 4.4 percent of the weights fire on each token. The full model still has to sit in memory: about 78 GB in FP8. Minimum hardware on the card is two A100 80 GB, two H100s, or one H200, B200, or B300. This is not a laptop model.
The tweet says up to 1 million tokens of context. The card is more precise. They validated quality and serving up to 1,048,576 tokens. The long-context training phase stopped at 262,144, and they recommend staying at or under that for serving speed and for hard tasks. Positional encoding sits only in the sliding-window layers, so the extra length is an extension, not a free lunch.
German is the point, and the three official pages do not use one share. The model card's pre-training mix is about 62.5 percent English, 23.9 percent German, and 13.6 percent code, on 20 trillion tokens. The blog says 21.3 percent of pre-training tokens are German, with translation used sparingly at 6 percent overall. The tech report says German is more than 20 percent of the mix, including more than 2 trillion German tokens they curated or wrote themselves. Mid-training adds 3.44 trillion tokens. The long-context extension adds 201 billion. The report sums the stages as 24 trillion. The card's three lines add to about 23.6 trillion. Close, not identical.
Training ran on 768 NVIDIA B200s, in Germany and Finland. Pre-training took 21 days, 392,000 GPU-hours. Mid-training was 5 days. The long-context phase was 13 hours. They estimate 950 MWh including data-centre overhead, and that figure leaves out supervised fine-tuning and reinforcement learning.
Their own post-training table, MoE models only, puts Kolibri at 75.5 overall in English and 70.8 in German. The tech-report speed plot, a different average, puts the post-trained model at 75.8 and 71.0. On that plot the base model is faster: 105,000 decoded bytes per second per GPU, tied with Nemotron Nano, at 81.5 in both languages. After post-training the same chart shows Kolibri at 31,000 bytes per second, while Nemotron Nano is still near 80,000 to 90,000 and scores lower. Do not read the base speed as the speed of the model you would actually serve.
A few rows, all from their tables:
- GPQA Diamond: 84.3 English, 81.3 German. On the same card, Qwen3.5 35B-A3B scores 84.2 in German.
- AIME 2026: 96.0 English, 90.0 German. The dense Qwen3.8 27B, which they grey out because it activates more parameters, scores 97.7 and 96.9.
- SWE-bench Verified: 66.4. TerminalBench 2.1: 27.7.
- Their AA-Omniscience index, from minus 100 to 100: Kolibri at minus 32.8. Qwen3.6 35B-A3B is at minus 15.3 on the same row.
The blog also shows internal customer-proxy scores climbing from Kolibri Origin to Kolibri: German public sector 0.54 to 0.75, aerospace 0.14 to 0.59. Those are their suites, not a public leaderboard.
Why this is interesting
- A German model you can actually hold - Apache 2.0 weights, trained in Germany and Finland, bilingual by design rather than an English model with a German patch. For a Swiss firm that cannot send files to a US API, that is the product.
- The 1 million token line is the ceiling - They tested it. They tell you to serve at 262,144 if the job is hard or the bill matters.
- Active parameters are the trick, memory is the bill - 3.46 billion compute per token, 78 billion in VRAM. Cheap to run a token, expensive to load.
- German is strong, not magic - Ahead of most open MoE peers on their German average. Not ahead of every peer on German GPQA, and not ahead of the dense 27 billion model they set aside.
- It is built to stop - They trained abstention, and they say the model should sit on the advisory side, with a person reviewing the output. The omniscience index is still negative. The intent and the score are both on the page.
What it is not
Not a chat app. Not a hosted API you can call today. Not a model that fits on a Mac. Not a win over every open model in German. Not an independent audit of the Pareto chart. The internal industry scores are Aleph Alpha grading Aleph Alpha. The EU AI Act and GDPR language in the report is a design claim, not a certificate that your deployment is compliant.
Bottom line
Kolibri is the first open European weights release in a while that is actually about German, and small enough in active compute that a serious on-prem box can serve it. Read the card for the hardware and the 262,144-token recommendation. Read the tables before you repeat the tweet. If the work is German documents on hardware you own, this is the one to try. If you needed a phone model, it is not that.
