Aleph Alpha Open-Weights Kolibri With 1M Context

Germany's Aleph Alpha released Kolibri, a 78.1B-parameter open-weight bilingual model with a 1M-token context, under Apache 2.0 on Unity Day.

Read as article

Aleph Alpha Open-Weights Kolibri With 1M Context

By @sharedot · · 6 pages

  • AI Frontier
  • Open Weight Models
  • Sovereign AI

Germany's Aleph Alpha released Kolibri, a 78.1B-parameter open-weight bilingual model with a 1M-token context, under Apache 2.0 on Unity Day.

What happened: Kolibri goes fully open

Aleph Alpha released Kolibri on October 3, 2026, an open-weight English-German Mixture-of-Experts model built for sovereign, mission-critical work in government and regulated industries. The model carries 78.1 billion parameters but activates only 3.46 billion per token, supports contexts up to 1,048,576 tokens, and ships with full weights on Hugging Face in BF16 and FP8 under the Apache 2.0 license, per TestingCatalog and ForkLog. Both outlets note organizations can deploy it on-premises rather than send internal data to a third-party inference service.

Why it's surprising: a 1M-context sovereign model

The surprise for builders is the combination: a fully downloadable model with a one-million-token context window, trained end-to-end inside Europe, is rare at any license tier. TestingCatalog reports the architecture uses 384 experts with six active per token, and full attention in only 10 of 50 layers — the other 40 use a 512-token sliding window to contain inference costs. Training ran on 768 Nvidia B200 GPUs, starting with 20 trillion tokens in 21 days and finishing near 24 trillion tokens total across mid-training and long-context adaptation, a scale and efficiency story both sources corroborate. German makes up 21.3% of pre-training tokens, supported by a bilingual 128,000-entry vocabulary and a tokenizer built to preserve German compounds — not machine translations, which ForkLog says accounted for only about 6% of the German data.

The evidence: benchmarks and grounding, vendor-run

Aleph Alpha's own results put Kolibri on the quality-versus-serving-cost Pareto frontier in English and German, claiming it can match models with up to four times as many active parameters on math, code, grounding, agentic, and long-context tasks — but both outlets stress the numbers are vendor-run and not independently verified. The two sources report slightly different scores: TestingCatalog lists 96.9 on AIME 2025 and 85.9 on LiveCodeBench v6, while ForkLog cites 96.7 on AIME 2026, 90.2 on the German-language version, 84.8 on GPQA Diamond, and 85.1 on LiveCodeBench v6. Grounding is the other headline claim: trained with Aleph Alpha's Merlin-Arthur abstention procedure, Kolibri withheld a wrong answer on 44% of AA-Omniscience items versus 14.8% for its predecessor Kolibri Origin, per TestingCatalog. ForkLog adds that Aleph Alpha built the model with the European AI Act, the GPAI Code of Practice, and GDPR in mind.

Stakes and what comes next

For practitioners, Kolibri is a concrete sovereignty option: weights you can run on your own hardware under Apache 2.0, with four reasoning settings (none, low, medium, high) plus tool calling, deployed via Aleph Alpha's inference package and a Kolibri-specific vLLM plugin, as TestingCatalog describes. ForkLog reports the FP8 version occupies about 78 GB of memory, with a minimum of one Nvidia H200, B200, or B300, or two A100 80GB/H100 SXM5 units; BF16 weights need roughly 156 GB. Aleph Alpha also shipped sector-specific evaluation suites for public administration, automotive, semiconductors, industrial technology, and aerospace that use no customer data, per TestingCatalog. The open questions now are independent verification of the Pareto-frontier claims and whether European institutions adopt a fully open model where closed frontier APIs previously dominated.

Sources

  1. testingcatalog.com › Aleph Alpha releases open-weight Kolibri with 1M context
  2. forklog.com › Germany Unveils Open AI Model for Public Sector

More on AI Frontier