Aleph Alpha's Kolibri: what you really save with 3.46B active parameters (and why you still need 78 GB of VRAM)

Aleph Alpha published Kolibri on October 3, 2026: a mixture of experts reasoning model in German and English with 78B parameters, of which only 3.46B are active per token, with open weights under Apache 2.0. That ratio makes each token cheap to compute, but all 78B parameters must be in GPU memory: the minimum is two 80 GB A100s, not a laptop.

What is Aleph Alpha’s Kolibri?

Kolibri is an open-weight reasoning model trained from scratch by Aleph Alpha, the Heidelberg lab behind the Pharia models. The model card declares no dependencies on other models: it’s not a fine-tune of Llama, Qwen, or anyone else.

What it brings:

  • explicit reasoning mode with adjustable effort
  • tool calling
  • native context window of 262,144 tokens, validated up to 1,048,576

Aleph Alpha positions it for multi-step reasoning, retrieval-augmented generation (RAG), agents with tools, coding, and assistants in German and English. Its knowledge extends to June 18, 2026.

There are two official repositories:

  • Aleph-Alpha/Kolibri-1 is the FP8 version, the one you’re most likely to deploy.
  • Aleph-Alpha/Kolibri-1-BF16 contains the weights in full precision.

What is a mixture-of-experts model and what do 3.46B active parameters mean?

It means Kolibri computes like a 3B model and occupies memory like a 78B one.

Kolibri is a mixture of experts (MoE) model. Each of its 50 layers contains 384 small “expert” networks plus a shared expert. For each token, a router selects 6 of the 384 experts, and only those (plus the shared expert and attention layers) do work. That adds up to 3.46B active parameters out of 78.1B total.

This design creates two costs that move in opposite directions:

  • Computation per token follows active parameters. Generating a token costs roughly what a dense 3 to 4B model costs. That’s where the throughput savings and lower serving cost come from.
  • Memory follows total parameters. The router can choose any expert for the next token, so all of them must be loaded and ready. Aleph Alpha’s own model card acknowledges this: memory is the price you pay.

Kolibri also saves on attention. Four of every five layers only look at the previous 512 tokens (sliding-window attention), and only one of every five attends to the full context. That’s what lets it serve very long contexts at reasonable cost.

How much VRAM does Kolibri need?

The FP8 version needs at minimum two 80 GB A100s, two H100s, or a single H200. Here are the figures from the official model cards:

Kolibri-1 (FP8) Kolibri-1-BF16
Weights in memory ~78 GB ~156 GB
Minimum 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200, or 1× B300 4× A100 80 GB, 4× H100 SXM5, 2× H200, 1× B200, or 1× B300
Recommended 2× H100 SXM5, 2× H200, 1× B200, or 1× B300 4× H100 SXM5, 2× H200, 2× B200, or 1× B300

FP8 is the intended deployment format, not a cutdown added later. During reinforcement learning, Aleph Alpha used quantization-aware training precisely so the model works well in low precision with vLLM. The FP8 version stores weights in FP8 and keeps embeddings, the output head, normalizations, and the MoE router in bfloat16.

On context length, the model card recommends not exceeding 262,144 tokens in latency- or throughput-sensitive deployments. The 1M window is available, but requires additional flags (see below).

At the time this note was published, Hugging Face listed several community quantizations of Kolibri labeled for llama.cpp, LM Studio, and Ollama. They’re not from Aleph Alpha and the model card doesn’t evaluate them, so their quality is unknown until you test them.

How to install and serve Kolibri with vLLM?

Install Aleph Alpha’s inference package and then run vllm serve with Kolibri’s own parsers. The package brings the Kolibri plugin for vLLM and installs the version of vLLM that supports it:

pip install 'aleph-alpha-inference>=1'

If you prefer not to install it, Aleph Alpha also publishes a container image: ghcr.io/aleph-alpha/aleph-alpha-inference.

Serve the FP8 model with reasoning and tool calling enabled:

vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 \
  --reasoning-parser kolibri1 \
  --tool-call-parser kolibri1 \
  --enable-auto-tool-choice

To exceed 262,144 tokens of context, add these flags:

--max-model-len 1048576 --hf-overrides '{"max_position_embeddings": 1048576}'

The recommended sampling parameters are temperature=1.0, top_p=0.97, and top_k=128.

How do you call Kolibri from code?

Point any OpenAI client at the server. vLLM exposes an OpenAI-compatible API, and the model card example uses http://localhost:8000/v1:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")

response = client.chat.completions.create(
    model="Aleph-Alpha/Kolibri-1",
    messages=[{"role": "user", "content": "Explain briefly what a mixture-of-experts model is."}],
    extra_body={
        "chat_template_kwargs": {
            "reasoning_effort": "high",
            "enable_thinking": True,
        }
    },
)
print(response.choices[0].message.content)

The reasoning_effort parameter accepts low, medium, and high. If you set it to none, or pass enable_thinking=false, the model responds immediately without reasoning.

Tool calling is Hermes-style: pass your function schemas in the standard tools field. You can combine it with reasoning mode.

Since the endpoint speaks the OpenAI API, any agent that lets you bring your own model can point to it. Goose is an example.

Kolibri vs Qwen: Is it really better?

In Aleph Alpha’s own benchmarks, Kolibri has the best overall score among the MoE models it compares against. But it clearly loses on programming agent tasks, and all these figures are self-reported. Aleph Alpha ran all models with its own eval-framework, and at the time this note was published there were no independent evaluations.

Figures from the FP8 card:

Benchmark Kolibri Qwen3.5 35B-A3B Qwen3.6 35B-A3B GPT-OSS 120B
Global (EN) 75.5 74.7 71.4 72.3
Global (DE) 70.8 69.8 67.3 70.2
AIME 2025 (EN) 96.9 88.1 84.6 90.6
LiveCodeBench v6 85.9 77.8 82.5 87.5
SWE-Bench Verified 66.4 71.6 73.8 –
TerminalBench 2.1 27.7 39.7 – 29.2
BFCL v4 (global) 61.4 70.5 67.2 57.3

The pattern: Kolibri is strong in math, reasoning, and RAG in German and English, and weaker on software engineering agent tasks and multi-turn tool calling.

The top comment in the Hacker News thread points out that the dense Qwen3.8 27B beats Kolibri in German (79.9 vs 70.8) with Kolibri’s own harness. It’s true, and the figure is in Aleph Alpha’s table. What the comment leaves out is that Qwen3.8 27B is a dense model: its 27B parameters work on every token, roughly eight times Kolibri’s compute per token. Aleph Alpha shows dense models in gray, as reference, not as direct rivals. The fair reading is that Kolibri trades some quality for far less compute per token.

The same goes for hallucinations. On the non-hallucination rate of AA-Omniscience, Kolibri scores 44.0 versus 56.7 for Qwen3.6.

Does Kolibri work in Spanish?

Not officially. Kolibri supports only German and English. The card describes this as a deliberate choice of depth over breadth, and Spanish doesn’t appear in any of its evaluations.

The pretraining mix was approximately 62.5% English, 23.9% German, and 13.6% code. You’ll get responses in Spanish, but nothing guarantees their quality. If your users write in Spanish, evaluate it with your own data before you commit.

Is Kolibri free for commercial use?

Yes, as far as the weights go: they’re published under Apache 2.0. The card adds a caveat: the license applies only to the weights and configuration files in the repository. The code, model architecture, parameters, and training methods are explicitly excluded, and Aleph Alpha retains those rights.

Grego’s note: what “sovereign” means when the lab gets bought

Aleph Alpha launched Kolibri as a “sovereign” model. The top comment on Hacker News asked the obvious question: in April 2026, Cohere, based in Toronto, agreed to acquire Aleph Alpha. According to reports, Aleph Alpha shareholders would end up with roughly 10% of the combined company, which will operate under the name Cohere. At the time this note was published, I couldn’t confirm that the deal had closed.

I don’t think it detracts from the launch, but it does show where sovereignty really lives. It doesn’t belong to a vendor: a vendor gets bought, merged, or repositioned in a quarter. It belongs to the artifact: trained weights from scratch, under Apache 2.0, that you can run on hardware you control, in the jurisdiction you choose. Once those weights are on your server, who owns Kolibri doesn’t change anything for you.

There are two more details that matter to anyone selling in Europe, including our readers in Spain. Aleph Alpha signed the EU Code of Good Practice for general-purpose AI models, and publishes a summary of its training data using the European Commission’s template. If your compliance team asks where a model’s data came from, this is one of the few open-weight releases that comes with the paperwork.

Who should use Kolibri?

Kolibri makes sense if you already have, or can rent, a node with an H200 or two H100s, and your workload is RAG over long documents, reasoning, or bilingual German/English processing. On that hardware, a model with 3.46B active parameters gives you a lot of throughput per GPU.

It’s not the right choice in three cases:

  • Programming agents: according to Aleph Alpha’s own figures, Qwen MoEs perform better on SWE-Bench and TerminalBench.
  • Products in Spanish: Spanish isn’t evaluated.
  • Running it on your laptop: for that, look at a small dense model like Ornith-1.5-9B, or the Gemma 4 family MoEs, which also appear in Kolibri’s comparison table.

Where do you run your open-weight models?
  • On my laptop or PC
  • On my own server or rented GPUs
  • I don’t run them: I use third-party APIs
  • I prefer another option (tell us which)
0 votantes

Related: 5.63 GB on Your Laptop and 70.6 on SWE-bench Verified: Installing Ornith-1.5-9B