Published: October 5, 2026 | By Vito Ruocco
Germany just released a free AI model, and on paper it beats Qwen, Nvidia and Mistral. It is called Kolibri, it comes from Aleph Alpha in Heidelberg, it has seventy-eight billion parameters and a context window of up to a million tokens, and the weights are yours under Apache 2.0. Every headline this weekend repeated the benchmark chart. I wanted to see what happens when you actually use it, so I gave Kolibri a game and an app to write, and then I gave the exact same prompts, word for word, to four of the models it claims to beat.
This article is the long, written version of my latest VitoMind video. It explains what Kolibri is, why a European open model matters, how a 78-billion-parameter model manages to run like a 3-billion one, what the benchmark table really says when you read all of it, the one thing almost nobody mentions before telling you to “download it and run it at home”, and where you can try it today for free with no GPU at all. Then we get to the part that matters most: Kolibri against Qwen3.6, Nemotron 3 Super, Mistral Small 4 and Qwen3.8 on the same two prompts, with all ten results published on my site so you can use them yourself.
If you prefer to read, everything from the video is below, with the numbers, the sources and a few details I could not fit into eight minutes.
The Short Version
- Kolibri-1 is Aleph Alpha’s new open-weight model, released on October 3, 2026: a mixture-of-experts model with 78 billion parameters in total and only 3.46 billion active per token, up to 1 million tokens of context, German and English only, Apache 2.0.
- On its own model card it leads the mixture-of-experts models of its size, ahead of Qwen3.6 35B-A3B, Nvidia’s Nemotron 3 Super and Mistral Small 4. The dense Qwen3.8 27B still beats it almost everywhere.
- It needs about 78 GB of GPU memory and Aleph Alpha’s own plugin for vLLM. It does not run on a gaming PC, it is not on OpenRouter, and no inference provider on Hugging Face serves it yet.
- You can try it for free in Tesseracted Labs’ browser chat (ten messages a day as a guest, a thousand with a free account), and they offer a seven-day API pilot on request.
- In my test, writing a small game and an interactive app, Kolibri was the weakest of the five. The Qwen model of the same size did better on both prompts.
- That is not what Kolibri was built for. For German documents that must stay on your own servers, it is one of the most interesting open models you can get right now.
- All ten demos are free to play at ruocco.it/ai-demos/kolibri-1-vs-rivals. The exact prompts and the full code are in the Ruocco Academy.
What Aleph Alpha Released
The launch post from Aleph Alpha opened with a line that is also the whole pitch: “Small bird, fast wings, Kolibri is here.” Kolibri is the German word for hummingbird, and the name is chosen carefully. A hummingbird is tiny and incredibly fast, and that is exactly what the company is promising: a big model that moves like a small one.
The details are on the official model card on Hugging Face. Kolibri-1 has 78.1 billion parameters in total, but only 3.46 billion of them are active for each token it reads or writes. The context window has been validated up to 1,048,576 tokens, which is about a million, although Aleph Alpha recommends staying around 262,000 tokens for everyday serving, because that is the length the model was natively trained on and the sweet spot for speed and quality. It has an explicit reasoning mode with effort levels from none to high, it supports tool calling, and the license is Apache 2.0, which means you can use it commercially, modify it and redistribute it.
It speaks two languages, German and English, and that is a deliberate choice. Aleph Alpha says supporting two languages instead of many is “depth over breadth”. Its knowledge cutoff is June 18, 2026, and it was released on October 3, which in Germany is the Day of German Unity. That was not a coincidence: this is a model built to make a point about European independence.
Why a European Model Matters
Europe has very few labs that build large language models from scratch. Mistral is the famous exception, and after that the list gets short very quickly. Aleph Alpha’s own post after the launch said the company started the year “with the conviction that Europe needs its own sovereign AI”, built in Germany from scratch and on infrastructure in Europe.
The company also says it built Kolibri “compliance first”: following the EU AI Act and GDPR, with copyright safeguards that go beyond the EU’s Code of Practice for general-purpose AI, which Aleph Alpha has signed. For most people watching AI videos that sounds like paperwork. For a ministry, a bank or an aerospace supplier, it is the difference between “interesting” and “allowed”.
The blog post that accompanied the release is very clear about who the model is for. In Aleph Alpha’s words, Kolibri is “a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace”. These are places where data simply cannot leave the building. The company trained the model on 768 Nvidia B200 GPUs, in Germany and Finland, with 21 days of pre-training on 20 trillion tokens, followed by 3.44 trillion tokens of mid-training and a long-context phase. Roughly 21 percent of the pre-training data is German, which is far more than any general-purpose model would ever use.
How a 78-Billion-Parameter Model Runs Like a 3-Billion One
The clearest explanation I found was written by Tejas Kumar, an AI engineer at IBM, on his blog. If you want the full technical picture, his article and Aleph Alpha’s 200-page technical report are the places to go. Here is the plain-English version.
384 experts, and each token only meets six of them
In a normal, dense model, every word goes through every parameter. Kolibri is a mixture-of-experts model. Each of its 50 layers contains 384 small specialist networks, called experts, plus one shared expert. A small router looks at each token and sends it to just six of the 384 experts, and the shared expert always joins in. So for every token, seven experts do the work and everything else stays switched off. That is how 78 billion parameters turn into about 3.46 billion parameters of actual work per token, and that is why the model is fast and cheap to run per token.
To make this easier to picture, one of the demos from my test is an interactive page that shows exactly this: a sentence is split into tokens, and each token travels through the layers while its seven experts light up on a grid of 384 cells. Ironically, the best version of that page was not written by Kolibri, as you will see later.

The catch: memory
There is a catch, and it is a big one. Even though only a few experts work at any moment, all of them have to sit in GPU memory, because the router can pick any of them for the next token. So Kolibri thinks like a small model but needs the memory of a large one. Keep this in mind, because it comes back in the most important practical section of this article.
Most layers only look nearby
The second trick is about long context. According to Tejas’s write-up, 40 of Kolibri’s 50 layers use sliding-window attention: each token only looks at the 512 tokens right before it. Every fifth layer looks at everything. It is like reading a very long contract while mostly paying attention to the sentence you are on, and stopping every few pages to think about the whole thing. That is how a million tokens of context stays affordable instead of exploding the cost.
A tokenizer that reads long German words
My favourite detail is the tokenizer. Models do not read letters or words, they read chunks called tokens, and German is famous for very long compound words. Aleph Alpha trained a bilingual tokenizer that respects the way German words are built. Tejas ran it on the German constitution, the Basic Law, and Kolibri needed about 15 percent fewer tokens than GPT-5’s tokenizer for the same text. His favourite example is the word Bundesverfassungsgericht, the Federal Constitutional Court: GPT-5’s tokenizer splits it into six pieces, Kolibri’s into just two. Fewer tokens means cheaper and faster German, and more German text fitting into the context window.
The Benchmarks, Read Whole
Now the numbers everyone has been sharing. The evaluation table on the model card is large, and most posts only quote the top of it. Here is what it says if you read all of it.

Against the other mixture-of-experts models with a similar number of active parameters, Kolibri leads overall: 75.5 in English, ahead of Qwen3.5 35B-A3B (74.7), Nemotron 3 Super (73.0), Qwen3.6 35B-A3B (71.4) and Mistral Small 4 (63.1). In German the lead is similar, 70.8 overall.
The math results are genuinely impressive: 96.9 on AIME 2025 in English. And on a banking customer-service agent benchmark, Tau3-Bench Banking, Kolibri scores 38.1 while most of its rivals stay under 16. If you build agents for regulated industries, that line is worth looking at.
But read the whole table. The dense Qwen3.8 27B, which activates several times more parameters per token, beats Kolibri almost everywhere, with 80.2 overall in English. The model card itself greys the dense models out for that reason, and it is a fair choice, but it is also the part that most posts left out. On SWE-bench Verified, the coding benchmark, the Qwen models are ahead: Kolibri scores 66.4, Qwen3.6 35B-A3B 73.8 and Qwen3.8 27B 72.6. Multi-turn tool calling is another weaker area.
One more thing I really like: Kolibri is trained to say “I don’t know”. Aleph Alpha used a special training game in which parts of the documents are sometimes hidden, and the model has to tell whether it actually has the evidence. On the AA-Omniscience test, when Kolibri did not know an answer it admitted it 44 percent of the time. Qwen3.5 did so 11 percent of the time. For a company that wants a model to answer from its own documents and not invent things, that is a very important number.
The One Thing Nobody Tells You: Hardware
A lot of videos this weekend said something like “download it and run it on your own computer”. Before you start a 78 GB download, look at the hardware line on the model card.
The model’s memory footprint is about 78 GB, just for the weights in FP8. The minimum Aleph Alpha lists is two Nvidia A100s with 80 GB each, two H100s, or a single H200, B200 or B300. A gaming PC, even with the best consumer graphics card you can buy, is nowhere near that. Remember the catch from the experts section: only 3.46 billion parameters are active, but all 78 billion have to be in memory.
It is also not a one-click install. Kolibri needs Aleph Alpha’s own inference package, aleph-alpha-inference, which installs a plugin for vLLM and the vLLM version it supports. The usual desktop apps do not run it yet. The official command on the model card looks like this: vllm serve Aleph-Alpha/Kolibri-1 --kv-cache-dtype fp8 --reasoning-parser kolibri1 --tool-call-parser kolibri1 --enable-auto-tool-choice.
So can you rent it somewhere? I checked. On Hugging Face, the model page says it is not deployed by any inference provider. On OpenRouter, a search for Kolibri returns no matching models. As of October 5, 2026, there is no hosted API from Aleph Alpha either.
Where to Try Kolibri Today, for Free
There is one easy way. A German company, Tesseracted Labs, put Kolibri online as a gift for the launch. On X, Konark Modi wrote for the company that they care about sovereign AI and wanted anyone to be able to try it, with no GPU and no setup.
The chat is at tesseracted.com/kolibri-1-chat. It has example prompts, one of them in German, a button for extended thinking, and optional web search powered by Brave. As a guest you get ten messages a day per IP address; with a free account you get a thousand. Tesseracted is clear that this is an independent demo and not an official Aleph Alpha service.
If you are building a product, they also offer API access on request. Their developer page describes an OpenAI-compatible endpoint, and the default approved pilot gives you 100 million input and output tokens per day, 300 requests per minute and 20 concurrent requests, for seven days. One detail from that same page explains part of my test results: by default an answer gets a budget of 4,096 tokens, and when thinking is on, the reasoning counts against that budget.
There is even a compare page, where you ask one question and get two answers side by side, Kolibri next to Qwen3.8 27B. It is a smart way to judge the model on your own questions, and I recommend trying it in German.
This free chat is exactly where I ran my two tests with Kolibri.
Every Demo on My Site, and the Academy
All the demos from this video live on my site, ruocco.it. The idea is simple: every new model I test, I test live, and every demo is yours to play in the browser for free, no account needed.
If you want to go further, there is the Ruocco Academy. Members get every prompt I used, word for word, with the model, settings, time and cost, step-by-step notes on how each video was made and what broke, the full source code of every demo as a ZIP, and videos only members can watch. For this video that includes Kolibri’s untouched original answers and the exact commands to run Kolibri on your own GPU.
The Test: Kolibri vs Four Rivals on the Same Prompts
Here is how I set it up. I gave Kolibri two prompts in Tesseracted’s free chat. Then I sent the exact same text, word for word, to four models through OpenRouter: Qwen3.6 35B-A3B, which is the same size class as Kolibri with about three billion active parameters; Nvidia’s Nemotron 3 Super; Mistral Small 4; and the dense Qwen3.8 27B. Five models, two prompts, ten answers, one request each. The four rivals cost about ten cents in total.
Two prompts are not a benchmark, and I want to be clear about that. This is a small test of one practical skill, writing working code for a browser, which is also what many people try first with any new model. Every file is published exactly as the model wrote it, and where I had to change something to make it run, I say so and publish the untouched original next to it.

Prompt one: Kolibri Rush
The first prompt asks for a complete browser game in one HTML file, called Kolibri Rush: a small glowing hummingbird flying through a neon city at night, collecting nectar from flowers and dodging patrol drones, with a start screen, a game-over screen, a parallax skyline with lit windows, a trail behind the bird, a score and a best score. The prompt also asks every model to expose the game’s state as window.game, which is what allowed one small bot to play all five versions on the side-by-side page.
Kolibri-1. Kolibri wrote the whole file in about twenty seconds, which is fast. But the game froze on its very first frame, because it calls ctx.circle, a canvas function that does not exist. I changed that single call to ctx.arc, and then it runs. Unfortunately, the city skyline the prompt asked for is missing, and the flowers move at eight pixels per second. The code mixes per-frame and per-second speeds, so the game is technically running but you cannot really play it.
Qwen3.6 35B-A3B. Same size class, three billion active parameters. It starts on the first try, with a skyline full of lit windows, flowers, drones and a speed-up over time. This is what the prompt asked for. It has one bug: with the keyboard the bird never stops climbing, because the key release resets one flag but not the other. With the mouse it plays perfectly, so that is how the bot plays it.
Nemotron 3 Super. As delivered, it could not start at all: the start screen was never connected to the space bar or the mouse. Two lines fixed that, and then it plays, with a bright city behind it.
Mistral Small 4. It runs out of the box, but its idea of a city is rows of grey towers across the whole screen.
Qwen3.8 27B. The dense model needed one line: it never called its resize function at start, so the canvas stayed at the default 300 by 150 pixels. With that one call added, it is the most polished of the five, with a proper title screen, glowing flowers, drones and hearts for lives.
Prompt two: an interactive explainer of Kolibri’s own architecture
The second prompt asks for an interactive page that visually explains how a mixture-of-experts model routes tokens, using Kolibri’s real architecture: 50 layers, 384 experts per layer, one shared expert, top-six routing, 78 billion parameters in total and about 3.46 billion active. The page needed a sentence split into tokens, an animated grid of 384 experts, a layer slider, a play button, live counters and a heatmap of expert usage.
Kolibri-1. I ran this one with extended thinking on, and Kolibri hit the demo’s output limit. The chat showed the message “Output limit reached. The answer may be incomplete.” The file stops halfway through a function, so the page shows its controls and an empty canvas. To be fair, that is the free demo’s cap, a 4,096-token budget that includes the thinking, and not necessarily the model’s limit. With your own server, or by asking it to continue, it could finish. But this is what you get there today.

Qwen3.6 35B-A3B. It built the whole thing: tokens as clickable chips, the expert grid, live statistics with active experts and active parameters, and a usage heatmap.
Nemotron 3 Super. Its file has a syntax error, so the page shows nothing. I left it untouched.
Mistral Small 4. It draws the token and its connections to the chosen experts, but not the grid of 384 experts the prompt asked for.
Qwen3.8 27B. The best version of the five. You route the tokens and watch each one travel through the layers, with the shared expert and six routed experts lighting up and the heatmap filling in. It is the visualizer I used earlier in this article to explain how Kolibri works.
My Verdict
On paper, Kolibri beats these models. In my test, writing a game and an interactive app, the Qwen model of the same size, Qwen3.6 35B-A3B, did better on both prompts, and the larger dense Qwen3.8 27B produced the most polished results. Kolibri was the fastest to answer and the weakest result in both cases, partly because of the free demo’s output budget.
But that is not what Kolibri was built for, and its own model card says so: it lists multi-step reasoning, retrieval-augmented generation, agentic tool calling and German- and English-language assistance as its strengths, and the benchmarks agree that coding is not its best area. If you work with German documents that have to stay on your own servers, with a model that admits when it does not know, Kolibri is one of the most interesting open models you can get right now. For writing code, I would still pick Qwen.
If you want to test it properly, there are two good options. You can rent a single H200 in the cloud for a few dollars an hour and run Aleph Alpha’s container, or you can ask Tesseracted Labs for an API pilot. If enough of you want it, I will rent the GPU and test Kolibri the way it was meant to be used, on real German documents.
Try Every Demo, and Get the Code
- All ten demos, free in your browser: ruocco.it/ai-demos/kolibri-1-vs-rivals, including the side-by-side page where one bot plays all five games and Kolibri’s original answers as text.
- The exact prompts, the full source as a ZIP and how to run Kolibri yourself: Ruocco Academy.
- Every other model I tested this way: ruocco.it/ai-demos.
Sources
- Kolibri-1 model card on Hugging Face (overview, evaluation tables, hardware requirements, getting started)
- Aleph Alpha: Kolibri Has Landed, A Sovereign Open-Weight Model
- Tejas Kumar: Aleph Alpha Kolibri, How the Sovereign German LLM Works
- Tesseracted Labs’ free Kolibri-1 chat and its developer documentation
- Posts on X by @Aleph__Alpha, @TejasKumar_ and @konarkmodi
A note on honesty: Kolibri-1’s answers in this test come from Tesseracted Labs’ free demo, which is an independent service and not an official Aleph Alpha product, and which limits each answer to a 4,096-token budget. The four rivals ran through OpenRouter with medium reasoning. Two prompts are a small test of code writing, not a benchmark. Every change I made to make a file run is listed on its card in the gallery, and the untouched originals are published next to the fixed versions.
If you want every new AI model tested hands-on like this, subscribe to VitoMind on YouTube and turn on notifications. And tell me in the comments of the video which model I should test next.