Kolibri: A Sovereign Open-Weight Model
https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.
One of humanity’s biggest problems here is we don’t know how to do moderation.
We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.
The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.
I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.
We as many other’s were curious to try and benchmark it.
On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.
No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/
No, I'm not gonna do that, I asked you to do that Mr Kolibri.
Also feedback on that interface: It's very annoying while answering. It almost immediately shows a list of sources, which on my screen fill up all the space and then when it starts answering it keeps those in view but also scrolls down the tiny part of actual text its outputting but I can't scroll up to start reading from the top, coz it keeps scrolling. I have to wait until it's completely done generating its output.
And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.
Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.
Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?
Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.
So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.
If the second model is cheap and fast enough, there is a business model.
You don’t even need to audit all the intermediate steps, just tool calls and end results.
I get that doesn't invalidate the real "point" of the model, but...
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
This thing is worse than a Qwen3.8 27B.
Good to see Europe adding toe what Mistral is doing. +100
I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.
"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.
> intellectual-property safety
My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.
I wonder how far we are from this. How far are we from LLM's Debian moment?
How is a 27B dense model bigger than a 78B MoE?
If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.
Then what does it good at? Sending faxes?
But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.
We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).
This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.