Hacker News new | past | comments | ask | show | jobs | submit

Kolibri: A Sovereign Open-Weight Model

https://aleph-alpha.com/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/
The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.
loading story #49946129
loading story #49946420
In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.
loading story #49946466
This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.
loading story #49945886
Hopefully this becomes the new standard.

It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.

loading story #49945632
His point about regulation and innovation is great and I wish more people thought like that.

One of humanity’s biggest problems here is we don’t know how to do moderation.

We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.

The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.

I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.

loading story #49946138
Thank you Aleph Alpha team for making it open.

We as many other’s were curious to try and benchmark it.

On that note, as a small gesture of support, we’ve hosted and made Kolibri-1 free for anyone to try for the next few days.

No GPU. No setup. Just try it. tesseracted.com/kolibri-1-chat/

https://x.com/konarkmodi/status/2106373678589960260?s=46

loading story #49945353
Not impressed. I asked it how to run itself (giving it the Huggingface link) on limited RAM i.e. less than stated as needed and on llama.cpp and true to what we read about "it will tell you when it doesn't know" that's almost all I got: It doesn't know, it told me I should go click on tabs in the Huggingface interface for more information. This was with extended thinking on.

No, I'm not gonna do that, I asked you to do that Mr Kolibri.

Also feedback on that interface: It's very annoying while answering. It almost immediately shows a list of sources, which on my screen fill up all the space and then when it starts answering it keeps those in view but also scrolls down the tiny part of actual text its outputting but I can't scroll up to start reading from the top, coz it keeps scrolling. I have to wait until it's completely done generating its output.

loading story #49946281
Thank you for the feedback on UI, improved the streaming to make it less frustrating.
loading story #49946275
For a post to make such a big deal about sovereignty it is a bit misleading to not mention that the company is slated to be merged with Cohere, a Canadian company.

And that is a good thing - no need to hide it. Given the growing cost of keeping up, these few non-US, non-Chinese companies really need to do more sharing of efforts and costs.

Canada too is very much in need of sovereign AI options, but funding that on its own would be pretty much a waste of money. Would love to see this new German Canadian company cooperate with Mistral too, or maybe one of the Korean AI companies.

Qwen3.8 27B beats Kolibri 79.9 vs 70.8 in German in Kolibri's harness on Kolibri's benchmark.

Also, once the Cohere takeover is complete will they still be able to use this "sovereign" claim despite being 90% owned and 100% operated out of Toronto?

loading story #49945268
loading story #49945340
loading story #49944879
I think at the moment the main thing a sovereign AI model needs to be good at is auditing the results of other models.

Right now one could run an open model for most government applications and it would be good enough, you just cannot trust any of these.

So having a sovereign controlled model audit the first one would basically act like a “trust adapter”.

If the second model is cheap and fast enough, there is a business model.

You don’t even need to audit all the intermediate steps, just tool calls and end results.

loading story #49944417
loading story #49944770
The absence of any comparison to Qwen3.8 Flash, another MoE model with a small-ish (6B) number of active parameters, is pretty striking. Instead, it's compared with Qwen3-Next 80B-A3B, a model released almost a full year ago.

I get that doesn't invalidate the real "point" of the model, but...

I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach

Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8

loading story #49943876
loading story #49944108
loading story #49946354
I love how the mere mention of a "sovereign" in LLM's announcement is the declaration of defeat.

This thing is worse than a Qwen3.8 27B.

Looks interesting. I just went to download from HF, but they only have fp16 which won't fit on my Mac.

Good to see Europe adding toe what Mistral is doing. +100

loading story #49947703
"Languages: German and English" – this is odd. That means their dataset is limited. In my understanding, frontier models are trained on multilingual datasets and can combine knowledge no matter what language it was written in.
loading story #49947364
loading story #49947108
the sovereignty topic needs more attention in general so great to see. self-hosting the model is one piece of sovereignty, but how do we handle the rest of the agent stack - embeddings, retrieval, memory, etc. Has anyone put together a practical agent stack that's 100% sovereign, where they control it all?
For those sick of “Pareto frontier” talk, just shorthand it as “it’s the best at some very particular thing”. Obviously that one thing/tradeoff it’s good at may not necessarily be compelling, but it is either a loose sign of quality, or a sign that they’ve chased some tiny edge into the ground.

I’ll be curious to see which it becomes in the next year - nba “very narrow record”, or a sign you can hang with the big boys.

loading story #49947937
Was expecting that a "sovereign" AI model would at least use their own sovereign language (German) on the website as one of the options. Anyways, all the best and happy reunification day.
loading story #49944198
> The second was to rephrase German documents we already had. An LLM rewrites an organic German document in the style of an encyclopedia entry, a Q&A dialogue or a text passage, preserving its content.

"an LLM" -- does that mean they are effectively learning from that LLM the German encyclopedic style? makes me wonder which LLM and how that is really sovereign.

loading story #49947037
I wish nothing but luck for an EU model, but:

> intellectual-property safety

My suspicion is that you simply can't build an even slightly competitive model without liberally stealing your training data, in 2026, as much as I'd like it to be otherwise. You can get to the point that I suspect most of the frontier labs are at, where you've laundered the initially stolen data through the creation of huge amounts of derivative synthetic data, but still. Anyone who isn't comfortable stealing their training data is bringing a knife to a gun fight, and is going to die a noble but inevitable death.

loading story #49944412
loading story #49944700
loading story #49945070
loading story #49944473
loading story #49944463
loading story #49944439
loading story #49946618
i wonder if the custom tokenizer is better in practice, the examples look interesting though
3.5b active params sounds cheap until you remember all 78b still has to fit in memory. curious what the smallest practical self-hosted setup looks like for german docs.
The article talks about what setup is needed.
The benchmarks are impressive given the problem space they are working in.
loading story #49944096
German here. We are cheering for Mistral, which is making some good moves before the year is over, and Black Forest Labs for non-coding. But that's about it.
loading story #49944112
Im surprised by how well this works. What is the difference from this and union alpha (other than the fact that it is open weights)?
loading story #49945744
I wonder if we can start having LLM distros: community led distributed training runs with periodic releases, open weights, FOSS code, the whole shebang. Maybe the public training sets can reach a level where an LLM trained on them can be good enough for most things, such as web search and aggregation, coding, etc.

I wonder how far we are from this. How far are we from LLM's Debian moment?

Is this a truely open source model or also open weights?
> A bigger dense model beats it. Qwen3.8 27B ...

How is a 27B dense model bigger than a 78B MoE?

loading story #49944185
I got distracted by that scroll-wheel UI component on the page. Neat!
loading story #49944246
At least Germany is moving smarter than UK government...
loading story #49944306
Nice to see public goods in this space.
calling qwen 27B a bigger model.. I don't know man. My vram says otherwise.
{"deleted":true,"id":49944116,"parent":49943034,"time":1791034372,"type":"comment"}
Interesting they recommended high end software without considering quant 4 or 8 and still used A3B which should give good throughput on cheap hardware.

If they can follow Qwen3.8-Flash-Next, the could draft off the huge reduction in VRAM requirements.

loading story #49946768
The Sega 32X game??
loading story #49945725
It still steals my IP without attribution. Now we have state sanctioned sovereign theft instead of foreign theft.
Aleph Alpha is just a sad joke by now. The talent isn't there anymore, they never managed to catch up to the other labs, failed to deliver on several projects, and by now are just a cash grab for the investors.
loading story #49944080
loading story #49944379
> It knows less from memory, Multi-turn tool calling is weaker, It’s not the best coding agent

Then what does it good at? Sending faxes?

loading story #49943877
loading story #49944226
loading story #49943840
loading story #49943872
loading story #49944039
loading story #49943866
loading story #49943989
The ignorant, hostile, negative, and, frankly, kinda racist comments here really are just a sad showing for the currently online crowd.

But anyway. I think the main oversight when dismissing this is that not every use-case is coding a SV-style startup app. That market is quite saturated, so it would make sense to create something locally for the use-cases currently underserved by LLMs.

We will probably learn more about what this can really do once quants become available that can be run by people without an SV salary (and the biases that come with that).

does it have GDPR compliance?
Is Germany sovereign? It outsourced its defence onto the USA. Recently Trump wanted more diesel; Germany insta-submitted, also because oddly enough Macron submitted before Germany (Macron is suspicious). Before that, Leyen committed to insta-submission with a deal that made europeans poorer (and perhaps Leyen benefits from that). Canada shows the way. Many of the smaller countries in the EU too, such as Netherlands, Denmark, Finland, to some extent Sweden as well. Every time I read "sovereign" here I have to object. Nothing is sovereign here. The whole hardware is definitely not sovereign. Perhaps some of the software is, but that's about it. Plus, who gets all the data? The big US mega-corporations sniff non-stop. Remember how Facebook sniffed Libgen and Anna's Archive dry etc..., then suddenly libgen went down. The US corporations act as huge global leeches on every step of the stair. And lobbyists benefit from this too.
> 4. It thinks in German

This means that it’s always on time, it uses acronyms for everything and when there’s a decision to be made, it sets up a committee.

loading story #49944229
loading story #49945212
loading story #49944606
loading story #49944352
loading story #49944695
loading story #49944323
loading story #49944107
loading story #49944446
loading story #49946157
loading story #49947615