Hacker News new | past | comments | ask | show | jobs | submit
I would hope that Qwen 3.8 is better. It's been 4 months, and we've seen almost no progress in this space.

As people have called out, Glimmer appears to be a trade-off rather than a clear winner.

And from what I've been reading, no one is expecting Qwen 3.8's model in this space to be a clear winner, but just slightly and marginally better.

That's a little concerning as DeepSeek v4 Flash proved at it larger sizes there's a ton of room left to compress knowledge.

If we don't see something that's substantially better in the ~30B param space soon - it would appear we might've saturated that size with knowledge.

> If we don't see something that's substantially better in the ~30B param space soon - it would appear we might've saturated that size with knowledge.

I wouldn't be quite so pessimistic. We may have saturated the current approach, but I think there's a lot still left in terms of compression, attention, active parameters, caching etc. etc.

I don’t think four months without a major breakthrough is cause to abandon all hope just yet. ;) The wild pace of LLM development is highly atypical, and we’re still in the ‘initial rush’ phase of development.

For contrast, the Newcomen steam engine (widely considered the first commercially useful engine) was used for over 60 years before the next major improvements. Now, 300 years later, we’re still finding ways to significantly improve heat engines.

> For contrast, the Newcomen steam engine (widely considered the first commercially useful engine) was used for over 60 years before the next major improvements. Now, 300 years later, we’re still finding ways to significantly improve heat engines.

Off-topic, but I stumbled upon the first Newcomen engine imported into Australia in a museum in Sydney and I was unexpectedly charmed (not an Engine Guy). It's large, but nothing like the awe of "mega-engineering", it's crude, but it clearly has such amazing utility (when compared to a reality without it) and it changed the world

I honestly expect that major advances in the open 30B dense space will take about a year, but expect incremental advances every couple of months from different developers in the meantime.

Qwen 3.6 27B was already a massive gift to smaller homelabs around the world; anything more is just a delightful surprise.