Hacker News new | past | comments | ask | show | jobs | submit
Architecture thread! Afaict they continue to use gated attention + delta net, which was also adopted+adapted by K3, but im surprised theres no improvements to the residual stream (deepseek are using manifold hyper-connections, kimi have attention residuals) ?

Perf improvements seem to all come from training?

As was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training