It is pretty nifty. I spend some time over last weekend implementing fused TQ to allow for 1m context lengths on a 128 gb MacBook M5 Max when using Qwen 3.8 flash next (https://github.com/antirez/ds4/pull/1115 if you are interested). If I get bored I might port over the Metal kernels from oMLX -- the speed increase they have for the v0.7.0 release is amazeballs.
loading story #49939337