i definitely will be doing some drop of some faster attention kernels in the next few weeks.
like i can do all sorts of memory layout of tensors/matrices etc tricks that if you dont have the abstractions for it would just never happen. so i can optimize the kernel flops