I just tried to play around with it on my RTX pro 6000 setup, it spends way too many tokens on overthinking stuff even if it’s able to catch the correct approach
Its speed is pretty good on the other hand with only 3B active parameters I am getting around 170 tkn/s on fp8
> it spends way too many tokens on overthinking stuff
Yeah. It’s a German model.
loading story #49944024
loading story #49946314
loading story #49943998
Did you try lower reasoning modes as well?