Hacker News new | past | comments | ask | show | jobs | submit
That's very interesting, I wonder if this applies also to models quantized to ints like (-1,0,1), and I wonder if the labs could maintain frontier performance if they removed floating points but arbitrarily scaled up the parameters.

Edit: the Thinking Machines article in the other comment gets into this a bit