Breaking the 1.58-bit Barrier for Ternary LLMs
https://arxiv.org/abs/2609.16338> We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layout
I honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't you design it like that from the start (talking about the adaptive, not the measure part; just sacrifice a few bits to clarify your encoding and save a ton of bits)?
loading story #49734352
loading story #49737749
loading story #49734042
loading story #49734840
loading story #49734583
loading story #49735618
loading story #49735097
loading story #49734027
loading story #49734021
loading story #49736892
loading story #49733795
loading story #49735278
loading story #49734191
loading story #49733869