PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.
If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.
loading story #49737156
loading story #49734510
loading story #49735414