70 points brainless 6 hours ago 12 comments
badatnames 3 hours ago | parent
nico 3 hours ago | parent
big-chungus4 3 hours ago | parent
GaggiX 3 hours ago | parent
tcdent 2 hours ago | parent
So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.
Trellis quantization is a far more approachable enhancement right now, but it doesn't cross the 1-bit barrier (and perhaps doesn't intend to).
nbutton762 2 hours ago | parent
augment_me 2 hours ago | parent
I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance
bArray 2 hours ago | parent
cpldcpu 1 hour ago | parent
It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.
²As to why, I have seen few explanations. But the empirical evidence is there.
tempoponet 45 minutes ago | parent
They would be better served with smaller models that can reliably call tools and generate structured outputs without looping or totally hallucinating. These projects exist, but aren't getting amplified.
gobdovan 36 minutes ago | parent
Name them
ted_dunning 4 minutes ago | parent
https://arxiv.org/abs/2603.00042
It is true that block-headed quantization of everything doesn't work. But, as I am sure will read, if you remove the spiky parts of the parameter values, the residue can be dramatically quantized and compressed while retaining performance.
This is a multi-modal sort of compression where you use different techniques for different phenomena. Simply compression all of the weights, each in isolation, ignores the benefits of compressing them collectively.