70 points brainless 6 hours ago 12 comments

badatnames 3 hours ago | parent

Their paper shows this comes with huge quality loss, but that doesn't make it a negative result by any means

nico 3 hours ago | parent

Has anyone tried this on apple silicon M1-5? Any benchmarks/comps?

big-chungus4 3 hours ago | parent

Can this produce a useful model? So far 1 bit quants have been less useful than smaller models that use the same memory

GaggiX 3 hours ago | parent

I recently found this 1.58-bit model for ASR and it's surprising good (and very fast), that being said it's not a LLM.

https://huggingface.co/moondream/parakeet-redux

tcdent 2 hours ago | parent

I don't think it's trying to be a useful implementation, but the significance they do provide is that they are able to improve on the relative loss at lower quants.

So, not something anyone would want to run currently, but an indicator that there is still more to squeeze out of lower precisions.

Trellis quantization is a far more approachable enhancement right now, but it doesn't cross the 1-bit barrier (and perhaps doesn't intend to).

nbutton762 2 hours ago | parent

Thought this was going to be on the original Little Bit paper, always nice to find out about a surprise sequel!

augment_me 2 hours ago | parent

Perf goes from 80% to 47% on Wikitext-2. Also no comparisons to FP4 solutions that are able to maintain or exceed perf on the same dataset 80% perf with a 4.25-4.5 big budget.

I think more meaningful thing here would be a hybrid solution that went down to sub-bit representations when the informational representation does not need it (for example later layers) that still maintains task performance

bArray 2 hours ago | parent

Has anybody tested this? Are there any available computed models to test?

cpldcpu 1 hour ago | parent

I understand the obsession with low bit quantization, but it is empirically quite evident that it is not possible to compress models to less than 4 bit per weight without severe loss of capabilities².

It may be nice as an experiment, but it is obviously a very inefficient route for model training: spending all the flops on a saturated model only to prune its capabilties.

²As to why, I have seen few explanations. But the empirical evidence is there.

tempoponet 45 minutes ago | parent

Influencers have latched onto the pitch that everyday people can run frontier models on an 8gb GPU while sticking it to the labs. There's a large audience of people who haven't had the hardware to test larger quants to see the difference.

They would be better served with smaller models that can reliably call tools and generate structured outputs without looping or totally hallucinating. These projects exist, but aren't getting amplified.

gobdovan 36 minutes ago | parent

> These projects exist, but aren't getting amplified.

Name them

ted_dunning 4 minutes ago | parent

Counter evidence:

https://arxiv.org/abs/2603.00042

It is true that block-headed quantization of everything doesn't work. But, as I am sure will read, if you remove the spiky parts of the parameter values, the residue can be dramatically quantized and compressed while retaining performance.

This is a multi-modal sort of compression where you use different techniques for different phenomena. Simply compression all of the weights, each in isolation, ignores the benefits of compressing them collectively.