Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

100% Positive

Analyzed from 376 words in the discussion.

Trending Topics

#bits#per#quantization#ternary#llms#bit#weights#better#model#fast

Discussion (13 Comments)Read Original on HackerNews

infogulchabout 1 hour ago
So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.

If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.

kadushkaabout 1 hour ago
By “work out” you mean no accuracy degradation? That’s a big ask - currently we can barely quantize to dynamic fp4 with small block size - still not completely lossless on all benchmarks.
Vetch15 minutes ago
QAT, which bitnet training is a form of, helps a ton in preserving accuracy at such low bits per parameter. There are also better quantization approaches that try to preserve the most sensitive weights† but are computationally expensive and so not typically done. Another complementary option is, if the model is fast enough, we should be able to push up correctness by self-consistency voting at close to T=1. Smart/fast Zero-shot classifiers like the recent Jev could help with aggregation across answers too, extending applicability.

†Every paper I've read estimates the average information content of transformer LLMs at about 3-4 bits per parameter. Curiously, biological synapses are also estimated to be about 4-5 bits per synapse, possibly a bit lower.

montroser37 minutes ago
Well, you could train directly at this bitrate.
yalok26 minutes ago
sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.

And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...

0. https://arxiv.org/pdf/2402.17764 - The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits

om8about 2 hours ago
Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
janalsncmabout 2 hours ago
PTQ and vector quantization aren’t used for this because part of the point of ternary LLMs is to make them faster. In a ternary LLM every weight is an add, subtract, or no-op so it is fast on CPU.

If you’re just using a code book to reconstruct a f16 model the only savings you can get are in sending it over the wire.

mitxelaabout 1 hour ago
which is important though since sending it across the wire over and over and over is actually the main bottleneck.
om8about 2 hours ago
If you want sub-2 bit llm, get one that’s already trained in higher precision, and compress it with something like YAQA/QTIP with finetuning or PV-tuning + AQLM/HIGGS
wgdabout 1 hour ago
Only a presence bitmap? If we're contemplating packing schemes I'm tempted to write a paper that uses arithmetic coding to squeeze out a few more centi-bits.
plqbfbvabout 2 hours ago
Very interesting, I was just exploring this to hopefully fit one of the latest quantized models in 16GB of VRAM.
NooneAtAll3about 2 hours ago
This is the only time "1.58 bit" phrase makes more sense than "1 trit"

Who knew that if you actually look at information entropy you can pack stuff better!

Kevcmkabout 2 hours ago
Woah. Good science.