Paper surfaces on Hacker News frontpage
2 Sep 16 4:59 PM · 11d ago · 1 article · 3 posts · 3 sources · development 2 of 2
The arXiv preprint reaches Hacker News, gaining 141 points and sparking discussion among the community about implications for model deployment and compression techniques.
Evangelos Georganas Researcher, paper authorAlexander Heinecke Researcher, paper authorPradeep Dubey Researcher, paper author
The whole story articlesposts the bright band is this development · numbered dots are the others · click one to jump
Reported in the same hours no headline names this development itself — these 1 claim were published in its stretch
-
first by HN Best, 11d ago · also HN Frontpage, arXiv cs.AI
What people said 8 voices · verbatim
-
sounds like a perfect fit for ASIC-optimized models (where matrix ops could be supported directly in BITCOS format, potentially) & achieving record power efficiency for on-device inference.And it looks like per [0], a model needs only ~30% more weights to be at comparable quality, if quantization-aware training is done...0…
-
> We measure the actual symbol distribution of 29 ternary LLM models and find that zeros account for up to 51.5% of all weights. Motivated by this finding, we introduce BITCOS, a simple distribution-adaptive layoutI honestly assumed that's how they already work. I have to admit that I even explained it like that to a friend. Why on earth wouldn't…
-
So they get down from 1.58 to 1.48 bits per weight by exploiting the fact that actual weights in practice are 0 51% of the time. Neat.If ternary llms work out and are baked into hardware as custom silicon I bet they'll be shockingly efficient.
-
Only a presence bitmap? If we're contemplating packing schemes I'm tempted to write a paper that uses arithmetic coding to squeeze out a few more centi-bits.
-
I'm surprised that a variable length encoding like this is usable directly as in memory format and not just as storage/transfer format.
-
So this compression is only pertinent to the LLM file format? In memory it'd have to be expanded into the 1.58-bit form - 5 trits per byte.
-
Ternary quantization does not make any sense. Vector quantization and trellis based methods are better in this region for PTQ.
-
Pushing past log2(3) for real. This could drastically shrink LLMs for embedded systems, making them truly portable.
All 2 developments of Researchers break 1.58-bit barrier for ternary LLM… →
NewswiresHacker NewsMastodon