Minimizing the Hidden Cost of Scales: Graph-Guided Ultra-Low-Bit Quantization for Large Language Models
Signal
78
Hype
15
In three linesSAGE-PTQ, an ultra-low-bit quantization method for LLMs, reduces hidden scaling cost by separating salient and non-salient weights via distributional statistics and graph modeling. On LLaMA-3-8B: 6.74 WikiText2 perplexity vs 55.8 for BiLLM, using 50% less GPU memory. On LLaMA-2-70B: 1.5x faster decoding on NVIDIA L40.Read source
Your take?
Summary generated by Claude — human-verified