Back to feed
Reddit r/LocalLLaMA·

Why doesn’t 4-bit GPTQ wreck a model’s perplexity? I derived the compensation math from scratch

Signal
72
Hype
15
In three linesDeep mathematical breakdown of GPTQ: 4-bit quantization preserves perplexity by treating weights as correlated. Author derives the update rule using Lagrange multipliers, explains 1% Hessian dampening, Cholesky decomposition over raw inverse, and C-contiguous memory optimization.
Read source
Your take?
Fine-tuningOpen source

Summary generated by Claude — human-verified