Back to feed
arXiv cs.LG·

ScaleSweep: Accurate NVFP4 Post-Training Quantization of LLMs via Block Scale Initialization

Signal
72
Hype
15
In three linesScaleSweep optimizes NVFP4 quantization (hardware-supported FP4 4-bit format) of LLMs through block scale candidate sweeping. Theoretically bounded to reduce search space, the method preserves >93% full-precision performance on Llama and Qwen under end-to-end quantization (weights, activations, KV cache, query states).
Read source
Your take?
LlamaQwenBenchmarksPapers

Summary generated by Claude — human-verified