KVarN: Variance-Normalized KV-Cache Quantization [R]
KVarN is a KV-Cache quantization method combining Hadamard rotations with variance-normalization on K and V matrices. Achieves 3-4x compression with 0-1% accuracy drop on AIME24 and speedup over fp16 baseline in vLLM. Optimized for decode-heavy settings (reasoning, code-gen, agents).