I implemented KVarN in my llama.cpp fork and ran KLD benchmarks. It's promising!
Signal
72
Hype
35
In three linesKVarN (Huawei's KV-cache quantization) implemented in llama.cpp fork. 3-5× KV cache compression with speed-up. KLD benchmarks show q5 quality at 4-bit, q4 quality at 3.5-bit. Available in BeeLlama.cpp v0.3.2 with --cache-type-k/v kvarn4 flags.Read source
Your take?
Summary generated by Claude — human-verified