Do Transformers Need Three Projections? Systematic Study of QKV Variants
Signal
78
Hype
15
In three linesSystematic study of QKV variants in transformers. Authors test three projection sharing constraints (Q-K=V, Q=K-V, Q=K=V) on synthetic tasks, vision, and language models (300M-1.2B parameters). Q-K=V reduces KV cache by 50% with 3.1% perplexity degradation. Combined with GQA/MQA, achieves 87.5-96.9% cache reduction for edge inference.Read source
Your take?
Summary generated by Claude — human-verified