Back to feed
arXiv cs.LG·

LLM Compression with Jointly Optimizing Architectural and Quantization choices

Signal
78
Hype
22
In three linesDifferentiable NAS framework for LLM compression jointly optimizing architecture and mixed-precision quantization of linear layers. Results: 1.4x faster inference than sequential NAS-then-quantization baseline, or 6% higher average accuracy across seven reasoning tasks at equivalent latency.
Read source
Your take?
ReasoningBenchmarksFine-tuning

Summary generated by Claude — human-verified