Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression
Signal
72
Hype
18
In three linesStructural pruning framework for Mixture-of-Experts models operating at channel level rather than expert level. Attribution-based method reformulates pruning as channel-score coverage maximization. Experiments on DeepSeek and Qwen models achieve 50% structured pruning with 4-bit quantization, 5.27× memory reduction on Qwen3-30B-A3B.Read source
Your take?
Summary generated by Claude — human-verified