SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs
Signal
78
Hype
15
In three linesSHAPE is a pruning method for sparse MoE models that evaluates experts via observed coalitions rather than independently. Using Shapley attribution over top-k routings, it identifies experts essential to collaborations. Tested on Qwen3-30B-A3B, GPT-OSS-20B, and DeepSeek-V2-Lite, SHAPE maintains accuracy with 20-40% expert pruning without retraining and reduces peak GPU memory.Read source
Your take?
Summary generated by Claude — human-verified