Back to feed
arXiv cs.LG·

Less is MoE: Trimming Experts in Domain-Specialist Language Models

Signal
78
Hype
15
In three linesFisher-MoE proposes compressing Mixture-of-Experts models by targeting intermediate FFN dimensions rather than entire experts. On Qwen1.5-MoE, removing just 12 of 1.35M critical dimensions (identified via Fisher importance) preserves performance while reducing memory by ~45% and improving inference throughput by 21%.
Read source
Your take?
QwenBenchmarksFine-tuningInfrastructure

Summary generated by Claude — human-verified