REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
Abstract
Router-weighted Expert Activation Pruning (REAP) outperforms expert merging for generative tasks in SMoE models, achieving near-lossless compression.
Sparsely-activated Mixture-of-Experts (SMoE) models offer efficient pre-training and low latency but their large parameter counts create significant memory overhead, motivating research into expert compression. Contrary to recent findings favouring expert merging on discriminative benchmarks, we demonstrate that expert pruning is a superior strategy for generative tasks. We prove that merging introduces an irreducible error by causing a "functional subspace collapse", due to the loss of the router's independent, input-dependent control over experts. Leveraging this insight, we propose Router-weighted Expert Activation Pruning (REAP), a novel pruning criterion that considers both router gate-values and expert activation norms. Across a diverse set of SMoE models ranging from 20B to 1T parameters, REAP consistently outperforms merging and other pruning methods on generative benchmarks, especially at 50% compression. Notably, our method achieves near-lossless compression on code generation and tool-calling tasks with Qwen3-Coder-480B and Kimi-K2, even after pruning 50% of experts.
Community
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment (2026)
- Residual Sparsification via Output Importance for Compressing Mixture-of-Experts LLMs (2026)
- Shape Mutating Expert Compression:LorExperts and BTExperts (2026)
- When Load-Balancing Goes Too Far: Expert Pruning in Over-Dispersed Mixture-of-Experts Models (2026)
- Benchmarking Composable Compression Techniques in Mixture-of-Experts LLMs (2026)
- Half the Experts, All the Code: One-Shot Domain Pruning of Mixture-of-Experts LLMs for Coding (2026)
- Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2510.13999 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 294
cerebras/GLM-4.5-Air-REAP-82B-A12B
Datasets citing this paper 20
0xSero/qwen35-reap-layerwise-observations
0xSero/glm-5.3-reap-fidelity-study
Spaces citing this paper 0
No Space linking this paper