Back to feed
arXiv cs.LG·

MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs

Signal
78
Hype
15
In three linesMODE is an expert-level mixed-precision quantization framework for MoE multimodal LLMs. It decomposes expert selection frequency by modality (vision/text) and filters redundant vision tokens to correct estimation biases. Results: <2.9% performance loss at W3A16.
Read source
Your take?
VisionBenchmarksPapers

Summary generated by Claude — human-verified