Knowledge-Guided Structured Pruning for Multimodal Language Models
摘要
Multimodal Large Language Models (MLLMs) demonstrate strong reasoning capabilities by integrating visual and textual information with external knowledge, but their computational demands limit practical deployment. We propose a knowledge-guided structured pruning approach that leverages external knowledge graphs to inform compression decisions. Our method achieves a favorable trade-off between model size and performance: at 30% compression, we retain 89.2% of original accuracy while reducing inference time by 1.4x. Experiments on knowledge-grounded visual question answering show modest improvements over magnitude-based pruning baselines, though we observe increased hallucination rates typical of compressed models. Our approach provides a practical framework for deploying MLLMs in resource-constrained environments while maintaining reasonable performance on knowledge-intensive tasks.