A Multi-Expert Vision Adapter Framework for End-to-End Vision-Language Integration
摘要
We propose MEVE, a novel multimodal large language model (MLLM) framework that integrates multiple vision experts with a powerful language backbone. MEVE effectively responds to diverse visual inputs, such as natural images, text documents, and diagrams, by using a dynamic expert-routing mechanism that selectively engages specific vision encoders for each modality. We subsequently fuse these complementary features into a unified representation for enhanced downstream reasoning. This paper addresses the gap in efficient and adaptive vision encoding within MLLMs, which is crucial for handling the diverse nature of visual information. MEVE achieves state-of-the-art or competitive performance against open-source MLLMs across a wide range of benchmarks, including general VQA, text-oriented VQA, chart analysis, and mathematical visual tasks, while maintaining a parameter-efficient model design. Our ablation studies and visualization analyses demonstrate that the multi-expert paradigm excels at capturing both local and global image features, leading to strong performance across various tasks. MEVE also exhibits robust zero-shot generalization, effectively adapting to new domains without extensive fine-tuning.