We propose MEVE, a novel multimodal large language model (MLLM) framework that integrates multiple vision experts with a powerful language backbone. MEVE effectively responds to diverse visual inputs, such as natural images, text documents, and diagrams, by using a dynamic expert-routing mechanism that selectively engages specific vision encoders for each modality. We subsequently fuse these complementary features into a unified representation for enhanced downstream reasoning. This paper addresses the gap in efficient and adaptive vision encoding within MLLMs, which is crucial for handling the diverse nature of visual information. MEVE achieves state-of-the-art or competitive performance against open-source MLLMs across a wide range of benchmarks, including general VQA, text-oriented VQA, chart analysis, and mathematical visual tasks, while maintaining a parameter-efficient model design. Our ablation studies and visualization analyses demonstrate that the multi-expert paradigm excels at capturing both local and global image features, leading to strong performance across various tasks. MEVE also exhibits robust zero-shot generalization, effectively adapting to new domains without extensive fine-tuning.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Multi-Expert Vision Adapter Framework for End-to-End Vision-Language Integration

  • Xiaoyu Kang,
  • Zhixin Shi,
  • Degang Sun,
  • Tengfan Weng

摘要

We propose MEVE, a novel multimodal large language model (MLLM) framework that integrates multiple vision experts with a powerful language backbone. MEVE effectively responds to diverse visual inputs, such as natural images, text documents, and diagrams, by using a dynamic expert-routing mechanism that selectively engages specific vision encoders for each modality. We subsequently fuse these complementary features into a unified representation for enhanced downstream reasoning. This paper addresses the gap in efficient and adaptive vision encoding within MLLMs, which is crucial for handling the diverse nature of visual information. MEVE achieves state-of-the-art or competitive performance against open-source MLLMs across a wide range of benchmarks, including general VQA, text-oriented VQA, chart analysis, and mathematical visual tasks, while maintaining a parameter-efficient model design. Our ablation studies and visualization analyses demonstrate that the multi-expert paradigm excels at capturing both local and global image features, leading to strong performance across various tasks. MEVE also exhibits robust zero-shot generalization, effectively adapting to new domains without extensive fine-tuning.