Recent developments in large language models (LLMs) have demonstrated significant progress in both language understanding and generation. In particular, GPT-4, a model capable of processing both text and images, presents challenges in understanding its internal operations. To investigate this, we propose VisionLingua, a model that integrates a fixed visual encoder with the Vicuna language model through a projection layer. Our study shows that aligning a visual module with a state-of-the-art language model unlocks a broad range of capabilities, including detailed image description, conversion of hand-drawn sketches into digital formats, interpretation of complex visual phenomena, and detection and resolution of issues in images. In early experiments, VisionLingua’s performance was hindered by the limited size of the training dataset, resulting in inconsistent text outputs. To improve performance, we implemented a second training phase using an expanded dataset with detailed image descriptions, which significantly enhanced the model’s ability to produce coherent and accurate text outputs.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

VisionLingua: Empowering Multimodal Understanding with Cutting-Edge Language Models

  • Deepika Kamboj,
  • Kamal Kumar Gola,
  • Shikha Arya

摘要

Recent developments in large language models (LLMs) have demonstrated significant progress in both language understanding and generation. In particular, GPT-4, a model capable of processing both text and images, presents challenges in understanding its internal operations. To investigate this, we propose VisionLingua, a model that integrates a fixed visual encoder with the Vicuna language model through a projection layer. Our study shows that aligning a visual module with a state-of-the-art language model unlocks a broad range of capabilities, including detailed image description, conversion of hand-drawn sketches into digital formats, interpretation of complex visual phenomena, and detection and resolution of issues in images. In early experiments, VisionLingua’s performance was hindered by the limited size of the training dataset, resulting in inconsistent text outputs. To improve performance, we implemented a second training phase using an expanded dataset with detailed image descriptions, which significantly enhanced the model’s ability to produce coherent and accurate text outputs.