<p>Achieving precise alignment between visual and semantic representations in multimodal large models is critical for advancing medical image analysis. This paper proposes a Multi-Dimensional Visual-Language Alignment (MDVLA) framework that leverages multi-dimensional image features and various forms of textual prompt representations to enhance visual–semantic alignment in medical images. Our method incorporates a Visual Enhancement Module (VEM) that ensures robust representation of critical medical patterns by leveraging multi-angle rotated image feature enhancement. Simultaneously, the Language Enhancement Module (LEM) leverages various forms of expert-screened medical prompts combined with learnable text embeddings to bridge the gap between medical image features and corresponding textual descriptions. Our framework relies on high-performance computing (HPC) resources to meet the intensive demands of large-scale medical data and multi-dimensional modeling. Extensive evaluations on Derm7pt and Pneumonia benchmarks demonstrate that MDVLA consistently outperforms state-of-the-art CLIP-based methods, achieving improvements of +0.49% and +1.26% in classification accuracy on Derm7pt and Pneumonia, respectively. The results validate the effectiveness of the proposed method in aligning visual and semantic features, enabling more accurate predictions and better interpretability in medical contexts. MDVLA is modular and imaging-domain agnostic, with VEM and LEM designed to generalize across modalities and tasks, enabling straightforward extension to multimodal and cross-task benchmarks.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing visual and semantic alignment of multimodal large models in medical images

  • Qifan Liu,
  • Yuxuan Fang,
  • Xiyuan Liu

摘要

Achieving precise alignment between visual and semantic representations in multimodal large models is critical for advancing medical image analysis. This paper proposes a Multi-Dimensional Visual-Language Alignment (MDVLA) framework that leverages multi-dimensional image features and various forms of textual prompt representations to enhance visual–semantic alignment in medical images. Our method incorporates a Visual Enhancement Module (VEM) that ensures robust representation of critical medical patterns by leveraging multi-angle rotated image feature enhancement. Simultaneously, the Language Enhancement Module (LEM) leverages various forms of expert-screened medical prompts combined with learnable text embeddings to bridge the gap between medical image features and corresponding textual descriptions. Our framework relies on high-performance computing (HPC) resources to meet the intensive demands of large-scale medical data and multi-dimensional modeling. Extensive evaluations on Derm7pt and Pneumonia benchmarks demonstrate that MDVLA consistently outperforms state-of-the-art CLIP-based methods, achieving improvements of +0.49% and +1.26% in classification accuracy on Derm7pt and Pneumonia, respectively. The results validate the effectiveness of the proposed method in aligning visual and semantic features, enabling more accurate predictions and better interpretability in medical contexts. MDVLA is modular and imaging-domain agnostic, with VEM and LEM designed to generalize across modalities and tasks, enabling straightforward extension to multimodal and cross-task benchmarks.