Despite the success of prompt-based vision-language models (VLMs) in general visual tasks, their direct application to medical image classification remains suboptimal due to two key challenges: strict data privacy constraints and limited annotated medical datasets. To address this, we propose two novel strategies for osteoporosis classification: (1) A text cue-assisted training strategy that replaces generic class names with detailed, LLM-generated visual descriptions to improve feature alignment; and (2) A Text-Image Feature Fusion Module (TIFM) that integrates visual and textual cues in high-dimensional space, enabling deeper multimodal synergy. Experiments show our approach significantly boosts accuracy and robustness in osteoporosis classification.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Osteoporosis Image Classification Algorithm Based on Textual Prompts

  • Yang Peng,
  • Yuchen Chen,
  • Jun Liang,
  • Ruihua Nie

摘要

Despite the success of prompt-based vision-language models (VLMs) in general visual tasks, their direct application to medical image classification remains suboptimal due to two key challenges: strict data privacy constraints and limited annotated medical datasets. To address this, we propose two novel strategies for osteoporosis classification: (1) A text cue-assisted training strategy that replaces generic class names with detailed, LLM-generated visual descriptions to improve feature alignment; and (2) A Text-Image Feature Fusion Module (TIFM) that integrates visual and textual cues in high-dimensional space, enabling deeper multimodal synergy. Experiments show our approach significantly boosts accuracy and robustness in osteoporosis classification.