Multimodal Osteoporosis Image Classification Algorithm Based on Textual Prompts
摘要
Despite the success of prompt-based vision-language models (VLMs) in general visual tasks, their direct application to medical image classification remains suboptimal due to two key challenges: strict data privacy constraints and limited annotated medical datasets. To address this, we propose two novel strategies for osteoporosis classification: (1) A text cue-assisted training strategy that replaces generic class names with detailed, LLM-generated visual descriptions to improve feature alignment; and (2) A Text-Image Feature Fusion Module (TIFM) that integrates visual and textual cues in high-dimensional space, enabling deeper multimodal synergy. Experiments show our approach significantly boosts accuracy and robustness in osteoporosis classification.