<p>This study addressed the challenge of limited labeled medical imaging data and weak cross-modal alignment between clinical narratives and visual content by proposing Text-2-Diagnosis, a transformer-based text-to-image framework for medical diagnostic imaging. The proposed approach combined self-attention and cross-attention mechanisms with CLIP-guided image-text alignment to improve semantic consistency between textual prompts and generated images while preserving anatomical structure. We evaluated computed tomography (CT) and ultrasound subsets of RadImageNet and the multi-modal ROCO dataset using Fréchet Inception Distance (FID) for perceptual realism and Dice for pixel-level fidelity. Across GAN baselines, Binary Cross-Entropy (BCE) yielded the strongest FID on every dataset: CT 7.33, Ultrasound 0.59, ROCO 5.43. Under the same setting, peak Dice was 0.94 for CT, 0.93 for Ultrasound, and 0.92 for ROCO. With relativistic loss, CT reached FID 27.87/Dice 0.83; Ultrasound achieved FID 34.09/Dice 0.31; ROCO attained FID 48.29/Dice 0.68. Text-2-Diagnosis transformer-based model that integrates self- and cross-attention achieved their strongest results on CT but did not surpass the BCE-trained GAN and struggled on Ultrasound and ROCO. This underscored modality and dataset challenges for attention-centric architectures. We conclude with targeted directions to further stabilize and adapt generative models, with the long-term goal of supporting clinically relevant image synthesis research; however, clinical utility remains to be established through expert and task-based evaluation.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Text-2-Diagnosis: an exploration of text-to-image generation in medical diagnostic imaging

  • Arwa H. Alshanbari,
  • Salha M. Alzahrani

摘要

This study addressed the challenge of limited labeled medical imaging data and weak cross-modal alignment between clinical narratives and visual content by proposing Text-2-Diagnosis, a transformer-based text-to-image framework for medical diagnostic imaging. The proposed approach combined self-attention and cross-attention mechanisms with CLIP-guided image-text alignment to improve semantic consistency between textual prompts and generated images while preserving anatomical structure. We evaluated computed tomography (CT) and ultrasound subsets of RadImageNet and the multi-modal ROCO dataset using Fréchet Inception Distance (FID) for perceptual realism and Dice for pixel-level fidelity. Across GAN baselines, Binary Cross-Entropy (BCE) yielded the strongest FID on every dataset: CT 7.33, Ultrasound 0.59, ROCO 5.43. Under the same setting, peak Dice was 0.94 for CT, 0.93 for Ultrasound, and 0.92 for ROCO. With relativistic loss, CT reached FID 27.87/Dice 0.83; Ultrasound achieved FID 34.09/Dice 0.31; ROCO attained FID 48.29/Dice 0.68. Text-2-Diagnosis transformer-based model that integrates self- and cross-attention achieved their strongest results on CT but did not surpass the BCE-trained GAN and struggled on Ultrasound and ROCO. This underscored modality and dataset challenges for attention-centric architectures. We conclude with targeted directions to further stabilize and adapt generative models, with the long-term goal of supporting clinically relevant image synthesis research; however, clinical utility remains to be established through expert and task-based evaluation.