Training generalist models capable of addressing diverse medical imaging tasks is challenging, and often requires large datasets with detailed annotations. This is particularly problematic in healthcare, where expert-labelled data are limited and costly to obtain. To address this, we present MOSAIC, a self-supervised transformer-based architecture designed for efficient learning in multimodal medical-diagnosis tasks. MOSAIC leverages advances in self-supervised learning by combining image trunks, text embeddings, multimodal fusion, and task-specific heads to substantially outperform existing approaches. Experiments on benchmark datasets, including CheXpert, RSNA Pneumonia and MIMIC-CXR, demonstrated an overall improvement in lesion detection, achieving a 91.87% mAUC, surpassing state-of-the-art models. The robust design of MOSAIC enables it to generalize across diverse clinical scenarios with minimal reliance on annotated datasets, highlighting its potential to transform medical diagnostics by improving the efficiency and accuracy in real-world applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MOSAIC: A Multimodal Self-supervised Transformer for Medical Image Diagnosis

  • Farah Oubelkas,
  • Zakaria Benhaili,
  • Lahcen Moumoun,
  • Abdellah Jamali

摘要

Training generalist models capable of addressing diverse medical imaging tasks is challenging, and often requires large datasets with detailed annotations. This is particularly problematic in healthcare, where expert-labelled data are limited and costly to obtain. To address this, we present MOSAIC, a self-supervised transformer-based architecture designed for efficient learning in multimodal medical-diagnosis tasks. MOSAIC leverages advances in self-supervised learning by combining image trunks, text embeddings, multimodal fusion, and task-specific heads to substantially outperform existing approaches. Experiments on benchmark datasets, including CheXpert, RSNA Pneumonia and MIMIC-CXR, demonstrated an overall improvement in lesion detection, achieving a 91.87% mAUC, surpassing state-of-the-art models. The robust design of MOSAIC enables it to generalize across diverse clinical scenarios with minimal reliance on annotated datasets, highlighting its potential to transform medical diagnostics by improving the efficiency and accuracy in real-world applications.