<p>In clinical practice, it is time-consuming for experts to write reports. Existing methods focus more on abnormal regions when there is an imbalance between normal and abnormal samples and fail to adequately capture the correlations in medical observations. The transformer exhibits performance degradation when processing long sequences. To address these challenges, we propose a hybrid Transformer architecture with Mamba for medical image report generation. The architecture integrates the strengths of both Transformer and Mamba: it leverages global and local image information extracted by visual extractors and introduces prior knowledge in the method of streaming, jointly assisting the encoder to focus more on abnormal regions for mitigating data bias. Additionally, Mamba is integrated into the decoder to enhance multi-modal information processing capabilities and enhance the ability to capture long sequence dependencies and be context-aware. On the one hand, the experimental results on two public datasets: IU X-Ray, COV-CTR show that the proposed model has significant improvements on evaluation metrics, especially in the BLEU scores on the IU X-Ray, surpassing the state-of-the-art models, indicating an obvious advantage in phrase overlap; on the other hand, the experimental results on the Bladder Whole Slide Datasets show that the proposed model has a certain generalization ability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A hybrid transformer architecture with Mamba for medical image report generation

  • Ruipeng Yang,
  • Fang Li

摘要

In clinical practice, it is time-consuming for experts to write reports. Existing methods focus more on abnormal regions when there is an imbalance between normal and abnormal samples and fail to adequately capture the correlations in medical observations. The transformer exhibits performance degradation when processing long sequences. To address these challenges, we propose a hybrid Transformer architecture with Mamba for medical image report generation. The architecture integrates the strengths of both Transformer and Mamba: it leverages global and local image information extracted by visual extractors and introduces prior knowledge in the method of streaming, jointly assisting the encoder to focus more on abnormal regions for mitigating data bias. Additionally, Mamba is integrated into the decoder to enhance multi-modal information processing capabilities and enhance the ability to capture long sequence dependencies and be context-aware. On the one hand, the experimental results on two public datasets: IU X-Ray, COV-CTR show that the proposed model has significant improvements on evaluation metrics, especially in the BLEU scores on the IU X-Ray, surpassing the state-of-the-art models, indicating an obvious advantage in phrase overlap; on the other hand, the experimental results on the Bladder Whole Slide Datasets show that the proposed model has a certain generalization ability.