Artificial Intelligence (AI) is transforming the field of healthcare, which brings advancements to the diagnosis and treatment of various conditions, especially breast cancer. It is considered the most dangerous illness among women. Thus, AI technologies including machine learning, deep learning, and advanced data analytics offer promising solutions to this sickness. The primary goal of this study is to develop systems called BCCNetAttention (Breast Cancer Captioning with ResNet50 and Cross Attention) that can automatically interpret and describe the content of images in a way that is meaningful and coherent to humans. In more detail, this research employed the Convolutional Neural Network (CNN) model for feature extraction and the Transformer architecture with Future Masked Language Modeling (FMLM) to train the model in the HisBreast dataset. As a result, the proposed approach reached a surprise result based on two calculations with the highest scores from BLEU-1 to BLEU-4 and Rouge-L at 0.65, 0.69, 0.72, 0.73, and 0.71. The study also demonstrated the power of CNN and Transformer with Cross Attention architecture in breast cancer image captioning in the INBreast dataset, creating an effective pipeline in the breast cancer diagnosis process.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BCCNetAttention: Enhancing Breast Cancer Report Through Image Captioning with Convolution Neural Network and Transformers Architecture

  • Huong Hoang Luong,
  • Nguyen Thai-Nghe,
  • Hai Thanh Nguyen

摘要

Artificial Intelligence (AI) is transforming the field of healthcare, which brings advancements to the diagnosis and treatment of various conditions, especially breast cancer. It is considered the most dangerous illness among women. Thus, AI technologies including machine learning, deep learning, and advanced data analytics offer promising solutions to this sickness. The primary goal of this study is to develop systems called BCCNetAttention (Breast Cancer Captioning with ResNet50 and Cross Attention) that can automatically interpret and describe the content of images in a way that is meaningful and coherent to humans. In more detail, this research employed the Convolutional Neural Network (CNN) model for feature extraction and the Transformer architecture with Future Masked Language Modeling (FMLM) to train the model in the HisBreast dataset. As a result, the proposed approach reached a surprise result based on two calculations with the highest scores from BLEU-1 to BLEU-4 and Rouge-L at 0.65, 0.69, 0.72, 0.73, and 0.71. The study also demonstrated the power of CNN and Transformer with Cross Attention architecture in breast cancer image captioning in the INBreast dataset, creating an effective pipeline in the breast cancer diagnosis process.